Analysis method and device of EAP hearing analyzer and analyzer
By using the EAP auditory analysis instrument to conduct psychological assessments using voice information, identifying phoneme boundaries and extracting features, and combining them with an emotion classification model, a scientific psychological analysis report is generated. This solves the problem of strong subjectivity in existing psychological assessment methods and achieves more accurate psychological assessment results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA PETROLEUM & CHEMICAL CORP
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-15
AI Technical Summary
Existing psychological assessment methods mainly rely on answering questions and scoring, which are highly subjective and easily influenced by self-judgment and time constraints, resulting in assessment results that are not objective and accurate enough.
Using an EAP auditory analysis instrument, the system acquires speech information, identifies phoneme boundaries, extracts the duration and amplitude features of speech segments, and combines them with an emotion classification model to generate a scientific and objective psychological analysis report. The system utilizes the fusion feature vector of speech segments and preset keywords for psychological assessment, reducing subjectivity and improving the accuracy of the evaluation.
It achieves objectivity and accuracy in psychological assessment results, reduces subjective interference from assessors, and provides more accurate personalized psychological analysis guidance.
Smart Images

Figure CN122030968A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent assessment technology for mental health, and more specifically, to an analysis method, device, and analyzer for an EAP auditory analysis instrument. Background Technology
[0002] Employee Assistance Program (EAP), also known as Employee Assistance Program or Total Employee Psychological Management Technique (hereinafter referred to as EAP), is a systematic and long-term welfare and support program established by companies for their employees. Through professional diagnosis and advice to the organization, and by providing professional guidance, training, and counseling to employees and their immediate family members, EAP aims to help resolve various psychological and behavioral problems of employees and their family members, thereby improving employee performance within the company.
[0003] Current psychological assessments mostly rely on question-and-answer scoring, a method highly subjective and susceptible to interference from self-judgment and time constraints, leading to significant biases in the assessment results. In contrast, professional diagnosis and recommendations to organizations, as well as professional guidance, training, and counseling for employees and their immediate family members, require highly qualified professionals and are less efficient. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing psychological assessments, which mostly rely on question-and-answer scoring. This method is highly subjective and easily influenced by self-judgment and time constraints, resulting in subjective psychological assessment results. This invention provides an analysis method, device, and analyzer for an EAP auditory perception analyzer, mainly used to generate scientific, objective, and accurate psychological analysis reports. This provides individuals with more accurate guidance and targeted solutions to their psychological problems.
[0005] In a first aspect, the present invention provides an analysis method for an EAP hearing analyzer, which employs a psychological assessment model. The psychological assessment model has several dimensions, each containing several quantitative questions. Each quantitative question is assigned at least one preset keyword. The analysis method includes the following steps:
[0006] S1. Acquire speech information; identify the phoneme boundaries of the speech information, and divide the speech information into several speech segments according to the phoneme boundaries;
[0007] S2. Extract the duration and amplitude features of each speech segment from the speech information, combine the duration and amplitude features of each speech segment to form a fusion feature vector of each speech segment, input the fusion feature vector of the speech segment into the emotion classification model, and obtain the emotion quantification value of each speech segment from the emotion classification model.
[0008] S3. Based on the duration characteristics of each speech segment, convert the speech information into text segments corresponding to the speech segments, and combine all text segments to form a text.
[0009] S4. Traverse the text, find the text fragments corresponding to the preset keywords, and then classify them into the corresponding sub-pattern of the quantitative problem corresponding to the preset keywords.
[0010] S5. Match the text fragments in each sub-pattern with the emotion quantification value;
[0011] S6. Perform sentiment analysis and semantic correlation analysis on the text fragments in each sub-pattern to determine whether the quantitative problem corresponding to the sub-pattern is positive sentiment, negative sentiment, or neutral sentiment.
[0012] If the quantitative question corresponding to the sub-pattern is determined to have a positive sentiment tendency, the score for the quantitative question corresponding to the sub-pattern is 1. If the quantitative question corresponding to the sub-pattern is determined to have a negative sentiment tendency, the score for the quantitative question corresponding to the sub-pattern is -1. If the quantitative question corresponding to the sub-pattern is determined to have a neutral sentiment tendency, then based on the matching results in S5, it is determined whether the emotional quantitative value corresponding to the text fragment in the sub-pattern is higher than the defined level. If the emotional quantitative value corresponding to the text fragment in the sub-pattern is higher than the defined level, the score for the quantitative question corresponding to the sub-pattern is 1. If the emotional quantitative value corresponding to the text fragment in the sub-pattern is equal to the defined level, the score for the quantitative question corresponding to the sub-pattern is 0. If the emotional quantitative value corresponding to the text fragment in the sub-pattern is lower than the defined level, the score for the quantitative question corresponding to the sub-pattern is -1.
[0013] S7. By judging the sum of the scores of all quantitative questions for each dimension, if the sum of the scores for that dimension is positive, then the dimension performs well; if the sum of the scores for that dimension is 0, then the dimension performs well; if the dimension is negative, then the dimension performs poorly.
[0014] S8. Automatically generate psychological reports based on performance in each dimension.
[0015] This invention provides an analysis method for an EAP (Employment Assessment Program) auditory analysis instrument. The voice information can be dialogue information, which the person being evaluated (e.g., an employee) is unaware of during the evaluation. After acquiring the voice information, by identifying the phoneme boundaries, several voice segments can be determined. Processing these segments yields the duration and amplitude features of each segment. Combining these features into a fused feature vector, and using an emotion classification model, a quantitative emotion value for each voice segment can be obtained. This serves as an auxiliary tool for quantifying issues within the dimensions of a psychological assessment model. Furthermore, after converting the voice information into text segments corresponding to the voice segments and forming text, the text is traversed, and text segments corresponding to preset keywords are associated with sub-patterns. This yields the text segments and their emotional quantification values within each sub-pattern. This process ensures that text segments, voice segments, sub-patterns, and emotional quantification values are all associated with the quantification problem. Together, they form a corresponding relationship. By performing sentiment analysis and semantic correlation analysis on the text fragments in each sub-pattern, it is possible to initially determine whether the quantitative question corresponding to the sub-pattern has a positive, negative, or neutral sentiment tendency. When the sub-pattern has a neutral sentiment tendency, it can be further judged through the emotional quantification value, thus making the sentiment tendency judgment of the quantitative question corresponding to the sub-pattern more accurate. After obtaining the accurate sentiment tendency of all sub-patterns in each dimension, the sentiment tendency of that dimension can be obtained as the expression of that dimension. This method can directly analyze and evaluate mental health based on the voice information of the personnel (employees) who need psychological assessment, without requiring the personnel to answer subjective questions. This can reduce the subjectivity of the personnel who need psychological assessment and make the obtained psychological assessment results more objective. Furthermore, by using preliminary text judgment and voice emotion to assist in judgment, the obtained psychological assessment results are more accurate.
[0016] Preferably, step S1 includes:
[0017] S1.1 Directional radio deployment: Multiple directional radios are deployed in the radio reception area to capture voice information.
[0018] S1.2 Audio signal preprocessing: Digital filtering technology is used to sequentially reduce noise, adjust gain, and cancel echo in the captured audio information.
[0019] S1.3 Data standardization and formatting: Transform the pre-processed audio signal data into a unified standard format;
[0020] S1.4 Identify the phoneme boundaries of speech information and divide the speech information into several speech segments based on the phoneme boundaries;
[0021] This also includes data privacy protection, which involves storing all voice information data obtained in S1.1 to S1.3 and encrypting it during storage and transmission.
[0022] Preferably, step S2 is as follows:
[0023] S2.1 Extract the duration and amplitude features of each speech segment from the speech information;
[0024] The duration characteristics of each speech segment include the average duration of the speech segment, the standard deviation of the speech segment duration, and the average duration of the pause. The duration of each speech segment is determined based on the phoneme boundaries, and the average duration of all speech segments is obtained by averaging the durations of all speech segments. The standard deviation of the speech segment duration is calculated based on the duration of all speech segments and the average duration of the speech segment. The duration of silent segments is determined based on the phoneme boundaries, and the average duration of all silent segments is obtained by averaging the durations of all silent segments.
[0025] The amplitude features of each speech segment include the average amplitude, standard deviation of the amplitude, and duration of energy distribution above a threshold. The short-time energy of each frame is calculated, and the average amplitude of the speech segment is obtained by averaging the values. The amplitude envelope is analyzed, and peaks and troughs are extracted to calculate the standard deviation of the amplitude. Segments with short-time energy above a threshold are detected, and their durations are recorded to obtain the duration of energy distribution above the threshold, where the threshold is [value missing]. Full audio clip Twice the average short-term energy ;
[0026] S2.2 Combine the six features of speech segment average duration, speech segment duration standard deviation, pause average duration, speech segment average amplitude, speech segment amplitude standard deviation, and speech segment energy distribution duration above the threshold into a fusion feature vector in order.
[0027] S2.3. Input the fused feature vector of the speech segments into the emotion classification model, and obtain the emotion quantification value of each speech segment from the emotion classification model, including the following steps:
[0028] The fused feature vectors are standardized and normalized.
[0029] The fused feature vectors after standardization and normalization are divided into training and testing sets;
[0030] The emotion classification model is trained and evaluated using training and test sets. The emotion classification model includes support vector machines, K-nearest neighbors, and neural networks.
[0031] The standardized and normalized fused feature vectors are input into the trained emotion classification model. The emotion classification model classifies the standardized and normalized fused feature vectors and gives the emotion quantification value of each speech segment.
[0032] Preferably, in step S1, the number of times the signal crosses the zero axis within each frame is calculated, the zero-crossing rate (ZCR) is calculated based on the number of times the signal crosses the zero axis within each frame, and the boundaries of phonemes in the speech signal are determined based on the ZCR. The formula for calculating the ZCR is:
[0033]
[0034] Where, x t Represents the amplitude at the t-th sampling point, 1{sign(x t )≠sign(x t+1 )} is the indicator function, T is the number of sampling points in the frame, and sign function sign(x) t This is used to determine the sign of the signal at the t-th sampling point;
[0035] In step S2.1,
[0036] The Short Time Energy (STE) of each frame of a speech segment is obtained. The formula for calculating the Short Time Energy (STE) is as follows:
[0037]
[0038] Where w(n) is the window function, x(n) is the signal sample, and N is the window length, which represents the number of sampling points of the signal. N is a natural number. n is the index of the signal sampling point;
[0039] The signal of each frame of the speech segment is converted to the frequency domain to obtain the amplitude X(k) of the frequency component. The amplitude spectrum is extracted based on the amplitude X(k) of the frequency component. The formula for obtaining the amplitude X(k) of the frequency component is as follows:
[0040]
[0041] Where X(k) is the amplitude of the frequency component, k is the frequency index, and Y is the number of points in the FFT. n is the index of the signal sampling point, x(n) is the signal sample, e is the complex exponent kernel, j is the imaginary unit, and 2πkn / Y represents the angle or phase.
[0042] In step S2.1, peaks and troughs are identified through amplitude envelope detection. Dilation and erosion operations in morphological calculations are applied to calculate the duration and amplitude differences between peaks. The formula for the dilation operation is:
[0043]
[0044] The formula for corrosion calculation is:
[0045]
[0046] in, This represents the expansion operation. This represents the erosion operation, where f is the input signal, b is the structuring element, and B is the domain of the structuring element.
[0047] Preferably, step S4 includes the following steps:
[0048] S4.1 Text preprocessing: Remove stop words and punctuation marks from the text, and simplify the text using lemmatization or stemming.
[0049] S4.2. Use TF-IDF (Term Frequency-Inverse Document Frequency) or word embedding models (such as Word2Vec or BERT) to calculate the similarity between the text fragment and the preset keyword list. The similarity calculation formula is as follows:
[0050]
[0051] Where A is a text fragment and B is a preset keyword vector;
[0052] S4.3 For each sub-pattern, find the text segment that best matches the preset keyword and classify it into the sub-pattern corresponding to the preset keyword.
[0053] Preferably, in step S3, when forming the text, the timestamp information of each text segment is retained;
[0054] Step S5 includes the following steps:
[0055] By using the timestamp information of the text segment in step S3 to associate with the timestamp information of the corresponding audio segment in step S2, the emotion quantization value data of the audio segment under the same timestamp information is obtained as the emotion quantization value data of the text segment, thus forming a match.
[0056] Preferably, step S6 includes the following steps:
[0057] S6.1. Text fragment preprocessing for each sub-pattern: including removing redundant information and part-of-speech tagging;
[0058] Remove redundant information: Use NLTK or spaCy to segment the text fragments of each sub-pattern and remove stop words, and perform lemmatization and stemming to unify the word forms;
[0059] Part-of-speech tagging: The part-of-speech tagging of each word in the text fragment of each subpattern is marked using the NLTK's pos_tag() function;
[0060] S6.2, Preset Keyword Recognition and Matching: Includes sentiment dictionary matching and context analysis:
[0061] Sentiment dictionary matching: Using a predefined sentiment dictionary, the preset keywords of each sub-pattern are matched with their corresponding sentiment scores, and the similarity between the preset keywords and sentiment words is calculated, using cosine similarity as the metric. Where A and B are the vectors of the text fragments and preset keywords corresponding to the sub-patterns, respectively;
[0062] Contextual analysis: Using VADER or TextBlob sentiment analysis tools, the sentiment tendency of preset keywords is further judged in the context of the text fragments in the corresponding sub-patterns, and the sentiment scores of sentences are extracted from the text fragments. The overall sentiment tendency of the sub-pattern is calculated by combining the weights of the preset keywords.
[0063] S6.3, Assigning Sentiment Scores: Assign a sentiment score to each preset keyword: positive keywords are assigned +1, negative keywords are assigned -1, and neutral keywords are assigned 0; if a preset keyword appears in multiple contexts, all sentiment tendencies are considered comprehensively.
[0064] S6.4. Based on the assignment of sentiment scores of the preset keywords in S6.3, combined with the quantitative questions of the corresponding sub-patterns in the psychological assessment model, the sentiment tendency of the quantitative questions corresponding to the sub-patterns is judged. If the quantitative questions corresponding to the sub-patterns are judged to have a positive sentiment tendency, the quantitative questions corresponding to the sub-patterns are judged to have a score of 1. If the quantitative questions corresponding to the sub-patterns are judged to have a negative sentiment tendency, the quantitative questions corresponding to the sub-patterns are judged to have a score of -1. If the quantitative questions corresponding to the sub-patterns are judged to have a neutral sentiment tendency, then according to the matching results in S5, it is judged whether the sentiment quantitative value corresponding to the text fragments in the sub-patterns is higher than the defined level. If the sentiment quantitative value corresponding to the text fragments in the sub-patterns is higher than the defined level, the quantitative questions corresponding to the sub-patterns are judged to have a score of 1. If the sentiment quantitative value corresponding to the text fragments in the sub-patterns is equal to the defined level, the quantitative questions corresponding to the sub-patterns are judged to have a score of 0. If the sentiment quantitative value corresponding to the text fragments in the sub-patterns is lower than the defined level, the quantitative questions corresponding to the sub-patterns are judged to have a score of -1.
[0065] The horizontal line is defined as (highest emotional quantification value + lowest emotional quantification value) / 2.
[0066] Preferably, in step S4, if the number of sub-patterns containing text fragments in each dimension accounts for 60% or more of all sub-patterns in that dimension, then the performance evaluation of that dimension is valid in step S7; otherwise, it is invalid.
[0067] In S8, the validity of the dimensional performance evaluation is presented in writing.
[0068] In a second aspect, the present invention provides an EAP hearing analysis device, comprising at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the analysis method of any of the above-described EAP hearing analyzers.
[0069] Thirdly, the present invention provides an EAP hearing analyzer, comprising:
[0070] The data acquisition module acquires voice information;
[0071] The speech intonation analysis module extracts the speech information collected by the data acquisition module, divides the speech information into several speech segments, and obtains the emotion quantification value of each speech segment based on the fundamental frequency, energy and morphological characteristics of the speech segments.
[0072] The speech-to-text module converts speech information into text segments corresponding to the speech segments, and all text segments form a text file.
[0073] The keyword recognition module traverses the text and matches the text fragments corresponding to preset keywords with the quantitative questions in various dimensions of the psychological assessment model;
[0074] The psychological report generation module matches text fragments with the emotional quantification values of corresponding audio fragments; then, based on the text fragments and emotional quantification values that match the quantitative questions, it obtains the answer tendencies for the quantitative questions; it integrates the answer tendencies of all quantitative questions for that dimension to obtain the performance of that dimension; it automatically generates psychological reports based on the performance of each dimension; and it provides data visualization interpretation of the generated psychological reports.
[0075] The data storage and management module saves all data generated by the data acquisition module, speech-to-text module, speech intonation analysis module, and psychological report generation module.
[0076] This invention provides an EAP (Employee Assistance Program) auditory analysis instrument that can directly analyze and assess mental health based on the voice information of employees (those requiring psychological evaluation), without requiring subjective answers. This reduces subjectivity during the evaluation process, resulting in more objective and accurate assessment results. Furthermore, by using textual information and voice emotion to judge the quantitative questions corresponding to each sub-mode, the accuracy of the judgment results is improved. A data storage and management module is employed to save all data generated by the data acquisition module, speech-to-text module, voice intonation analysis module, and psychological report generation module. This provides storage and encryption services for these modules, facilitating confidentiality and subsequent verification.
[0077] Preferably, the data storage and management module locally saves all data generated by the data acquisition module, speech-to-text module, speech intonation analysis module, and psychological report generation module;
[0078] The data storage and management module uses the non-relational database MongoDB to store different types of data; and employs the AES-256 encryption algorithm in all data storage processes.
[0079] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0080] 1. This invention provides an analysis method for an EAP (Employee Assistance Program) auditory analysis instrument, which can directly analyze and evaluate mental health based on the voice information of the personnel (employees) who need psychological assessment. It eliminates the need for the personnel to answer questions subjectively, reduces the subjectivity of the psychological assessment, and makes the obtained psychological assessment results more objective. Furthermore, by using preliminary textual judgment and voice emotion-assisted judgment, the obtained psychological assessment results are more accurate.
[0081] 2. The present invention provides an EAP hearing analysis device capable of performing the analysis method of an EAP hearing analyzer.
[0082] 3. This invention provides an EAP (Employee Assistance Program) auditory analysis instrument that can directly analyze and assess mental health based on the voice information of employees (those requiring psychological evaluation), without requiring subjective answers from the participants. This reduces subjectivity during the evaluation process, resulting in more objective and accurate assessment results. Furthermore, by using textual information and voice emotion to judge the quantitative questions corresponding to each sub-mode, the accuracy of the judgment results is improved. A data storage and management module is employed to save all data generated by the data acquisition module, speech-to-text module, voice intonation analysis module, and psychological report generation module. This provides storage and encryption services for these modules, facilitating confidentiality and subsequent verification. Attached Figure Description
[0083] Figure 1 This is a schematic diagram of the components of the EAP hearing analyzer. Detailed Implementation
[0084] The present invention will be further described in detail below with reference to experimental examples and specific embodiments. However, this should not be construed as limiting the scope of the above-mentioned subject matter of the present invention to the following embodiments; all technologies implemented based on the content of the present invention fall within the scope of the present invention.
[0085] Example 1
[0086] This embodiment provides an analysis method for an EAP hearing analyzer, employing a psychological assessment model. The model has several dimensions, each containing several quantitative questions. Each quantitative question is assigned at least one preset keyword. The analysis method includes the following steps:
[0087] S1. Acquire speech information; identify the phoneme boundaries of the speech information, and divide the speech information into several speech segments based on the phoneme boundaries; among which, acquiring speech information can be done by recording, and the specific recording method can be selected according to the actual situation; the speech segment can be a word, phrase, sentence or paragraph.
[0088] Optionally, step S1 includes:
[0089] S1.1 Directional radio deployment: Multiple directional radios are deployed in the radio reception area to capture voice information.
[0090] S1.2 Audio signal preprocessing: Digital filtering technology is used to sequentially reduce noise, adjust gain, and cancel echo in the captured audio information.
[0091] S1.3 Data standardization and formatting: Transform the pre-processed audio signal data into a unified standard format;
[0092] S1.4 Identify the phoneme boundaries of speech information and divide the speech information into several speech segments based on the phoneme boundaries;
[0093] This also includes data privacy protection, which involves storing all voice information data obtained in S1.1 to S1.3 and encrypting it during storage and transmission.
[0094] The above methods can obtain clearer voice information and can also encrypt the voice information.
[0095] Optionally, in step S1, the number of times the signal crosses the zero axis within each frame is calculated, the zero-crossing rate (ZCR) is calculated based on the number of times the signal crosses the zero axis within each frame, and the boundaries of phonemes in the speech signal are determined based on the ZCR. The formula for calculating the ZCR is:
[0096]
[0097] Where, x t Represents the amplitude at the t-th sampling point, 1{sign(x t )≠sign(x t+1 )} is the indicator function, T is the number of sampling points in the frame, and sign function sign(x) t The zero-crossing rate (ZCR) is used to determine the symbol of the signal at the t-th sampling point. The ZCR is obtained by calculating the number of times the signal crosses the zero axis within each frame. A high ZCR is usually associated with silent segments, while a low ZCR is associated with sound segments. A high ZCR indicates a pause, meaning the segment is silent.
[0098] High zero-crossing rate (ZCR) refers to the frequency with which a signal crosses the zero axis within a short period of time. Specifically, the zero-crossing rate is a measure of the rate of change of a signal, representing the number of times a signal changes from a positive to a negative value or vice versa within a frame. Typically, silent or noisy segments of a speech signal have a higher zero-crossing rate, while speech or phoneme segments usually have a lower zero-crossing rate.
[0099] The threshold for zero crossover rate is not fixed and is typically determined based on the specific application scenario, the signal sampling rate, and the type of speech to be processed. However, a typical threshold range might be as follows:
[0100] • For segments with sound (such as vowels, voiced sounds, etc.): the zero crossover rate is usually low, possibly between 10% and 20%.
[0101] • For silent or noisy segments: the zero crossover rate is high, possibly between 30% and 50% or even higher.
[0102] • The specific threshold will be adjusted based on experimental data. In this speech processing module, if the zero crossover rate exceeds the set value of 30%, the signal segment is determined to be a silent segment or noise.
[0103] The start and end of a speech segment can be determined using the zero crossover rate. Furthermore, the accuracy of factor boundaries can be adjusted using dynamic thresholding. Moreover, phoneme boundaries can be determined collaboratively based on the zero crossover rate and short-time energy (STE) to extract the duration of each speech segment.
[0104] Optional, a simplified process for determining phoneme boundaries and extracting duration using zero cross-rate (ZCR) and short-time energy (STE) in collaboration:
[0105] Zero cross rate (ZCR) detection:
[0106] (1) Calculate the zero cross rate for each frame to identify the frequency of signal changes.
[0107] (2) Here, a set value of 30% is used as the threshold. The part with a lower ZCR usually corresponds to a speech segment, while the part with a higher ZCR may be a silent segment or high-frequency noise (high noise is also equivalent to silent and is processed).
[0108] Short-Time Energy (STE) Detection:
[0109] (1) Calculate the short-time energy of each frame to distinguish between audio and silent segments.
[0110] (2) Here, twice the average value of STE is used as the threshold - the high-energy part is usually a speech segment, and the low-energy part is usually a silent segment.
[0111] Collaborative judgment of phoneme boundaries:
[0112] (1) High STE energy + low ZCR: determined as the start or continuation of a speech segment.
[0113] (2) Low STE energy + high ZCR: This indicates the end of a silent segment or a speech segment.
[0114] (3) Combine the changes in STE and ZCR to determine the start and end points of the phonemes.
[0115] Combined with the specific process of judgment:
[0116] 1) Initial detection: High-energy regions are initially identified through STE detection, and ZCR is used to determine whether these high-energy regions are continuous speech segments or short unvoiced segments.
[0117] 2) Boundary adjustment (initial determination of starting point index): For continuous high STE+low ZCR regions, they can be marked as the starting point of speech segments or continuous vocalization parts.
[0118] 3) Boundary adjustment (initial determination of the end point index): When the transition from high STE + low ZCR to low STE + high ZCR is marked as the end point of the speech segment.
[0119] 4) Boundary refinement (high noise included in speech): Combine high STE+high ZCR segments to further identify and separate unvoiced sounds and noise, avoiding misjudging high-frequency noise as speech segments.
[0120] 5) Boundary refinement (stable silence): The double low case can usually be stably judged as a silent segment, which helps to adjust and refine the boundary judgment and ensure accurate identification of the start and end points of speech.
[0121] Duration of extracted speech segments:
[0122] (1) Calculate the duration of each speech segment based on the detected phoneme boundaries.
[0123] (2) Formula for calculating the duration of a speech segment: The duration of each speech segment can be obtained by subtracting the start frame index from the end frame index of the phoneme segment, and then dividing by the sampling rate: that is...
[0124] Speech segment duration = (end frame index - start frame index) / sampling rate.
[0125] S2. Extract the duration and amplitude features of each speech segment from the speech information, combine the duration and amplitude features of each speech segment to form a fusion feature vector of each speech segment, input the fusion feature vector of the speech segment into the emotion classification model, and obtain the emotion quantification value of each speech segment from the emotion classification model.
[0126] The duration feature can include the duration of each speech segment to reflect the speech rate and rhythm of the speech flow; the amplitude feature can analyze the amplitude changes of the speech segment signal to assess the dynamic range and intensity changes of the speech signal of the speech segment. The purpose of fusing the duration and amplitude features of each speech segment is to combine the duration and amplitude features extracted from the speech segment into a multi-dimensional vector so that it can be input into the emotion classification model.
[0127] Optionally, step S2 is:
[0128] S2.1 Extract the duration and amplitude features of each speech segment from the speech information;
[0129] The duration features of each speech segment include the average duration of the speech segment, the standard deviation of the speech segment duration, and the average duration of pauses. The duration of each speech segment is determined based on phoneme boundaries, and the average duration of all speech segments is obtained by averaging the durations of all speech segments. The standard deviation of the speech segment duration is calculated based on the duration of all speech segments and the average duration of the speech segments. The duration of silent segments is determined based on phoneme boundaries, and the average duration of all silent segments is obtained by averaging the durations of all silent segments. This allows for the detection of the duration and distribution of silent segments.
[0130] The amplitude characteristics of each speech segment include the average amplitude of the speech segment, the standard deviation of the speech segment amplitude, and the duration of the energy distribution above the threshold. The short-time energy of each frame is calculated, and the average amplitude of the speech segment is obtained by averaging. The amplitude envelope is analyzed, and the peaks and troughs are extracted to calculate the standard deviation of the speech segment amplitude, which is the amplitude fluctuation. Segments with short-time energy above the threshold are detected, and their duration is counted to obtain the duration of the energy distribution above the threshold of the speech segment. The threshold is twice the average short-time energy of the complete speech segment.
[0131] Optionally, in step S2.1, the Short-Time Energy (STE) of each frame of the speech segment is obtained. The formula for calculating the Short-Time Energy (STE) is as follows:
[0132]
[0133] Where w(n) is the window function, x(n) is the signal sample, and N is the window length [the window length represents the number of signal sampling points; under normal circumstances, N is a natural number, therefore...]. [(Set of Natural Numbers)], where n is the index of the signal sampling point;
[0134] Short-time energy (STE) measures the signal energy within each frame to determine sound intensity and dynamic changes. STE calculations are used to measure the energy of each frame of the signal to detect changes in speech intensity. By acquiring the STE of each frame of a speech segment, the amplitude value for each frame can be obtained; variations in amplitude are related to the degree of emotional arousal, with larger amplitude changes potentially indicating emotional excitement.
[0135] The speech segment's signal frame is converted to the frequency domain using Short Time Fourier Transform (STFT) to obtain the amplitude X(k) of the frequency component. The amplitude spectrum is then extracted from the amplitude X(k). The formula for obtaining the amplitude X(k) of the frequency component is as follows:
[0136]
[0137] Where X(k) is the amplitude of the frequency component, and Y is the number of points in the FFT. (Set of Natural Numbers):
[0138] Specifically: X(k): the amplitude of the frequency component (often also called the coefficient of the Fourier transform), representing the strength of the signal at the k-th frequency.
[0139] k: Frequency index, representing the kkk-th frequency component in the frequency domain, with a value range of k = 0, 1, ..., Y-1.
[0140] Y: The number of points in the Fourier transform (also known as the number of points in the FFT), which is the number of frequencies segmented in the frequency domain. Y∈N is the set of natural numbers.
[0141] n: The index of the signal sampling point, representing the nth sample in the time domain.
[0142] x(n): Signal sample, that is, the amplitude of the signal at the nth sampling point, i.e. the signal value in the time domain.
[0143] e: Complex exponent kernel, representing the frequency components of the signal in the frequency domain, transformed from the time domain to the frequency domain by Fourier transform. Essentially, it is a complex number representing a point on the unit circle with a phase of -2πkn / Y.
[0144] j: is the imaginary unit.
[0145] 2πkn / Y: Represents the angle or phase, which determines the frequency of the sine wave. Specifically, it is the phase of the k-th frequency component at the n-th time point.
[0146] By obtaining the amplitude spectrum, i.e. the amplitude line shape, all amplitude values can be concatenated into a line graph; the amplitude information in the frequency domain is extracted, providing a foundation for subsequent emotion quantification analysis.
[0147] In step S2.1, peaks and troughs are identified through amplitude envelope detection. Dilation and erosion operations in morphological calculations are applied to calculate the duration and amplitude differences between peaks. The formula for the dilation operation is:
[0148]
[0149] The formula for corrosion calculation is:
[0150]
[0151] Where f is the input signal and b is the structuring element.
[0152] In the above dilation and auxiliary operations, 1. the dilation operation is used to identify the peaks in the signal and mark the positions of these peaks.
[0153] 2. Erosion operation is used to identify troughs in the signal and mark the locations of these troughs.
[0154] 3. The duration between peaks is obtained by calculating the time difference between adjacent peaks.
[0155] 4. Amplitude difference is obtained by calculating the amplitude difference between adjacent peaks and troughs.
[0156] The symbols are explained in detail below:
[0157] Circle plus sign Dilation is a morphological operation that expands the peak region of an input signal by sliding a structuring element b across the signal f(x) and taking the maximum value within its local neighborhood. Dilation enlarges the local maxima in a signal and is typically used to enhance or amplify bright spots or peak regions in images or signals.
[0158] Circle minus sign Erosion is a morphological operation that shrinks low-value regions of an input signal by sliding a structuring element b across the input signal f(x) and taking the minimum value within a local neighborhood. Erosion reduces the size of local minimum regions in a signal and is typically used to weaken or reduce dark spots or valleys in images or signals.
[0159] (x) or (x) refers to the processing of f(x) through dilation. Erosion operation We performed morphological processing on the signal f(x). f(x) is the input signal, representing its amplitude at position x.
[0160] (1) Expansion operation (x): Calculate the maximum value (i.e., max) of the signal f(x) at position x and its neighborhood (determined by the structuring element b), thereby expanding the peak region of the signal.
[0161] (2) Erosion operation (x): Calculate the minimum value (i.e., min) of the signal f(x) at position x and its neighborhood, thereby shrinking the low-value region of the signal.
[0162] b: Structural element, the same as the structural element used in the dilation operation.
[0163] B: The domain of the structuring element.
[0164] For example:
[0165] At a sampling rate of 96kHz (program hardware setting), we chose 5 sampling points as the appropriate symmetric window size for the structuring element b. This preserves the details of the signal while smoothing out possible noise. That is, b∈B and B=[-2,-1,0,1,2].
[0166] In the definition of the structuring element b, the numbers -2, -1, 0, 1, and 2 are not actual numerical values, but rather represent the relative position of the structuring element during signal processing. These numbers indicate the offset of the structuring element relative to the current processing point. For example, when b = [-2, -1, 0, 1, 2], it represents the range of the current sampling point and the two sampling points before and after it.
[0167] Practical applications: -2 indicates shifting forward by two sampling points (i.e., an earlier point in time). -1 indicates shifting forward by one sampling point (i.e., slightly earlier than the current point). 0 represents the current sampling point. 1 indicates shifting backward by one sampling point (i.e., slightly later than the current point). 2 indicates shifting backward by two sampling points (i.e., a later point in time).
[0168] Example: Suppose the signal value corresponding to the current sampling point x is f(x). In signal processing, if the structuring element b is defined as [-2, -1, 0, 1, 2], then for f(x), we will simultaneously consider the values of five sampling points: f(x-2), f(x-1), f(x), f(x+1), and f(x+2). That is, the dilation operation will look for the maximum value within the range of these five points, while the erosion operation will look for the minimum value within the range of these five points.
[0169] Corrosion calculation process:
[0170] (1) For each position x in the signal f(x), the erosion operation slides the structuring element b in the neighborhood of x and calculates the minimum value of f(x+b).
[0171] (2)x+b represents the position of the structuring element b on the signal f.
[0172] (3) The result is that the local minimum of the original signal is extended to every point in the neighborhood, thereby “eroding” the local valley of the signal.
[0173] The sharp areas of the amplitude line graph are smoothed by dilation and erosion operations, and morphological features are extracted.
[0174] S2.2 Combine the six features of speech segment average duration, speech segment duration standard deviation, pause average duration, speech segment average amplitude, speech segment amplitude standard deviation, and speech segment energy distribution duration above the threshold into a fusion feature vector in order.
[0175] For example, the duration and amplitude features of each speech segment are combined to form a fused feature vector for each speech segment as F = [tmean, tstd, pmean, amean, astd, ehigh], where each symbol represents the average duration of the speech segment, the standard deviation of the speech segment duration, the average duration of the pause, the average amplitude of the speech segment, the standard deviation of the speech segment amplitude, and the duration of the energy distribution of the speech segment above the threshold, respectively. Furthermore, the feature dimensions are standardized to eliminate dimensional differences.
[0176] S2.3. Input the fused feature vector of the speech segments into the emotion classification model, and obtain the emotion quantification value of each speech segment from the emotion classification model, including the following steps:
[0177] The fused feature vectors are standardized and normalized.
[0178] The fused feature vectors after standardization and normalization are divided into training and test sets to evaluate the performance of the sentiment classification model.
[0179] The emotion classification model is trained and evaluated using training and test sets. The emotion classification model includes support vector machines, K-nearest neighbors, and neural networks.
[0180] Support Vector Machines (SVMs) are suitable for binary or multi-class classification problems with small datasets. Their advantages include the ability to handle non-linear data through kernel functions (such as the RBF kernel), resulting in good classification performance.
[0181] K-Nearest Neighbors (KNN), Applicability: Suitable for pattern recognition and classification tasks, especially when the amount of training data is small. Advantages: Simple and easy to use, no model training is required, and classification can be performed directly using similarity metrics (such as Euclidean distance).
[0182] Neural networks (such as CNN and RNN) are suitable for classification problems involving large-scale datasets and complex features. Their advantages include the ability to capture complex nonlinear relationships through deep structures, making them suitable for multimodal sentiment analysis.
[0183] For SVM: Choose an appropriate kernel function (such as the RBF kernel) and adjust hyperparameters (such as the penalty coefficient C and the kernel parameter gamma) to improve classification accuracy. For KNN: Choose an appropriate number of neighbors K and use cross-validation to select the optimal K value.
[0184] The standardized and normalized fused feature vectors are input into the trained emotion classification model. The emotion classification model classifies the standardized and normalized fused feature vectors and gives the emotion quantification value of each speech segment.
[0185] Emotion Recognition and Quantification: Input the standardized fused feature vector F into the trained emotion classification model, and output the emotion classification result through the emotion classification model, mapping each emotion to a quantization value. The emotion quantization value can be a continuous value between 0 and 1, which can correspond to discrete emotion labels (such as anger, joy, sadness, etc., in this embodiment, only the quantization data value is used to represent). Use the emotion quantization value as the final output to identify the intensity of different emotion states. For example, an output emotion quantization value of 0.75 indicates that the emotion state is biased towards "agreement, pleasure", and the closer the value is to 1, the stronger the emotion.
[0186] In addition, a visualization chart of the emotion quantization value has been generated here, which will be shown in the finally exported EAP work assistance plan report to facilitate understanding of the emotion change trend. For example: Generate a visualization chart of emotion quantization through Matplotlib or Plotly.
[0187] S3. Convert the voice information into text segments corresponding to the voice segments according to the duration characteristics of each voice segment, and form a text after integrating all text segments;
[0188] Optionally, convert the preprocessed voice signal into overall text information that can distinguish multiple-person conversations in real time, display it in the software window, and allow EAP ideological and political workers to manually correct possible incorrect information;
[0189] Optionally, when forming the text in step S3, retain the timestamp information of each text segment for subsequent matching according to the timestamp information;
[0190] S4. Traverse the text, find the text segments corresponding to the preset keywords, and then classify them into the corresponding sub-patterns of the quantization questions corresponding to the preset keywords;
[0191] Optionally, step S4 includes the following steps:
[0192] S4.1. Preprocess the text: Remove the stop words and punctuation marks in the text, and simplify the text using word form reduction or stemming. In addition, the text can also be tokenized, and natural language processing techniques are used to decompose the text into words or phrases. Then remove the stop words: Remove common but meaningless words (such as "de", "shi", etc.). For example: Use NLTK or spaCy in Python for text preprocessing and preset keyword recognition.
[0193] S4.2. Calculate the similarity between the text segments and the preset keyword list using TF-IDF (Term Frequency - Inverse Document Frequency) or word embedding models (such as Word2Vec or BERT). The similarity calculation formula is:
[0194]
[0195] Where A is a text fragment and B is a preset keyword vector;
[0196] In the above similarity calculation formula:
[0197] Vector A: Represents the vector representation of the text segment. In steps S4.1 and S4.2, the text segment is represented through preprocessing (removing stop words and punctuation) and using Term Frequency-Inverse Document Frequency (TF-IDF) with a word embedding model (such as Word2Vec or BERT). Ultimately, the text segment is represented as a vector A, where each element corresponds to the importance or embedding value of a word in the vocabulary.
[0198] Vector B: Represents the vector representation of the preset keyword list. Similarly, the preset keyword list is also represented as a vector B using the same preprocessing and vectorization methods.
[0199] A·B represents the dot product of vectors A and B, which is the sum of the corresponding elements multiplied together.
[0200] ∥A∥ and ∥B∥ represent the norms (i.e., the lengths of vectors) of vectors A and B, respectively, which are obtained by taking the square root of the sum of the squares of each element.
[0201] The significance of similarity calculation: The value of cosine similarity is between -1 and 1, representing the degree of similarity between two vectors.
[0202] (1) 1: indicates that the direction is exactly the same, meaning that the text fragment is very similar to the list of keywords.
[0203] (2) 0: indicates no similarity (i.e., the vectors are orthogonal).
[0204] (3)-1: indicates the exact opposite direction.
[0205] (4) In text analysis, cosine similarity is used to evaluate the similarity between a text segment (A) and a predefined list of keywords (B). The closer the value is to 1, the higher the relevance of the text to the keyword list.
[0206] The preset keywords are predefined, such as a list of preset keywords for each of the 10 sub-patterns in each psychological dimension. The similarity between a text fragment and the preset keyword list of a sub-pattern is calculated to determine whether the text fragment belongs to that sub-pattern.
[0207] For example, using machine learning libraries (such as scikit-learn) for sentiment classification and similarity calculation.
[0208] S4.3 For each sub-pattern, find the text segment that best matches the preset keyword and classify it into the sub-pattern corresponding to the preset keyword.
[0209] S5. Match the text fragments in each sub-pattern with the emotion quantification value;
[0210] After the timestamp information of the text fragment is preserved in step S3, step S5 includes the following steps:
[0211] By using the timestamp information of the text segment in step S3 to associate with the timestamp information of the corresponding audio segment in step S2, the emotion quantization value data of the audio segment under the same timestamp information is obtained as the emotion quantization value data of the text segment, thus forming a match.
[0212] Specifically, the method for obtaining the sentiment quantification value data of text fragments:
[0213] The emotional quantification information is extracted as follows:
[0214] 1) Input: Obtain the emotional quantification value of all audio segments, along with the corresponding timestamp information.
[0215] Sub-emotion quantification information extraction: Based on the timestamp information of text segments identified by preset keywords, extract the emotion quantification value corresponding to the timestamp information from the emotion quantification values of all speech segments. NumPy or Pandas is used to process the data corresponding to the timestamps.
[0216] 2) Define the horizontal line: Use the average sentiment quantification value of all text fragments in each sub-pattern as the baseline, which is used to evaluate the sentiment tendency of the response results in step S6.
[0217] For example, in the optimal state, the output text information is as follows: 60 text fragments are generated, each corresponding to a sub-pattern, containing the identified preset keywords and sentiment quantification values. The timestamp information of each fragment, along with its corresponding sentiment quantification value and a defined horizontal line, are output to indicate sentiment changes.
[0218] In cases of insufficient recognition, the output text information is as follows: If there are fewer than 60 paragraphs, the effective recognition information in the psychological report will be displayed as follows: Dimensions with a recognition rate below 60% will be marked as having low reference value. Effective recognition information will be output to the EAP work support program report. If the recognition rate for both questioning and response in all 10 sub-patterns of this dimension is above 60%, then this dimension will be marked as having reference value. Dimensions that do not meet the criteria will be listed as items for improvement and output to the report.
[0219] S6. Perform sentiment analysis and semantic correlation analysis on the text fragments in each sub-pattern to determine whether the quantitative problem corresponding to the sub-pattern is positive sentiment, negative sentiment, or neutral sentiment.
[0220] If the quantitative question corresponding to the sub-pattern is determined to have a positive sentiment tendency, the score for the quantitative question corresponding to the sub-pattern is 1. If the quantitative question corresponding to the sub-pattern is determined to have a negative sentiment tendency, the score for the quantitative question corresponding to the sub-pattern is -1. If the quantitative question corresponding to the sub-pattern is determined to have a neutral sentiment tendency, then based on the matching results in S5, it is determined whether the emotional quantitative value corresponding to the text fragment in the sub-pattern is higher than the defined level. If the emotional quantitative value corresponding to the text fragment in the sub-pattern is higher than the defined level, the score for the quantitative question corresponding to the sub-pattern is 1. If the emotional quantitative value corresponding to the text fragment in the sub-pattern is equal to the defined level, the score for the quantitative question corresponding to the sub-pattern is 0. If the emotional quantitative value corresponding to the text fragment in the sub-pattern is lower than the defined level, the score for the quantitative question corresponding to the sub-pattern is -1.
[0221] Optionally, step S6 includes the following steps:
[0222] S6.1. Text fragment preprocessing for each sub-pattern: including removing redundant information and part-of-speech tagging;
[0223] Remove redundant information: Use NLTK or spaCy to segment the text fragments of each sub-pattern and remove stop words, and perform lemmatization and stemming to unify the word forms;
[0224] Part-of-speech tagging: The part-of-speech tagging of each word in the text fragment of each subpattern is marked using the NLTK's pos_tag() function;
[0225] S6.2, Preset Keyword Recognition and Matching: Includes sentiment dictionary matching and context analysis:
[0226] Sentiment dictionary matching: Using a predefined sentiment dictionary, the preset keywords of each sub-pattern are matched with their corresponding sentiment scores, and the similarity between the preset keywords and sentiment words is calculated, using cosine similarity as the metric. Where A and B are the vectors of the text fragments and preset keywords corresponding to the sub-patterns, respectively;
[0227] Contextual analysis: Using VADER or TextBlob sentiment analysis tools, the sentiment tendency of preset keywords is further judged in the context of the text fragments in the corresponding sub-patterns, and the sentiment scores of sentences are extracted from the text fragments. The overall sentiment tendency of the sub-pattern is calculated by combining the weights of the preset keywords.
[0228] Predefined sentiment dictionaries, for example:
[0229] The general sentiment dictionary is as follows;
[0230] (1) Positive sentiment vocabulary (+1)
[0231] Happiness: joy, delight, happiness, bliss; Satisfaction: contentment, fulfillment, satisfaction; Joy: pleasure, excitement, delight; Gratitude: gratitude, thanks, deep appreciation; Optimism: positive, sunny, cheerful, open-minded; Trust: reliance, trust, reliability, confidence; Liking: love, affection, admiration, appreciation; Positivity: initiative, ambition, striving; Peace: tranquility, serenity, serenity; Contentment: satisfaction, contentment, satisfaction; Agreement: agreement, support, recognition, acceptance, acceptance.
[0232] (2) Negative emotional vocabulary (-1)
[0233] Sadness: grief, pain, sorrow, grief; anger: anger, annoyance, irritability, rage; disappointment: despair, frustration, frustration, disappointment; dejection: discouragement, despondency, loss, low spirits; fear: fear, timidity, terror, dread; unease: anxiety, worry, tension, worry; jealousy: envy, admiration, jealousy; hatred: hatred, aversion, loathing, loathing.
[0234] Doubt: suspicion, doubt, disbelief, questioning; despair: hopelessness, despair, loss of hope; denial: opposition, rejection, criticism, questioning, rebuttal.
[0235] (3) Neutral emotional vocabulary (0)
[0236] Curiosity: exploring, seeking knowledge, investigating, asking questions; Thinking: reflecting, considering, contemplating, pondering; Analysis: analyzing, dissecting, researching, discussing; Neutrality: impartial, fair, balanced; Understanding: comprehending, experiencing, feeling, comprehending;
[0237] Supplementing the custom sentiment dictionary (taking two dimensions as an example)
[0238] (1)(1) Self-acceptance:
[0239] ① Explanation of terms related to keyword recognition pattern:
[0240] Acceptance: This includes accepting your strengths and weaknesses. Imperfection: Acknowledging that imperfection is part of growth. Learning: Learning from mistakes and experiences. Emotions: Maintaining an openness to all emotions. Experience: Actively participating in life experiences. Curiosity: Maintaining an interest in self-exploration. Forgiveness: Being able to forgive your own mistakes. Expression: Daring to express your true feelings and needs. Challenge: Facing life's challenges with courage. Respect: Respecting your own feelings and needs. Vulnerability: Accepting and showing your vulnerable side.
[0241] ② Expanding the Emotional Dictionary:
[0242] +1 positive vocabulary: acceptance, growth, learning, curiosity, courage, confidence.
[0243] Negative words - 1: denial, rejection, fear, avoidance.
[0244] Neutral vocabulary 0: observation, experience, reflection.
[0245] (2) Degree of ideological bias:
[0246] ① Keyword recognition mode:
[0247] Respect: Respect others' opinions and ideas. Trust: Trust others. Cooperation: Be willing to cooperate with others. Compromise: Find balance in relationships. Goodwill: Believe others act with good intentions. Independence: Respect others' independent decisions. Listening: Listen attentively to others' opinions. Understanding: Strive to understand others' feelings. Differences: Accept differences between people. Helping: Believe others can offer help.
[0248] ② Expanding the Emotional Dictionary:
[0249] +1 positive vocabulary: inclusiveness, cooperation, understanding, trust, openness.
[0250] Negative words - 1: prejudice, distrust, hostility, stubbornness.
[0251] Neutral word 0: neutral, consider, analyze.
[0252] S6.3, Assigning Sentiment Scores: Assign a sentiment score to each preset keyword: positive keywords are assigned +1, negative keywords are assigned -1, and neutral keywords are assigned 0; if a preset keyword appears in multiple contexts, all sentiment tendencies are considered comprehensively.
[0253] S6.4. Based on the assignment of sentiment scores of the preset keywords in S6.3, combined with the quantitative questions of the corresponding sub-patterns in the psychological assessment model, the sentiment tendency of the quantitative questions corresponding to the sub-patterns is judged. If the quantitative questions corresponding to the sub-patterns are judged to have a positive sentiment tendency, the quantitative questions corresponding to the sub-patterns are judged to have a score of 1. If the quantitative questions corresponding to the sub-patterns are judged to have a negative sentiment tendency, the quantitative questions corresponding to the sub-patterns are judged to have a score of -1. If the quantitative questions corresponding to the sub-patterns are judged to have a neutral sentiment tendency, then according to the matching results in S5, it is judged whether the sentiment quantitative value corresponding to the text fragments in the sub-patterns is higher than the defined level. If the sentiment quantitative value corresponding to the text fragments in the sub-patterns is higher than the defined level, the quantitative questions corresponding to the sub-patterns are judged to have a score of 1. If the sentiment quantitative value corresponding to the text fragments in the sub-patterns is equal to the defined level, the quantitative questions corresponding to the sub-patterns are judged to have a score of 0. If the sentiment quantitative value corresponding to the text fragments in the sub-patterns is lower than the defined level, the quantitative questions corresponding to the sub-patterns are judged to have a score of -1.
[0254] The horizontal line is defined as (highest sentiment quantification value + lowest sentiment quantification value) / 2. Furthermore, NumPy can be used to calculate the median average of the sentiment quantification values to ensure the accuracy of the defined horizontal line data.
[0255] Specifically, emotion quantification is used to assist in judgment: when the initial emotion judgment obtained through emotion dictionary matching and context analysis is unclear, the emotion quantification value of tone information is analyzed to determine the employee's response tendency.
[0256] A. Preliminary identification criteria (if the employee's response in the context of the preset keywords cannot be identified, or the sentiment score is 0):
[0257] a. No pre-defined keyword association: If no pre-defined keywords related to the quantitative question can be identified in the text content, the sentiment tendency cannot be directly determined.
[0258] b. Sentiment score of 0: After analyzing the sentiment tendency of the sentence, a preliminary judgment of a sentiment score of 0 indicates that the emotion is neutral or ambiguous.
[0259] B. Implementation details: The emotional tendency is determined by calculating the difference between the emotional quantification value of each speech segment and the horizontal line.
[0260] The score is determined by comparing the emotional quantification value with the defined level using an if-else structure.
[0261] a. If the emotional quantification value is higher than the defined level: it is judged as agreement, and the score is 1.
[0262] b. If the emotional quantification value is below the defined level: it is judged as disagreement, and the score is -1.
[0263] c. If the emotional quantification value equals the defined horizontal line: it is judged as unclear, and the score is 0.
[0264] The above judgment method can obtain a more accurate sentiment tendency for each sub-pattern, which can improve the accuracy of subsequent analysis and evaluation results.
[0265] S7. By judging the sum of the scores of all quantitative questions for each dimension, if the sum of the scores for that dimension is positive, then the dimension performs well; if the sum of the scores for that dimension is 0, then the dimension performs well; if the dimension is negative, then the dimension performs poorly.
[0266] Dimensional Score Calculation: By summing the scores of each sub-pattern, the total score for each dimension is calculated, and a psychological state assessment is performed. Implementation Steps:
[0267] 1) Score Integration: The scores of the emotion quantification judgment are integrated into the final quantitative question score to form a complete evaluation result for each sub-mode.
[0268] Technical Implementation:
[0269] a. Use a Pandas data frame to store the sentiment judgment result and sentiment quantification value judgment result for each sub-pattern.
[0270] b. For example, df.loc['subpattern','score'] = sentiment quantification value to determine the score and update the data.
[0271] 2) Accumulation of scores for quantification issues corresponding to sub-patterns
[0272] Accumulated calculation: The scores for the quantification questions within each core dimension are summed. Pandas is used for data processing and accumulation, simplifying the operation:
[0273]
[0274] 3) Interpretation of Assessment Results
[0275] Results analysis:
[0276] A positive score indicates excellent performance in this dimension.
[0277] A score of 0 indicates good performance in this dimension.
[0278] A negative score indicates a basic level of performance in this dimension.
[0279] Visualization: Use Matplotlib or Plotly to generate visual charts of the core dimension scores to facilitate understanding and analysis of the results.
[0280] S8. Automatically generate psychological reports based on performance in each dimension. Output results: Set up a custom template (example description); fill in the corresponding results and generate a report: Output an analysis report containing scores for each dimension, detailing performance in each dimension and related suggestions.
[0281] As one implementation, in step S4, if the sub-patterns containing text fragments account for 60% or more of all sub-patterns in each dimension, then the performance evaluation of that dimension is valid in step S7; otherwise, it is invalid.
[0282] In S8, the validity of the dimensional performance evaluation is presented in writing.
[0283] This embodiment provides an analysis method for an EAP (Employment Aptitude Test) auditory analysis instrument. The voice information can be dialogue information, and the person being evaluated (e.g., an employee) is unaware that it is for evaluation purposes. After acquiring the voice information, by identifying the phoneme boundaries, several voice segments can be determined. Processing these voice segments yields the duration and amplitude features of each segment. Combining these features into a fused feature vector, and under the analysis of an emotion classification model, the emotion quantification value of the voice segment can be obtained. This serves as an auxiliary judgment tool for sub-patterns within the dimensions of the psychological assessment model. After converting the voice information into text segments corresponding to the voice segments and forming text, the text is traversed, and the text segments corresponding to preset keywords are matched with sub-patterns. This yields the text segments and emotion quantification values within each sub-pattern. By conducting sentiment analysis and semantic correlation analysis, it is possible to initially determine whether the quantitative question corresponding to a sub-pattern has a positive, negative, or neutral sentiment tendency. When the sub-pattern has a neutral sentiment tendency, it is possible to make a further judgment based on the emotional quantification value, thereby making the sentiment tendency judgment of the quantitative question corresponding to the sub-pattern more accurate. After obtaining the accurate sentiment tendency of all sub-patterns for each dimension, the sentiment tendency of that dimension can be obtained as the expression of that dimension. This method can directly analyze and evaluate mental health based on the voice information of the personnel (employees) who need psychological assessment, without requiring them to answer questions subjectively. This reduces the subjectivity of the psychological assessment and makes the obtained psychological assessment results more objective. Furthermore, by using preliminary judgment based on text and auxiliary judgment based on voice emotion, the obtained psychological assessment results are more accurate.
[0284] Example 2
[0285] This embodiment provides an EAP hearing analysis device, including at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the analysis method of the EAP hearing analyzer in Embodiment 1.
[0286] Example 3
[0287] This embodiment provides an EAP hearing analyzer, such as Figure 1 As shown, it includes: a data acquisition module, a speech-to-text module, a keyword recognition module, a speech intonation analysis module, a psychological report generation module, and a data storage and management module.
[0288] The data acquisition module acquires voice information; for example, it can be the voice acquisition function built into an electronic device such as an iPad, which can record audio; the voice information can be real-time recording or non-real-time recording.
[0289] In one or more embodiments, the data acquisition module captures voice information using a microphone. By using a separate microphone for voice information capture, the voice information can be obtained more clearly and accurately. Furthermore, the data acquisition module captures voice information using a directional microphone, which provides better directional sound reception and makes the obtained voice information more accurate.
[0290] The speech intonation analysis module extracts the speech information collected by the data acquisition module, divides the speech information into several speech segments, and obtains the emotion quantification value of each speech segment based on the fundamental frequency, energy, and morphological features of the speech segments. Optionally, it extracts the speech information collected by the data acquisition module, divides the speech information into several speech segments, obtains the duration and amplitude features of the speech segments, combines the duration and amplitude features of each speech segment to form a fusion feature vector for each speech segment, and inputs the fusion feature vector of the speech segments into the emotion classification model, which then obtains the emotion quantification value of each speech segment.
[0291] The speech-to-text module converts speech information into text segments corresponding to the speech segments, and all text segments form a text; optionally, the speech information collected by the data acquisition module is converted into text segments corresponding to the speech segments based on the duration characteristics of the speech segments obtained by the speech intonation analysis module, and all text segments form a text.
[0292] In one or more embodiments, the speech-to-text module employs multi-model fusion technology, processing the same speech segment in parallel using multiple different models and combining the results of each model through voting or weighted averaging to effectively reduce the bias of a single model. Multi-model fusion technology (also known as ensemble learning) processes the same speech segment by running multiple different models in parallel and combining the results of each model through voting or weighted averaging to improve overall accuracy and robustness. The core idea of this technology is to leverage the advantages and characteristics of different models, reducing the bias and overfitting of a single model by combining the results of multiple models, thereby improving the overall performance of the system. In the field of speech processing, multi-model fusion technology can significantly improve the performance of tasks such as speech recognition, speech synthesis, and speech emotion analysis. By combining the advantages of different models, the bias of a single model can be effectively reduced, improving the robustness of the system. The text module introduces dynamically adaptive speech recognition technology to analyze the user's speech characteristics (including syllables and tones) in real time and dynamically adjust the speech recognition model parameters to adapt to differences in accents and speech rates among different regions and populations. During dynamic adjustment, those skilled in the art can adjust the parameters in the introduced recognition technology as needed.
[0293] DynamicTime speech recognition technology is a method of identifying text contained in speech based on dynamic time warping. Compared to traditional static speech recognition technology, dynamic speech recognition technology is more suitable for recognizing speech that changes frequently and at a fast pace, and it also has a higher accuracy rate.
[0294] In one or more embodiments, the speech-to-text conversion has generalization capabilities.
[0295] In one or more embodiments, the speech-to-text module incorporates a real-time feedback mechanism. When a user discovers an error in the transcription result, the user can immediately annotate and provide feedback, and the speech-to-text module can automatically learn and adjust. When a user discovers a transcription error, they can provide feedback by manually annotating or inputting the correct text. This feedback is used to improve future transcription results. After the user inputs a correct transcription, the system stores this feedback data in a database for subsequent model training and optimization. The system periodically or in real-time uses this data for retraining to improve the model's accuracy. When computational resources are sufficient, user feedback can be used in real-time to fine-tune the model. Through online learning technology, the system can quickly adjust parameters and improve its performance after receiving feedback. When it is determined that the user's input content needs to be learned: the system identifies the feedback, verifies the data, performs contextual analysis, and finally adjusts the weights.
[0296] (1) The system first needs to recognize user feedback. This is usually done through the user interface, where the user explicitly points out transcription errors and enters the correct text. The system uses this explicit feedback to identify what needs to be learned.
[0297] (2) To ensure the accuracy of the feedback, the system may use a variety of strategies for verification. For example, the accuracy of the data may be verified by comparing user feedback with existing data in the system, or by using redundancy mechanisms (e.g., consistency of feedback from multiple users).
[0298] (3) The system can also determine the validity of feedback through context analysis. If the user's feedback is inconsistent with the current voice context, the system can mark the feedback for further review or ignore it.
[0299] (4) The system will dynamically adjust the weights of the speech recognition model based on the frequency and consistency of the feedback data. Frequent erroneous feedback will cause the speech recognition model to make more significant adjustments, while occasional feedback may only cause minor adjustments.
[0300] The keyword recognition module traverses the text and matches the text fragments corresponding to preset keywords with the quantitative questions in each dimension of the psychological assessment model; optionally, it traverses the text and assigns the text fragments corresponding to preset keywords to the sub-patterns of the corresponding preset keywords.
[0301] These text snippets reflect users' tendencies regarding these sub-patterns. Sentiment analysis of these text snippets can correlate with the answer tendencies of corresponding questions within the psychological scale of the psychological assessment model. In other words, by obtaining the sentiment tendencies of these text snippets, the answers to the corresponding questions within the psychological scale can be obtained. This leads to the score on the psychological scale of the psychological assessment model, which can then be used for evaluation and the generation of psychological reports. The analysis of these text snippets can help improve the accuracy and practicality of EAP (Educational Aptitude Test) auditory analysis instruments. Instead of directly asking questions from within the psychological scale, it obtains text snippets through conversation and then further derives the answers to the questions within the psychological scale. This method reduces the subjectivity of the assessed individuals when answering questions within the psychological scale, thus making the results more objective and accurate.
[0302] The psychological report generation module matches text fragments with the emotional quantification values of corresponding audio fragments. Then, based on the text fragments and emotional quantification values matching the quantification question, it obtains the answer tendency for the quantification question. It then synthesizes the answer tendencies of all quantification questions for that dimension to obtain the performance of that dimension, and automatically generates a psychological report based on the performance of each dimension. The generated psychological report also provides data visualization interpretation. Optionally, it matches text fragments in each sub-mode with the emotional quantification values of corresponding audio fragments. It performs sentiment tendency analysis and semantic correlation analysis on the text fragments in each sub-mode to determine whether the quantification question corresponding to that sub-mode has a positive, negative, or neutral sentiment tendency. When the quantification question corresponding to that sub-mode is determined to have a neutral sentiment tendency, the sub-mode is used to generate the report. The emotional quantification value of the text fragment is used to reassess the emotional tendency of the corresponding quantitative question for the sub-mode, and then a score is assigned to the sub-mode based on the emotional tendency. The sum of scores of all sub-modes in each dimension is obtained through a psychological assessment model. If the sum of scores for a dimension is positive, the dimension performs well; if the sum of scores for a dimension is 0, the dimension performs well; if the dimension is negative, the dimension performs poorly. A psychological report is automatically generated based on the performance of each dimension. The data in the generated psychological report is then visualized and interpreted. Optionally, the results analyzed by the psychological assessment model are visualized in the report using bar charts, pie charts, radar charts, etc. The psychological report generation module analyzes the report results based on psychological research and practical experience and provides corresponding intervention suggestions and support strategies.
[0303] In this embodiment, the psychological assessment model comprehensively assesses the psychological state based on two core dimensions of the six dimensions of acceptance and commitment therapy.
[0304] The internal logic of the psychological assessment model is based on two core dimensions of the six dimensions of Acceptance and Commitment Therapy (ACBT) to comprehensively assess the psychological state of employees. The six dimensions of ACBT are as follows:
[0305] ① Self-acceptance (experiential avoidance): refers to the form and frequency of one's internal experiences (such as thoughts, emotions, bodily sensations, or memories) appearing in one's mind, or one's sensitivity to situations.
[0306] ② Thought bias (cognitive fusion): This refers to the tendency for human behavior to be limited by the content of thought. This tendency causes people's behavior to be excessively controlled by language rules and cognitive evaluation, thus making it impossible to guide behavior with here-and-now experience and direct experience.
[0307] ③ Affective Awareness (Present Awareness): This dimension emphasizes an individual's overall awareness of their emotions and experiences in the current situation. It involves attention and sensitivity to immediate experiences, including awareness of the environment, bodily sensations, emotions, and thoughts. Individuals with high affective awareness are able to clearly identify and accept their feelings, rather than avoid or downplay them.
[0308] ④ Value Orientation (Commitment to Action): Value orientation refers to the degree to which an individual's behavior is guided by their core values. This includes recognizing and clarifying one's personal values, and taking concrete actions based on these values. Individuals with high value orientation are able to consistently take actions consistent with their values in their daily lives.
[0309] ⑤ Self-awareness (self-observation): This dimension emphasizes an individual's awareness and understanding of their own inner psychological state. It involves self-reflection and self-consciousness, including understanding one's own thought patterns, emotional responses, and behavioral habits. Individuals with high self-awareness are better able to understand their inner world and make more informed choices accordingly.
[0310] ⑥ Relational Flexibility (Relational Framework): Relational flexibility focuses on an individual's adaptability and flexibility in interpersonal relationships. It involves identifying and adjusting roles and behaviors in interpersonal interactions, as well as the ability to maintain positive and healthy relationships. Individuals with high relational flexibility are able to flexibly adjust their behavior in different social situations to promote effective and harmonious interpersonal communication.
[0311] Two core dimensions of Acceptance and Commitment Therapy's six dimensions are self-acceptance and paranoia. The specific questioning and analysis logic (corresponding to quantitative questions) relates to the following:
[0312] (1) Taking self-acceptance as an example:
[0313] ① Are you able to accept your imperfections and learn from them?
[0314] ②Are you curious about your own emotions and experiences?
[0315] ③Are you willing to experience a variety of emotions, including those that are uncomfortable?
[0316] ④ When you make a mistake, are you able to forgive yourself and try to correct it?
[0317] ⑤ Do you believe you have the right to express your feelings and needs?
[0318] ⑥ Are you willing to face and accept your challenges and troubles?
[0319] ⑦ Do you understand and respect your own emotions?
[0320] ⑧ Do you pay attention to your inner feelings?
[0321] 9. Do you allow yourself to feel vulnerable at times?
[0322] ⑩ Do you consider yourself good at understanding and accepting your own emotions?
[0323] (2) Taking the degree of ideological bias as an example:
[0324] ① Do you respect other people's opinions and ideas?
[0325] ②Are you willing to trust others and cooperate with them to solve problems?
[0326] ③Are you willing to make some compromises in order to maintain the relationship?
[0327] ④ Do you believe that others are mostly acting out of goodwill?
[0328] ⑤ Do you respect other people's independence and allow them to make their own decisions freely?
[0329] ⑥ Do you believe that others can help you?
[0330] ⑦ Are you willing to listen to other people's suggestions and feedback?
[0331] ⑧ Do you try your best to understand other people's feelings?
[0332] 9. Do you think people should respect and understand each other's differences?
[0333] ⑩ Do you believe that most people are trustworthy?
[0334] The voice analysis results focus on the positive or negative response to each question: 1 for agreement, -1 for disagreement, and 0 for unclear responses.
[0335] (1) Self-acceptance: A positive total score indicates a higher level of self-acceptance, which is a more positive state.
[0336] ① Basic Self-Acceptance (Negative Total Score): Your low score on self-acceptance may indicate that you are overly critical of yourself in daily life and work, and may feel troubled and anxious about your own emotional experiences and frustrations. This situation may cause you to put excessive pressure on yourself at work, affecting your work efficiency and quality of life. It is recommended that you consider seeking professional psychological counseling to help you learn how to better accept yourself and understand your emotions and needs. In addition, at work, try to set some achievable goals for yourself, avoid overly expecting too much of yourself, try to discover and cultivate your strengths and talents, and face the challenges of work and life in a more positive way.
[0337] ② Good Self-Acceptance (Total Score: 0): Your self-acceptance score is average. You may have some understanding and acceptance of your emotions and needs when facing work pressure and life challenges, but there may be room for improvement. It is recommended that you further cultivate habits of self-care and self-love in your daily work and life. For example, you can use your lunch break to do some relaxing activities such as yoga, meditation, or reading to enhance your self-acceptance and understanding. At work, when encountering problems and setbacks, remember to give yourself some patience and space, try to understand and accept your emotions, and don't suppress yourself excessively.
[0338] ③ Excellent Self-Acceptance (Positive Score): Your high score on self-acceptance indicates that you understand and accept your own emotions and needs well, which is highly commendable. Building on this foundation, you can further translate this acceptance into more effective actions, such as maintaining a positive mood and work efficiency through effective time management and self-care in your daily work, thereby improving your quality of life and work.
[0339] (2) Degree of ideological bias: A positive total score indicates that the lower the degree of ideological bias, the more normal and positive the ability to establish relationships with others.
[0340] ① Basic Social Adaptability (Negative Total Score): Your thinking is highly biased, which may mean you harbor suspicion, distrust, or misunderstanding towards others in your work or interpersonal interactions. This could negatively impact your teamwork and workplace relationships. It is recommended that you try some social skills training to improve your interpersonal and communication abilities. At work, be more open to others' viewpoints and suggestions, and try to understand and respect their work methods. This will help you build more harmonious relationships and improve teamwork.
[0341] ② Good Social Adaptability (Total Score: 0): Your level of stubbornness is average. You may sometimes feel confused or conflicted when facing others and handling interpersonal relationships. In interactions with colleagues and teamwork, you may sometimes feel accepted and understood, and at other times feel doubtful and uneasy. It is recommended that you try more to understand and accept the viewpoints and feelings of others at work. This will help you build better connections with others, improve your teamwork skills, and thus increase work efficiency.
[0342] ③ Excellent social adaptability (positive score): Your level of bias is relatively low. In your interactions and collaborations with others, you are generally able to understand, accept, and respect them, which is highly commendable. Building on this foundation, you can further translate this understanding and acceptance into concrete actions. For example, at work, actively participate in teamwork, proactively listen to and accept others' suggestions, and collaboratively solve work-related problems. This will help improve your teamwork effectiveness and work efficiency.
[0343] The data storage and management module saves all data generated by the data acquisition module, speech-to-text module, speech intonation analysis module, and psychological report generation module. All data includes personnel information, recording data, converted text data, keyword recognition results, speech intonation analysis results, and psychological reports.
[0344] In one or more embodiments, the data storage and management module locally stores all data generated by the data acquisition module, speech-to-text module, speech intonation analysis module, and psychological report generation module. Local storage is more secure and avoids privacy leaks.
[0345] Furthermore, the data storage and management module uses the non-relational database MongoDB to store different types of data; and employs the AES-256 encryption algorithm in all data storage processes. Of course, the choice of database and encryption algorithm can be made according to the actual situation.
[0346] like Figure 1 As shown, this embodiment provides an EAP (Employee Assistance Program) hearing analyzer, including a data acquisition module. Data acquisition is the initial step in the core workflow of the EAP hearing analyzer, employing a high-precision directional microphone to capture speech information. This microphone is specifically designed to accurately capture the speech details in employee conversations, ensuring clear, noise-free speech signals in any environment. The collected speech information is further processed using speech-to-text technology for subsequent data analysis. The key to this technology lies in its real-time performance and accuracy, ensuring that the converted text data reflects the original speech information as faithfully as possible.
[0347] Sound recording is a crucial hardware component of the entire EAP (Earning and Assisted Audio) listening analysis instrument, determining the quality and accuracy of the initial data. To ensure high-quality captured speech information, this embodiment employs high-precision directional sound recording. This design is characterized by its extremely strong directionality, enabling precise collection of sound from the target source while effectively reducing background noise from other directions. The deployment location and angle of the directional sound recorder have been carefully calculated to ensure optimal sound recording performance in any environment, whether it's an open conference room or a small, enclosed discussion room.
[0348] Before further processing, the audio information undergoes a preprocessing stage. This stage uses digital filtering techniques to perform noise reduction, gain adjustment, and echo cancellation on the captured speech signal. This ensures that subsequent speech-to-text transcription can be performed on a clear, noise-free audio base, greatly improving the accuracy of the transcription.
[0349] The EAP auditory analyzer in this embodiment provides two data acquisition modes: real-time and non-real-time. In real-time mode, the analyzer can immediately capture and analyze data during a conversation, suitable for situations requiring immediate feedback. The non-real-time mode stores the raw voice information for subsequent analysis, making it more suitable for in-depth and detailed psychological analysis.
[0350] To ensure the stability and consistency of voice data in subsequent analysis, this embodiment also includes a data standardization and formatting process. Regardless of the environment or conditions from which the data is collected, this process transforms it into a unified standard format, laying a solid foundation for subsequent data processing and analysis. For example, there is an MP3 to WAV conversion algorithm: MP3 is a lossy compression format, and conversion to WAV requires decompression to restore the compressed data to the original audio sample data. WAV standard audio format – parameter settings: Sampling rate: 96kbps (ultra-high sampling rate for better audio detail capture); Bit depth: 24 bits (provides higher dynamic range and more accurate sound representation); Number of channels: 7.1 channels (surround sound system for a more spatial and immersive audio experience).
[0351] Considering the sensitivity and privacy of the conversation content, this embodiment employs a series of measures to ensure the security of data collection. All voice data is encrypted during storage and transmission to ensure data integrity and privacy. Furthermore, the analyzer is equipped with a multi-factor authentication mechanism, ensuring that only authorized personnel can access and use the device.
[0352] After voice data collection, a speech-to-text module converts the collected speech information into text. This technology primarily utilizes deep learning algorithms and a large amount of speech training data, and through multiple iterations and optimizations, provides users with a near-perfect speech-to-text experience. Furthermore, to better adapt to different accents and speaking speeds, this technology also incorporates multiple speech models to ensure optimal transcription results in various scenarios.
[0353] The breakthrough in speech-to-text technology achieved by the EAP auditory analyzer is attributed to the novel application of deep learning algorithms. Deep learning, as an algorithm that simulates the workings of the human brain, has demonstrated exceptional performance in speech recognition tasks. Through a multi-layered neural network structure, it can automatically extract features from speech signals and convert these features into corresponding text.
[0354] A successful speech-to-text system relies heavily on a large amount of high-quality speech training data. This system collected relevant survey recordings, as well as speech data from various accents and scenarios, covering diverse everyday, professional, and special environments. After annotation and cleaning, this data laid a solid foundation for training the neural network model.
[0355] Taking into account the differences in accents and speaking speeds among various regions and populations, the EAP auditory analyzer incorporates dynamic adaptation technology in its speech recognition. This technology can analyze the user's speech characteristics in real time and dynamically adjust model parameters to ensure optimal transcription results under any circumstances.
[0356] To further improve transcription accuracy, this embodiment employs multi-model fusion technology. By processing the same speech segment in parallel using multiple different models, and then combining the results of each model through voting or weighted averaging, the bias of a single model can be effectively reduced, resulting in a more accurate transcription result.
[0357] The EAP hearing analyzer also features a real-time feedback mechanism. When users find errors in the transcription results, they can immediately mark and provide feedback. The system will automatically learn and adjust to ensure that similar errors are avoided in future recognition. This mechanism not only improves the system's accuracy but also greatly enhances the user experience.
[0358] Keyword identification is crucial for in-depth analysis of conversation content. The EAP auditory analysis instrument, through natural language processing technology and machine learning algorithms, can accurately identify and extract keywords from conversations. This goes beyond simple word frequency statistics; it identifies keywords that decisively influence emotional and psychological states within the context of the dialogue. This provides vital data support for the generation of subsequent psychological reports.
[0359] For in-depth analysis of conversation content, keyword identification is particularly important. Keywords serve as a window into summarizing and capturing the core meaning, sentiment, and attitude of the conversation. Natural Language Processing (NLP) technology, part of the intersection of computer science, artificial intelligence, and linguistics, aims to enable computers to interpret, understand, and generate human language. In keyword identification, NLP technology first performs sentence segmentation, breaking down continuous text into individual words or phrases. Then, it uses specific algorithms, such as TF-IDF (Term Frequency-Inverse Document Frequency) and TextRank, to assign a weight to each word to determine its importance in the text.
[0360] Machine learning has provided a higher level of capability for natural language processing, particularly in terms of accuracy in keyword recognition and contextual understanding. Through extensive training data, machine learning models can capture patterns and relationships within text, thus identifying keywords more accurately. More advanced models, such as the Transformer in deep learning, can even understand complex contexts and semantic relationships within text.
[0361] Simply identifying keywords is insufficient; understanding the emotions and intentions behind them is also crucial. By analyzing the context, grammatical structure, and relationship of keywords with other words, we can discern the speaker's emotional inclination and psychological state. For example, words like "stress," "confusion," and "dissatisfaction" may suggest that the speaker is facing a predicament or emotional distress. This analysis goes beyond single keywords; it requires considering the entire sentence or paragraph's context. With accurate keyword identification and sentiment analysis, a comprehensive and in-depth psychological report can be generated for leaders. This report not only lists keywords but also demonstrates their distribution, frequency, and relevance within the text, providing leaders with a holistic perspective. Furthermore, based on the sentiment analysis results, the report offers specific suggestions and solutions to help leaders better understand and support their employees.
[0362] The keyword recognition module provides users with an intuitive and clear method to deeply understand the content of conversations. Leveraging advanced technology and algorithms, it is now possible to more accurately and profoundly capture employees' thoughts and emotions, reducing employee subjectivity.
[0363] Furthermore, the information contained in speech is not limited to textual content; intonation, speaking speed, and pauses are all important carriers of emotion. The EAP (Earning Aptitude Test) auditory analyzer utilizes advanced speech analysis technology to deeply analyze these non-verbal signals. It can extract features such as fundamental frequency, energy, and morphology of speech, forming a set of feature algorithms specifically customized for the organization. This analysis can capture subtle emotional changes, thus providing leaders with more in-depth psychological analysis.
[0364] Advanced speech analysis technology employs innovative High Dimension Emotional Voice Analysis (HDEA-VAT) technology. By adjusting different weights to extract the following design, this HDEA-VAT technology can provide more accurate and in-depth emotional and psychological analysis, helping leaders and decision-makers better understand and manage the psychological state of employees.
[0365] 1. Multi-level feature extraction
[0366] Frequency domain analysis: Using high-precision Fourier transform (FFT) and wavelet transform, frequency domain analysis is performed on speech signals to extract features such as fundamental frequency, harmonic structure, and duration of energy distribution above the threshold in speech segments.
[0367] Temporal analysis: Extracting instantaneous energy, amplitude envelope, and morphological features of speech through autoregressive (AR) model and time-frequency analysis.
[0368] Duration and rhythm analysis: Analyze pauses, speech rate changes and rhythm patterns in speech signals to reveal the speaker's emotions and psychological state.
[0369] 2. Integrating deep learning and psychological models
[0370] Deep Neural Networks (DNNs): Construct a multi-layer deep neural network model. The input layer receives speech feature data, the intermediate layers perform feature fusion and emotion classification, and the output layer provides emotion state prediction.
[0371] Emotional psychological model: Combining emotion theories in psychology (such as Plutchik's emotion wheel model), the output of DNN is mapped to specific emotion categories to achieve fine-grained analysis of complex emotional states.
[0372] 3. High-precision emotion detection
[0373] Micro-emotion recognition: By analyzing minute changes in speech signals, it captures subtle fluctuations in the speaker's emotions, providing high-precision emotion detection.
[0374] Emotion intensity quantification: An emotion intensity quantification model is introduced to assess the intensity of each emotion and provide quantitative data on emotion intensity, i.e., to obtain the emotion quantification value, which is convenient for further analysis and application.
[0375] When conveying information, speech goes beyond simple textual content. Often, it expresses emotions and attitudes through nonverbal signals such as intonation, speech rate, and pauses. This is why the same sentence can convey vastly different meanings depending on the intonation and speech rate. Therefore, this embodiment uses a speech intonation analysis module to extract the fundamental frequency, energy, and morphological features of speech, capturing the speaker's emotional changes. The fundamental frequency is the most basic sound frequency in speech, directly related to the speaker's pitch. By analyzing changes in the fundamental frequency, we can understand the speaker's emotional state. For example, when people are tense or anxious, their fundamental frequency may increase. When they are relaxed or frustrated, the fundamental frequency may decrease. The energy of speech reflects the strength of the speaker's vocalization. Generally, when emotions are strong, whether anger, excitement, or joy, the energy of speech increases. Conversely, in cases of melancholy or boredom, the energy may decrease. Therefore, by analyzing changes in speech energy, we can gain a preliminary understanding of the speaker's emotional intensity. Morphological analysis studies the shape and structure of sound waves, including their duration and amplitude. The speaker's personality, character, and current emotional state all affect the shape of the sound waves. For example, some people speak in a short, forceful style, with sharp sound waves; while others speak softly, with smooth sound waves.
[0376] Based on the preceding data analysis and algorithmic model, the EAP auditory analysis instrument can automatically generate a psychological report. This report is not merely a simple list of data, but rather combines a multi-dimensional psychological assessment model with various technical means such as keyword analysis and tone analysis to provide leaders with a comprehensive, in-depth, and intuitive psychological analysis result. Furthermore, the report also provides specific suggestions and solutions to help leaders better understand and support their employees.
[0377] In today's data-driven world, extracting, analyzing, and transforming data into meaningful insights is crucial. The EAP (Ear, Listening, and Perception) Analyzer precisely achieves this goal, especially in mental health assessment.
[0378] Before generating a psychological report, the first step is data integration and processing. Due to the diverse sources of data, such as voice, text, and tone of voice, this data first needs to be normalized and standardized. This embodiment uses data processing algorithms to ensure data quality and consistency, laying a solid foundation for subsequent analysis.
[0379] The psychological assessment model employed by the EAP auditory analysis instrument goes beyond traditional psychological theories. This embodiment combines modern artificial intelligence technology with traditional psychological knowledge to construct a multi-dimensional psychological assessment model. This model covers multiple dimensions, including emotion, cognition, and behavior, and comprehensively assesses employees' psychological state based on two core dimensions of the six dimensions of Acceptance and Commitment Therapy (ACRT), providing leaders with a 360-degree mental health assessment.
[0380] As mentioned earlier, keywords and tone of voice are important indicators for assessing employees' psychological state. EAP (Employee Assistance Program) auditory analysis systems can perform in-depth analysis of this data, such as sentiment analysis and semantic correlation analysis. This not only allows us to understand employees' current psychological state but also predicts their future psychological trends.
[0381] To ensure the report's intuitiveness and ease of understanding, this embodiment incorporates numerous visualization elements. These include, but are not limited to, bar charts, pie charts, and radar charts. These graphics visually represent employees' mental health status, helping leaders quickly obtain key information. Based on these analyses, the system will also automatically generate corresponding psychological support plans to provide substantial assistance to employees.
[0382] The EAP auditory data visualization feature presents psychological data through engaging graphics and interfaces. It utilizes various chart types, such as radar charts, bar charts, pie charts, and time series graphs, making complex psychological data concise and intuitive. This clear presentation not only helps leaders quickly grasp key information but also provides strong support for in-depth research.
[0383] The EAP (Emotional Aptitude Test) auditory analysis system combines fundamental psychological theories, models, and practical applications with technology to ensure the scientific accuracy of data interpretation. For example, when a report indicates that an employee scored low on the "emotional state" dimension, the system will cite relevant psychological theories to explain the possible reasons and underlying psychological mechanisms.
[0384] Beyond data analysis, the EAP (Employee Assistance Program) auditory analyzer can also provide specific suggestions and solutions based on the analysis results. These suggestions are based on extensive psychological research and practical experience, aiming to help leaders better understand and support their employees. For example, if the analysis results show that an employee has mild anxiety symptoms, the report will provide corresponding intervention suggestions and support strategies.
[0385] The EAP auditory analysis system, in its first phase design, comprehensively assessed employees' psychological state based on two core dimensions of the six dimensions of Acceptance and Commitment Therapy:
[0386] Based on the above psychological analysis, the EAP auditory analysis system automatically generates specific and practical support solutions. For example, if an employee's "psychological stress" score is high, the system may recommend providing stress management training or short-term psychological counseling. If the "interpersonal relationships" score is low, it may recommend team building activities or communication skills training. These recommendations are not only based on data but also incorporate the latest psychological research and practical experience, ensuring they are both scientific and effective.
[0387] The standardized psychological analysis report and EAP support plan are shown in the table below:
[0388]
[0389]
[0390]
[0391]
[0392] In this application, the psychological report generation template employs an innovative keyword analysis algorithm, combined with the logic of psychological scales, to achieve quantitative analysis and automatic report generation. Unlike traditional psychological scales that typically rely on direct input selection results, this application identifies and analyzes speech signals in everyday conversations, including non-verbal signals such as tone, speech rate, and pauses, converting them into quantifiable data. This method not only improves the accuracy of the analysis but also effectively reduces employee resistance, making psychological assessments more natural and stress-free. Furthermore, this application utilizes high-dimensional emotional speech analysis technology to deeply analyze the fundamental frequency, energy, and morphological features of speech, capturing subtle emotional changes. These multi-layered feature extractions and sentiment analyses allow leaders to obtain more in-depth and comprehensive psychological analysis reports, thereby better understanding and supporting employees. Through real-time adaptive learning and multimodal fusion analysis, this application not only provides immediate feedback but also continuously optimizes based on user feedback, adapting to different users and scenarios, significantly improving the accuracy and practicality of psychological assessments.
[0393] Finally, data security and privacy protection are paramount for the project. The EAP hearing analyzer stores all data locally, without transmitting or processing it through any third-party servers. This significantly reduces the risk of data leakage. The local server (after the full extension phase) employs the latest encryption technology and firewall system to ensure data integrity and security. Furthermore, the system provides robust data management functions, allowing leaders to query, statistically analyze, and perform other tasks as needed.
[0394] In this embodiment, a hierarchical data storage structure is adopted to ensure that the data is not only secure but also easy to manage and retrieve. The use of the non-relational database MongoDB provides specialized storage solutions for different types of data (such as text, audio clips, analysis results, etc.). Furthermore, this structure allows for flexible expansion of storage capacity in the future as needed.
[0395] To ensure data confidentiality, all data is encrypted during storage. This implementation uses the AES-256 encryption algorithm for data encryption. In addition, the system performs automatic backups daily to ensure data recovery even in extreme circumstances, such as physical damage or other unforeseen events.
[0396] To further enhance data security, this embodiment also establishes a strict data access permission management system. Each user's access permissions are assigned based on their role and needs. For example, only users with specific permissions can access and analyze raw data. This strategy aims to ensure that data can only be accessed by authorized personnel, significantly reducing the risk of data leakage.
[0397] Secondly, to meet the requirements for auditing and tracking data operations, this system also records the operation history of all data, including data creation, modification, and deletion. This is not only for security reasons, but also to verify the integrity and compliance of data operations.
[0398] The EAP hearing analyzer in this embodiment adopts a front-end and back-end separation design. The front-end uses the React framework to develop an app, providing a user-friendly and responsive interface. The back-end is developed using Node.js, ensuring efficient and stable data processing capabilities. For data storage and management, the system uses MongoDB, a non-relational database, which ensures efficient data reading and writing while flexibly handling various data formats.
[0399] In modern application development, front-end and back-end separation has become a popular trend. The main advantage of this design pattern is that it separates the user interface (UI) from data processing and storage functions. This allows development teams to work more flexibly, accelerates the iteration process, and ensures higher system stability and maintainability. The front-end focuses on user experience, while the back-end handles data logic and business processes.
[0400] This example uses React to build the user interface or UI components. React is an open-source JavaScript library developed and maintained by Facebook. React's main advantage lies in its virtual DOM mechanism, meaning that parts of the page are only updated when necessary, thus providing a smooth user experience. Furthermore, React supports component-based development, enhancing code reusability and significantly improving development efficiency. In this project, the React framework was chosen to ensure a consistent and high-performance experience across different devices.
[0401] This embodiment combines Express.js, a popular Node.js web application framework, to quickly build stable, high-performance backend services. Node.js is a JavaScript runtime based on the V8 engine. It allows developers to write JavaScript code on the server side, providing a unified development experience. Node.js's non-blocking I / O and event-driven design make it ideal for handling a large number of concurrent requests, ensuring efficient data processing.
[0402] In the EAP hearing analysis system project, MongoDB provides a highly flexible and scalable solution because the data format may change over time and with evolving needs. MongoDB is a document-oriented, non-relational database that stores data in a JSON-like document format, featuring a flexible data model and high read / write performance. Compared to traditional relational databases, MongoDB does not require a fixed table structure, meaning developers can add or modify data fields more flexibly.
[0403] To ensure seamless integration between the front-end and back-end, this implementation utilizes Docker containerization technology, enabling the application and its dependencies to run consistently in any environment. CI / CD tools, such as Jenkins, are also used to automate testing and deployment processes, ensuring stable and secure releases for every code update.
[0404] During the development of this embodiment, a series of specialized software programs were used to support data analysis and algorithm development. MATLAB and Python were the main development tools, providing powerful data processing and algorithm implementation capabilities.
[0405] Furthermore, TensorFlow and PyTorch support the design and training of deep learning models. To ensure software stability and security, this embodiment also uses Git and Docker for version control and software deployment.
[0406] The EAP auditory analysis device is not just a simple tool, but a standardized EAP device that uses a variety of technologies, such as semi-closed professional questioning guidance, speech-to-text conversion, keyword recognition, and voice tone analysis, to help monitor employees' emotional state.
[0407] The widespread adoption of this technology will not only improve communication between leaders and employees, but also provide a deeper understanding of employee psychology, increasing organizational satisfaction and a sense of belonging, thus bringing breakthroughs and innovations to the field of mental health research. Furthermore, the implementation of this project will not only enhance communication between leaders and employees, but also lead to greater employee satisfaction and a stronger sense of belonging within the organization.
[0408] Through the above design, this invention not only helps improve individual psychological well-being but also promotes a healthy workplace environment, thereby contributing to increased individual productivity and team cohesion. Simultaneously, this strategy provides company leaders and employees with a more constructive tool to better understand and support employees' psychological needs, thus contributing to building a harmonious work environment.
[0409] (1) The EAP hearing analysis instrument of the present invention, based on the psychological report generated by the EAP hearing analysis instrument, can effectively improve the communication effect between the two parties in the dialogue, so that both parties can better understand each other's mood and effectively improve the communication efficiency of both parties.
[0410] (2) When this invention is applied to EAP (Employee Assistance Program), leaders or relevant ideological and political workers can obtain a more scientific and objective psychological analysis report through this instrument. This provides more accurate guidance for the heart-to-heart talk with employees and can identify or prevent possible psychological balance or problems of employees in a targeted manner.
[0411] (3) The widespread application of this invention, by emphasizing prevention and enhancing individual psychological adaptability, can establish a more positive and acceptable mental health culture. This tool aims to prevent problems in advance, helping individuals improve their psychological adaptability and flexibility in coping with challenges. It enhances the overall mental well-being of individuals and prevents the development of potential problems.
[0412] (4) This invention not only helps improve individual psychological well-being but also promotes a healthy workplace environment, thereby contributing to increased individual productivity and team cohesion. Furthermore, this strategy provides company leaders and employees with a more constructive tool to better understand and support employees' psychological needs, thus contributing to the creation of a harmonious work environment.
[0413] Note: The use of voice information mentioned in this invention will be communicated to the person being evaluated (such as an employee) after recording, and consultation will be conducted with them. This complies with the Civil Code of the People's Republic of China and the Personal Information Protection Law of the People's Republic of China, and does not infringe on the privacy rights of the person being evaluated (such as an employee).
[0414] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An analysis method for an EAP (Earning Aptitude Test) auditory analyzer, employing a psychological assessment model. This model has several dimensions, each containing several quantitative questions. Its characteristics are as follows: Each quantitative question has at least one preset keyword, and the analysis method includes the following steps: S1. Acquire speech information; identify the phoneme boundaries of the speech information, and divide the speech information into several speech segments according to the phoneme boundaries; S2. Extract the duration and amplitude features of each speech segment from the speech information, combine the duration and amplitude features of each speech segment to form a fusion feature vector of each speech segment, input the fusion feature vector of the speech segment into the emotion classification model, and obtain the emotion quantification value of each speech segment from the emotion classification model. S3. Based on the duration characteristics of each speech segment, convert the speech information into text segments corresponding to the speech segments, and combine all text segments to form a text. S4. Traverse the text, find the text fragments corresponding to the preset keywords, and then classify them into the corresponding sub-pattern of the quantitative problem corresponding to the preset keywords. S5. Match the text fragments in each sub-pattern with the emotion quantification value; S6. Perform sentiment analysis and semantic correlation analysis on the text fragments in each sub-pattern to determine whether the quantitative problem corresponding to the sub-pattern is positive sentiment, negative sentiment, or neutral sentiment. If the quantitative question corresponding to the sub-pattern is determined to have a positive sentiment tendency, the score for the quantitative question corresponding to the sub-pattern is 1. If the quantitative question corresponding to the sub-pattern is determined to have a negative sentiment tendency, the score for the quantitative question corresponding to the sub-pattern is -1. If the quantitative question corresponding to the sub-pattern is determined to have a neutral sentiment tendency, then based on the matching results in S5, it is determined whether the emotional quantitative value corresponding to the text fragment in the sub-pattern is higher than the defined level. If the emotional quantitative value corresponding to the text fragment in the sub-pattern is higher than the defined level, the score for the quantitative question corresponding to the sub-pattern is 1. If the emotional quantitative value corresponding to the text fragment in the sub-pattern is equal to the defined level, the score for the quantitative question corresponding to the sub-pattern is 0. If the emotional quantitative value corresponding to the text fragment in the sub-pattern is lower than the defined level, the score for the quantitative question corresponding to the sub-pattern is -1. S7. By judging the sum of the scores of all quantitative questions for each dimension, if the sum of the scores for that dimension is positive, then the dimension performs well; if the sum of the scores for that dimension is 0, then the dimension performs well; if the dimension is negative, then the dimension performs poorly. S8. Automatically generate psychological reports based on performance in each dimension.
2. The analysis method of the EAP hearing analyzer according to claim 1, characterized in that, Step S1 includes: S1.1 Directional radio deployment: Multiple directional radios are deployed in the radio reception area to capture voice information. S1.2 Audio signal preprocessing: Digital filtering technology is used to sequentially reduce noise, adjust gain, and cancel echo in the captured audio information. S1.3 Data standardization and formatting: Transform the pre-processed audio signal data into a unified standard format; S1.4 Identify the phoneme boundaries of standard format speech information and divide the speech information into several speech segments based on the phoneme boundaries; This also includes data privacy protection, which involves storing all voice information data obtained in S1.1 to S1.3 and encrypting it during storage and transmission.
3. The analysis method of the EAP hearing analyzer according to claim 1, characterized in that, Step S2 is as follows: S2.1 Extract the duration and amplitude features of each speech segment from the speech information; The duration characteristics of each speech segment include the average duration of the speech segment, the standard deviation of the speech segment duration, and the average duration of the pause. The duration of each speech segment is determined based on the phoneme boundaries, and the average duration of all speech segments is obtained by averaging the durations of all speech segments. The standard deviation of the speech segment duration is calculated based on the duration of all speech segments and the average duration of the speech segment. The duration of silent segments is determined based on the phoneme boundaries, and the average duration of all silent segments is obtained by averaging the durations of all silent segments. The amplitude features of each speech segment include the average amplitude of the speech segment, the standard deviation of the speech segment amplitude, and the duration of the energy distribution of the speech segment above the threshold. The short-time energy of each frame is calculated, and the average amplitude of the speech segment is obtained by averaging the values. The amplitude envelope is analyzed, and the peaks and troughs are extracted to calculate the standard deviation of the speech segment amplitude. Segments with short-time energy above the threshold are detected, and their duration is counted to obtain the duration of the energy distribution of the speech segment above the threshold. The threshold is twice the average short-time energy of the complete speech segment. S2.2 Combine the six features of speech segment average duration, speech segment duration standard deviation, pause average duration, speech segment average amplitude, speech segment amplitude standard deviation, and speech segment energy distribution duration above the threshold into a fusion feature vector in order. S2.
3. Input the fused feature vector of the speech segments into the emotion classification model, and obtain the emotion quantification value of each speech segment from the emotion classification model, including the following steps: The fused feature vectors are standardized and normalized. The fused feature vectors after standardization and normalization are divided into training and testing sets; The emotion classification model is trained and evaluated using training and test sets. The emotion classification model includes support vector machines, K-nearest neighbors, and neural networks. The standardized and normalized fused feature vectors are input into the trained emotion classification model. The emotion classification model classifies the standardized and normalized fused feature vectors and gives the emotion quantification value of each speech segment.
4. The analysis method of the EAP hearing analyzer according to claim 3, characterized in that, In step S1, the number of times the signal crosses the zero axis within each frame is calculated. The zero-crossing rate (ZCR) is calculated based on the number of times the signal crosses the zero axis within each frame. The boundaries of phonemes in the speech signal are determined based on the ZCR. The formula for calculating the ZCR is: Where xt represents the amplitude of the t-th sampling point, 1{sign(xt)≠sign(xt+1)} is the indicator function, T is the number of sampling points in the frame, and the sign function sign(xt) is used to determine the sign of the signal at the t-th sampling point; In step S2.1, the Short Time Energy (STE) of each frame of the speech segment is obtained. The formula for calculating the Short Time Energy (STE) is as follows: Where w(n) is the window function, x(n) is the signal sample, and N is the window length, which represents the number of sampling points of the signal. N is a natural number. n is the index of the signal sampling point; The signal of each frame of the speech segment is converted to the frequency domain to obtain the amplitude X(k) of the frequency component. The amplitude spectrum is extracted based on the amplitude X(k) of the frequency component. The formula for obtaining the amplitude X(k) of the frequency component is as follows: Where X(k) is the amplitude of the frequency component, k is the frequency index, and Y is the number of points in the FFT. n is the index of the signal sampling point, x(n) is the signal sample, e is the complex exponent kernel, j is the imaginary unit, and 2πkn / Y represents the angle or phase. In step S2.1, peaks and troughs are identified through amplitude envelope detection. Dilation and erosion operations in morphological calculations are applied to calculate the duration and amplitude differences between peaks. The formula for the dilation operation is: The formula for corrosion calculation is: Where ⊕ represents the expansion operation, This represents the erosion operation, where f is the input signal, b is the structuring element, and B is the domain of the structuring element.
5. The analysis method of the EAP hearing analyzer according to claim 1, characterized in that, Step S4 includes the following steps: S4.1 Text preprocessing: Remove stop words and punctuation marks from the text, and simplify the text using lemmatization or stemming. S4.2 Calculate the similarity between the text fragment and the preset keyword list using TF-IDF or word embedding models. The similarity calculation formula is as follows: Where A is a text fragment and B is a preset keyword vector; S4.3 For each sub-pattern, find the text segment that best matches the preset keyword and classify it into the sub-pattern corresponding to the preset keyword.
6. The analysis method of the EAP hearing analyzer according to claim 1, characterized in that, In step S3, when generating text, the timestamp information of each text segment is retained; Step S5 includes the following steps: By using the timestamp information of the text segment in step S3 to associate with the timestamp information of the corresponding audio segment in step S2, the emotion quantization value data of the audio segment under the same timestamp information is obtained as the emotion quantization value data of the text segment, thus forming a match.
7. The analysis method of the EAP hearing analyzer according to claim 1, characterized in that, Step S6 includes the following steps: S6.
1. Text fragment preprocessing for each sub-pattern: including removing redundant information and part-of-speech tagging; Remove redundant information: Use NLTK or spaCy to segment the text fragments of each sub-pattern and remove stop words, perform lemmatization and stemming to unify the word form; Part-of-speech tagging: The part-of-speech tagging of each word in the text fragment of each subpattern is marked using the NLTK's pos_tag() function; S6.2, Preset Keyword Recognition and Matching: Includes sentiment dictionary matching and context analysis: Sentiment dictionary matching: Using a predefined sentiment dictionary, the preset keywords of each sub-pattern are matched with their corresponding sentiment scores, and the similarity between the preset keywords and sentiment words is calculated, using cosine similarity as the metric. Where A and B are the vectors of the text fragments and preset keywords corresponding to the sub-patterns, respectively; Contextual analysis: Using VADER or TextBlob sentiment analysis tools, the sentiment tendency of preset keywords is further judged in the context of the text fragments in the corresponding sub-patterns, and the sentiment scores of sentences are extracted from the text fragments. The overall sentiment tendency of the sub-pattern is calculated by combining the weights of the preset keywords. S6.3, Assigning Sentiment Scores: Assign a sentiment score to each preset keyword: positive keywords are assigned +1, negative keywords are assigned -1, and neutral keywords are assigned 0; if a preset keyword appears in multiple contexts, all sentiment tendencies are considered comprehensively. S6.
4. Based on the assignment of sentiment scores of the preset keywords in S6.3, combined with the quantitative questions of the corresponding sub-patterns in the psychological assessment model, the sentiment tendency of the quantitative questions corresponding to the sub-patterns is judged. If the quantitative questions corresponding to the sub-patterns are judged to have a positive sentiment tendency, the quantitative questions corresponding to the sub-patterns are judged to have a score of 1. If the quantitative questions corresponding to the sub-patterns are judged to have a negative sentiment tendency, the quantitative questions corresponding to the sub-patterns are judged to have a score of -1. If the quantitative questions corresponding to the sub-patterns are judged to have a neutral sentiment tendency, then according to the matching results in S5, it is judged whether the sentiment quantitative value corresponding to the text fragments in the sub-patterns is higher than the defined level. If the sentiment quantitative value corresponding to the text fragments in the sub-patterns is higher than the defined level, the quantitative questions corresponding to the sub-patterns are judged to have a score of 1. If the sentiment quantitative value corresponding to the text fragments in the sub-patterns is equal to the defined level, the quantitative questions corresponding to the sub-patterns are judged to have a score of 0. If the sentiment quantitative value corresponding to the text fragments in the sub-patterns is lower than the defined level, the quantitative questions corresponding to the sub-patterns are judged to have a score of -1. The horizontal line is defined as (highest emotional quantification value + lowest emotional quantification value) / 2.
8. The analysis method of the EAP hearing analyzer according to claim 1, characterized in that, In step S4, if the number of sub-patterns containing text fragments in each dimension accounts for 60% or more of all sub-patterns in that dimension, then the performance evaluation of that dimension is valid in step S7; otherwise, it is invalid. In S8, the validity of the dimensional performance evaluation is presented in writing.
9. An EAP (Early Access Point) auditory analysis device, characterized in that, It includes at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform an analysis method of an EAP hearing analyzer according to any one of claims 1 to 8.
10. An EAP hearing analyzer, characterized in that, include: The data acquisition module acquires voice information; The speech intonation analysis module extracts the speech information collected by the data acquisition module, divides the speech information into several speech segments, and obtains the emotion quantification value of each speech segment based on the fundamental frequency, energy and morphological characteristics of the speech segments. The speech-to-text module converts speech information into text segments corresponding to the speech segments, and all text segments form a text file. The keyword recognition module traverses the text and matches the text fragments corresponding to preset keywords with the quantitative questions in various dimensions of the psychological assessment model; The psychological report generation module matches text fragments with the emotional quantification values of corresponding audio fragments; then, based on the text fragments and emotional quantification values that match the quantitative questions, it obtains the answer tendencies for the quantitative questions; it integrates the answer tendencies of all quantitative questions for that dimension to obtain the performance of that dimension; it automatically generates psychological reports based on the performance of each dimension; and it provides data visualization interpretation of the generated psychological reports. The data storage and management module saves all data generated by the data acquisition module, speech-to-text module, speech intonation analysis module, and psychological report generation module.