Voice processing method, electronic device, storage medium and program product
By analyzing the voice and physiological characteristics of the object, evaluating psychological stress in real time and dynamically adjusting the voice output of the intelligent agent, the problem of the lack of stress situation simulation of the intelligent agent in customer service and education is solved, and the interactive intelligence and stress resistance are improved.
Patent Information
- Application Number
- CN202510876320.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies for the interaction between intelligent agents and objects in fields such as customer service and education lack the ability to simulate real-life stressful situations and are unable to dynamically adjust the difficulty of training. This leads to insufficient stress tolerance among agents, easily causing language conflicts and customer complaints. Furthermore, the evaluation dimensions are single and cannot fully reflect the agent's state under stress.
By analyzing the voice and physiological characteristics of the subject, evaluating the psychological stress index and changing trends in real time, and dynamically adjusting the voice output of the intelligent agent to adapt to different dialogue scenarios and psychological states, we can build a seated training system and online education and tutoring system for real stress scenarios.
It improves the intelligence level and dynamic adaptability of the intelligent agent in voice interaction, enhances the stress resistance and learning experience of the seats, and reduces the occurrence of language conflicts.
Smart Images

Figure CN120690197A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to artificial intelligence technology, and in particular to a speech processing method, electronic device, storage medium and program product. Background Art
[0002] With the rapid development of artificial intelligence (AI), intelligent agents are increasingly being used in customer service, psychological counseling, education, and other fields to interact with objects through voice. During these interactions, intelligent agents can generate corresponding voice responses based on preset rules or models to adapt to different conversation scenarios. Summary of the Invention
[0003] The embodiments of the present application provide a voice processing method, electronic device, storage medium and program product, which can improve the intelligence level of an intelligent agent in voice interaction.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] This embodiment of the present application provides a speech processing method, the method comprising:
[0006] determining a real-time psychological stress index of a subject in conversation with the agent based on first speech data of the subject;
[0007] determining a change trend of the subject's psychological stress based on the subject's historical psychological stress index;
[0008] Based on the psychological stress change trend and the real-time psychological stress index, the second voice data of the intelligent agent is adjusted to obtain third voice data, and the third voice data is used as the voice data for the intelligent agent to communicate with the object.
[0009] The present invention provides a speech processing device, including:
[0010] a first determining module for determining a real-time psychological stress index of a subject having a conversation with the agent based on first speech data of the subject;
[0011] a second determining module, determining a change trend of the subject's psychological stress based on the subject's historical psychological stress index;
[0012] An adjustment module is used to adjust the second voice data of the intelligent agent based on the psychological pressure change trend and the real-time psychological pressure index to obtain third voice data, and use the third voice data as the voice data for the intelligent agent to communicate with the object.
[0013] An embodiment of the present application provides an electronic device, comprising:
[0014] a memory for storing computer-executable instructions or computer programs;
[0015] The processor is used to implement the speech processing method provided in the embodiment of the present application when executing the computer-executable instructions or computer programs stored in the memory.
[0016] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the speech processing method provided in the embodiment of the present application when executed by a processor.
[0017] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the speech processing method provided in the embodiment of the present application is implemented.
[0018] The embodiments of the present application have the following beneficial effects:
[0019] The subject's first speech data is used to determine their real-time psychological stress index, and the trend of their stress is assessed in combination with their historical stress index. Based on this real-time stress index and stress trend, the agent's second speech data is intelligently adjusted to produce third speech data tailored to the current conversation. The agent can adjust its speech output based on the subject's stress, enhancing its intelligence and dynamic adaptability in voice interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a schematic diagram of the architecture of the speech processing system provided in an embodiment of the present application;
[0021] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;
[0022] Figure 3 This is a flow diagram of the voice processing method provided in the embodiment of the present application. Figure 1 ;
[0023] Figure 4 This is a flow diagram of the voice processing method provided in the embodiment of the present application. Figure 2 ;
[0024] Figure 5 This is a flow diagram of the voice processing method provided in the embodiment of the present application. Figure 3 ;
[0025] Figure 6 This is a flow diagram of the voice processing method provided in the embodiment of the present application. Figure 4 ;
[0026] Figure 7This is a module diagram of the speech processing system provided in an embodiment of the present application;
[0027] Figure 8 This is a flow chart of a pressure scenario construction module provided in an embodiment of the present application;
[0028] Figure 9 This is a flow chart of a real-time conversation pressure injection module provided in an embodiment of the present application;
[0029] Figure 10 This is a flow chart of the multidimensional data acquisition module provided in an embodiment of the present application;
[0030] Figure 11 is a flow chart of the pressure assessment engine provided in an embodiment of the present application;
[0031] Figure 12 It is a flow chart of the dynamic regulator provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0033] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0034] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0035] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0036] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0037] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0038] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0039] 1) In response to: used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0040] 2) Voice data: This refers to the sound signals emitted by users or intelligent robots, captured by microphones or other audio collection devices, typically stored and processed in digital form. Voice data contains voice characteristics such as pitch, speaking rate, loudness, and rhythm, and can be used to analyze the speaker's emotional state and psychological stress.
[0041] 3) Psychological Stress Index: A quantitative indicator used to reflect a user's psychological stress level. The psychological stress index can be calculated by analyzing speech characteristics (such as speech rate changes and intonation fluctuations) or combining other physiological characteristics, and is usually expressed in numerical form.
[0042] 4) Speech synthesis parameters: These are adjustable variables used to define the speech characteristics of speech data, including but not limited to pitch, speaking rate, volume, and timbre. By adjusting speech synthesis parameters, the presentation of the synthesized speech data can be changed to better meet specific needs.
[0043] 5) Time Window: This refers to a fixed or sliding time period set in data analysis, used to extract and analyze data within a specific time period. For example, in psychological stress trend analysis, time windows of varying lengths can be used to observe changes in a user's short-term or long-term psychological stress index.
[0044] 6) Speech features: characteristic parameters extracted from speech data, including spectral features (such as Mel-frequency cepstral coefficients), time domain features (such as speaking rate and pause time), and energy features (such as loudness).
[0045] 7) Physiological characteristics: refers to the measurable characteristics of the user at the physiological level, such as heart rate, skin conductance, respiratory rate, etc.
[0046] The agent training system mainly uses conventional conversation robots to train new agents, such as conventional multi-round conversation simulations and real-time robot voice conversation simulations, but lacks special stress training for agents and cannot effectively train the agents' ability to withstand pressure. The relevant technologies have the following technical problems: 1) Lack of stress scenario simulation: The system can only perform general conversation training and cannot simulate various stress situations in real customer service scenarios (such as customer anger, urgent issues, multi-tasking, etc.). 2) Single evaluation dimension: Performance is evaluated only by voice recognition accuracy and conversation completion, which cannot fully reflect the true state of the agent under pressure. 3) No dynamic difficulty adjustment: The training difficulty is fixed, and the stress level cannot be adjusted in real time according to the agent's performance. 4) Feedback lag: The analysis of training results is usually carried out after the end, and real-time guidance cannot be provided. 5) Easy to cause customer complaints: There is no targeted stress training for agents. If the agent's ability to withstand pressure is not good, it is easy to have language conflicts with customers, causing customer complaints.
[0047] Based on the problems existing in the related art, the embodiments of the present application provide a voice processing method, an electronic device, a storage medium and a program product, which can improve the intelligence level of the intelligent agent in voice interaction. The following describes an exemplary application of the voice processing device provided by the embodiment of the present application, which is an electronic device for implementing the voice processing method. The electronic device provided by the embodiment of the present application can be implemented as various types of terminals such as laptops, tablet computers, desktop computers, set-top boxes, smart phones, smart speakers, smart watches, smart TVs, and vehicle-mounted terminals, and can also be implemented as servers. Below, an exemplary application when the electronic device is implemented as a terminal or a server will be described.
[0048] See also Figure 1 , Figure 11 is a schematic diagram of the architecture of the voice processing system provided in the embodiment of the present application. In order to perform voice processing operations, a voice processing application can be provided. For example, the voice processing application can be an application dedicated to voice processing, or it can be a functional module in other applications (such as a voice processing module in a financial application, etc.). The voice processing system 100 in the embodiment of the present application includes at least a terminal 400, a network 300 and a server 200, wherein the server 200 is a server for the voice processing application. The server 200 can constitute the voice processing device in the embodiment of the present application, that is, the voice processing method in the embodiment of the present application is implemented by the server 200. The terminal 400 is connected to the server 200 via the network 300, and the network 300 can be a wide area network or a local area network, or a combination of the two.
[0049] See also Figure 1 , the user can perform interactive operations on the client side of the voice processing application through the terminal 400, and the interactive operations can be, for example, establishing a voice dialogue with the intelligent agent. After receiving the interactive operation of the user, the client sends a voice processing request to the server 200 through the network 300. After receiving the voice processing request, the server 200 responds to the voice processing request sent by the terminal and determines the real-time psychological stress index of the object based on the first voice data of the object in dialogue with the intelligent agent; the server 200 determines the psychological stress change trend of the object based on the historical psychological stress index of the object; the server 200 adjusts the second voice data of the intelligent agent based on the psychological stress change trend and the real-time psychological stress index to obtain third voice data, and uses the third voice data as the voice data for the dialogue between the intelligent agent and the object.
[0050] In some embodiments, the terminal 400 can also execute the voice processing method of the embodiment of the present application, that is, the terminal 400 can determine the real-time psychological stress index of the object based on the first voice data of the object communicating with the intelligent agent; the terminal 400 can determine the psychological stress change trend of the object based on the historical psychological stress index of the object; the terminal 400 can adjust the second voice data of the intelligent agent based on the psychological stress change trend and the real-time psychological stress index to obtain third voice data, and use the third voice data as the voice data for the intelligent agent to communicate with the object.
[0051] In some embodiments, in customer service scenarios, new agents need to undergo rigorous training to cope with complex and changing customer communication scenarios. To this end, a seat training system based on a conversational robot can be constructed. The agent training system records the first voice data of the agent (i.e., the object) by simulating real stress scenarios (such as angry customer complaints or emergency problem handling), and calculates the real-time psychological stress index based on the first voice data. At the same time, the agent training system analyzes the trend of psychological stress changes based on the agent's historical psychological stress index. On this basis, the conversational robot (i.e., the intelligent agent) adjusts the second voice data it outputs according to the agent's real-time psychological stress index and the trend of psychological stress changes (such as increasing the speaking speed to increase the sense of urgency or slowing down the speaking speed to provide guidance), and generates more targeted third voice data, thereby helping agents adapt to high-pressure work environments more quickly and improve their adaptability.
[0052] In some embodiments, in online education scenarios, an intelligent robot (i.e., an intelligent agent) acts as a virtual teacher to provide one-on-one tutoring services to students (i.e., subjects). When a student expresses tension or anxiety during the learning process, the features in the first voice data change, and the system calculates the student's real-time psychological stress index based on the first voice data. The intelligent robot will adjust its teaching strategy based on the real-time psychological stress index and the trend of psychological stress changes. For example, it may generate new third voice data by reducing the speed of explanation, using a gentler tone, or inserting motivational words, thereby creating a relaxed learning atmosphere and improving the student's learning experience.
[0053] In some embodiments, the electronic device may be Figure 1 Terminal 400 in Figure 2 , Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application, Figure 2 The electronic device shown includes: at least one processor 410, a memory 450, at least one network interface 420 and a user interface 430. The various components in the electronic device are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 440 is not shown in FIG. Figure 2 Various buses are labeled as bus system 440 .
[0054] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0055] The user interface 430 includes one or more output devices 431 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0056] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0057] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0058] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.
[0059] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;
[0060] A network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wi-Fi, and Universal Serial Bus (USB);
[0061] a presentation module 453 for enabling presentation of information via one or more output devices 431 (e.g., a display screen, a speaker, etc.) associated with the user interface 430 (e.g., a user interface for operating peripheral devices and displaying content and information);
[0062] The input processing module 454 is configured to detect one or more user inputs or interactions from one of the one or more input devices 432 and to translate the detected inputs or interactions.
[0063] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 The speech processing device 455 stored in the memory 450 is shown. The speech processing device 455 can be software in the form of a program or plug-in, and includes the following software modules: a first determination module 4551, a second determination module 4552, and an adjustment module 4553. These modules are logical and can be arbitrarily combined or further separated according to the functions implemented. The functions of each module will be described below.
[0064] In other embodiments, the apparatus provided in the embodiments of the present application may be implemented in hardware. As an example, the apparatus provided in the embodiments of the present application may be a processor in the form of a hardware decoding processor, which is programmed to execute the speech processing method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor may be one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0065] The following describes the speech processing method provided by the embodiment of the present application. As mentioned above, the electronic device that implements the speech processing method of the embodiment of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.
[0066] See also Figure 3 , Figure 3 This is a flow diagram of the voice processing method provided in the embodiment of the present application. Figure 1 , will combine Figure 3 The steps shown are explained as Figure 3 As shown, the voice processing method is described as an example in which the execution subject is a server. The method includes the following steps 101 to 103:
[0067] In step 101, a real-time psychological stress index of an object in conversation with an agent is determined based on first voice data of the object.
[0068] Here, the intelligent agent is the subject that has a conversation with the object, and the intelligent agent can be an intelligent robot or a voice interaction system. The intelligent agent generates second voice data and has a conversation with the object based on the second voice data. The object is the subject that has a conversation with the intelligent agent, and the object can be a user or other intelligent robot, etc. The first voice data is the voice signal generated by the object during the conversation with the intelligent agent (such as an intelligent robot). The real-time psychological stress index is a quantitative indicator obtained after real-time analysis of the first voice data output by the object at the current time during the current conversation, which is used to reflect the user's psychological stress level at the current time. The current time can be the current specific moment, such as xx hours xx minutes xx seconds, or it can be the current time period, such as within 1 minute.
[0069] In an embodiment of the present application, the first voice data of the object can be collected by an audio acquisition device (such as a microphone), and the first voice data is pre-processed (such as noise reduction and framing) to obtain pre-processed voice data. Voice features are extracted from the pre-processed voice data. The voice features may include fundamental frequency jitter features, amplitude jitter features, voice pause frequency, and speech rate change rate, etc. In an embodiment of the present application, there is no limitation on the method of determining the real-time psychological stress index of the object based on voice features. For example, the normalized voice features can be weighted and summed to obtain the real-time psychological stress index of the object. Alternatively, a machine learning model (such as a support vector machine or a deep neural network) is pre-trained, and the extracted voice features are input into the machine learning model to calculate the real-time psychological stress index. The machine learning model is trained with a large amount of labeled data and can accurately predict the psychological stress index based on the voice features.
[0070] For example, let's assume the subject is a new agent, and the agent is an intelligent robot. The new agent is undergoing customer service training and is simulating a complaint scenario involving an "angry customer." The new agent's first voice data is recorded. After analyzing and calculating the first voice data, the new agent's real-time psychological stress index is 78, indicating that the new agent is currently experiencing high psychological stress.
[0071] In some embodiments, before determining the subject's real-time psychological stress index based on first voice data of a subject in a conversation with an agent, in response to the subject selecting a business scenario, a conversation text and third speech synthesis parameters configured for the business scenario are determined. The conversation text is converted into fourth voice data, and the fourth voice data is adjusted based on the third speech synthesis parameters to obtain second voice data of the agent.
[0072] Here, the speech processing system can be pre-configured with multiple business scenarios, each corresponding to different dialogue text and speech synthesis parameters. Multiple business scenarios can include customer service, mental health counseling, educational counseling, etc. In the customer service scenario, business scenarios can also include customer complaints, emergency issues, multitasking, etc. The dialogue text is pre-defined text content configured according to the business scenario and is used to guide the conversation between the intelligent agent and the object. The third speech synthesis parameter is the speech synthesis parameter pre-set for the business scenario, such as speech rate, pitch, conversation interruption frequency, background noise level, etc., which is used to adjust the generated speech data to meet the needs of the specific business scenario.
[0073] In an embodiment of the present application, before interacting with an intelligent agent, the object selects a specific business scenario from a plurality of business scenarios. Based on the business scenario selected by the object, the corresponding dialogue text is extracted from a preset dialogue text library. The dialogue text is standardized content designed according to actual business needs, which can guide the dialogue process and provide reference information. Through text-to-speech technology, the selected dialogue text is converted into initial fourth voice data. According to the third voice synthesis parameters pre-configured for the business scenario, the fourth voice data is adjusted to generate second voice data that better meets the needs of the business scenario. The adjusted second voice data is used as the initial voice output of the intelligent agent for dialogue with the object.
[0074] For example, a new agent can select the "Handling Customer Complaints" business scenario to simulate handling an angry customer's complaint. A pre-set conversational text is loaded based on this business scenario, for example: "We apologize for the inconvenience. Please describe the issue in detail. We will resolve it as soon as possible." The agent then loads the third speech synthesis parameters that match the "Handling Customer Complaints" business scenario, for example: speaking rate: 140 words per minute (slightly slower to demonstrate patience); pitch: 3 (normalized value, indicating a gentle and non-aggressive tone); interruption frequency: 6 times per minute (with appropriate pauses to demonstrate listening and reflection); background noise level: level 2 (simulating a quiet but slightly distracting customer service environment). The conversational text is converted into initial fourth speech data, which is then adjusted based on the third speech synthesis parameters to generate the final second speech data. For example, the speaking rate can be adjusted to 140 words per minute, the pitch set to 3, six pauses inserted, and appropriate background noise (such as keyboard tapping) added. The adjusted second speech data serves as the initial output of the intelligent robot, which then begins a simulated conversation with the new agent.
[0075] The embodiment of the present application determines the conversation text and third speech synthesis parameters that match the business scenario selected by the response object, and generates the second speech data of the intelligent agent based on these parameters, which can generate speech output that is more in line with the needs of the specific scenario and ensure that the content and expression of the conversation conform to the actual application environment. The introduction of third speech synthesis parameters (such as speech speed, tone, conversation interruption frequency and background noise level) makes the generated speech data closer to the characteristics of real human conversation, enhancing the naturalness and immersion of the interaction. The speech synthesis parameters can be dynamically adjusted according to different business scenarios, without the need to fix a single speech output mode, so as to adapt to diverse conversation needs.
[0076] In some embodiments, see Figure 4 , Figure 4 It is shown that in step 101, determining the real-time psychological stress index of the object based on the first voice data of the object in dialogue with the intelligent agent can be achieved by the following steps 1011 to 1013:
[0077] In step 1011, speech features of the first speech data are extracted and normalized to obtain first features.
[0078] Here, the number of speech features can be one or more. The speech features can be fundamental frequency jitter features, amplitude jitter features, speech pause frequency, and speech rate change rate, etc. The fundamental frequency jitter feature reflects the range fluctuation of the sound frequency, the amplitude jitter feature reflects the range fluctuation of the sound amplitude, the speech pause frequency is the number of pauses during speaking per unit time, and the speech rate change rate is the rate of change of speech rate over time. Extract multiple speech features of the first speech data output by the object at the current moment, normalize each speech feature separately, and obtain the first feature corresponding to each speech feature, ensuring that all first features are in the same numerical range (for example, 0-100).
[0079] Exemplarily, normalization of the fundamental frequency jitter feature can be achieved in the following manner: calculating a first difference between the actual fundamental frequency jitter feature and the minimum fundamental frequency jitter feature, calculating a second difference between the maximum fundamental frequency jitter feature and the minimum fundamental frequency jitter feature, and taking the ratio of the first difference to the second difference as the normalized first feature corresponding to the fundamental frequency jitter feature.
[0080] In step 1012, the physiological characteristics of the object are normalized to obtain a second characteristic.
[0081] Here, physiological characteristics are physiological data acquired by sensors (e.g., heart rate monitors or galvanic skin response meters), such as heart rate, respiratory rate, and galvanic skin response. Secondary characteristics are normalized physiological characteristics converted to a value between 0 and 100. The secondary characteristics are obtained by collecting and normalizing the physiological characteristics of the subject at the current moment.
[0082] For example, the physiological characteristic is heart rate variability (HRV) data, which is collected via Bluetooth connection to a physiological characteristic collection device (such as a smart bracelet) worn by the user, with a sampling frequency of 1 Hz.
[0083] In step 1013, a preset operation is performed on the second feature and the first feature to obtain a real-time psychological stress index of the subject.
[0084] Here, performing a preset operation on the second feature and the first feature to obtain the subject's real-time psychological stress index can be achieved by performing a weighted summation of the second feature and multiple first features to obtain the subject's real-time psychological stress index. This embodiment of the application does not limit the weight of the second feature or the weight of each first feature, and can be set based on needs.
[0085] For example, if the second feature is the HRV index, and the multiple first features are the voice jitter index (the weighted sum of the first features corresponding to the fundamental frequency jitter feature and the amplitude jitter feature), speech rate change rate, and pause frequency, the psychological stress index = 0.4 × HRV index + 0.3 × voice jitter index + 0.2 × speech rate change rate + 0.1 × speech pause frequency.
[0086] The embodiment of the present application combines voice features and physiological features, making full use of the emotional information in the voice data and the objective indicators of the physiological data, overcoming the limitations that may exist in a single data source, and thus significantly improving the accuracy and reliability of psychological stress assessment. Normalization of voice features and physiological features of different dimensions or ranges ensures the comparability between the features and avoids calculation deviations caused by differences in data scales. Through preset operation methods such as linear weighted summation or machine learning models, a flexible and scalable psychological stress index calculation framework is provided, which can adapt to different application scenarios and needs.
[0087] In some embodiments, for each question included in the second voice data of the intelligent agent, the voice data of the subject's response to the question in the first voice data is converted into text, and the standard response script preset for the question is extracted from the knowledge base. The similarity between the text and the standard response script is calculated. If the similarity is less than a preset threshold, the subject's response is determined to be incorrect. The number of incorrect responses in the subject's first voice data is counted, and this number, the second feature, and the first feature are weighted and summed to obtain the subject's real-time psychological stress index.
[0088] In step 102 , based on the historical psychological stress index of the subject, a change trend of the subject's psychological stress is determined.
[0089] Here, the historical psychological stress index is the psychological stress index of the object during the past interaction with the intelligent agent. The historical psychological stress index can be a psychological stress index recorded in the past. The past time can be a certain moment in the past or a certain time period. For example, the current moment is the 5th minute, the real-time psychological stress index calculated in the 5th minute is 58, the historical psychological stress index recorded in the 4th minute is 60, and the historical psychological stress index recorded in the 3rd minute is 40. The psychological stress change trend is the result obtained by statistically analyzing multiple historical psychological stress indices within a preset time range, reflecting the change pattern of the user's psychological stress index over time (such as rising, falling or stable).
[0090] In some embodiments, see Figure 5 , Figure 5 It is shown that determining the change trend of the psychological stress of the subject based on the historical psychological stress index of the subject in step 102 can be achieved by the following steps 1021 to 1023:
[0091] In step 1021 , the multiple historical psychological stress indexes of the subject are divided into multiple psychological stress index sequences according to multiple time windows.
[0092] Each psychological stress index sequence includes historical psychological stress indices within a corresponding time window.
[0093] Here, the time window is a pre-set time period, for example, a time window can be 30 seconds, 60 seconds, etc. The embodiment of the present application does not specifically limit the time window, and can be set voluntarily. Starting from the most recent time point (the current moment), with the length of each time window as the step length, multiple historical psychological stress indices within each time window are sequentially intercepted to form a psychological stress index sequence.
[0094] For example, assuming that multiple historical psychological stress indices are arranged in chronological order as 60, 65, 70, 75, 80, 85, and 90, the corresponding times are [0s, 10s, 20s, 30s, 40s, 50s, 60s], and the time window is 30s, then the psychological stress index sequence corresponding to the first time window is [60, 65, 70, 75], and the psychological stress index sequence corresponding to the second time window is [75, 80, 85, 90].
[0095] In step 1022 , an average value of the historical psychological stress indexes included in each psychological stress index sequence is calculated.
[0096] Here, for each psychological stress index sequence, the sum of the multiple historical psychological stress indices included in the psychological stress index sequence is calculated, and the ratio of the sum to the number of historical psychological stress indices included in the psychological stress index sequence is taken as the average value.
[0097] For example, the psychological stress index sequence corresponding to the first time window is [60, 65, 70, 75], and the average value = (60+65+70+75) / 4 = 67.5.
[0098] In step 1023, the psychological stress change trend of the subject is determined based on the average values corresponding to the multiple psychological stress index sequences.
[0099] Here, a linear regression analysis may be performed on the average values corresponding to the multiple psychological stress index sequences to obtain the psychological stress change trend of the determined object.
[0100] The embodiment of the present application can comprehensively reflect the short-term and long-term changing trends of the subject's psychological stress by dividing the historical psychological stress index into multiple psychological stress index sequences according to multiple time windows and calculating the average value of each sequence. By dividing into different time windows, it is possible to capture short-term fluctuations and evaluate long-term trends, thereby improving the comprehensiveness of the analysis. The calculation method based on the average value is easy to implement and has high computational efficiency, making it suitable for real-time application scenarios. The time window length can be flexibly adjusted according to actual needs, and is suitable for different business scenarios and analysis objectives.
[0101] In some embodiments, determining a subject's psychological stress change trend based on the average values corresponding to multiple psychological stress index sequences can be achieved in the following manner: First, performing linear regression on the average values corresponding to the multiple psychological stress index sequences based on the time windows corresponding to the multiple psychological stress index sequences to obtain a slope representing changes in psychological stress. Then, if the slope is greater than or equal to a first threshold, an upward trend is determined as the psychological stress change trend. If the slope is greater than a second threshold and less than the first threshold, a stable trend is determined as the psychological stress change trend, where the first threshold is greater than the second threshold. If the slope is less than or equal to the second threshold, a downward trend is determined as the psychological stress change trend.
[0102] Here, for each psychological stress index sequence, the minimum moment in the time window corresponding to the psychological stress index sequence can be used as the timestamp corresponding to the psychological stress index sequence. For example, if the time window is [0s-30s], the timestamp is 0s. Based on the time windows corresponding to multiple psychological stress index sequences, a linear regression is performed on the average values corresponding to the multiple psychological stress index sequences to obtain a slope representing the change in psychological stress. This can be achieved by calculating a first average value of the timestamps corresponding to the multiple psychological stress index sequences and a second average value of the average values corresponding to the multiple psychological stress index sequences. For each psychological stress index sequence, the difference between the timestamp corresponding to the psychological stress index sequence and the first average value is multiplied by the difference between the average value corresponding to the psychological stress index sequence and the second average value to obtain a first value. The first values of the multiple psychological stress index sequences are then summed to obtain a second value. For each psychological stress index sequence, the difference between the timestamp corresponding to the psychological stress index sequence and the first average value is squared to obtain a third value. The third values of the multiple psychological stress index sequences are summed to obtain a fourth value. The ratio of the second value to the fourth value is used as the slope of the change in psychological stress.
[0103] For example, there are three psychological stress index sequences, with timestamps x being [0, 30, 60] and average values y being [67.5, 82.5, 95]. The first average value = (0 + 30 + 60) / 3 = 30, and the second average value = (67.5 + 82.5 + 95) / 3 = 81.67. The second value = (0-30) × (67.5-81.67) + (30-30) × (82.5-81.67) + (60-30) × (95.0-81.67) = 825. The fourth value = (0-30) 2 +(30-30) 2 +(60-30) 2 = 1800. Slope = second value / fourth value = 825 / 1800 = 0.46.
[0104] It should be noted that the embodiments of the present application do not limit the values of the first threshold and the second threshold, and can be set based on actual needs. For example, the first threshold is 0.5 and the second threshold is -0.5. When the slope is greater than or equal to 0.5, the trend of psychological stress change is an upward trend. When the slope is less than or equal to -0.5, the trend of psychological stress change is a downward trend. When the slope is greater than -0.4 and less than 0.5, the trend of psychological stress change is a stable trend.
[0105] The embodiments of the present application can accurately quantify the changing trend of psychological stress through linear regression analysis and threshold judgment, and divide it into three states: rising, stable, and declining. Based on the linear regression model, it can objectively reflect the changing pattern of psychological stress over time. By adjusting the first and second thresholds, it can adapt to different application scenarios and needs. By quantifying the slope, it is possible to determine not only the trend direction but also the speed of change, providing more refined data support for subsequent voice adjustment strategies.
[0106] In step 103, based on the psychological stress change trend and the real-time psychological stress index, the second voice data of the agent is adjusted to obtain third voice data, and the third voice data is used as voice data for the agent to communicate with the object.
[0107] Here, the second voice data is the voice data output by the intelligent agent. According to the trend of changes in psychological stress and the real-time psychological stress index, the adjustment strategy for the second voice data is determined. The adjustment strategies may include the following: an upward trend and a high real-time psychological stress index: This indicates that the psychological stress of the subject is increasing rapidly, and a soothing adjustment strategy needs to be adopted, such as reducing the speaking speed, using a soft tone, increasing the frequency of pauses, etc. A downward trend and a low real-time psychological stress index: This indicates that the psychological stress of the subject is easing, and an incentive adjustment strategy can be adopted, such as appropriately speeding up the speaking speed, raising the tone, reducing the frequency of pauses, etc. A stable trend: This indicates that the psychological stress of the subject remains stable, and a neutral voice output style can be maintained, but the voice parameters are fine-tuned according to the real-time psychological stress index to adapt to the current state. The adjusted third voice data is used as the real-time voice output of the intelligent agent for conversations with the subject.
[0108] In some embodiments, see Figure 6 , Figure 6 It is shown that in step 103, the second voice data of the agent is adjusted based on the psychological stress change trend and the real-time psychological stress index to obtain the third voice data, which can be achieved by the following steps 1031 to 1034:
[0109] In step 1031 , a target value interval of the real-time psychological stress index is determined from a plurality of value intervals, and a target psychological stress level corresponding to the target value interval is determined.
[0110] The multiple value intervals are obtained by dividing the value range of the psychological stress index, and the value intervals correspond to the psychological stress levels one by one.
[0111] Here, the multiple value intervals are obtained by dividing the value range of the psychological stress index. The embodiment of the present application does not limit the method of dividing the value intervals, nor the correspondence between the value intervals and the psychological stress level, and can be set based on actual needs. The value of the psychological stress index is positively correlated with the psychological stress level.
[0112] For example, the multiple value intervals are [0, 30), [30, 60], [60, 80], and [80, +∞], corresponding to the psychological stress levels of low pressure, medium pressure, high pressure, and excessive pressure, respectively. If the real-time psychological stress index is 85, the target value interval for the real-time psychological stress index is [80, +∞], and the target psychological stress level is excessive pressure.
[0113] In step 1032 , a first adjustment factor for adjusting speech synthesis parameters is determined based on the psychological stress change trend and the target psychological stress level.
[0114] Here, the first adjustment factor is a parameter determined based on the psychological stress change trend and the target psychological stress level, and is used to indicate whether the speech synthesis parameter should be increased or decreased according to the first adjustment factor. When the first adjustment factor is a positive number, the speech synthesis parameter is increased; when the first adjustment factor is a negative number, the speech synthesis parameter is decreased. When the first adjustment factor is 0, the speech synthesis parameter remains unchanged.
[0115] In the embodiment of the present application, a correspondence between the psychological stress change trend, the psychological stress level and the adjustment factor can be pre-established and stored in a mapping table. The mapping table is queried based on the psychological stress change trend and the target psychological stress level to obtain the corresponding first adjustment factor.
[0116] In some embodiments, determining a first adjustment factor for adjusting speech synthesis parameters based on a psychological stress change trend and a target psychological stress level can be achieved by: first, determining a second adjustment factor corresponding to the psychological stress change trend and the target psychological stress level based on the correspondence between the psychological stress change trend, the psychological stress level, and the adjustment factor. Then, determining a first rate of change based on a first average value of historical psychological stress indices within a first time window and a second average value of historical psychological stress indices within a second time window, and determining a second rate of change based on the second average value and a third average value of historical psychological stress indices within a third time window, where the first time window is the time window including the real-time psychological stress index, the second time window is the time window adjacent to the first time window, and the third time window is the time window adjacent to the second time window. Finally, if the difference between the first rate of change and the second rate of change is greater than a third threshold, increasing the second adjustment factor based on a preset ratio to obtain the first adjustment factor. If the difference is less than or equal to the third threshold, the second adjustment factor is used as the first adjustment factor.
[0117] Here, a correspondence between the psychological stress trend, the psychological stress level, and the adjustment factor can be pre-established and stored in a mapping table. Based on the psychological stress trend and the target psychological stress level, the mapping table is queried to obtain the corresponding second adjustment factor. For example, when the psychological stress trend is increasing and the target psychological stress level is low, the corresponding second adjustment factor can be +10%.
[0118] In the embodiment of the present application, the time window including the real-time psychological stress index is used as the first time window (e.g., the last 30 seconds), the time window adjacent to the first time window (e.g., the previous 30 seconds) is used as the second time window, and the time window adjacent to the second time window (e.g., the previous 30 seconds) is used as the third time window. Based on the first average value of the historical psychological stress index within the first time window and the second average value of the historical psychological stress index within the second time window, the first rate of change is determined. This can be achieved by calculating the difference between the first average value and the second average value, and the ratio of this difference to the second average value as the first rate of change. Similarly, the difference between the second average value and the third average value is calculated, and the ratio of this difference to the third average value is used as the second rate of change. If the difference between the first rate of change and the second rate of change is greater than a third threshold value, it is considered that the change in the psychological stress index is in an accelerated state and further regulation is required. Therefore, the second regulation factor is increased based on a preset ratio to obtain the first regulation factor. The embodiment of the present application does not limit the preset ratio, for example, it can be 10%. If the difference is less than or equal to the third threshold value, the original second regulation factor remains unchanged, that is, the second regulation factor is used as the first regulation factor.
[0119] It should be noted that if the second adjustment factor is negative, increasing the second adjustment factor based on the preset ratio will result in the first adjustment factor increasing only in absolute value, with the negative sign remaining unchanged. For example, if the second adjustment factor is -15% and the preset ratio is 10%, the first adjustment factor will be -25%.
[0120] In the embodiment of the present application, by calculating the first rate of change and the second rate of change, and combining the comparison of the difference with the third threshold value, the system can more carefully capture whether the rate of change of psychological stress has increased significantly. This multi-dimensional analysis not only takes into account the current level of psychological stress, but also focuses on the acceleration or deceleration of its changing trend, thereby improving the accuracy of psychological state assessment. The second adjustment factor is adjusted according to the significance of the rate of change of psychological stress, so that the system can flexibly respond to different situational needs. For example, when the rate of change of psychological stress increases significantly, the system will increase the adjustment strength (such as further reducing the speech speed or softening the tone) to better appease the user; and when the rate of change is stable, a moderate adjustment strategy is maintained to avoid excessive intervention. The embodiment of the present application improves the level of intelligence in human-computer voice interaction.
[0121] In step 1033, the first speech synthesis parameter corresponding to the second speech data is adjusted based on the first adjustment factor to obtain a second speech synthesis parameter.
[0122] Here, the first speech synthesis parameters are speech synthesis parameters used when generating the second speech data. The first speech synthesis parameters may include speech rate, pitch, conversation interruption frequency, background noise level, etc. For each first speech synthesis parameter, the first speech synthesis parameter may be adjusted based on the first adjustment factor to obtain a second speech synthesis parameter corresponding to the first speech synthesis parameter.
[0123] For example, the first adjustment factor is +10%, and the first speech synthesis parameters include speech rate = 160 words / minute, pitch = 5, conversation interruption frequency = 4 times / minute, and background noise level 0. For the first speech synthesis parameters with specific numerical values, such as speech rate, pitch, and conversation interruption frequency, the first adjustment factor can be multiplied by the first speech synthesis parameter to obtain the adjustment amount. For a speech rate of 160 words / minute × 10% = 16 words / minute, the speech rate adjustment amount is increased by 16, resulting in a second speech synthesis parameter of 176. For the background noise level, etc., the level can be mapped to a numerical value within a range, such as 0-100, and then multiplied by the first adjustment factor to obtain the adjustment amount. After obtaining the adjusted numerical value, it is reversely mapped back to the adjusted background noise level to obtain the second speech synthesis parameter.
[0124] In step 1034, the second speech data is synthesized based on the second speech synthesis parameters to obtain third speech data.
[0125] Here, the second speech data is resynthesized using the adjusted second speech synthesis parameters to generate the final third speech data. For example, if the original second speech data has a speaking rate of 160 words / minute, a pitch of 5, a conversation interruption frequency of 4 times / minute, and a background noise level of 0, the adjusted third speech data will have a speaking rate of 176 words / minute, a pitch of 6, a pause frequency of 6 times / minute, and a background noise level of 1.
[0126] In some embodiments, if the first adjustment factor is used to increase the speech synthesis parameter, the preset script speech data is inserted into the third speech data to obtain new third speech data.
[0127] The embodiment of the present application achieves refined assessment and response to psychological states through value interval division and psychological stress level mapping. A first adjustment factor is determined based on the psychological stress change trend and psychological stress level, and a first speech synthesis parameter in the second speech data output by the intelligent agent is adjusted to obtain a second speech synthesis parameter. The second speech data is then resynthesized based on the second speech synthesis parameter to generate third speech data. This allows the voice data output by the intelligent agent to be adjusted in real time based on the subject's psychological stress. By changing the speech synthesis parameter in the speech data, the psychological stress of the subject can be indirectly affected during subsequent conversations.
[0128] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.
[0129] An embodiment of the present application provides a speech processing method, which is a seat stress training method. Figure 7 This is a module diagram of the speech processing system provided by the embodiment of the present application. Figure 7 The speech processing system includes a stress scenario construction module 701, a real-time conversation pressure injection module 702, a multi-dimensional data acquisition module 703, a stress assessment engine 704 and a dynamic regulator 705. The stress scenario construction module 701 is used to construct a stress scenario library. After the agent selects a training scenario (corresponding to the business scenario in the above embodiment) from the stress scenario library, the speech processing system applies initial pressure through the real-time conversation pressure injection module 702. The multi-dimensional data acquisition module 703 collects multi-dimensional physiological and voice data, and the stress assessment engine 704 calculates the current stress value of the agent in real time based on the multi-dimensional physiological and voice data (corresponding to the real-time psychological stress index in the above embodiment). The dynamic regulator 705 adjusts the subsequent pressure injection intensity according to the evaluation result of the pressure value, that is, adjusts the real-time conversation pressure injection module 702 to form a closed-loop regulation.
[0130] The stress scenario building module is introduced below. Figure 8 It is a flow chart of the pressure scenario construction module provided in the embodiment of the present application.
[0131] Step 801: Administrator logs in.
[0132] Here, administrator login and permission verification are first performed: an authentication mechanism based on JSON WebToken (JWT) is adopted, and administrator role and permission information is stored in a distributed relational database (TiDB Database, Tidb) to ensure that only authorized personnel can configure scenarios.
[0133] Step 802: scene template selection.
[0134] Here, the voice processing system provides a basic scenario template library, including common business scenarios such as "customer complaints" and "emergency services". The template data of the business scenarios is stored in the JSON field of the distributed relational database Tidb.
[0135] Step 803: pressure parameter configuration.
[0136] Here, the stress parameters (corresponding to the speech synthesis parameters in the above embodiment) can include a speech speed adjustment range (80-200 words / minute), intonation, conversation interruption frequency (0-10 times / minute), background noise level (0-5 levels), and customer emotion index (1-10 levels). By adjusting the level of the customer emotion index, the speech speed and intonation can be uniformly adjusted. For example, the higher the level, the faster the speech speed and the higher the intonation.
[0137] Step 804: Publish the scene.
[0138] Here, the configured business scenarios are stored in a distributed relational database and synchronized to the Remote Dictionary Server (Redis) cache for fast loading on the agent side.
[0139] The real-time conversation pressure injection module is introduced below. Figure 9 This is a flow chart of the real-time dialogue pressure injection module provided in an embodiment of the present application.
[0140] Step 901: The agent selects a scene.
[0141] Here, the agent (corresponding to the object in the above embodiment) selects a business scenario from the business scenarios constructed by the stress scenario construction module.
[0142] Step 902: Load scene parameters.
[0143] Here, the agent terminal obtains the pressure parameter configured for the selected business scenario from Redis (corresponding to the third speech synthesis parameter in the above embodiment).
[0144] Step 903: Establish a voice connection.
[0145] Here, Web Real-Time Communication (WebRTC) technology is used to establish a real-time voice channel between the agent end and the real-time dialogue robot (corresponding to the intelligent agent in the above embodiment), supporting low-latency two-way communication.
[0146] Step 904: Receive an adjustment instruction.
[0147] Here, we use the open-source stream processing framework (Flink) to process commands sent by the dynamic regulator in real time. Based on these commands, we adjust speech synthesis parameters (such as speaking rate, intonation, and pauses (i.e., the frequency of conversation endpoints in the stress parameter)) and dynamically insert preset interference factors (such as background noise and sudden interruptions (simulating loud noises that cause call pauses)).
[0148] It should be noted that when the business scenario is loaded for the first time, there is no instruction sent by the dynamic regulator, and the audio stream is generated based on the pre-configured pressure parameters.
[0149] Step 905: Dynamically adjust the pressure.
[0150] At the speech level, the Fast Forward MPEG (FFmpeg) codec uses the adjusted pressure parameters to process the audio stream (pre-synthesized audio for different scenarios) in real time, adjusting the speech rate and adding noise. At the conversation level, based on a pre-set script (multiple rounds of dialogue), high-pressure dialogue content is inserted at specific moments (randomly selected insertion times) (triggered by the increased pressure adjustment strategy).
[0151] The multi-dimensional data acquisition module is introduced below. Figure 10 It is a flow chart of the multi-dimensional data acquisition module provided in an embodiment of the present application.
[0152] Step 1001: wristband data collection.
[0153] Here, physiological data (corresponding to the physiological characteristics in the above embodiment) is collected by connecting a smart bracelet via Bluetooth to collect heart rate variability (HRV) data. The sampling frequency is 1Hz, and Flink is used for real-time data stream processing.
[0154] Step 1002: Speech feature extraction.
[0155] Here, an open source speech processing library is used to extract the following speech features: fundamental frequency jitter, amplitude jitter (shimmer), speech pause frequency, and speech rate change rate.
[0156] Step 1003: data preprocessing.
[0157] Here, Flink is used to clean and normalize the collected physiological data and speech features, detect and filter outliers, and align time windows (for example, a 5-second sliding window).
[0158] Step 1004: Feature storage.
[0159] Here, the preprocessed feature data is stored in the time series table in Tidb. At the same time, the preprocessed feature data is written to Redis for real-time evaluation.
[0160] The following describes the stress assessment engine. Figure 11 It is a flow chart of the pressure assessment engine provided in an embodiment of the present application.
[0161] Step 1101, obtain real-time data.
[0162] Here, the latest multi-dimensional feature data, including voice feature data and physiological data, is read from Redis. Flink is used to aggregate the streaming data, that is, to calculate the stress index (corresponding to the real-time psychological stress index in the above embodiment).
[0163] Step 1102: Calculate the pressure index.
[0164] Here, the Psychological Stress Index (PSI) = 0.4 × HRV index + 0.3 × voice jitter index (the weighted sum of fundamental frequency jitter characteristics and amplitude jitter characteristics) + 0.2 × speech rate change rate + 0.1 × pause frequency. Each index in the formula is a normalized value between 0 and 100. Alternatively, the similarity between the agent's answer and the standard answer preset in the knowledge base can be calculated, and the number of incorrect answers can be counted and weighted summed to obtain the PSI.
[0165] Step 1103: Determine the pressure level.
[0166] Here, when the stress index is less than 30, the stress level (corresponding to the psychological stress level in the above embodiment) is low, indicating that the agent is relaxed. When 30 ≤ Stress Index ≤ 60, the stress level is medium, indicating that the agent is in an ideal training state. When 60 ≤ Stress Index ≤ 80, the stress level is high, indicating that the agent is nearing the limit of tolerance. When the stress index is greater than or equal to 80, the stress level is excessive, indicating that immediate intervention is required.
[0167] Step 1104: Output the evaluation result.
[0168] Here, the evaluation results are written to Tidb for later analysis, and the real-time stress level and stress index are pushed to the dynamic regulator through the message queue (Kafka).
[0169] The dynamic regulator is introduced below. Figure 12 It is a flow chart of the dynamic regulator provided in an embodiment of the present application.
[0170] Step 1201: Receive a stress assessment.
[0171] Here, a Kafka consumer group subscribes to the PSI (Stress Index) data stream output by the stress assessment engine in real time. Flink's windowing mechanism is used to maintain a 30-second sequence of stress indices (corresponding to the psychological stress index sequence in the above example). A sliding window (corresponding to the time window in the above example, with a window size of 30 seconds and a sliding interval of 5 seconds) is used to ensure data continuity. A state cache is maintained in the Flink operator, recording the average PSI value of the last three windows for trend analysis.
[0172] Step 1202: Pressure state analysis.
[0173] Here, anomaly detection is performed first: instantaneous interference data is filtered through a preset threshold (for example, a sudden change of PSI exceeding 15 points is considered a sensor abnormality) to obtain a filtered pressure index sequence. Trend calculation (corresponding to the trend of psychological stress changes in the above embodiment): Based on the pressure index sequence, the slope of the PSI change is calculated by linear regression. When the slope is greater than 0.5, it is judged as an upward trend, and less than -0.5 is a downward trend, otherwise it is a stable trend. Acceleration calculation: Compare the differences in the PSI change rates of adjacent windows. If the current window change rate (corresponding to the first change rate in the above embodiment) is 20% higher than the previous window (corresponding to the second change rate in the above embodiment), it is judged to be an accelerated state.
[0174] Step 1203: Select an adjustment strategy.
[0175] Here, a two-level decision-making mechanism is used to select the adjustment strategy. First, the primary strategy: Based on the current stress level (corresponding to the target psychological stress level in the above embodiment) and the preset rules of change trend, the basic adjustment amplitude (corresponding to the second adjustment factor in the above embodiment) is directly mapped from the mapping table (see Table 1).
[0176] Table 1 Mapping table
[0177] Current pressure level Changing trends Basic adjustment range low pressure rise +10% pressure low pressure decline +20% pressure Medium pressure rise maintain Medium pressure decline +5% pressure high pressure any -15% pressure Overpressure any -30% pressure
[0178] Then the strategy is corrected: when the acceleration state is detected, a 10% adjustment amount is added to the basic adjustment range (for example, the original basic adjustment range is -15% pressure, and after correction it is -25%); three consecutive adjustments in the same direction trigger a strategy review.
[0179] Step 1204: Generate an adjustment instruction.
[0180] Here, the adjustment command includes three dimensions: voice parameters: speech rate adjustment percentage (±5% to 20%), frequency of pause insertion (0-5 times / minute); conversation content: increase / decrease the conversation conflict level (1-5 levels); environmental interference: increase or decrease the background noise level (±10dB). The command is serialized into JSON format through the Kafka producer, including a timestamp, session ID, and adjustment details.
[0181] The embodiments of the present application significantly improve the ability of agents to cope with real high-pressure environments through special stress scenario training, thereby improving the targeted nature of training. Combined with multi-dimensional data such as physiological indicators, stress assessment is made more comprehensive and accurate, enhancing the objectivity of the assessment. The dynamic adjustment mechanism avoids training discomfort or insufficient effect caused by fixed stress levels, and optimizes the training experience. By simulating real stress scenarios, the trial and error costs in on-site training are reduced, thereby reducing training costs. It supports full process coverage from basic business training to high-pressure special training, expanding application scenarios.
[0182] The following further describes the exemplary structure of the voice processing device 455 provided in the embodiment of the present application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the speech processing device 455 of the memory 450 may include:
[0183] The first determination module 4551 is used to determine the real-time psychological stress index of the object based on the first voice data of the object in dialogue with the intelligent agent.
[0184] The second determining module 4552 determines the change trend of the subject's psychological stress based on the subject's historical psychological stress index.
[0185] The adjustment module 4553 is used to adjust the second voice data of the intelligent agent based on the psychological stress change trend and the real-time psychological stress index to obtain third voice data, and use the third voice data as the voice data for the intelligent agent to communicate with the object.
[0186] In some embodiments, the second determination module 4552 is further used to divide the subject's multiple historical psychological stress indices into multiple psychological stress index sequences according to multiple time windows, wherein each psychological stress index sequence includes the historical psychological stress index within the corresponding time window; calculate the average value of the historical psychological stress index included in each psychological stress index sequence; and determine the subject's psychological stress change trend based on the average values corresponding to the multiple psychological stress index sequences.
[0187] In some embodiments, the second determination module 4552 is further used to perform linear regression processing on the average values corresponding to multiple psychological stress index sequences based on the time windows corresponding to the multiple psychological stress index sequences to obtain a slope representing the change in psychological stress; if the slope is greater than or equal to a first threshold, the upward trend is determined as the trend of change in psychological stress; if the slope is greater than a second threshold and less than the first threshold, the steady trend is determined as the trend of change in psychological stress, wherein the first threshold is greater than the second threshold; if the slope is less than or equal to the second threshold, the downward trend is determined as the trend of change in psychological stress.
[0188] In some embodiments, the adjustment module 4553 is also used to determine the target value interval of the real-time psychological stress index from multiple value intervals, and determine the target psychological stress level corresponding to the target value interval, wherein the multiple value intervals are obtained by dividing the value range of the psychological stress index, and the value intervals correspond to the psychological stress levels one by one; based on the psychological pressure change trend and the target psychological stress level, determine the first adjustment factor for adjusting the speech synthesis parameter; adjust the first speech synthesis parameter corresponding to the second speech data based on the first adjustment factor to obtain the second speech synthesis parameter; synthesize the second speech data based on the second speech synthesis parameter to obtain the third speech data.
[0189] In some embodiments, the adjustment module 4553 is also used to determine a second adjustment factor corresponding to the psychological stress change trend and the target psychological stress level based on the correspondence between the psychological stress change trend, the psychological stress level and the adjustment factor; determine a first change rate based on a first average value of the historical psychological stress index in the first time window and a second average value of the historical psychological stress index in the second time window, and determine a second change rate based on the second average value and a third average value of the historical psychological stress index in the third time window, the first time window being a time window including the real-time psychological stress index, the second time window being a time window adjacent to the first time window, and the third time window being a time window adjacent to the second time window; if the difference between the first change rate and the second change rate is greater than a third threshold, the second adjustment factor is increased based on a preset ratio to obtain the first adjustment factor; if the difference is less than or equal to the third threshold, the second adjustment factor is used as the first adjustment factor.
[0190] In some embodiments, the first determination module 4551 is also used to extract voice features of the first voice data, and normalize the voice features to obtain a first feature; normalize the physiological features of the object to obtain a second feature; perform a preset operation on the second feature and the first feature to obtain the real-time psychological stress index of the object.
[0191] In some embodiments, the first determination module 4551 is also used to determine the dialogue text and third speech synthesis parameters configured for the business scenario in response to the object selecting the business scenario; convert the dialogue text into fourth speech data, and adjust the fourth speech data based on the third speech synthesis parameters to obtain the second speech data of the intelligent agent.
[0192] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the speech processing method described in the present invention.
[0193] The embodiment of the present application provides a computer-readable storage medium in which computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the speech processing method provided in the embodiment of the present application, for example, Figure 3 The speech processing method shown.
[0194] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0195] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0196] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).
[0197] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.
[0198] To sum up, through the embodiments of the present application, the speech speed, pitch, interruption frequency, etc. in the voice data output by the dialogue robot can be adjusted in real time based on the user's psychological stress state, thereby indirectly affecting the user's psychological stress state and realizing intelligent human-computer voice communication.
[0199] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.
Claims
1. A speech processing method, characterized in that: The method comprises: determining a real-time psychological stress index of a subject in conversation with the agent based on first speech data of the subject; determining a change trend of the subject's psychological stress based on the subject's historical psychological stress index; Based on the psychological stress change trend and the real-time psychological stress index, the second voice data of the intelligent agent is adjusted to obtain third voice data, and the third voice data is used as the voice data for the intelligent agent to communicate with the object.
2. The method according to claim 1, characterized in that Determining the change trend of the subject's psychological stress based on the subject's historical psychological stress index includes: Dividing the multiple historical psychological stress indexes of the subject into multiple psychological stress index sequences according to multiple time windows, wherein each of the psychological stress index sequences includes the historical psychological stress index within the corresponding time window; Calculating an average value of the historical psychological stress indexes included in each of the psychological stress index sequences; The psychological stress change trend of the subject is determined based on average values corresponding to the multiple psychological stress index sequences.
3. The method according to claim 2, characterized in that The determining of the change trend of the subject's psychological stress based on the average values corresponding to the multiple psychological stress index sequences includes: performing linear regression processing on average values corresponding to the multiple psychological stress index sequences based on time windows corresponding to the multiple psychological stress index sequences to obtain a slope representing a change in psychological stress; If the slope is greater than or equal to a first threshold, the upward trend is determined as the psychological stress change trend; If the slope is greater than a second threshold and less than the first threshold, a steady trend is determined as the psychological stress change trend, wherein the first threshold is greater than the second threshold; If the slope is less than or equal to the second threshold, the downward trend is determined as the psychological stress change trend.
4. The method according to claim 1, wherein The step of adjusting the second voice data of the agent based on the psychological stress change trend and the real-time psychological stress index to obtain third voice data includes: Determining a target value interval for the real-time psychological stress index from a plurality of value intervals, and determining a target psychological stress level corresponding to the target value interval, wherein the plurality of value intervals are obtained by dividing the value range of the psychological stress index, and the value intervals correspond one-to-one to the psychological stress level; determining a first adjustment factor for adjusting a speech synthesis parameter based on the psychological stress change trend and the target psychological stress level; adjusting a first speech synthesis parameter corresponding to the second speech data based on the first adjustment factor to obtain a second speech synthesis parameter; The second speech data is synthesized based on the second speech synthesis parameter to obtain third speech data.
5. The method according to claim 4, characterized in that The determining, based on the psychological stress change trend and the target psychological stress level, a first adjustment factor for adjusting a speech synthesis parameter includes: Determining a second adjustment factor corresponding to the psychological stress change trend and the target psychological stress level based on a correspondence between the psychological stress change trend, the psychological stress level, and the adjustment factor; determining a first change rate based on a first average value of a historical psychological stress index within a first time window and a second average value of the historical psychological stress index within a second time window, and determining a second change rate based on the second average value and a third average value of the historical psychological stress index within a third time window, wherein the first time window is a time window including the real-time psychological stress index, the second time window is a time window adjacent to the first time window, and the third time window is a time window adjacent to the second time window; If the difference between the first change rate and the second change rate is greater than a third threshold, increasing the second adjustment factor based on a preset ratio to obtain the first adjustment factor; If the difference is less than or equal to the third threshold, the second adjustment factor is used as the first adjustment factor.
6. The method according to claim 1, wherein The step of determining the real-time psychological stress index of the subject based on first voice data of the subject in dialogue with the agent comprises: Extracting speech features of the first speech data and performing normalization processing on the speech features to obtain first features; Normalizing the physiological feature of the subject to obtain a second feature; A preset operation is performed on the second feature and the first feature to obtain a real-time psychological stress index of the subject.
7. The method according to claim 1, characterized in that The method further comprises: In response to the object selecting a business scenario, determining a dialogue text and a third speech synthesis parameter configured for the business scenario; The dialogue text is converted into fourth voice data, and the fourth voice data is adjusted based on the third voice synthesis parameter to obtain the second voice data of the agent.
8. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the speech processing method according to any one of claims 1 to 7 when executing the computer-executable instructions or computer program stored in the memory.
9. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the speech processing method according to any one of claims 1 to 7 is implemented.
10. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the speech processing method according to any one of claims 1 to 7 is implemented.