Soothing interaction method, system, device and storage medium for intelligent voice terminal

By adopting the emotional recognition and soothing interaction model on the smart voice terminal, the emotional state in the sound signal is recognized in real time and the soothing audio output is solved, and the problem of smart voice terminals failing to consider the use of the object's emotions during interaction is improved, improving the interactive experience and humanization.

CN114120985BActive Publication Date: 2025-05-16SHANGHAI SHENSILICON SEMICON CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111370015.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2025-05-16
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

When existing smart voice terminals interact with the use object, they fail to effectively consider the emotions of using the object, resulting in poor intelligence and poor experience.

Method used

The emotion recognition model and comfort interaction model are used to receive sound signals in real time and perform semantic recognition and emotional discrimination. In case of poor network communication, the pre-configured comforting interaction model is used to output the comforting audio corresponding to the emotional level and output during the waiting interval.

Benefits of technology

Before the problem of the interactive object cannot be solved immediately, provide emotional comfort, improve interactive experience, and enhance the humanization of smart products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114120985B_ABST
    Figure CN114120985B_ABST
Patent Text Reader

Abstract

The present invention discloses a soothing interaction method, system, device and storage medium for an intelligent voice terminal, the method comprising: receiving a collected sound signal; processing the sound signal for semantic recognition and emotion discrimination; accessing the Internet through the recognition result; determining the emotion level to which the sound signal belongs; receiving the Internet access status and emotion level in real time; when the network communication is poor and the access is in a waiting state, using a pre-configured soothing interaction model to directly output the first soothing audio corresponding to the emotion level according to the access status and emotion level, and then outputting the soothing audio corresponding to the emotion level in sequence at intervals corresponding to the emotion level. The soothing interaction method implemented by the present invention does not make the interactive object feel ignored before the problem of the interactive object cannot be solved, thereby maximizing the interactive experience of the interactive object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to voice interaction technology, and in particular to a soothing interaction method, system, device and storage medium for an intelligent voice terminal. Background Art

[0002] Existing voice interaction products include fully online, fully offline, and semi-online and semi-offline types. These products use some neural network emotion algorithms to acquire, classify, recognize, and respond to the real-time voice of the interactive object, making the human-computer interaction friendly and accurate.

[0003] The usual interaction process is that the voice product captures the key information in the corpus of the interaction object, identifies the emotional state based on the constructed recognition model, distinguishes the emotional changes of the interaction object, and judges the expected emotions of the interaction object after the emotional changes, activates the corresponding database, and actively feeds back the required new information to the interaction object.

[0004] However, there are some problems with the use of such voice products. For example, if the emotional algorithm is set in the cloud, when the network is poor, and the problem asked by the interactive object is not in the terminal instruction set, the data will not be transmitted to the cloud, resulting in the voice product being unable to solve the problem asked by the interactive object, and there is no other response on the terminal to soothe the emotions, which makes the interactive object's experience extremely poor. Even when the network is normal, the data feedback from the cloud will be delayed, resulting in the interactive object not being able to get a response in a short time, making the voice product technology less humane. For example, the emotional algorithm is set on the terminal side. Although it is not affected by the network status, if the problem asked does not belong to the terminal instruction set, the problem asked will still not be fed back, which still makes the voice product less humane. Summary of the invention

[0005] The embodiments of the present application provide a soothing interaction method, system, device and storage medium for an intelligent voice terminal, thereby solving the problem that in the prior art, the intelligent voice terminal does not consider the emotions of the user during the interaction with the user, resulting in poor intelligence effect and poor experience. It achieves emotional soothing before feedback of the response result, thereby improving the user experience of the intelligent product.

[0006] In a first aspect, the present application provides a soothing interaction method for an intelligent voice terminal, the method comprising:

[0007] S100, receiving the collected sound signal;

[0008] S200, processing the sound signal for semantic recognition and emotion discrimination; wherein, after semantic recognition of the sound signal is performed using a pre-configured semantic recognition model, the Internet is accessed through the semantic recognition result; emotion discrimination is performed on the semantic recognition result and the sound signal using a pre-configured emotion recognition model, and the emotion level to which the sound signal belongs is determined according to multiple levels of emotional states preset in the emotion recognition model;

[0009] S300, receives the Internet access status and emotional level in real time; when the network communication is stable and the access is completed, directly outputs the access result; when the network communication is poor and the access is in a waiting state, a pre-configured soothing interaction model is used to directly output the first soothing audio corresponding to the emotional level according to the access status and emotional level, and then outputs the soothing audio corresponding to the emotional level in sequence at intervals corresponding to the emotional level.

[0010] Furthermore, in the step S200, it also includes, when using the emotion recognition model to perform emotion discrimination on the continuous sound signals, when the continuous sound signals are adapted to the same level of emotional state and the same semantic recognition results, determining that the emotional state of the sound signal discriminated last is one level higher than that of the sound signal discriminated last.

[0011] Furthermore, the step S300 further includes configuring a plurality of soothing audios and their output time intervals adapted to the emotion level in the soothing interaction model, so that when network communication is poor, the first output of the soothing audio corresponding to the emotion level is performed directly according to the emotion level, and during the waiting period, the output continues according to its preset output time interval.

[0012] Furthermore, after step S300, when the network communication returns to normal, if the received emotional state includes the highest emotional level, human-computer voice interaction is performed to provide manual channel selection to solve technical problems related to the intelligent voice terminal.

[0013] In a second aspect, the present application provides a soothing interaction system for an intelligent voice terminal, using any method of the first aspect, the system comprising:

[0014] A sound receiving module, configured to receive the collected sound signal;

[0015] A signal processing module is configured to process the sound signal for semantic recognition and emotion discrimination; wherein, after the sound signal is semantically recognized using a pre-configured semantic recognition model, the Internet is accessed through the semantic recognition result; the semantic recognition result and the sound signal are emotionally discriminated using a pre-configured emotion recognition model, and the emotion level to which the sound signal belongs is determined according to multiple levels of emotional states preset in the emotion recognition model;

[0016] The soothing output module is configured to receive the Internet access status and emotional level in real time; when the network communication is stable and the access is completed, the access result is directly output; when the network communication is poor and the access is in a waiting state, a pre-configured soothing interaction model is used to directly output the first soothing audio corresponding to the emotional level according to the access status and emotional level, and the soothing audio corresponding to the emotional level is output in sequence at intervals corresponding to the emotional level.

[0017] In a third aspect, the present application provides a computer device, characterized in that the computer device includes a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by the processor to implement the soothing interaction method of the intelligent voice terminal as described in any one of the first aspects.

[0018] In a fourth aspect, the present application provides a computer-readable storage medium, characterized in that the storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by a processor to implement the soothing interaction method of the intelligent voice terminal as described in any one of the first aspects.

[0019] The technical solution provided in the embodiments of the present application has at least the following technical effects:

[0020] Since the present invention adopts an emotion recognition model, the emotional state carried in the sound signal can be obtained in time, and the interactive object will not feel ignored before the problem of the interactive object can be solved, thereby maximizing the interactive experience of the interactive object.

[0021] Due to the adoption of the soothing interaction model, soothing audio can be fed back according to emotional states of different emotional levels, and soothing audio of emotional states with high emotional levels is output preferentially, which can make the interacting objects feel that they are not being ignored, but are waiting for the problem to be solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is a flow chart of the soothing interaction method of the intelligent voice terminal in the first embodiment of the present application;

[0023] Figure 2 This is a module diagram of the soothing interaction system of the intelligent voice terminal in Example 2 of the present application. DETAILED DESCRIPTION

[0024] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0025] Embodiment 1

[0026] The present application embodiment provides a soothing interaction method for an intelligent voice terminal, the method comprising:

[0027] Step S100: receiving a collected sound signal.

[0028] Step S200, processing the sound signal for semantic recognition and emotion discrimination; wherein, after semantic recognition of the sound signal is performed using a pre-configured semantic recognition model, the Internet is accessed through the semantic recognition result; emotion discrimination is performed on the semantic recognition result and the sound signal using a pre-configured emotion recognition model, and the emotion level to which the sound signal belongs is determined according to the multiple levels of emotional states preset in the emotion recognition model.

[0029] Step S300, receiving the Internet access status and emotional level in real time; when the network communication is stable and the access is completed, directly output the access result; when the network communication is poor and the access is in a waiting state, a pre-configured soothing interaction model is used to directly output the first soothing audio corresponding to the emotional level according to the access status and the emotional level, and then output the soothing audio corresponding to the emotional level in sequence at intervals corresponding to the emotional level.

[0030] In step S100 of this embodiment, the receiving subject of the collected sound signal can be any intelligent voice terminal, such as "Tmall Genie" and "Xiaodu at Home" available in the market, or a smart phone or tablet computer that can perform semantic recognition on a daily basis, or an application program that sets semantic recognition, which is not limited in this embodiment. Since the semantic recognition technology is loaded on the intelligent terminal, the intelligent voice terminal in this embodiment can be any terminal device that uses the semantic recognition technology.

[0031] Further explanation, in one embodiment, the semantic recognition model in step S200 processes the semantic recognition process, and also includes preprocessing the sound signal. Specifically, it includes: obtaining a sound signal and text information containing sentence markers; it includes identifying the language type to which the sound signal belongs, performing sentence processing on the sound signal according to the language type, and obtaining a sound signal containing sentence markers; translating the sound signal containing sentence markers into text information. The text information containing sentence markers can be used to access the Internet, so as to query the relevant content corresponding to the recognized text information through the Internet.

[0032] For the emotion recognition of the sound signal, this embodiment also identifies the language type to which the sound signal belongs. The intelligent voice terminal in this embodiment is not limited to the use of Chinese people, and the expression methods or rhythmic formats of different languages ​​are different. For example, Chinese is an ideographic language, which has a high degree of generalization and simplicity, and high expression efficiency; English is a phonetic language, and sentences need to be expressed through certain external morphological markers. English emphasizes that only one thing is said in one sentence, and this thing is either a "subject-verb-object" structure or a "subject-predicate-object" structure, that is, "what to do" or "what is what", while Chinese does not have such a particularity. In this implementation scheme, it is necessary to understand the emotional state in the sound signal, so it is impossible to obtain the emotional state by only understanding its meaning, so it is necessary to pre-identify the language type of the sound signal, and there are many languages ​​in the world, and many languages ​​are similar. The intelligent voice terminal in this embodiment pre-configures a language type recognition model, and is a language type recognition model that is trained and perfected in the service background. The language type recognition model in this embodiment uses the ngrams concept in artificial intelligence to build a model architecture, and establishes a 4-grams language model based on the existing corpus. Regarding the recognition of the language type, the present embodiment is not limited to the described technology, and it is sufficient to obtain the language type of the sound signal.

[0033] Furthermore, the sound signal is processed into sentences according to the language type to obtain a sound signal containing sentence marks. The sound signal is processed into sentences based on the language type. Sentence processing means setting pause marks between sound signals. In text expression, punctuation marks are usually used to indicate pauses between semantics. In this embodiment, the sound signal is processed into sentences so that there are sentence marks between the sound signals. The sound signal between each sentence mark has a complete semantics. The pause format between sound signals of different language types is different. Therefore, in this embodiment, sentence processing needs to be performed based on the language type.

[0034] Thus, the sound signal containing the sentence mark is translated into text information. The "translation" in this step is the conversion of the expression mode, which converts the sound wave of the sound signal into semantic text, rather than converting one language text into another language text. Further explanation, the sound signal in this embodiment has been segmented according to the language type, and at this time, it is only necessary to convert the sound signal with the sentence mark into text information.

[0035] In one embodiment, the method of converting sound to text includes: preprocessing: removing silence at the beginning and end to reduce interference; the operation of removing silence is generally called VAD; sound framing, cutting the sound into small segments, each small segment is called a frame, using a moving window function to achieve, and there is overlap between frames. Feature extraction: using linear prediction cepstral coefficients (LPCC) and Mel cepstral coefficients (MFCC), each frame waveform is converted into a multidimensional vector containing sound information. Acoustic model (AM): obtained by training speech data, the input is a feature vector, and the output is phoneme information. Dictionary: the correspondence between characters or words and phonemes, while Chinese is the correspondence between pinyin and Chinese characters, and English is the correspondence between phonetic symbols and words. Language model (LM): by training a large amount of text information, the probability of the correlation between single characters or words is obtained. Decoding: that is, the audio data after feature extraction is output as text through acoustic models, dictionaries, and language models.

[0036] In one embodiment, the identification of emotional state further includes the following:

[0037] Receiving the sound signal and text information containing the sentence mark. In this embodiment, the sound signal and text information received can be through two channels, one channel directly receives the text information, and the other channel receives the sound signal and text information.

[0038] In this embodiment, the query response of text information and the emotion recognition in the sound signal are processed. That is to say, in the same time sequence, after the semantic recognition result is completed, both the query response of text information and the emotion recognition of the sound signal are performed. Since the query response time is limited by the network status, when the network status is not smooth or the queried data resources are limited, the feedback response is relatively slow, and the emotion recognition in the process has obtained the emotional state. At this time, the response can be soothed according to the emotional state, which is the purpose of this embodiment.

[0039] Query and respond to text messages. Query and respond to text messages include identifying characteristic keywords in text messages and content to be responded to. Characteristic keywords can be verbs fed back by the intelligent voice terminal, such as playing a certain singer's song, "play" is the characteristic keyword, and "a certain singer's song" is the content to be queried. When the intelligent voice terminal has stored the content to be queried locally, it can respond directly, but when it needs to be searched from the Internet, it may take more time to feedback. For interactive objects with bad emotions, waiting affects emotions more. Based on this, setting an emotion-based soothing design is the purpose of this embodiment.

[0040] The pre-configured emotion recognition model is used to perform emotion discrimination on the sound signals and text information containing sentence markers to determine the emotional state of each sound signal.

[0041] Of course, in some embodiments, the emotion recognition model includes extracting language feature values ​​from a sound signal containing language markers, and extracting rhythmic features including intonation, falling tone, accent, and stress from the sound signal. In some embodiments, the emotion recognition model also includes calculating the probability of the emotional state of the sound signal in each sentence using a vector segmentation Mahalanobis distance discriminant method or a principal component analysis method or a hidden Markov model based on the language markers and language feature values, and determining the emotional state when the probability value of the emotional state exceeds a threshold. In some embodiments, the emotion recognition model also includes using an emotion lexicon to perform emotion retrieval on text information, obtaining the emotional state contained in the text information, counting the number of corresponding emotional states in each sentence of text information, and determining the emotional state with the highest number of times.

[0042] Further explanation, for the recognized emotional state, the emotion recognition model of this embodiment presets multiple emotion levels, that is, the recognized emotional state matches the corresponding emotion level when output. The emotional state of each sound signal output by the emotion recognition model in this embodiment has a corresponding emotion level.

[0043] Therefore, the emotion recognition model of this embodiment performs emotion discrimination on the sound signal. The emotion recognition model presets multiple levels of emotional states. After identifying the emotional state, the emotion level can be obtained, thereby achieving the emotion level of each sound signal through the emotion recognition model.

[0044] In one embodiment, step S200 also includes, when using the emotion recognition model to perform emotion discrimination on continuous sound signals, when the continuous sound signals are adapted to the same level of emotional state and the same semantic recognition result, determining that the emotional state of the sound signal discriminated the second time is one level higher than that of the sound signal discriminated the first time. In other words, for two consecutive sound signals with the same semantics, the first sound signal may not be fed back in time by the intelligent voice terminal, and even if the emotional state of the second sound signal does not change during the emotion judgment, the intelligent voice terminal, in order to be more humane, upgrades the emotional level of the latter time, so that the intelligent voice terminal pays more attention to the interactive object, thereby providing a deeper degree of comfort.

[0045] Step S300 in this embodiment further includes that a variety of soothing audios and their output time intervals adapted to the emotional level are configured in the soothing interaction model, so that when the network communication is poor, the corresponding soothing audio is directly output for the first time according to the emotional level, and during the waiting period, the output continues according to its preset output time interval. In other words, the output strategy is set for the emotional level in the soothing interaction model. For example, if the emotional level is Class A, the soothing audio output is Class A audio, and the output interval time of the corresponding Class A audio is also Class A time period. It is further explained that the higher the level of the emotional level, the higher the soothing degree of the soothing audio, and the shorter the time interval of the output audio. When the Internet access state has been in a waiting state, no response result has been given. In this state, this embodiment outputs the corresponding soothing audio through the identified emotional level, thereby improving the humanized experience of the intelligent voice terminal.

[0046] It is further explained that in the present embodiment, priorities can also be set according to the emotion levels, and soothing audio for emotional states with high priority levels can be output. That is, soothing audio is output according to emotional states with high priority levels of emotion levels. For example, people have six types of emotions, namely happiness, sadness, anger, surprise, fear, and disgust, and each emotion type is prioritized for soothing feedback, so as to give priority to outputting soothing audio corresponding to high-level emotional states. It is further explained that in the present embodiment, each emotion type also includes multiple sub-levels, and each sub-level combination of each emotion type generates emotional states of different emotion levels. In the present embodiment, soothing audio is preset according to emotional states of different emotion levels, so that soothing audio is fed back according to the highest level of emotional state determined by the received sound signal within the time threshold, thereby alleviating the emotions of the interactive object in the waiting state.

[0047] In addition, in order to better improve the humanized settings of the intelligent voice terminal, after step S300, it also includes that when the network communication returns to normal, if the received emotional state includes the highest emotional level, human-computer voice interaction is performed to provide manual channel selection to solve technical problems related to the intelligent voice terminal. That is to say, when it comes to access delays caused by network failures of intelligent voice terminals, this embodiment directly provides manual soothing channel selection, which can, on the one hand, manually negotiate and soothe, help the interactive object to relieve emotions, thereby increasing the service impression of the intelligent voice terminal supplier, and on the other hand, it can help the supplier of the intelligent voice terminal collect relevant data on fault access delays, and then make corresponding technical improvements.

[0048] From the above, it can be seen that in this embodiment, when the network conditions are good, since a soothing interaction method is pre-set in this embodiment, the emotional state transmitted by the interactive object is determined according to the sound signal, and whether the interactive object is anxious is felt. Since the storage space of the intelligent voice terminal is limited, the various computing models used in this embodiment are trained models. No matter what the network conditions are, the emotional state is obtained through each model in the soothing interaction method and the soothing audio is output in time according to the soothing mechanism to soothe the emotions of the interactive object, thereby making the interaction between the intelligent voice terminal and the interactive object (person) more humane. For example, when the query response is "querying, please wait", the intelligent voice terminal outputs the corresponding soothing audio according to the emotional state of the determined emotional level, and outputs the soothing audio at a preset time interval.

[0049] Since different emotional states correspond to different soothing audios, if the emotional state level is not high, that is, if you are not very anxious, then the soothing audio output may be a speech with a lower soothing degree, and the interval between soothing speeches can be relatively long. On the contrary, if the emotional state is excited and anxious, the soothing audio output belongs to a speech with a higher soothing degree, and the time interval between the soothing audio output should be relatively short, so that the interactive object can feel that he is not ignored, but is waiting for the problem to be solved. And based on the statistics of emotional states, in this embodiment, query results corresponding to a large number of emotional states are output preferentially for query responses. As a result, even if the problem of the interactive object cannot be solved directly, the interactive object does not feel ignored, thereby maximizing the interactive experience of the interactive object.

[0050] Embodiment 2

[0051] The embodiment of the present application provides a soothing interaction system for an intelligent voice terminal, which adopts the soothing interaction method for an intelligent voice terminal in the first embodiment. The system includes:

[0052] The sound receiving module 100 is configured to receive the collected sound signal.

[0053] The signal processing module 200 is configured to process the sound signal for semantic recognition and emotion discrimination; wherein, after the sound signal is semantically recognized using a pre-configured semantic recognition model, the Internet is accessed through the semantic recognition result; the semantic recognition result and the sound signal are emotionally discriminated using a pre-configured emotion recognition model, and the emotion level to which the sound signal belongs is determined according to the multiple levels of emotional states preset in the emotion recognition model.

[0054] The soothing output module 300 is configured to receive the Internet access status and emotional level in real time; when the network communication is stable and the access is completed, the access result is directly output; when the network communication is poor and the access is in a waiting state, a pre-configured soothing interaction model is used to directly output the first soothing audio corresponding to the emotional level according to the access status and the emotional level, and then output the soothing audio corresponding to the emotional level in sequence at intervals corresponding to the emotional level.

[0055] Embodiment 3

[0056] The present embodiment provides a computer device, characterized in that the computer device includes a processor and a memory, the memory stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by the processor to implement the soothing interaction method of the intelligent voice terminal as any one of the first embodiments.

[0057] The present embodiment provides a computer-readable storage medium, characterized in that the storage medium stores at least one instruction, at least one program, code set or instruction set, and the at least one instruction, at least one program, code set or instruction set is loaded and executed by a processor to implement the soothing interaction method of an intelligent voice terminal as described in any one of the first embodiments.

[0058] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0059] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0060] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0061] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0062] Although the preferred embodiments of the present invention have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0063] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A soothing interaction method for an intelligent voice terminal, characterized in that: The method comprises: S100, receiving the collected sound signal; S200, performing semantic recognition and emotion discrimination on the sound signal; wherein, after performing semantic recognition on the sound signal using a pre-configured semantic recognition model, accessing the Internet through the semantic recognition result; performing emotion discrimination on the semantic recognition result and the sound signal using a pre-configured emotion recognition model, and determining the emotion level to which the sound signal belongs according to multiple levels of emotional states preset in the emotion recognition model; S300, receives the Internet access status and emotional level in real time; when the network communication is stable and the access is completed, directly outputs the access result; when the network communication is poor and the access is in a waiting state, a pre-configured soothing interaction model is used to directly output the first soothing audio corresponding to the emotional level according to the access status and emotional level, and then outputs the soothing audio corresponding to the emotional level in sequence at intervals corresponding to the emotional level.

2. The soothing interaction method of the intelligent voice terminal according to claim 1, characterized in that: In the step S200, it also includes, when using the emotion recognition model to perform emotion discrimination on the continuous sound signals, when the continuous sound signals are adapted to the same level of emotional state and the same semantic recognition results, determining that the emotional state of the sound signal discriminated last is one level higher than that of the sound signal discriminated last.

3. The soothing interaction method of the intelligent voice terminal according to claim 1, characterized in that: The step S300 further includes configuring a plurality of soothing audios and their output time intervals adapted to the emotion level in the soothing interaction model, so that when network communication is poor, the first output of the soothing audio corresponding to the emotion level is performed directly according to the emotion level, and during the waiting period, the output continues according to the preset output time interval.

4. The soothing interaction method of the intelligent voice terminal according to claim 1, characterized in that: After step S300, it also includes, when the network communication returns to normal, if the received emotional state includes the highest emotional level, performing human-computer voice interaction to provide manual channel selection to solve technical problems related to the intelligent voice terminal.

5. A soothing interaction system for an intelligent voice terminal, using the method of any one of claims 1 to 4, characterized in that: The system comprises: A sound receiving module, configured to receive the collected sound signal; A signal processing module is configured to process the sound signal for semantic recognition and emotion discrimination; wherein, after the sound signal is semantically recognized using a pre-configured semantic recognition model, the Internet is accessed through the semantic recognition result; the semantic recognition result and the sound signal are emotionally discriminated using a pre-configured emotion recognition model, and the emotion level to which the sound signal belongs is determined according to multiple levels of emotional states preset in the emotion recognition model; The soothing output module is configured to receive the Internet access status and emotional level in real time; when the network communication is stable and the access is completed, the access result is directly output; when the network communication is poor and the access is in a waiting state, a pre-configured soothing interaction model is used to directly output the first soothing audio corresponding to the emotional level according to the access status and emotional level, and the soothing audio corresponding to the emotional level is output in sequence at intervals corresponding to the emotional level.

6. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the soothing interaction method of the intelligent voice terminal as described in any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the soothing interaction method of the intelligent voice terminal as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Mobile internet intelligent wearable device

    CN107049280A

  • Senior citizen emotion monitoring system based on wearable equipment

    CN110881987A