Voice interaction method and device, storage medium and electronic device

By comprehensively analyzing the semantics of the elderly's input speech, identifying and matching emotional tags, the problem of speech recognition errors caused by differences in the elderly's pronunciation is solved. This enables accurate command recognition and emotional interaction even offline, making it particularly suitable for elderly care devices for empty-nest and solitary elderly people.

CN114495930BActive Publication Date: 2025-12-12SUZHOU KEYI-SKY SEMITECH INC +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210105815.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-28
Publication Date
2025-12-12
Estimated Expiration
2042-01-28

AI Technical Summary

Technical Problem

In elderly care settings, speech recognition systems may fail to recognize or misrecognize emergency calls due to differences in pronunciation among the elderly, and the lack of emotional interaction may prevent timely assistance, especially when network conditions are insufficient.

Method used

Through semantic-based comprehensive analysis, the system identifies acoustic features and sentiment tags of input speech by utilizing speech frequency, tone, volume, and speech rate. It then searches for command words with high similarity and sentiment matching, and obtains response words from a special vocabulary database for elderly care, supporting sentiment analysis and alarm processing.

Benefits of technology

It improves the accuracy and efficiency of speech recognition, enabling it to recognize commands from elderly people with different pronunciations even offline, and provides personalized emotional interaction. It is especially suitable for elderly people living alone or in empty nests, and can handle emergencies in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495930B_ABST
    Figure CN114495930B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a voice interaction method, device, storage medium and electronic equipment. The method comprises: in response to a situation of entering an activated state, acquiring an input voice; performing semantic-based comprehensive analysis on the input voice to determine a first emotional label corresponding to the input voice and a first acoustic feature vector of a keyword in the input voice, the comprehensive analysis being further based on at least one of the following factors: tone, intonation, volume and speed; searching for a first command phrase from a preset command phrase set, the first command phrase having a second acoustic feature vector with a similarity to the first acoustic feature vector greater than a first preset value, and the first command phrase having a second emotional label identical to the first emotional label; and acquiring a corresponding reply from a preset special vocabulary for the elderly according to the first command phrase, the reply comprising a child reply. The present disclosure can identify the emotion of the old person's pronunciation, respond to the old person accordingly, and use the child's voice to chat with the old person, thereby comforting the old person's emotion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, and in particular, to a voice interaction method and device, a storage medium and an electronic device. BACKGROUND

[0002] At present, various cloud voice recognition technologies have been successfully applied to household appliances. When a user speaks a wake-up word and a command word, continuous voice recognition of a small vocabulary word group can be realized. However, the existing technology must rely on the Internet and a cloud recognition server. Since many old people in the elderly care scene lack network conditions, an offline voice recognition scheme becomes the first choice. The offline voice recognition method in the related file generally has a built-in fixed command word group, and can recognize the pronunciation of commonly used Chinese characters in Putonghua and a small amount of dialects in specific regions. For example, in the elderly care scene, due to the differences in voice tone, dialect, and habitual expression of old people, and the problem of unclear pronunciation, the pronunciation of the wake-up word and the command word of different users for a command may be quite different from the normal pronunciation set by default. In particular, when an urgent help is needed, the voice tone of the old people is significantly different from that in the daily state, such as "help". This may cause the voice recognition system of the device to fail to recognize or incorrectly recognize the command word of the user, and thus the old people cannot be helped in time.

[0003] In addition, the elderly living alone lack chat and confidants, and their emotions are prone to be pessimistic. Therefore, an offline voice emotional interaction supporting emotional analysis is also very important for the spiritual comfort of the old people. SUMMARY

[0004] To solve at least one of the technical problems mentioned above, the present disclosure provides a voice interaction method, device, storage medium and electronic device.

[0005] According to an aspect of the present disclosure, a voice interaction method is provided, which includes:

[0006] In response to a situation of entering an activated state, an input voice is acquired;

[0007] A semantic-based comprehensive analysis is performed on the input voice to determine a first emotional label corresponding to the input voice and a first acoustic feature vector of a key word in the input voice. The comprehensive analysis is also based on at least one of the following factors: frequency, tone, volume, and speed.

[0008] A first command word is searched for in a preset command word group. A second acoustic feature vector of the first command word has a similarity greater than a first preset value with the first acoustic feature vector, and a second emotional label of the first command word is the same as the first emotional label.

[0009] According to the first command word, a corresponding reply is obtained from a preset pension special vocabulary, and the reply includes a child reply, a music reply, or a program reply.

[0010] In some possible implementation manners, the method further includes:

[0011] If the reply is the child reply, and the child reply includes a preset sensitive word, an alarm processing corresponding to the preset sensitive word is performed.

[0012] In some possible implementation manners, the method further includes constructing the preset command word group, and the constructing the preset command word group includes:

[0013] A second command word input by a user is obtained based on a simulated scene, the comprehensive analysis is performed on the second command word, a third acoustic feature vector and a third emotion label corresponding to the second command word are obtained, and the simulated scene includes a falling scene, a fire scene, an illegal intrusion scene, or a chat scene.

[0014] The second command word, the third acoustic feature vector, and the third emotion label are stored in the command word group.

[0015] In some possible implementation manners, the pension special vocabulary includes a child voice library, and the method further includes constructing the child voice library, and the constructing the child voice library includes:

[0016] A chat reply input by a user is obtained.

[0017] A fourth emotion label of the chat reply is set.

[0018] A reply keyword of the chat reply is obtained.

[0019] According to the reply keyword, a mapping relationship between the chat reply and a third command word in the command word group is determined, and a fifth emotion label of the third command word is the same as the fourth emotion label.

[0020] The chat reply, the fourth emotion label, and the mapping relationship are stored in the child voice library.

[0021] In some possible implementation manners, the method of entering the active state includes:

[0022] A wake-up voice input by a user is obtained.

[0023] The comprehensive analysis is performed on the wake-up voice, and a fourth acoustic feature vector is obtained.

[0024] If the fourth acoustic feature vector has a similarity to an acoustic feature vector of a wake-up word in the wake-up word group greater than a second preset value, an active state is entered.

[0025] In some possible implementations, the method further includes constructing the wake-up word group, and the constructing the wake-up word group includes:

[0026] obtaining a first wake-up word inputted for the first time;

[0027] performing the comprehensive analysis on the first wake-up word to obtain a fifth acoustic feature vector of the first wake-up word, and storing the first wake-up word and the fifth acoustic feature vector into the wake-up word group;

[0028] obtaining at least one second wake-up word;

[0029] For each of the second wake-up words, performing the comprehensive analysis on the second wake-up word to obtain a sixth acoustic feature vector of the second wake-up word, calculating a similarity of each wake-up word in the wake-up word group to the second wake-up word based on the sixth acoustic feature vector, and if a maximum similarity is greater than a third preset value, storing the second wake-up word into the wake-up word group;

[0030] If a number of wake-up words in the wake-up word group reaches a fourth preset value, prompting that the wake-up word inputting is completed.

[0031] According to a second aspect of the present disclosure, a voice interaction device is provided, and the device includes:

[0032] An input voice obtaining module is configured to obtain input voice in response to a case of entering an active state.

[0033] An emotion label obtaining module is configured to perform semantic-based comprehensive analysis on the input voice to determine a first emotion label corresponding to the input voice and a first acoustic feature vector of a keyword in the input voice, and the comprehensive analysis is further based on at least one of the following factors: tone, intonation, volume, and speed.

[0034] A command word searching module is configured to search for a first command word from a preset command word group, a second acoustic feature vector of the first command word has a similarity to the first acoustic feature vector greater than a first preset value, and a second emotion label of the first command word is the same as the first emotion label.

[0035] A reply obtaining module is configured to obtain a corresponding reply from a preset special word library for the elderly according to the first command word, and the reply includes a child reply, a music reply, or a program reply.

[0036] In some possible implementations, the device further includes:

[0037] The alarm module is configured to, if the reply is the child reply and the child reply includes a preset sensitive word, perform an alarm process corresponding to the preset sensitive word.

[0038] According to a third aspect of the present disclosure, an electronic device is provided, comprising at least one processor, and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements the voice interaction method according to any one of the first aspect by executing the instructions stored in the memory.

[0039] According to a fourth aspect of the present disclosure, a computer readable storage medium is provided, and the computer readable storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by a processor to implement the voice interaction method according to any one of the first aspect.

[0040] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, but not limiting the present disclosure.

[0041] The implementation of the present disclosure has the following beneficial effects:

[0042] The present disclosure provides a personalized offline dialect voice recognition and emotional interaction method and device based on scenario guidance and a special pension vocabulary, improves the efficiency and accuracy of semantic recognition of various types of dialects, avoids losses caused by recognition errors, and comprehensively judges and analyzes the user's speech frequency, tone, volume, and speed to recognize the current emotion of the human body, select appropriate reply based on the special pension vocabulary, simulate the voice of children, and carry out emotional interaction. It is especially suitable for empty nest, solitary and other pension scenarios, and can modify and edit the special pension vocabulary according to the application scenario, and solve the problem that specific command words cannot be recognized and interacted due to different pronunciation of the elderly in the family pension scenario.

[0043] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the drawings. BRIEF DESCRIPTION OF DRAWINGS

[0044] In order to more clearly illustrate the technical solutions and advantages of the embodiments or prior art in the specification, the drawings needed in the embodiment or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the specification, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0045] Figure 1 A flowchart of a voice interaction method according to an embodiment of the present disclosure is shown.

[0046] Figure 2 Fig. 1 shows a structural schematic diagram of a voice terminal device according to an embodiment of the present disclosure;

[0047] Figure 3 Fig. 2 shows a device schematic diagram of a voice interaction method according to an embodiment of the present disclosure;

[0048] Figure 4 Fig. 3 shows a block diagram of an electronic device according to an embodiment of the present disclosure;

[0049] Figure 5 Fig. 4 shows a block diagram of another electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0050] The technical solutions in the embodiments of the present disclosure will be clearly and completely described in the embodiments of the present disclosure in combination with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present disclosure.

[0051] It should be noted that the terms "first", "second", and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or server including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0052] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numbers in the drawings represent functionally identical or similar elements. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless specifically indicated.

[0053] The word "exemplary" is used herein in the sense of being an example, illustration, or demonstration. Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments.

[0054] The term "and / or", as used herein, merely describes association between associated objects, and can indicate that there are three cases, for example, A and / or B can indicate that there are three cases of A alone, A and B, and B alone. In addition, the term "at least one" herein indicates any one of a plurality or any combination of at least two of a plurality, for example, at least one of A, B and C includes any one or more elements selected from the set consisting of A, B and C.

[0055] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the specific embodiments below. Those skilled in the art should understand that the present disclosure can also be implemented without some specific details. In some examples, methods, means, elements and circuits well known to those skilled in the art are not described in detail, in order to highlight the main ideas of the present disclosure.

[0056] The implementation scenario of the present disclosure includes a voice terminal device, which can be used offline or online.

[0057] Figure 1 A flowchart of a voice interaction method according to an embodiment of the present disclosure is shown as follows, Figure 1 The method includes the following steps:

[0058] S101, in response to the situation of entering the active state, acquiring the input voice;

[0059] When the voice terminal device receives the wake-up word input by the user, it enters the active state. In other embodiments, the voice terminal device can also be in the active state all the time and does not need a wake-up word. When the voice terminal device is in the active state, it can receive the input voice input by the user at any time;

[0060] S102, performing semantic-based comprehensive analysis on the input voice to determine the first emotional label corresponding to the input voice and the first acoustic feature vector of the keyword in the input voice, and the comprehensive analysis is also based on at least one of the following factors: tone, intonation, volume and speed;

[0061] The input voice is subjected to semantic-based comprehensive analysis, wherein the comprehensive analysis comprises emotion analysis of the voice and extraction of an acoustic feature vector. The emotion analysis of the input voice can be performed by inputting the input voice into a preset emotion analysis model. The preset emotion analysis model quantifies and grades the frequency, tone, volume and speed of the input voice, performs emotion analysis by using a convolutional neural network (CNN), and infers the current user's emotional state as anger, happiness, fear, sadness, surprise or neutrality. The received input voice is input into the preset emotion analysis model to obtain a first emotion label corresponding to the input voice. The first emotion label comprises an anger label, a happiness label, a fear label, a surprise label or a neutral label.

[0062] S103. A first command word is searched from the preset command word group. A second acoustic feature vector of the first command word has a similarity to the first acoustic feature vector greater than a first preset value, and a second emotion label of the first command word is the same as the first emotion label.

[0063] The keyword of the input voice can be extracted by using an unsupervised algorithm, such as term frequency-inverse document frequency (TF-IDF) based keyword extraction. In other embodiments, other keyword extraction algorithm models can also be used. The keyword voice of the input voice is obtained, the first acoustic feature vector of the keyword voice is extracted, the command word with the same emotion label as the emotion label of the input voice is searched from the pre-stored command word group, and the second acoustic feature vector with a similarity to the first acoustic feature vector greater than the first preset value is searched from the command word with the same emotion label. The first command word corresponding to the second acoustic feature vector is obtained, i.e., the second emotion label of the obtained first command word is the same as the first emotion label of the input voice, and the second acoustic feature vector of the first command word is similar to the second acoustic feature vector of the keyword of the input voice.

[0064] S104. A corresponding reply is obtained from a preset pension special vocabulary according to the first command word. The reply comprises a child reply, a music reply or a program reply.

[0065] The first command word is input into a preset pension special vocabulary, and through a mapping relationship between the command word and the pension special vocabulary, a corresponding reply is found. For example, when the first command word is "play music", the corresponding music reply is played; when the first command word is "listen to a radio program", the corresponding program reply is output; and when the first command word is "what to eat", the corresponding child reply in the child reply library is output. The semantic, frequency, tone, volume and speed of the input voice are comprehensively judged and analyzed to identify the current emotion of the user, select a suitable reply based on the pension special vocabulary, simulate the voice of the child, and carry out emotional interaction, which is particularly suitable for empty nest, single and other pension scenarios.

[0066] In one embodiment, the method further comprises:

[0067] If the reply is the child reply, and the child reply includes a preset sensitive word, an alarm corresponding to the preset sensitive word is performed.

[0068] In one embodiment, as shown in Figure 2 The voice terminal device includes a control center, a wake-up module, an emotion recognition and analysis model, a pension special vocabulary, and a remote alarm module. The pension special vocabulary includes a child voice library, a music voice library, and a program voice library, and the voice terminal device is connected to the cloud.

[0069] When the voice terminal device is connected to the network, the cloud module includes an alarm module, which can be published through a WeChat public number module or alarmed through a telephone message. For example, when the first command word is "call an ambulance", the corresponding child reply includes "don't be afraid, I will call the police", and the voice terminal device directly dials 120 and the child's phone number; when the voice terminal device is not connected to the network, the first command word is "help", and the corresponding child reply includes "I have called the police", and the voice terminal device directly dials the emergency number. In the event of an emergency, the old person can be calmed down in time, and the old person can be helped to dial an emergency phone number, so that the old person can dial an emergency phone number when he or she falls down and cannot move, and thus the best treatment time is not missed.

[0070] In one embodiment, the method further comprises constructing the preset command word group, and the constructing the preset command word group comprises:

[0071] Based on the simulated scenario, a second command word input by the user is obtained, the second command word is comprehensively analyzed, a third acoustic feature vector and a third emotion label corresponding to the second command word are obtained, and the simulated scenario includes a fall scenario, a fire scenario, an illegal intrusion scenario, or a chat scenario.

[0072] The second command word, the third acoustic feature vector, and the third emotion label are stored in the command word group.

[0073] The voice terminal device is set with a command word group, the old people are guided in a scene, and some command words of simulated falling scene, fire scene, illegal intrusion scene or chat scene are simulated, such as the normal reaction of the old people in the falling scene is "ah…!", and the sound of heavy object falling and pain groan, at this time the voice terminal device collects these sounds as a second command word, when the second command word is input, a third acoustic feature vector of the second command word is obtained, the second command word is input to the emotion analysis model, a third emotion label of the second command word is recognized, for example, the third emotion label is "pain label" and "frightened label", etc., the second command word, the third acoustic feature vector and the third emotion label are stored in the command word group. According to the specific application, the command word group can be modified and customized. Through the voice terminal device, various scenes are simulated, such as fire emergency scene, and the old people are guided to say "help" and other alarm command words with voice tone in the scene, so as to facilitate voice help in real scene.

[0074] In one embodiment, the special vocabulary for the elderly includes a child voice library; the method further includes constructing the child voice library, and the constructing the child voice library includes:

[0075] Obtaining a chat reply speech input by a user based on the simulated scene;

[0076] Setting a fourth emotion label of the chat reply speech;

[0077] Obtaining a reply keyword of the chat reply speech;

[0078] According to the reply keyword, determining a mapping relationship between the chat reply speech and a third command word in the command word group, wherein a fifth emotion label of the third command word is the same as the fourth emotion label;

[0079] Storing the chat reply speech, the fourth emotion label and the mapping relationship to the child voice library.

[0080] In an embodiment, the pension special vocabulary includes a child voice library, a music voice library, and a program voice library. The music voice library is mainly related to pre-downloaded offline music or music obtained from a cloud disk after the voice terminal device is connected to the network. When a user wants to listen to music, a related command word can be input to obtain the music. The program voice library stores related programs such as stand-up comedy. The child voice library is customized according to the needs of the user. The child voice library of the voice terminal device can be the voice recorded by a close person, such as a parent whose children are not nearby. The parent wants to hear the voice of the child at any time. A command word in a command word group is input, and a voice is input, such as a command word "what to eat for dinner". A chat reply voice can be input, such as "eat light in the evening, and cook a green vegetable and egg noodle". A fourth emotional tag of the chat reply voice is set as a neutral tag. The fourth emotional tag is mainly related to the emotional tag of the command word responded by the chat reply voice, and is not the emotion of the chat reply voice. For example, when an old person is angry, the emotional tag of the command word spoken by the old person is an angry tag, but the corresponding child reply voice should have an emotional tag such as a gentle tag or a comfort tag. The child is guided to speak a reply word in a specific situation, and the voice is stored in the voice terminal device. The old person can chat and interact with the voice terminal device, and the emotion of the old person can be soothed.

[0081] In an embodiment, the method of entering the active state includes:

[0082] obtaining a wake-up voice input by a user;

[0083] performing the comprehensive analysis on the wake-up voice to obtain a fourth acoustic feature vector;

[0084] if the similarity between the fourth acoustic feature vector and the acoustic feature vector of the wake-up word in the wake-up word group is greater than a second preset value, entering the active state.

[0085] The voice terminal device enters the active state, which can be woken up by a wake-up word. The voice terminal device obtains a wake-up voice input by a user, performs feature recognition on the wake-up voice, obtains a fourth acoustic feature vector, finds a wake-up word with the greatest similarity to the wake-up voice from a wake-up word group through the fourth acoustic feature vector, obtains the greatest similarity, and judges whether the similarity is greater than a second preset value. If it is greater, it means that the wake-up word is correct, and the voice terminal device enters the active state. At this time, the voice terminal device can communicate and interact with the user. The voice terminal device is woken up by the wake-up word, which avoids noise interference of the voice terminal device.

[0086] In an embodiment, the method further includes constructing the wake-up word group, and the constructing the wake-up word group includes:

[0087] obtaining a first wake-up word input for the first time;

[0088] performing the comprehensive analysis on the first wake-up word to obtain a fifth acoustic feature vector of the first wake-up word, and storing the first wake-up word and the fifth acoustic feature vector into the wake-up word group;

[0089] obtaining at least one second wake-up word;

[0090] For each of the second wake-up words, performing the comprehensive analysis on the second wake-up word to obtain a sixth acoustic feature vector of the second wake-up word, calculating a similarity between each wake-up word in the wake-up word group and the second wake-up word based on the sixth acoustic feature vector, and if the maximum similarity is greater than a third preset value, storing the second wake-up word into the wake-up word group;

[0091] If the number of wake-up words in the wake-up word group reaches a fourth preset value, it is prompted that the wake-up word input is completed. The command words in the command word group can be multiple, including various forms of voice, such as English wake-up words, emergency wake-up words, dialect wake-up words, etc., which are input by the user. When the user sets the wake-up word, the voice terminal device prompts the input of the wake-up word. When the voice terminal device prompts the input of the wake-up word for the first time, the user speaks the wake-up word. The wake-up word can be in English, dialect or Mandarin form. The content is set by the user according to his own situation. It can be "ah, ah, ah" for the convenience of the old people, or it can be a name. When the user inputs the wake-up word for the first time, that is, the first wake-up word, the acoustic feature vector A21 of the first wake-up word is extracted, and the first wake-up word and the acoustic feature vector A21 of the first wake-up word are stored in the wake-up word group. When the voice terminal device prompts the input of the wake-up word for the second time, the user repeats the content of the first wake-up word to obtain the second wake-up word. The acoustic feature vector A22 of the second wake-up word is extracted, and the acoustic feature vector A22 of the second wake-up word is compared with the acoustic feature vector A21 of the first wake-up word. If the similarity reaches a third preset value, the acoustic feature vector A22 of the second wake-up word is stored in the wake-up word group. When the voice terminal device prompts the input of the wake-up word for the third time, the user repeats the content of the first wake-up word to obtain the third wake-up word. The acoustic feature vector A23 of the third wake-up word is extracted, and the acoustic feature vector A23 of the third wake-up word is compared with the acoustic feature vector A21 of the first wake-up word and the acoustic feature vector A22 of the second wake-up word in the wake-up word group. The maximum similarity value is obtained. If the maximum similarity value is greater than the third preset value, the third wake-up word and the acoustic feature vector A23 are stored in the wake-up word group. Each time the wake-up word is input, the wake-up word is compared with the wake-up words already stored in the wake-up word group. If the similarity is greater than the third preset value, it is stored in the wake-up word group. If the similarity is less than the third preset value, it is discarded. When the number of wake-up words in the wake-up word group is greater than the fourth preset value, which can be 3, it is prompted that the input is successful, and the input is stopped. Different wake-up words are defined, the commonly used command words in the word library are updated according to the voice tone of the old people, arbitrary personalized dialect recognition is realized, the problem that the old people cannot recognize specific command words and interact due to different pronunciation is solved, and the universality is strong.

[0092] Please refer to Figure 3 According to a second aspect of the present disclosure, a voice interaction device is provided, which comprises:

[0093] The input voice acquisition module 10 is configured to acquire input voice in response to the entering of the active state.

[0094] The emotion tag obtaining module 20 is configured to perform semantic-based comprehensive analysis on the input voice to determine a first emotion tag corresponding to the input voice and a first acoustic feature vector of a keyword in the input voice, and the comprehensive analysis is further based on at least one of the following factors: tone, intonation, volume, and speed.

[0095] The command word searching module 30 is configured to search for a first command word from a preset command word set, the first command word has a second acoustic feature vector that is similar to the first acoustic feature vector by more than a first preset value, and the first command word has a second emotion tag that is the same as the first emotion tag.

[0096] The reply obtaining module 40 is configured to obtain a corresponding reply from a preset special word bank for the elderly according to the first command word, and the reply includes a child reply, a music reply, or a program reply. In an embodiment, the device further includes:

[0097] The alarm module is configured to, if the reply is the child reply and the child reply includes a preset sensitive word, perform an alarm corresponding to the preset sensitive word.

[0098] In some embodiments, the device provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and specific implementations can refer to the descriptions of the above method embodiments. For brevity, they will not be repeated here.

[0099] The embodiments of the present disclosure also propose a computer-readable storage medium, the computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the above method. The computer-readable storage medium can be a non-volatile computer-readable storage medium.

[0100] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method.

[0101] The electronic device can be provided as a terminal, a server, or other forms of devices.

[0102] Figure 4 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. For example, the electronic device 800 can be a terminal such as a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0103] Reference Figure 4The electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0104] The processing component 802 usually controls overall operations of the electronic device 800, such as operations associated with displaying, making phone calls, data communications, camera operations, and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or part of steps of the above methods. In addition, the processing component 802 can include one or more modules to facilitate

[0105] The memory 804 is configured to store various types of data to support operations of the electronic device 800. Examples of these data include instructions for any application or method operating on the electronic device 800, contact data, phonebook data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic or optical disk.

[0106] The power supply component 806 provides power for the various components of the electronic device 800. The power supply component 806 can include a power supply management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0107] The multimedia component 808 includes a screen to provide an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and intensity of the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a camera mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zooming capability.

[0108] The audio component 810 is configured to output and / or input an audio signal. For example, the audio component 810 includes a microphone (MIC) to receive an external audio signal when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker to output an audio signal.

[0109] The I / O interface 812 provides an interface for the processing component 802 and a peripheral interface module, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0110] The sensor component 814 includes one or more sensors to provide various state assessments for the electronic device 800. For example, the sensor component 814 can detect an open / closed state of the electronic device 800, relative positioning of components, such as a display and a keypad of the electronic device 800, a change in position of the electronic device 800 or a component of the electronic device 800, presence or absence of user contact with the electronic device 800, an orientation or acceleration / deceleration of the electronic device 800, and a temperature change of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect presence of a nearby object without any physical touch. The sensor component 814 can further include a light sensor, such as a CMOS or CCD image sensor, for use in an imaging application. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0111] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G, 5G, or a combination thereof. In an example embodiment, the communication component 816 receives a broadcast signal or broadcast related information from an external broadcast management system via a broadcast channel. In an example embodiment, the communication component 816 can further include a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) techniques, infrared data association (IrDA) techniques, ultra-wideband (UWB) techniques, Bluetooth (BT) techniques, and other techniques.

[0112] In an example embodiment, the electronic device 800 can be implemented with one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements, for performing the above-described methods.

[0113] In an example embodiment, a non-transitory computer-readable storage medium, such as the memory 804 including computer program instructions, is also provided, which can be executed by the processor 820 of the electronic device 800 to complete the above-described methods.

[0114] Figure 5 A block diagram of another electronic device according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server. Referring to Figure 5 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932, for storing instructions executable by the processing component 1922, such as an application program. The application program stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described methods.

[0115] The electronic device 1900 can also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output (I / O) interface 1958. The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, or the like.

[0116] In example embodiments, a non-transitory computer-readable storage medium, e.g., memory 1932 including computer program instructions, is also provided that can be executed by processing component 1922 of electronic device 1900 to implement the above-described methods.

[0117] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0118] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a

[0119] The computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0120] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0121] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0122] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0123] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0124] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and

[0125] Embodiments of the present disclosure have been described above, with examples of the specification being exemplary and not exhaustive, and not limited to the disclosed embodiments. Many modifications and variations to the described embodiments will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The choice of words in the specification is intended to best explain the principles of the embodiments, practical application, or improvement to the art in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A voice interaction method, characterized in that, The method comprises: in response to entering an active state, obtaining an input voice; performing semantic-based comprehensive analysis on the input voice to determine a first emotional label corresponding to the input voice and a first acoustic feature vector of a keyword in the input voice, the comprehensive analysis being further based on at least one of the following factors: tone, intonation, volume and speed, and the first emotional label comprising an angry label, a happy label, a scared label, a surprised label or a neutral label; finding a first command word from a preset command word group, the first command word having a second acoustic feature vector similar to the first acoustic feature vector by more than a first preset value, and the first command word having a second emotional label identical to the first emotional label, the command word group being constructed based on comprehensive analysis of command words corresponding to simulated scenarios, the simulated scenarios comprising a fall scenario, a fire scenario, an illegal intrusion scenario or a chat scenario; obtaining a corresponding reply from a preset elderly-care-specific vocabulary according to the first command word, the reply comprising a child reply, a music reply or a program reply, the child reply being a reply simulating a child's voice, and the elderly-care-specific vocabulary comprising a child voice library; if the reply is the child reply and the child reply includes a preset sensitive word, performing alarm processing corresponding to the preset sensitive word; obtaining a chat reply input by a user based on a simulated scenario; setting a fourth emotional label of the chat reply; obtaining a reply keyword of the chat reply; determining a mapping relationship between the chat reply and a third command word in the command word group according to the reply keyword, wherein the third command word has a fifth emotional label identical to the fourth emotional label; storing the chat reply, the fourth emotional label and the mapping relationship to the child voice library.

2. The method of claim 1, wherein, The method further comprises constructing the preset command word group, and the constructing the preset command word group comprises: obtaining a second command word input by a user based on a simulated scenario, and performing the comprehensive analysis on the second command word to obtain a corresponding third acoustic feature vector and a third emotional label; storing the second command word, the third acoustic feature vector and the third emotional label to the command word group.

3. The method of claim 1, wherein, The method of entering an active state comprises: obtaining a wake-up voice input by a user; performing the comprehensive analysis on the wake-up voice to obtain a fourth acoustic feature vector; if the fourth acoustic feature vector has a similarity to an acoustic feature vector of a wake-up word in a wake-up word group by more than a second preset value, entering an active state.

4. The method of claim 3, wherein, The method further comprises constructing the wake-up word group, and the constructing the wake-up word group comprises: obtaining a first wake-up word input for the first time; performing the comprehensive analysis on the first wake-up word to obtain a fifth acoustic feature vector of the first wake-up word, and storing the first wake-up word and the fifth acoustic feature vector to the wake-up word group; obtaining at least one second wake-up word; For each of the second wake-up word, the comprehensive analysis is performed on the second wake-up word to obtain a sixth acoustic feature vector of the second wake-up word, a similarity between each wake-up word in the wake-up word group and the second wake-up word is calculated based on the sixth acoustic feature vector, and if the maximum similarity is greater than a third preset value, the second wake-up word is stored in the wake-up word group; If the number of wake-up words in the wake-up word group reaches a fourth preset value, it is prompted that the wake-up word input is completed.

5. A voice interaction device, characterized by The device comprises: An input voice acquisition module configured to acquire input voice in response to entering an active state; An emotion label acquisition module configured to perform semantic-based comprehensive analysis on the input voice to determine a first emotion label corresponding to the input voice and a first acoustic feature vector of a keyword in the input voice, the comprehensive analysis being further based on at least one of the following factors: tone, intonation, volume, and speed, and the first emotion label comprising an angry label, a happy label, a scared label, a surprised label, or a neutral label; A command word searching module configured to search for a first command word from a preset command word group, the similarity between a second acoustic feature vector of the first command word and the first acoustic feature vector being greater than a first preset value, and the second emotion label of the first command word being the same as the first emotion label, the command word group being constructed based on a comprehensive analysis of command words corresponding to simulated scenarios, the simulated scenarios comprising a fall scenario, a fire scenario, an illegal intrusion scenario, or a chatting scenario; A reply speech acquisition module configured to acquire corresponding reply speech from a preset special word library for the elderly according to the first command word, the reply speech comprising child reply speech, music reply speech, or program reply speech, the child reply speech being reply speech simulating the voice of a child, and the special word library for the elderly comprising a child voice library; An alarm module configured to perform alarm processing corresponding to a preset sensitive word if the reply speech is the child reply speech and the child reply speech comprises the preset sensitive word; A simulated reply speech acquisition module configured to acquire chatting reply speech input by a user based on a simulated scenario; An emotion label setting module configured to set a fourth emotion label of the chatting reply speech; A keyword acquisition module configured to acquire a reply keyword of the chatting reply speech; A mapping relationship determination module configured to determine a mapping relationship between the chatting reply speech and a third command word in the command word group according to the reply keyword, wherein a fifth emotion label of the third command word is the same as the fourth emotion label; A storage module configured to store the chatting reply speech, the fourth emotion label, and the mapping relationship in the child voice library.

6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction or at least one program, which is loaded and executed by the processor to implement the voice interaction method according to any one of claims 1-4.

7. An electronic device, comprising: The application discloses a voice interaction method and device, and a computer readable storage medium. The voice interaction device comprises at least one processor and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements the voice interaction method according to any one of claims 1-4 by executing the instructions stored in the memory.

Citation Information

Patent Citations

  • Smart home voice control system and method based on voice fuzzy recognition technology

    CN106847281A

  • Intelligent elderly care service system and method

    CN112086091A

  • Voice generation method and device, storage medium and electronic equipment

    CN112185389A