Parent-child communication guiding method and device, and electronic equipment
By using voice acquisition and recognition technology, the voice characteristics and emotions of parents and children are identified, and guiding voice messages are output to improve parent-child communication, thus solving the problem of parent-child relationship tension caused by negative emotions and improving parent-child relationship.
Patent Information
- Application Number
- CN202210499189.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2042-05-09
AI Technical Summary
When parents accompany their children to read and study, the atmosphere of parent-child communication may become tense due to the child's slow comprehension or the parents' loss of patience, which may affect the parent-child relationship.
Voice signals are collected by a voice acquisition device, and acoustic feature recognition and voice recognition are performed to identify the speaker's identity and emotions, and output communication guidance voice to improve parent-child relationships.
When parents and children experience negative emotions, promptly remind and guide them to communicate in a positive manner to improve parent-child relationships and prevent the communication atmosphere from deteriorating further.
Smart Images

Figure CN114944166B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of information processing, and in particular to a parent-child communication guiding method and device and electronic equipment. BACKGROUND
[0002] In the growth stage of children, accompanying children to read and learn by parents is an important part of family education. In the process of accompanying children to read and tutoring children's homework by parents, emotional changes will inevitably occur due to the absence of children, slow understanding of knowledge, loss of patience of parents, etc., so that the communication atmosphere between children and parents becomes tense and unsmooth, affecting the parent-child relationship. How to use intelligent means to help maintain a good relationship between parents and children during the process of accompanying children to read is a problem to be solved at present. SUMMARY
[0003] The present application provides a parent-child communication guiding method, device and electronic equipment to realize intelligent guiding of communication mode between parents and children.
[0004] The present application provides a parent-child communication guiding method, comprising:
[0005] obtaining a voice signal collected by a voice collection device;
[0006] performing acoustic feature recognition on the voice signal, and determining speaker identity information based on the recognized acoustic feature;
[0007] performing voice recognition on the voice signal to obtain voice content corresponding to the speaker identity information;
[0008] performing volume detection on the voice signal to obtain volume information corresponding to the speaker identity information;
[0009] identifying the emotion of the speaker corresponding to the speaker identity information based on the voice content and the volume information to obtain a target emotion;
[0010] in response to the target emotion being an undesirable emotion, outputting a communication mode guiding voice.
[0011] According to the parent-child communication guiding method provided by the present application, after obtaining the voice signal collected by the voice collection device, the parent-child communication guiding method further comprises:
[0012] performing speech rate detection on the voice signal to obtain speech rate information corresponding to the speaker identity information;
[0013] The identification of the emotion of the speaker corresponding to the speaker identity information based on the voice content and the volume information to obtain a target emotion comprises:
[0014] identify an emotion of a speaker corresponding to the speaker identity information based on the voice content, the speech speed information and the volume information, and obtain a target emotion.
[0015] According to the parent-child communication guiding method provided by the application, the outputting of the communication mode guiding voice in response to the target emotion being an undesirable emotion comprises:
[0016] determining the content of the communication mode according to the speaker identity information and the target emotion in response to the target emotion being an undesirable emotion.
[0017] converting the content of the communication mode into a communication mode guiding voice.
[0018] outputting the communication mode guiding voice.
[0019] According to the parent-child communication guiding method provided by the application, the determining of the content of the communication mode according to the speaker identity information and the target emotion comprises:
[0020] matching a corresponding communication mode from a communication mode information table according to the speaker identity information and the target emotion to obtain the content of the communication mode, wherein the communication mode information table stores a corresponding relationship among speaker identity, emotion and communication mode.
[0021] According to the parent-child communication guiding method provided by the application, the communication mode information table is created based on a good communication mode of psychology.
[0022] According to the parent-child communication guiding method provided by the application, in response to the speaker identity information indicating that the speaker corresponding to the speaker identity information is a child, the target emotion comprises at least one of collapse, crying and temper; and the outputting of the communication mode guiding voice comprises:
[0023] outputting the communication mode guiding voice in a soothing tone.
[0024] According to the parent-child communication guiding method provided by the application, in response to the speaker identity information indicating that the speaker corresponding to the speaker identity information is a parent, the target emotion comprises at least one of impatience, temper and loud scolding; and the outputting of the communication mode guiding voice comprises:
[0025] outputting the communication mode guiding voice in a child's tone.
[0026] The application further provides a parent-child communication guiding device, comprising:
[0027] an acquisition module, configured to acquire a voice signal collected by a voice collection device;
[0028] The first identification module is configured to perform acoustic feature identification on the voice signal, and determine the speaker identity information based on the identified acoustic feature.
[0029] The second identification module is configured to perform voice recognition on the voice signal, and obtain voice content corresponding to the speaker identity information.
[0030] The volume detection module is configured to perform volume detection on the voice signal, and obtain volume information corresponding to the speaker identity information.
[0031] The third identification module is configured to identify the emotion of the speaker corresponding to the speaker identity information based on the voice content and the volume information, and obtain a target emotion.
[0032] The output module is configured to output a communication mode guiding voice in response to the target emotion being an undesirable emotion.
[0033] The present application also provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the parent-child communication guiding method according to any one of the above when executing the computer program.
[0034] The present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the parent-child communication guiding method according to any one of the above.
[0035] The present application also provides a computer program product, which comprises a computer program, and the computer program is executable on a processor to implement the parent-child communication guiding method according to any one of the above.
[0036] The parent-child communication guiding method, device and electronic device provided by the present application can collect a voice signal, perform acoustic feature identification, voice recognition and volume detection on the voice signal, obtain speaker identity information and corresponding voice content and volume information, identify the emotion of the speaker corresponding to the speaker identity information based on the voice content and the volume information, and output a communication mode guiding voice when the emotion is an undesirable emotion, so as to remind and guide the speaker to communicate in a good communication mode by using the communication mode guiding voice, and to help improve undesirable emotions and communication atmosphere and prevent further deterioration of the communication atmosphere, thereby realizing intelligent guiding of the communication mode between parents and children and providing help for maintaining a good parent-child relationship. BRIEF DESCRIPTION OF DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required by the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0038] Figure 1 is one of the process schematic diagrams of the parent-child communication guiding method provided by the present application;
[0039] Figure 2 is a process schematic diagram of the method for outputting the communication mode guiding voice when the target emotion is an undesirable emotion;
[0040] Figure 3 is the second process schematic diagram of the parent-child communication guiding method provided by the present application;
[0041] Figure 4 is a structural schematic diagram of the parent-child communication guiding device provided by the present application;
[0042] Figure 5 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION
[0043] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below in combination with the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0044] In the process of reading and learning with the accompaniment of parents, due to the influence of various factors such as the different cognitive levels of children and parents, the personalities of children and parents, etc., the situation that children cannot understand the content of parents' speech and parents lose patience and start to get angry or have impulsive emotions may occur when parents guide children to do homework. The situation that children cry and get angry after being reminded by parents may occur when children do not do homework seriously. No matter parents or children, the communication atmosphere between parents and children becomes tense when these undesirable emotions occur, and it will be further deteriorated with the development of undesirable emotions, which affects the parent-child relationship. If parents and children can be reminded and guided to communicate in a good communication mode when undesirable emotions occur, it will help to improve the parent-child relationship.
[0045] Based on this, in the embodiment of the present application, a parent-child communication guidance scheme is provided, which can monitor the emotional changes of children and parents in the process of parents accompanying reading with the aid of intelligent technology, guide parents and children to control emotions in a benign guidance manner in a timely manner from the perspective of children, and improve the parent-child relationship.
[0046] The parent-child communication guidance method of the present application is described below. The parent-child communication guidance method can be applied to servers, mobile phones and other electronic devices, and can also be applied to special intelligent devices such as companion robots, and can also be applied to a parent-child communication guidance device provided in a server, a mobile phone or other electronic device or a special intelligent device. The parent-child communication guidance device can be realized by software, hardware or a combination of both. Taking the case where the parent-child communication guidance method is applied to a special intelligent device (such as a companion robot) or an electronic device, the intelligent device or the electronic device can include a voice collection device and a voice output device. The voice collection device can be a microphone, for example. The intelligent device can collect the voice signals of parents and children through the microphone, and then perform speaker identification, voice content recognition and volume detection on the voice signals. If it is identified that the speaker is a child and the child has adverse emotions such as collapse, crying or temper, the child can be reminded to express and communicate well with the parents in a gentle tone; if it is identified that the speaker is a parent and the parent has adverse emotions such as impatience, temper or shouting at the child, the parent can be reminded to control emotions in the child's tone, and communicate with the child in a calm tone to alleviate the tense communication atmosphere and improve the parent-child relationship.
[0047] Figure 1 An exemplary flowchart of the parent-child communication guidance method provided by the embodiment of the present application is shown, referring to Figure 1 The parent-child communication guidance method can include the following steps 110-160.
[0048] Step 110: Acquire the voice signal collected by the voice collection device.
[0049] The electronic device or the intelligent device can collect the voice signal in the environment where the voice collection device is located through the voice collection device. For example, the electronic device or the intelligent device can include the voice collection device, such as collecting the voice signal through its own microphone. For example, the electronic device or the intelligent device can also be in communication connection with an external voice collection device to acquire the voice signal collected by the voice collection device.
[0050] Step 120: Perform acoustic feature recognition on the voice signal, and determine the speaker identity information based on the recognized acoustic features.
[0051] In an example embodiment, after the electronic device or the smart device acquires the voice signal, the voice of the monitored child and the parent can be registered first, and the acoustic features of the child and the parent are extracted to obtain the first reference acoustic features of the child and the second reference acoustic features of the parent. For example, among family members, the parent can be multiple parent identities, such as father, mother, grandfather, grandmother, etc., and the monitored child can also be multiple, such as child 1 and child 2. When registering the voice, the registration can be performed by role. For example, in the parent identity category, the voices of the father, mother, grandfather, and grandmother can be registered respectively to obtain the sub-reference acoustic features of each parent, and the sub-reference acoustic features of each parent can carry an identity identification, by which the parent role can be identified. These sub-reference acoustic features constitute the second reference acoustic features corresponding to the parent identity category; in the child category, the voice of each child can also be registered, such as the voice of child 1 and child 2, to obtain the sub-reference acoustic features of child 1 and child 2, and the sub-reference acoustic features of each child can carry an identity identification, by which the child identity can be identified. These sub-reference acoustic features of the children constitute the first reference acoustic features corresponding to the child category.
[0052] During the accompanying reading process of the parent, the electronic device or the smart device can collect the voice signal through the voice collection device, extract the acoustic features of the voice signal, match the extracted acoustic features with the registered first reference acoustic features and second reference acoustic features, and if the matching with the first reference acoustic features is successful, it is determined that the speaker identity information is the child; if the matching with the second reference acoustic features is successful, it is determined that the speaker identity information is the parent identity.
[0053] For example, the extracted acoustic features can be matched with all the sub-reference acoustic features in the registered first reference acoustic features and all the sub-reference acoustic features in the second reference acoustic features, and the specific role of the speaker can be determined, such as the extracted acoustic features matching with the sub-reference acoustic features of child 2 in the first reference acoustic features, it is determined that the speaker identity information is child 2; such as the extracted acoustic features matching with the sub-reference acoustic features of the mother in the second reference acoustic features, it is determined that the speaker identity information is the mother.
[0054] Exemplarily, the speaker identity information can be recognized through a voiceprint feature in the acoustic feature. After the speech signal is acquired, a speaker recognition (SR) technology can be used to perform voiceprint recognition on the speech signal to extract a voiceprint feature in the speech signal. The SR can also be referred to as voiceprint recognition (VPR), which is a biometric recognition technology for recognizing the speaker identity according to the speaker individual information in the speech signal, and uses the theory that each voice has unique features. By performing voiceprint feature extraction on the collected speech signal, the voiceprint feature of the speaker can be obtained, and the speaker identity can be identified based on the voiceprint feature.
[0055] Step 130: performing speech recognition on the speech signal to obtain speech content corresponding to the speaker identity information.
[0056] Automatic speech recognition (ASR) can convert the speech signal of a person into corresponding text to obtain speech content corresponding to the speech signal. After the electronic device or the smart device acquires the speech signal, a speech recognition technology can be used to perform speech recognition on the speech signal to obtain speech content corresponding to the speaker identity information.
[0057] Step 140: performing volume detection on the speech signal to obtain volume information corresponding to the speaker identity information.
[0058] Exemplarily, a volume detection plug-in can be arranged in the electronic device or the smart device, and after the speech signal is acquired, the volume detection plug-in can be used to detect the volume of the speech signal to obtain volume information corresponding to the speaker identity information. Exemplarily, the volume detection plug-in can be in communication connection with the speech collection device, and can detect the volume of the speech signal collected by the speech collection device.
[0059] Exemplarily, the volume of the speech represents the intensity of the speech, and after the electronic device or the smart device acquires the volume signal, the volume signal can be divided into frames according to the time granularity, and the amplitude of each audio frame signal can be calculated to obtain the volume.
[0060] Step 150: identifying the emotion of the speaker corresponding to the speaker identity information based on the speech content and the volume information to obtain a target emotion.
[0061] The volume of the voice can reflect the emotion of a person. For example, when a person is angry, scolds others or is impatient, the volume of the voice is relatively high, and the volume can also fluctuate greatly. When the person is in a stable emotion, the volume of the voice is relatively stable and moderate. Exemplarily, the emotion of the speaker can be determined by a volume threshold. For example, when the volume is greater than the volume threshold, it indicates that the speaker has an undesirable emotion, and when the volume is less than the volume threshold, it indicates that the speaker has a good emotion. Exemplarily, different volume thresholds can be set for different emotions. For example, when the volume is less than a first volume threshold, it indicates that the speaker has a good emotion, when the volume is greater than the first volume threshold and less than a second volume threshold, it indicates that the speaker has an impatient emotion, and when the volume is greater than the second volume threshold and less than a third volume threshold, it indicates that the speaker has an emotion of scolding, shouting or being angry, and the like. It should be noted that the above is only an example and is not used to limit the present application.
[0062] The voice content of a person can also reflect the emotion of the person. For example, if the voice content includes the sentence "How many times have I told you, why did you make a mistake again!", it can be determined that the emotion of the speaker is scolding, shouting or being angry. Exemplarily, the voice content can be processed by word segmentation, and the word segmentation representing the emotion can be recognized to determine the emotion of the speaker. Exemplarily, only the undesirable emotion can be recognized by the voice content. For example, an undesirable emotion word segmentation library can be established, the word segmentation in the voice content is matched with the undesirable emotion word segmentation library, and if the word segmentation is matched, it is determined that the speaker has an undesirable emotion.
[0063] In an example embodiment, the volume information of the voice and the voice content can be combined to determine the emotion of the speaker. For example, for the voice content "read it clearly", if the volume of the speaker is moderate, it indicates that the speaker only gently reminds the child and does not have an undesirable emotion. If the speaker speaks this sentence with a relatively large volume, it indicates that the speaker has an undesirable emotion such as impatience or scolding. In this way, by combining the volume information of the voice and the voice content to determine the emotion of the speaker, the accuracy of emotion recognition can be improved.
[0064] In an example embodiment, the speech content and volume information can be recognized based on a pre-constructed first speech emotion recognition model to obtain the emotion of the speaker. In an example, an initial neural network model can be selected, and the initial neural network model can be trained using a first training sample set to obtain the speech emotion recognition model. The first training sample set can include speech sample data and first emotion type data corresponding to the speech sample data. The speech sample data can be recognized to obtain text data, and text features can be extracted from the text data, and volume features can be extracted from the speech sample data. The first emotion type data is annotation data of the text features and the volume features. The extracted text features and volume features are used as input data of the initial neural network model, and the first emotion type data is used as output data of the initial neural network model. The initial neural network model is trained to obtain the first speech emotion recognition model.
[0065] In an example, only keywords of bad emotions can be annotated when annotating the text features, and only bad emotions can be recognized when recognizing the speech emotions.
[0066] Step 160: outputting a communication mode guiding voice in response to the target emotion being a bad emotion.
[0067] After obtaining the target emotion, the speaker can be guided in communication mode according to the target emotion. In an example, after obtaining the target emotion, the target emotion can be classified using a classifier to determine whether the target emotion is a good emotion or a bad emotion.
[0068] In an example embodiment, Figure 2 A flowchart of a method for outputting a communication mode guiding voice when the target emotion is a bad emotion is shown, and the method is provided by the present application. Referring to Figure 2 As shown, the method can include steps 161-163.
[0069] Step 161: determining the content of the communication mode according to the speaker identity information and the target emotion in response to the target emotion being a bad emotion.
[0070] In an example embodiment, determining the content of the communication mode according to the speaker identity information and the target emotion can include: matching a corresponding communication mode from a communication mode information table according to the speaker identity information and the target emotion to obtain the content of the communication mode, and the communication mode information table stores the corresponding relationship between the speaker identity, the emotion, and the communication mode.
[0071] In an example, the communication mode information table can be created based on good communication modes of psychology. For example, a good communication mode prompt template can be provided by a psychologist, and the communication mode information table can be created based on the prompt template.
[0072] Step 162: convert the content of the communication mode into a communication mode guiding voice.
[0073] The content of the communication mode can be text information, and the content of the communication mode can be converted from text information into voice information to obtain a communication mode guiding voice.
[0074] Step 163: output the communication mode guiding voice.
[0075] For example, when the speaker identity information indicates that the speaker corresponding to the speaker identity information is a child, the target emotion identified can include at least one of collapse, crying, and temper, but is not limited thereto. Accordingly, the output communication mode guiding voice can include outputting the communication mode guiding voice in a soothing tone. For example, the communication mode guiding voice can be output in a tone that is easy for children to accept, or a tone that can adjust the atmosphere and output the communication mode guiding voice in a soothing and interesting tone. In this way, not only can the child be reminded to communicate with the parent in a good attitude in a tone that is easy for the child to accept, but the child's emotions can also be soothed, helping the child to calm down as soon as possible. When the child's emotions are not good, the child can be reminded and guided to control the emotions, prevent the parent-child relationship from deteriorating further, and improve the parent-child relationship.
[0076] For example, when the speaker identity information indicates that the speaker corresponding to the speaker identity information is a parent, the target emotion identified can include at least one of impatience, temper, and loud scolding, but is not limited thereto. Accordingly, the output communication mode guiding voice can include outputting the communication mode guiding voice in a tone of a child.
[0077] For example, when the speaker identity information indicates that the speaker corresponding to the speaker identity information is a parent, the corresponding communication mode guiding voice can be output according to the specific role of the parent identity. For example, when the parent identity identified is a mother, the title "Mom" can be added when the communication mode guiding voice is output in the tone of a child. For example, when the parent identity identified is a father, the title "Dad" can be added when the communication mode guiding voice is output in the tone of a child. In this way, the specific identity of the parent can be distinguished, the pertinence and interactivity are stronger, and the communication mode guiding is more conducive.
[0078] The parent-child communication guiding method provided by the example embodiment can collect a voice signal, perform acoustic feature recognition, voice recognition and volume detection on the voice signal, obtain speaker identity information and corresponding voice content and volume information, identify the emotion of a speaker corresponding to the speaker identity information based on the voice content and volume information, and output a communication mode guiding voice when the emotion is an undesirable emotion. The communication mode guiding voice can be used to remind and guide the speaker to communicate in a good communication mode, can remind and guide the speaker when the parent and the child just have undesirable emotions, help to improve undesirable emotions and a communication atmosphere, prevent the communication atmosphere from further deteriorating, realize intelligent guiding of a communication mode between parents and children, and provide help for maintaining a good parent-child relationship.
[0079] Figure 3 An example shows a second flowchart of the parent-child communication guiding method provided by the embodiment of the application. As shown in the second flowchart, the parent-child communication guiding method can include the following steps 310 to 370. Figure 3
[0080] Step 310: Obtain a voice signal collected by a voice collection device.
[0081] Step 320: Perform acoustic feature recognition on the voice signal, and determine speaker identity information based on the recognized acoustic features.
[0082] Step 330: Perform voice recognition on the voice signal, and obtain voice content corresponding to the speaker identity information.
[0083] Step 340: Perform volume detection on the voice signal, and obtain volume information corresponding to the speaker identity information.
[0084] Step 350: Perform speech rate detection on the voice signal, and obtain speech rate information corresponding to the speaker identity information.
[0085] The speech rate reflects the speed of a person speaking. For example, the number of phonemes in a set time period in the voice signal can be counted, and then divided by the set time period to obtain the speech rate information of the voice signal.
[0086] Step 360: Identify the emotion of a speaker corresponding to the speaker identity information based on the voice content, volume information and speech rate information, and obtain a target emotion.
[0087] In addition to the volume of the voice and the voice content of the person, the speaking speed of the person under different emotions is also different. For example, when the person is angry, scolds others or is impatient, the speaking speed is relatively fast, and when the emotion is stable, the speaking speed is relatively slow. Exemplarily, the emotion of the speaker can be determined by a speaking speed threshold. For example, when the speaking speed exceeds the speaking speed threshold, it indicates that the speaker has an adverse emotion, and when the speaking speed is less than the speaking speed threshold, it indicates that the speaker has a good emotion. Exemplarily, different speaking speed threshold ranges can be set for different emotions, and the specific emotion category can be determined by judging the speaking speed threshold range in which the detected speaking speed falls.
[0088] In an example embodiment, the volume information, the voice content and the speaking speed information of the voice can be combined to jointly determine the emotion of the speaker, and the accuracy of emotion recognition is further improved. The voice content, the volume information and the speaking speed information can be recognized based on a second pre-constructed voice emotion recognition model to obtain the emotion of the speaker. Exemplarily, an initial neural network model can be selected, and the initial neural network model can be trained by using a second training sample set to obtain the second voice emotion recognition model. The second training sample set can include voice sample data and second emotion type data corresponding to the voice sample data. The voice sample data can be subjected to voice recognition to obtain text data, from which text features can be extracted, and volume features and speaking speed features can be extracted from the voice sample data. The second emotion type data is the labeled data of the text features, the volume features and the speaking speed features. The extracted text features, the volume features and the speaking speed features are used as the input data of the initial neural network model, and the second emotion type data is used as the output data of the initial neural network model. The initial neural network model is trained to obtain the second voice emotion recognition model.
[0089] Step 370: outputting a communication mode guiding voice in response to the target emotion being an adverse emotion.
[0090] The parent-child communication guiding method provided by the example embodiment can improve the accuracy of emotion recognition by collecting voice signals, performing acoustic feature recognition, voice recognition, volume detection and speaking speed detection on the voice signals, obtaining speaker identity information and corresponding voice content, volume information and speaking speed information, and recognizing the emotion of the speaker corresponding to the speaker identity information based on the voice content, the volume information and the speaking speed information. When the emotion is an adverse emotion, a communication mode guiding voice is outputted, which can be used to remind and guide the speaker to communicate in a good communication mode. The parent-child communication guiding method can remind and guide the parent and the child when the adverse emotion just occurs, help to improve the adverse emotion and the communication atmosphere, prevent the further deterioration of the communication atmosphere, and realize the intelligent guiding of the communication mode between the parent and the child, thereby providing help for the maintenance of a good parent-child relationship.
[0091] The parent-child communication guidance device provided by the present invention is described below. The parent-child communication guidance device described below can be referred to in correspondence with the parent-child communication guidance method described above.
[0092] Figure 4 An exemplary schematic diagram of the parent-child communication guidance device provided in an embodiment of the present invention is shown, with reference to... Figure 4 As shown, the parent-child communication guidance device 400 may include an acquisition module 410, a first recognition module 420, a second recognition module 430, a volume detection module 440, a third recognition module 450, and an output module 460. Specifically: the acquisition module 410 is used to acquire the voice signal collected by the voice acquisition device; the first recognition module 420 is used to perform acoustic feature recognition on the voice signal and determine the speaker's identity information based on the recognized acoustic features; the second recognition module 430 is used to perform speech recognition on the voice signal to obtain the voice content corresponding to the speaker's identity information; the volume detection module 440 is used to perform volume detection on the voice signal to obtain the volume information corresponding to the speaker's identity information; the third recognition module 450 is used to identify the speaker's emotion corresponding to the speaker's identity information based on the voice content and volume information to obtain the target emotion; and the output module 460 is used to output communication guidance voice in response to the target emotion being a negative emotion.
[0093] In one example embodiment, the parent-child communication guidance device 400 may further include a speech rate detection module, which can be used to detect the speech rate of the speech signal to obtain the speech rate information corresponding to the speaker's identity information; the third recognition module 450 may be specifically used to recognize the speaker's emotion corresponding to the speaker's identity information based on the speech content, speech rate information and volume information to obtain the target emotion.
[0094] In one example embodiment, the output module 460 may include a determining unit, a conversion unit, and an output unit. Specifically: the determining unit can be used to determine the content of the communication method based on the speaker's identity information and the target emotion in response to the target emotion being negative; the conversion unit can be used to convert the content of the communication method into guiding speech; and the output unit can be used to output the guiding speech.
[0095] In one example embodiment, the determining unit can be specifically used to match the corresponding communication method from the communication method information table based on the speaker's identity information and the target emotion, and obtain the content of the communication method. The communication method information table stores the correspondence between the speaker's identity, emotion and communication method.
[0096] In one example embodiment, the communication style information sheet can be created based on psychological principles of good communication.
[0097] In an example embodiment, when the speaker identity information indicates that the speaker corresponding to the speaker identity information is a child, the target emotion can include at least one of a breakdown, crying, and a temper tantrum; and the output unit can be specifically configured to output the communication mode guiding voice in a soothing tone.
[0098] In an example embodiment, when the speaker identity information indicates that the speaker corresponding to the speaker identity information is a parent identity, the target emotion can include at least one of impatience, a temper tantrum, and a loud scolding; and the output unit can be specifically configured to output the communication mode guiding voice in a tone of a child.
[0099] Figure 5 An example of a schematic diagram of an entity structure of an electronic device is shown in Figure 5 As shown, the electronic device 500 can include a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 can complete mutual communication through the communication bus 540. The processor 510 can invoke a logical instruction in the memory 530 to execute the steps of the parent-child communication guiding method provided by the above-mentioned embodiments, for example, can include: acquiring a voice signal collected by a voice collection device; performing acoustic feature recognition on the voice signal, and determining speaker identity information based on the recognized acoustic feature; performing voice recognition on the voice signal to obtain voice content corresponding to the speaker identity information; performing volume detection on the voice signal to obtain volume information corresponding to the speaker identity information; identifying the emotion of the speaker corresponding to the speaker identity information based on the voice content and the volume information to obtain a target emotion; and in response to the target emotion being an undesirable emotion, outputting a communication mode guiding voice.
[0100] In addition, the logical instruction in the memory 530 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in the embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0101] In another aspect, the present application also provides a computer program product comprising a computer program, which can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer can execute the steps of the parent-child communication guiding method provided by the above embodiments, for example, can comprise: acquiring a voice signal collected by a voice collection device; performing acoustic feature recognition on the voice signal, determining speaker identity information based on the recognized acoustic feature; performing voice recognition on the voice signal to obtain voice content corresponding to the speaker identity information; performing volume detection on the voice signal to obtain volume information corresponding to the speaker identity information; identifying the emotion of the speaker corresponding to the speaker identity information based on the voice content and the volume information to obtain a target emotion; and outputting a communication mode guiding voice in response to the target emotion being an undesirable emotion.
[0102] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the parent-child communication guiding method provided by the above embodiments are implemented, for example, can comprise: acquiring a voice signal collected by a voice collection device; performing acoustic feature recognition on the voice signal, determining speaker identity information based on the recognized acoustic feature; performing voice recognition on the voice signal to obtain voice content corresponding to the speaker identity information; performing volume detection on the voice signal to obtain volume information corresponding to the speaker identity information; identifying the emotion of the speaker corresponding to the speaker identity information based on the voice content and the volume information to obtain a target emotion; and outputting a communication mode guiding voice in response to the target emotion being an undesirable emotion.
[0103] The device embodiments described above are only schematic, wherein the units illustrated as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement without creative labor.
[0104] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus necessary general hardware platforms, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in the various embodiments or some parts of the embodiments.
[0105] Finally, it should be noted that the above examples are only used to illustrate the technical solutions of the present application, and are not intended to limit the same; although the present application has been described in detail with reference to the foregoing examples, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A parent-child communication guiding method, characterized by, The method comprises the following steps: During the process of accompanying reading by parents, a voice signal collected by a voice collecting device is acquired; Speech rate detection is performed on the voice signal to obtain speech rate information corresponding to the speaker identity information; Acoustic feature recognition is performed on the voice signal, and the speaker identity information is determined based on the recognized acoustic feature; Speech recognition is performed on the voice signal to obtain voice content corresponding to the speaker identity information; Volume detection is performed on the voice signal to obtain volume information corresponding to the speaker identity information; The emotion of the speaker corresponding to the speaker identity information is recognized based on the voice content and the volume information, and a target emotion is obtained; The emotion of the speaker corresponding to the speaker identity information is recognized based on the voice content, the speech rate information and the volume information, and a target emotion is obtained; The emotion of the speaker corresponding to the speaker identity information is recognized based on the voice content, the speech rate information and the volume information, and a target emotion is obtained, which comprises: The emotion of the speaker corresponding to the speaker identity information is recognized based on the voice content, the speech rate information and the volume information, and a target emotion is obtained; The emotion of the speaker corresponding to the speaker identity information is recognized based on the voice content, the speech rate information and the volume information, and a target emotion is obtained, which comprises: The voice content is subjected to word segmentation processing, and a target emotion is recognized by a pre-established bad emotion word segmentation library from the word segmentation representing the emotion in the voice content; The volume information and the voice content are combined, and a target emotion is obtained based on the volume information; The voice content and the volume information are recognized based on a pre-constructed first voice emotion recognition model, and a target emotion is obtained; 2. The parent-child communication guiding method according to claim 1, characterized by, In response to the target emotion being a bad emotion, a communication mode guiding voice is outputted, and the communication mode guiding voice comprises a communication mode guiding voice outputted in a soothing tone and a communication mode guiding voice outputted in a child's tone. The response to the target emotion being a bad emotion and the output of the communication mode guiding voice comprises: In response to the target emotion being a bad emotion, the content of the communication mode is determined according to the speaker identity information and the target emotion; The content of the communication mode is converted into a communication mode guiding voice; 3. The parent-child communication guiding method according to claim 2, characterized by, The communication mode guiding voice is outputted. The determination of the content of the communication mode according to the speaker identity information and the target emotion comprises:
4. The parent-child communication guiding method according to claim 2, characterized by, The content of the communication mode is obtained by matching the corresponding communication mode from a communication mode information table according to the speaker identity information and the target emotion, and the communication mode information table stores the corresponding relationship among the speaker identity, the emotion and the communication mode. In response to the speaker identity information indicating that the speaker corresponding to the speaker identity information is a child, the target emotion comprises at least one of collapse, crying and temper tantrums; The output of the communication mode guiding voice comprises:
5. The parent-child communication guiding method according to claim 2, characterized by, The communication mode guiding voice is outputted in a soothing tone. In response to the speaker identity information indicating that the speaker corresponding to the speaker identity information is a parent, the target emotion comprises at least one of impatience, temper tantrums and loud scolding; The output of the communication mode guiding voice comprises:
6. A parent-child communication guiding device, characterized by, The communication mode guiding voice is outputted in a child's tone. The method comprises the following steps: An acquisition module is configured to acquire a voice signal collected by a voice collecting device during the process of accompanying reading by parents; The speech signal is subjected to speech rate detection to obtain speech rate information corresponding to the speaker identity information; The first recognition module is configured to recognize acoustic features of the speech signal, and determine the speaker identity information based on the recognized acoustic features; The second recognition module is configured to recognize the speech signal to obtain speech content corresponding to the speaker identity information; The volume detection module is configured to detect the volume of the speech signal to obtain volume information corresponding to the speaker identity information; The third recognition module is configured to recognize the emotion of the speaker corresponding to the speaker identity information based on the speech content and the volume information to obtain a target emotion; The third recognition module is configured to recognize the emotion of the speaker corresponding to the speaker identity information based on the speech content and the volume information to obtain a target emotion; The output module is configured to output a communication mode guiding voice in response to the target emotion being an undesirable emotion, wherein the communication mode guiding voice includes a communication mode guiding voice output in a soothing tone and a communication mode guiding voice output in a child's tone.
7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the parent-child communication guiding method of any one of claims 1-5.
8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the parent-child communication guiding method of any one of claims 1-5.
9. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the parent-child communication guiding method of any one of claims 1-5.
Citation Information
Patent Citations
Family emotion management device and method
CN105280187A
Emotion recognition processing method and device, medium and electronic device
CN112102850A
Interaction method, device and equipment of anthropomorphic robot and storage medium
CN113580166A