Doll animation expression display method and device, electronic equipment and readable medium

By collecting and processing environmental audio and image data in the doll, accurate user mood information is generated, which solves the problem of voice distortion caused by background sound interference and improves the matching of the doll's animated expressions and user experience.

CN120823291APending Publication Date: 2025-10-21SHANGHAI SHIKONG TECHNOLOGY DEVELOPMENT CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510894280.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

In the existing technology, the method for displaying doll animation expressions does not remove the background sound of the collected audio, resulting in distortion of the voice signal, low accuracy in predicting user moods, inconsistency between the animation expressions and the user's actual mood, and poor user experience.

Method used

The microphone collects the ambient audio frame sequence, performs voice detection and background sound removal processing, and combines it with the doll's front-view scene image to generate user mood information. Based on the user's mood and the doll's personality data, animated expressions are generated and displayed on the doll's display screen.

Benefits of technology

It improves the accuracy and interactivity of user mood information, enhances the matching degree between the doll animation expression and the user mood, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823291A_ABST
    Figure CN120823291A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a doll animation expression display method and device, electronic equipment and a readable medium. A specific embodiment of the method comprises the following steps: acquiring updated doll character data; collecting an environment audio frame sequence through a microphone arranged in the doll; performing voice detection processing on the environment audio frame sequence; collecting a voice audio frame sequence within a preset time period after the timestamp through a microphone, and collecting a foresight scene image of the doll through a camera device built in the doll; carrying out background sound removal processing on the voice audio frame sequence; generating user mood information based on the target voice audio frame sequence and the foresight scene image; generating doll animation expression information based on the user mood information and the updated doll character data; acquiring excitation feedback information associated with the expression information of the doll animation; and displaying the animation expression corresponding to the animation expression information of the doll and the excitation feedback information in a display screen of the doll. According to the embodiment, the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method, device, electronic device, and readable medium for displaying animated expressions of dolls. Background Art

[0002] Doll animated expression display is a technology that displays animated expressions on the doll's display. Currently, the common method for displaying animated expressions is to directly predict the user's current mood based on the collected audio and display an animated expression (such as a slightly narrowed smile) that suits the user's current mood (for example, happiness).

[0003] However, when using the above method to display animated expressions on the doll's display screen, the following technical problems often occur:

[0004] The current user's mood is directly predicted based on the collected audio, and animated expressions that are appropriate to the current user's mood are displayed. The collected audio is not processed for background sound removal. Background sound (such as TV sound, traffic noise, other people's conversations, etc.) will mix with the user's voice, causing voice signal distortion, which in turn leads to low accuracy in predicting the current user's mood. The animated expressions displayed by the doll may not match the user's actual mood, resulting in a poor user experience.

[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the inventive concept and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention

[0006] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0007] Some embodiments of the present disclosure provide a method, device, electronic device, and computer-readable medium for displaying animated expressions of a doll to solve one or more of the technical problems mentioned in the above background technology section.

[0008] In a first aspect, some embodiments of the present disclosure provide a method for displaying animated expressions of a doll, the method comprising: acquiring and updating the personality data of the doll; collecting an ambient audio frame sequence through a microphone built into the doll; performing voice detection processing on the collected ambient audio frame sequence to determine the timestamp corresponding to the audio frame of the voice starting point; collecting a voice audio frame sequence within a preset time period after the above timestamp through the microphone, and simultaneously collecting a front-view scene image of the doll through a camera device built into the doll; performing background sound removal processing on the above voice audio frame sequence to obtain a target voice audio frame sequence; generating user mood information based on the above target voice audio frame sequence and the above front-view scene image; generating doll animated expression information based on the above user mood information and the above updated doll personality data; acquiring incentive feedback information associated with the above doll animated expression information; and displaying the animated expression corresponding to the doll animated expression information and the above incentive feedback information on the display screen of the doll.

[0009] In a second aspect, some embodiments of the present disclosure provide a doll animation expression display device, the device comprising: a first acquisition unit, configured to acquire and update doll personality data; a first acquisition unit, configured to acquire an environmental audio frame sequence through a microphone built into the doll; a voice detection processing unit, configured to perform voice detection processing on the acquired environmental audio frame sequence to determine the timestamp corresponding to the voice starting point audio frame; a second acquisition unit, configured to acquire a voice audio frame sequence within a preset time period after the above timestamp through the microphone, and at the same time acquire a front-view scene image of the doll through a camera device built into the doll; background sound The removal processing unit is configured to perform background sound removal processing on the above-mentioned voice audio frame sequence to obtain a target voice audio frame sequence; the first generation unit is configured to generate user mood information based on the above-mentioned target voice audio frame sequence and the above-mentioned forward-looking scene image; the second generation unit is configured to generate doll animation expression information based on the above-mentioned user mood information and the above-mentioned updated doll personality data; the second acquisition unit is configured to acquire incentive feedback information associated with the above-mentioned doll animation expression information; the display unit is configured to display the animated expression corresponding to the doll animation expression information and the above-mentioned incentive feedback information on the display screen of the above-mentioned doll.

[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation of the first aspect is implemented.

[0012] The above-mentioned embodiments of the present disclosure have the following beneficial effects: The method for displaying animated expressions on dolls in some embodiments of the present disclosure improves the user experience. Specifically, the poor user experience is caused by directly predicting the current user mood based on collected audio and displaying animated expressions that match the current user mood without performing background noise removal on the collected audio. Background sounds (such as television sound, traffic noise, and other people's conversations) can mix with the user's voice, causing distortion of the voice signal. This in turn leads to low accuracy in predicting the current user's mood. The animated expressions displayed by the doll may not match the user's actual mood, resulting in a poor user experience. Based on this, the method for displaying animated expressions on dolls in some embodiments of the present disclosure first obtains updated doll personality data. This allows the updated doll personality data used to generate doll animated expression information to be obtained. Then, a microphone built into the doll collects a sequence of ambient audio frames. Voice detection is performed on the collected ambient audio frame sequence to determine the timestamp corresponding to the audio frame at the start of the voice. This allows the timestamp of the voice start to be determined. Next, a microphone is used to capture a sequence of speech and audio frames within a preset time period after the timestamp, while a camera built into the doll captures an image of the doll's forward-looking scene. Consequently, after determining the speech start time, the speech and audio frames within the preset time period after the timestamp and the doll's forward-looking scene image are simultaneously captured. Next, background sound removal is performed on the speech and audio frames to obtain a target speech and audio frame sequence. This removes the background sound contained in the speech and audio frame sequence. Next, user mood information is generated based on the target speech and audio frame sequence and the forward-looking scene image. This can be combined with the background-removed target speech and audio frame sequence and the doll's forward-looking scene image to generate user mood information. The doll's forward-looking scene image may capture the user's facial expressions, and combining the target speech and audio frame sequence with the doll's forward-looking scene image to generate user mood information may further improve the accuracy of the user mood information. Next, based on the user mood information and the updated doll personality data, animated expression information for the doll is generated. Thus, the doll's personality, i.e., updated doll personality data, and the user's current mood, i.e., user mood information, can be combined to generate doll animated expression information representing the doll's animated expressions. Next, motivational feedback information associated with the doll animated expression information is obtained. Finally, the animated expression corresponding to the doll animated expression information and the motivational feedback information are displayed on the doll's display screen. Thus, the motivational feedback information can be displayed simultaneously with the animated expression, enhancing the doll's interactivity. Furthermore, because the background sound contained in the speech audio frame sequence is removed before generating the user mood information, the interference of the background sound on the speech signal is reduced. Generating the user mood information based on the target speech audio frame sequence after background sound removal improves the accuracy of the user mood information, ensuring that the animated expression displayed on the doll's display screen is more consistent with the user's actual mood, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0014] Figure 1 is a flow chart of some embodiments of the method for displaying doll animation expressions according to the present disclosure;

[0015] Figure 2 1 is a schematic structural diagram of some embodiments of the doll animation expression display device according to the present disclosure;

[0016] Figure 3 It is a structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0018] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0023] Figure 1 The process 100 of some embodiments of the method for displaying doll animation expressions according to the present disclosure is shown. The method for displaying doll animation expressions comprises the following steps:

[0024] Step 101: Get and update the character data of the doll.

[0025] In some embodiments, the execution body of the doll animation expression display method (eg, a microcontroller inside the doll) can obtain and update the doll personality data.

[0026] In some optional implementations of some embodiments, the execution entity may obtain and update the character data of the doll through the following steps:

[0027] The user information and acquisition request information input by the user through the above-mentioned display screen are sent to the doll personality configuration server to obtain the updated doll personality data corresponding to the above-mentioned acquisition request information from the above-mentioned doll personality configuration server. Among them, the above-mentioned user information includes a user name and password. The above-mentioned acquisition request information is the demand or instruction for obtaining the doll personality input by the user through the display screen, and the above-mentioned acquisition request information may include the updated doll personality type (for example, lively, cute, etc.). The above-mentioned updated doll personality data includes each doll animation expression data under each mood corresponding to the doll personality type included in the acquisition request information. Each doll animation expression data in each doll animation expression data corresponds to emotion category information (for example, happy, angry, sad, etc.), and can represent the doll's animated expression (for example, animated expression in GIF format). The above-mentioned user mood information includes emotion category information, which can represent the user's mood (for example, happy, angry, sad, etc.). As an example, the updated doll personality data can include animated expression data corresponding to a lively doll personality type. For example, animated expression data corresponding to a happy emotion type can include a smiling expression with wide eyes and a mouth that is raised to the ears. For animated expression data corresponding to a sad emotion type, it can include a frowning expression with a drooping mouth. The doll personality configuration server can be a server that stores updated doll personality data. The doll personality configuration server can verify user information and, if verification is successful, authorize the doll's internal microcontroller to access the updated doll personality data.

[0028] Step 102: Collect an ambient audio frame sequence through a microphone built into the doll.

[0029] In some embodiments, the execution entity may collect the ambient audio frame sequence using a microphone built into the doll. The microphone may be integrated within the doll and used to detect ambient sounds. The ambient audio frame sequence may be a sequence of audio frames collected in real time by the microphone over a period of time.

[0030] Step 103: Perform voice detection processing on the collected environmental audio frame sequence to determine the timestamp corresponding to the voice starting point audio frame.

[0031] In some embodiments, the execution entity may perform speech detection processing on the collected ambient audio frame sequence to determine a timestamp corresponding to the speech starting point audio frame.

[0032] In some optional implementations of some embodiments, the execution entity may perform speech detection processing on the collected ambient audio frame sequence to determine the timestamp corresponding to the speech starting point audio frame through the following steps:

[0033] In the first step, for the first ambient audio frame in the above ambient audio frame sequence, the following speech detection process is performed:

[0034] The first sub-step is to perform speech detection processing on the ambient audio frame to obtain detection information.

[0035] The second sub-step is, in response to determining that the detection information indicates that the ambient audio frame contains a speech signal, determining the ambient audio frame as a speech starting point audio frame, and determining the timestamp corresponding to the above ambient audio frame as the timestamp corresponding to the speech starting point audio frame.

[0036] The third sub-step is, in response to determining that the detection information indicates that the ambient audio frame does not contain a speech signal, performing the following updating steps:

[0037] Sub-step 1: Deleting the first ambient audio frame from the ambient audio frame sequence to update the ambient audio frame sequence.

[0038] Based on the first ambient audio frame in the updated ambient audio frame sequence, the above-mentioned voice detection process is performed again.

[0039] In some optional implementations of some embodiments, the execution entity may perform speech detection processing on the ambient audio frame to obtain detection information through the following steps:

[0040] The first step is to perform feature extraction processing on the ambient audio frame to obtain ambient audio feature information. In practice, the execution entity may use Mel-Frequency Cepstral Coefficient (MFCC) technology to perform feature extraction processing on the ambient audio frame to obtain the ambient audio feature information. The ambient audio feature information may be an MFCC feature vector representing the audio acoustic characteristics of the ambient audio frame.

[0041] In the second step, the ambient audio feature information corresponding to the above-mentioned ambient audio frame is input into a pre-trained wake-up word detection model to obtain a wake-up word detection label. The wake-up word detection model can be a neural network model (for example, a deep neural network (DNN) model, a hidden Markov model (HMM)) for judging whether the ambient audio frame corresponding to the input ambient audio feature information contains a specific wake-up word. The wake-up word detection label can include a positive label or a negative label, wherein a positive label can indicate that the ambient audio frame contains a wake-up word. A negative label can indicate that the ambient audio frame does not contain a wake-up word.

[0042] In the third step, in response to determining that the above-mentioned wake-up word detection tag represents that the environmental audio frame does not contain the wake-up word, the information representing that the environmental audio frame does not contain the voice signal is determined as the detection information.

[0043] In the fourth step, in response to determining that the wake-up word detection tag represents that the ambient audio frame contains the wake-up word, information representing that the ambient audio frame contains a voice signal is determined as detection information.

[0044] Optionally, the wake-up word detection model can be trained by the following steps:

[0045] Obtain a sample set, wherein the samples in the sample set include sample environment audio feature information and a sample target wake-up word detection label corresponding to the sample environment audio feature information.

[0046] In the first step, the following training steps are performed based on the sample set:

[0047] In the first sub-step, the sample environment audio feature information of at least one sample in the sample set is input into the initial neural network to obtain a sample predicted wake-up word detection label corresponding to each sample in the at least one sample.

[0048] In a second sub-step, the sample predicted wake-up word detection label corresponding to each of the at least one sample is compared with the corresponding sample target wake-up word detection label. In practice, the execution subject may use a cross-entropy loss function to compare and determine the difference between the sample predicted wake-up word detection label corresponding to each of the at least one sample and the corresponding sample target wake-up word detection label.

[0049] The third sub-step is to determine whether the initial neural network has achieved a preset optimization goal based on the comparison result. The optimization goal may be that the cross entropy loss function value is lower than a preset threshold.

[0050] In a fourth sub-step, in response to determining that the initial neural network achieves the above-mentioned optimization goal, the initial neural network is used as a trained wake-up word detection model;

[0051] In a second step, in response to determining that the initial neural network does not achieve the above-mentioned optimization goal, the network parameters of the initial neural network are adjusted, and a sample set is formed using unused samples. The adjusted initial neural network is used as the initial neural network, and the above-mentioned training steps are performed again. As an example, the network parameters of the above-mentioned initial neural network can be adjusted using a back propagation algorithm (BP algorithm) and a gradient descent method (such as a mini-batch gradient descent algorithm).

[0052] In some optional implementations of some embodiments, the execution entity may further perform speech detection processing on the ambient audio frame to obtain detection information through the following steps:

[0053] The first step is to perform voiceprint feature extraction on the above-mentioned environmental audio frame to obtain voiceprint feature information corresponding to the above-mentioned environmental audio frame. In practice, the above-mentioned execution entity can use voice activity detection (VAD) or a deep learning model (such as CRNN) to separate potential speech segments from the environmental audio frame. Then, the above-mentioned execution entity can use perceptual linear prediction (PLP) technology to perform feature extraction on the above-mentioned speech segment to obtain speech segment PLP feature information corresponding to the above-mentioned speech segment. The above-mentioned speech segment PLP feature information can represent the PLP feature of the speech segment. Finally, the speech segment PLP feature information is input into a pre-trained voiceprint model (for example, an X-Vector model) to obtain voiceprint feature information (i.e., a voiceprint embedding vector).

[0054] The second step is to obtain a pre-stored user audio feature information set. Each user audio feature information in the user audio feature information set may represent a user's voiceprint feature. The user audio feature information set may represent the voiceprint features of each user (e.g., each member of a family). The user audio feature information set may also represent the voiceprint features of the same user in different states (e.g., sadness, excitement, etc.).

[0055] In the third step, for each user audio feature information in the user audio feature information set, the similarity between the user audio feature information and the voiceprint feature information is determined as the audio feature similarity. The similarity between the user audio feature information and the voiceprint feature information can be represented by cosine similarity.

[0056] In the fourth step, the determined audio feature similarities are determined as an audio feature similarity set.

[0057] In a fifth step, in response to determining that the audio feature similarity set contains an audio feature similarity that satisfies a preset condition, information indicating that the ambient audio frame contains a speech signal is determined as detection information. The preset condition may be that the audio feature similarity with the greatest similarity in the audio feature similarity set is greater than a preset similarity.

[0058] In step 6, in response to determining that there is no audio feature similarity that meets the preset condition in the audio feature similarity set, information representing that the ambient audio frame does not contain a speech signal is determined as detection information.

[0059] Step 104 : The microphone collects a sequence of voice audio frames within a preset time period after the timestamp, and the camera device built into the doll collects a front-view scene image of the doll.

[0060] In some embodiments, the execution entity may use a microphone to capture a sequence of voice and audio frames within a preset time period after the timestamp, and simultaneously use a camera built into the doll to capture an image of the doll's forward-facing scene. The preset time period may begin after the timestamp and last for a preset duration. The camera may be a camera. The camera may be mounted at a fixed position on the doll's head to capture images of the doll's forward-facing scene.

[0061] Step 105: Perform background sound removal processing on the speech audio frame sequence to obtain a target speech audio frame sequence.

[0062] In some embodiments, the execution entity may perform background sound removal processing on the speech audio frame sequence to obtain a target speech audio frame sequence. In practice, a Wiener filter technique may be used to perform background sound removal processing on the speech audio frame sequence to obtain a target speech audio frame sequence. Each speech audio frame in the speech audio frame sequence corresponds to a corresponding target speech audio frame in the target speech audio frame sequence. The target speech audio frame may be an audio frame obtained by removing background sound from the speech audio frame.

[0063] In some optional implementations of some embodiments, the execution entity may perform background sound removal processing on the speech audio frame sequence to obtain a target speech audio frame sequence through the following steps:

[0064] In the first step, a preset number of environmental audio frames in the environmental audio frame sequence before the voice starting point audio frame are determined as environmental noise audio frames to obtain an environmental noise audio frame sequence.

[0065] In the second step, for each ambient noise audio frame in the ambient noise audio frame sequence, the following processing is performed.

[0066] The first sub-step is to convert the ambient noise audio frame to obtain frequency-domain complex spectrum information. In practice, the execution entity may employ discrete Fourier transform (DFT) technology to convert the ambient noise audio frame to obtain frequency-domain complex spectrum information. The frequency-domain complex spectrum information may represent the complex spectrum of the ambient noise audio frame.

[0067] The second sub-step is to perform a modulo operation on the complex spectrum corresponding to the frequency domain complex spectrum information to obtain amplitude spectrum information, wherein the amplitude spectrum information may represent the amplitude spectrum of the ambient noise audio frame.

[0068] The third step is to generate an amplitude spectrum matrix based on the obtained amplitude spectrum information. In practice, the amplitude spectrum information of each ambient noise audio frame can be used as a horizontal column vector, and the horizontal column vectors corresponding to the ambient noise audio frame sequence can be arranged in time from top to bottom (that is, according to the order of the ambient noise audio frames corresponding to the amplitude spectrum information in the ambient noise audio frame sequence) to obtain the amplitude spectrum matrix. As an example, the above-mentioned ambient noise audio frame sequence can be {ambient noise audio frame 1, ambient noise audio frame 2, ambient noise audio frame 3}. The amplitude spectrum information corresponding to the above-mentioned ambient noise audio frame 1 can be [0.583, 0.583, 0.580]. The amplitude spectrum information corresponding to the above-mentioned ambient noise audio frame 2 can be [0.586, 0.583, 0.583]. The amplitude spectrum information corresponding to the above-mentioned ambient noise audio frame 3 can be [0.583, 0.586, 0.583]. Then the amplitude spectrum matrix can be:

[0069] The fourth step is to generate noise mean information based on the above-mentioned amplitude spectrum matrix. In practice, the above-mentioned execution entity can calculate the mean of each column of the amplitude spectrum matrix (it should be noted that if it cannot be divided evenly, the mean can be rounded to three decimal places) to obtain a mean vector as the noise mean information. As an example, the mean of the first column can be (0.583+0.583+0.580)÷3=0.582. The mean of the second column can be (0.586+0.583+0.583)÷3=0.584. The mean of the third column can be (0.583+0.586+0.583)÷3=0.584. Then the noise mean information can be {0.582, 0.584, 0.584}.

[0070] Step 5: For each speech audio frame in the speech audio frame sequence, perform the following processing:

[0071] Sub-step 1: Convert the speech audio frame to obtain target frequency domain complex spectrum information. In practice, the execution entity may use discrete Fourier transform (DFT) technology to convert the speech audio frame to obtain target frequency domain complex spectrum information. The target frequency domain complex spectrum information may represent the complex spectrum of the speech audio frame.

[0072] Sub-step 2: modulo the complex spectrum corresponding to the target frequency domain complex spectrum information to obtain target amplitude spectrum information.

[0073] Sub-step three, based on the above-mentioned target amplitude spectrum information and the above-mentioned noise mean information, generate the background sound amplitude spectrum information. In practice, the above-mentioned execution entity can determine the vector difference (such as "element-by-element subtraction") between the target amplitude spectrum information and the noise mean information as the background sound amplitude spectrum information. As an example, the above-mentioned target amplitude spectrum information can be {1.102, 1.500, 1.400}, and the above-mentioned noise mean information can be {0.582, 0.584, 0.584}, then {1.102, 1.500, 1.400}-{0.582, 0.584, 0.584}={0.520, 0.916, 0.816}, then the background sound amplitude spectrum information can be {0.520, 0.916, 0.816}.

[0074] In the sixth step, the obtained background sound amplitude spectrum information is sorted according to the order of the corresponding target speech audio frames in the above target speech audio frame sequence to obtain a background sound amplitude spectrum information sequence.

[0075] The seventh step is to perform audio signal reconstruction processing on each background-removed sound amplitude spectrum information in the above background-removed sound amplitude spectrum information sequence to obtain a target speech audio frame sequence. In practice, the Griffin-Lim algorithm can be used to perform audio signal reconstruction processing on each background-removed sound amplitude spectrum information in the background-removed sound amplitude spectrum information sequence to obtain a target speech audio frame sequence. Optionally, in practice, each background-removed sound amplitude spectrum information in the above background-removed sound amplitude spectrum information sequence can be input into the WaveGlow model to obtain a target speech audio frame sequence. Among them, each background-removed sound amplitude spectrum information in the above background-removed sound amplitude spectrum information sequence corresponds to a corresponding target speech audio frame in the above target speech audio frame sequence. The above target speech audio frame can be an audio frame constructed based on the amplitude spectrum corresponding to the background-removed sound amplitude spectrum information.

[0076] Step 106 : Generate user mood information based on the target speech audio frame sequence and the front-view scene image.

[0077] In some embodiments, the execution entity may generate user mood information based on the target speech audio frame sequence and the forward-looking scene image.

[0078] In the process of adopting technical solutions to solve the problems mentioned in the background technology, the following problems often arise:

[0079] Because the puppet's front-view scene image is captured after voice detection, it may not contain the user's face. When predicting the user's mood information based on the front-view scene image without a face and the target voice audio frame sequence, the system attempts to extract features from the front-view scene image without a face that have no actual correspondence with mood, mistakenly associating them with certain mood states. This results in poor accuracy in the generated user mood information. Furthermore, extracting features from the front-view scene image without a face that have no actual correspondence with mood wastes computer computing resources.

[0080] Faced with the above technical problems, the inventors decided to adopt the following solutions:

[0081] In some optional implementations of some embodiments, the execution entity may generate user mood information based on the target speech audio frame sequence and the front view scene image through the following steps:

[0082] In the first step, for each target speech audio frame in the target speech audio frame sequence, the following steps are performed:

[0083] In the first sub-step, audio signal feature extraction processing is performed on the target speech audio frame to obtain speech audio feature information corresponding to the target speech audio frame. In practice, the Librosa library can be called to perform audio signal feature extraction processing on the target speech audio frame using Mel-Frequency Cepstral Coefficient (MFCC) technology to obtain speech audio feature information corresponding to the target speech audio frame. The speech audio feature information can be the MFCC feature vector of the target speech audio frame.

[0084] In the second sub-step, the speech audio feature information is input into a pre-trained acoustic model to obtain predicted text information. The acoustic model may be a DNN-HMM model. The predicted text information may be the semantic text represented by the target speech audio frame.

[0085] The second step is to splice the obtained prediction text information to obtain the speech prediction text information corresponding to the above-mentioned target speech audio frame sequence. In practice, the above-mentioned prediction text information can be arranged and spliced ​​according to the order of the corresponding target speech audio frames in the target speech audio frame sequence to obtain the speech prediction text information. As an example, the above-mentioned target speech audio frame sequence can be "{target speech audio frame 1, target speech audio frame 2, target speech audio frame 3}". The prediction text information corresponding to the target speech audio frame 1 can be "I am today", the prediction text information corresponding to the target speech audio frame 2 can be "very", and the prediction text information corresponding to the target speech audio frame 3 can be "happy", then the speech prediction text information can be "I am very happy today".

[0086] The third step is to input the above-mentioned speech prediction text information into a pre-trained mood prediction model to obtain mood prediction information, wherein the above-mentioned mood prediction information includes various mood prediction sub-information, and each mood prediction sub-information in the above-mentioned various mood prediction sub-information includes emotion category information and category score. The above-mentioned mood prediction model can be a BERT model or a text sentiment analysis and emotion recognition model that takes speech prediction text information as input and mood prediction information as output. As an example, the speech prediction text information can be "I am very happy today." The various mood prediction sub-information can be ""emotion category information: happy, category score: 10 points", "emotion category information: angry, category score: 0 points", "emotion category information: sad, category score: 0 points".

[0087] The fourth step is to perform preliminary facial region positioning processing on the front-view scene image to obtain facial positioning information. In practice, a Haar cascade classifier can be used to perform preliminary facial region positioning processing on the front-view scene image to obtain facial positioning information. The facial positioning information includes one of the following: empty information (e.g., NULL indicates that no facial image is located) or at least one positioning region information. Each positioning region information in the at least one positioning region information represents the regional position of the face in the front-view scene image, which can be represented by various coordinate points (i.e., various facial edge coordinate points).

[0088] In step 6, in response to determining that the face location information indicates that a face image has not been located, the mood prediction sub-information that satisfies a preset screening condition among the mood prediction sub-information is determined as the target mood prediction sub-information. The preset screening condition may be that the category score included is the largest among the category scores included in the mood prediction sub-information.

[0089] In the seventh step, the emotion category information included in the target mood prediction sub-information is determined as the user mood information.

[0090] In step 8, in response to determining that the face positioning information includes at least one positioning area information, user mood information is generated based on the at least one positioning area information, the forward-viewing scene image and the mood prediction information.

[0091] The above technical solution and its related contents, as an inventive point of an embodiment of the present disclosure, solve the technical problem of "poor accuracy of generated user mood information and waste of computer computing resources." The factors that lead to poor accuracy of generated user mood information and waste of computer computing resources are often as follows: Since the front view scene image of the doll is collected after the voice is detected, the front view scene image of the doll may not contain the user's face image. Based on the front view scene image without the face and the target voice audio frame sequence, when predicting the user's mood information, the system will attempt to extract features that have no actual correspondence with the mood from the front view scene image without the face image, mistakenly associating them with certain mood states, thereby resulting in poor accuracy of the generated user mood information. At the same time, when extracting features that have no actual correspondence with the mood from the front view scene image without the face image, it results in a waste of computer computing resources. If the above factors are resolved, the accuracy of the generated user mood information can be improved and the waste of computer computing resources can be reduced. To achieve this effect, first, for each target speech audio frame in the target speech audio frame sequence, the following steps are performed: First, audio signal feature extraction is performed on the target speech audio frame to obtain speech audio feature information corresponding to the target speech audio frame. Second, the speech audio feature information is input into a pre-trained acoustic model to obtain predicted text information. Then, the obtained individual predicted text information is concatenated to obtain speech prediction text information corresponding to the target speech audio frame sequence. This allows the speech represented by the target speech audio frame sequence to be converted into corresponding text information, namely, speech prediction text information. Then, the speech prediction text information is input into a pre-trained mood prediction model to obtain mood prediction information, wherein the mood prediction information includes various mood prediction sub-information, each of which includes emotion category information and a category score. Thus, mood prediction information can be generated using the mood prediction model based on the speech prediction text information. Next, initial facial region localization is performed on the front-view scene image to obtain face location information. Thus, a preliminary facial region location process can be performed on the forward-view scene image to detect whether the forward-view scene image contains a facial image. Next, in response to determining that the facial location information indicates that no facial image has been located, the mood prediction sub-information that satisfies a preset screening condition among the various mood prediction sub-information is determined as target mood prediction sub-information. The emotion category information included in the target mood prediction sub-information is then determined as the user's mood information.Thus, through the above steps, when the forward-view scene image does not include a face image, user mood information can be generated based solely on the mood prediction sub-information generated based on the target speech audio frame sequence. This reduces processing of the forward-view scene image (e.g., extracting features that have no actual correspondence with mood), reduces interference with the generation of user mood information caused by the forward-view scene image that does not include the user's face image (e.g., performing mood recognition on an image without a face may cause the system to misjudge patterns or objects in the background as expressions, thereby providing an erroneous mood prediction result), improves the accuracy of the generated user mood information, and reduces waste of computer computing resources. In response to determining that the above-mentioned face positioning information includes at least one positioning area information, user mood information is generated based on the at least one positioning area information, the above-mentioned forward-view scene image, and the above-mentioned mood prediction information. Thus, user mood information can be generated when the forward-view scene image includes a face image. Also, before generating the user mood information, the face area of ​​the front-view scene image is initially located and processed, and when the front-view scene image does not include a face image, the user mood information is generated only based on the various mood prediction sub-information. This reduces the processing of the front-view scene image (for example, extracting features that have no actual correspondence with the mood), reduces the interference of the front-view scene image that does not include the user's face image on the generation of the user mood information (for example, performing mood recognition on an image without a face may cause the system to misjudge patterns or objects in the background as expressions, thereby giving an erroneous mood prediction result), improves the accuracy of the generated user mood information, and reduces the waste of computer computing resources.

[0092] In the process of adopting technical solutions to solve the problems mentioned in the background technology, the following problems often arise:

[0093] Since the front-view scene image containing facial images may contain multiple people (such as at parties and public places), the collected front-view scene image not only contains the user's facial image, but also includes facial images of non-users other than the user. If emotion recognition is performed directly based on the front-view scene image and combined with various mood prediction sub-information to generate user mood information, the non-user's expression (such as the smile of a bystander) may be incorrectly associated with the user's mood prediction (for example, the user says "I am angry", but another person in the image is smiling, but the system mistakenly judges the user's emotion as "happy"), resulting in poor accuracy of the generated user mood information. As a result, the animated expressions displayed based on the user mood information may not match the user's actual mood, resulting in a poor user experience.

[0094] Faced with the above technical problems, the inventors decided to adopt the following solutions:

[0095] In some optional implementations of some embodiments, the execution entity may generate user mood information based on the at least one positioning area information, the forward-viewing scene image, and the mood prediction information through the following steps:

[0096] The first step is to determine at least one area image corresponding to the at least one positioning area information in the front view scene image as at least one candidate face image;

[0097] The second step is to obtain the pre-stored user face image. In practice, the execution subject can obtain the pre-stored user face image from the SD card in the doll.

[0098] In the third step, for each candidate facial image in the at least one candidate facial image, the candidate facial image is compared with the user facial image to obtain a comparison similarity. In practice, the execution entity may input the user facial image into a pre-trained facial feature extraction model to obtain a high-dimensional feature vector corresponding to the user facial image as the user facial image feature information. For each candidate facial image in the at least one candidate facial image, the execution entity may input the candidate facial image into a pre-trained facial feature extraction model (such as a ResNet model, a VGGFace model, etc.) to obtain a high-dimensional feature vector corresponding to the facial image as the candidate facial image feature information. Afterwards, the execution entity may determine the similarity between the user facial image feature information and the candidate facial image feature information as the comparison similarity. The comparison similarity may be represented by cosine similarity.

[0099] In the third step, the alignment similarity with the greatest similarity among the obtained alignment similarities is determined as the target alignment similarity.

[0100] In the fourth step, in response to determining that the target comparison similarity is greater than a preset similarity, the candidate facial image corresponding to the target comparison similarity in the at least one candidate facial image is determined as the facial image to be recognized.

[0101] In the fifth step, the facial image to be recognized is input into the input layer of a pre-trained facial emotion recognition model to obtain facial image information. The facial emotion recognition model includes the input layer, a facial feature extraction layer, and a facial emotion recognition layer. The input layer may convert the facial image to be recognized into a tensor format that the model can process. The facial image information may be a tensor representing the facial image to be recognized.

[0102] In the sixth step, the facial image information is input into a facial feature extraction layer to obtain facial feature information. The facial feature extraction layer may be a feature extraction layer (e.g., a convolutional layer or a convolutional neural network) that takes the facial image information as input and outputs the facial feature information. The facial feature information may be a feature vector that extracts features such as facial texture, shape, and expression from the facial image information.

[0103] In the seventh step, the facial feature information is input into the facial emotion recognition layer to obtain facial emotion recognition information, wherein the facial emotion recognition information includes various facial emotion recognition sub-information, and each facial emotion recognition sub-information in the various facial emotion recognition sub-information includes emotion category information and category score. The facial emotion recognition layer can be a classification layer (such as a fully connected layer) that takes facial feature information as input information and facial emotion recognition information as output information. The emotion category information can represent the category of emotion (for example, happiness, anger, sadness, etc.). The category score can be a score that represents whether the emotion represented by the facial feature information belongs to the category represented by the emotion category information.

[0104] Step 8: Generate user mood information based on the facial emotion recognition information and the mood prediction information. In practice, for each facial emotion recognition sub-information included in the facial emotion recognition information, the execution entity may determine the emotion category information included in the facial emotion recognition sub-information as target emotion category information. Thereafter, the execution entity may determine the mood prediction sub-information in the mood prediction information that includes the target emotion category information as the target mood prediction sub-information. Next, the execution entity may determine the sum of the category score included in the facial emotion recognition sub-information and the category score included in the target mood prediction sub-information as the target category score. Thereafter, the execution entity may determine the target emotion category information and the target category score as initial user mood information. The execution entity may determine each mood prediction sub-information in the mood prediction information other than the target mood prediction sub-information as initial user mood information. Finally, the execution entity may determine the emotion category information included in the initial user mood information with the highest category score among the determined initial user mood information as the user mood information. It should be noted that the highest category score indicates the highest category score among the category scores included in the initial user mood information. As an example, the initial user mood information may include "emotion category information: happy, category score: 15 points," "emotion category information: angry, category score: 2 points," and "emotion category information: sad, category score: 4 points." In this case, "emotion category information: happy, category score: 10 points" represents the user mood information.

[0105] The above technical solution, combined with steps 107 to 109 and their related contents, serves as an inventive point of an embodiment of the present disclosure and solves the technical problem of "poor user experience." Factors that lead to a poor user experience are often as follows: Since the front-view scene image containing a facial image may contain multiple people (such as at a party or public place), the captured front-view scene image contains not only the user's facial image, but also facial images of non-users other than the user. If emotion recognition is performed directly based on the front-view scene image and combined with various mood prediction sub-information to generate user mood information, non-user expressions (such as a bystander's smile) may be incorrectly associated with the user's mood prediction (for example, if the user says "I'm angry," but another person in the image is smiling, the system may mistakenly judge the user's emotion as "happy"). This results in poor accuracy of the generated user mood information. Furthermore, the animated expression displayed based on the user mood information may not match the user's actual mood, resulting in a poor user experience. If the above factors are resolved, the effect of improving the user experience can be achieved. To achieve this effect, first, at least one area image corresponding to the at least one positioning area information in the front-view scene image is determined as at least one candidate face image. Thus, all regional images that may contain the user's face are extracted from the front-view scene image. Then, the pre-stored user face image is obtained. Thus, the user's pre-stored user face image can be obtained. Afterwards, for each candidate face image in the at least one candidate face image, the candidate face image is compared with the user face image to obtain a comparison similarity. Thus, the degree of similarity between the candidate face image and the user's face can be quantified, and each comparison similarity used to generate a target comparison similarity can be obtained. Then, the comparison similarity with the greatest similarity among the obtained comparison similarities is determined as the target comparison similarity. Thus, the target comparison similarity used to determine the face image most likely to belong to the user can be obtained. Then, in response to determining that the target comparison similarity is greater than the preset similarity, the candidate face image corresponding to the target comparison similarity in the at least one candidate face image is determined as the face image to be identified. Thus, only when the target comparison similarity is greater than the threshold value, the corresponding candidate face image can be determined as the face image to be identified, thereby filtering out non-user face images, ensuring that subsequent emotion recognition is only for the user face, reducing misjudgment (i.e., the non-user expression (such as the smile of the bystander) may be mistakenly associated with the user's mood prediction (for example, the user says "I am angry", but another person in the image is laughing, but the system mistakenly judges the user's emotion as "happy")). Afterwards, the above-mentioned face image to be identified is input into the input layer of the pre-trained facial emotion recognition model to obtain facial image information, wherein the above-mentioned facial expression recognition model includes the above-mentioned input layer, facial feature extraction layer, and facial emotion recognition layer. Next, the above-mentioned facial image information is input into the facial feature extraction layer to obtain facial feature information.The facial feature information is then input into the facial emotion recognition layer to obtain facial emotion recognition information, wherein the facial emotion recognition information includes various facial emotion recognition sub-information, each of which includes emotion category information and a category score. Thus, the facial emotion recognition model can predict the facial emotion recognition information based on the face image to be recognized that is most similar to the user in the front-view scene image. Finally, user mood information is generated based on the facial emotion recognition information and the mood prediction information. Thus, the facial emotion recognition information and mood prediction information can be combined to generate more accurate user mood information. Furthermore, by extracting all region images that may contain the user's face, namely candidate face images, from the front-view scene image and filtering out non-user face images from each candidate face image, subsequent emotion recognition is performed solely on the user's face, reducing misjudgments (i.e., the possibility of incorrectly associating non-user expressions (such as a bystander's smile) with the user's mood prediction (e.g., if the user says "I'm angry," but another person in the image is smiling, the system may mistakenly identify the user's emotion as "happy")), thereby improving the accuracy of the facial emotion recognition information and user mood information. In combination with steps 107 to 109, animated expressions associated with the user's mood information can be displayed based on the user's mood information with higher accuracy, so that the animated expressions displayed on the doll's display screen are more consistent with the user's actual mood, thereby improving the user experience.

[0106] Step 107: Generate doll animation expression information based on the user mood information and the updated doll personality data.

[0107] In some embodiments, the execution entity may generate doll animated expression information based on the user mood information and the updated doll personality data. In practice, the execution entity may determine the emotion category information included in the user mood information. The execution entity may then determine the doll animated expression data corresponding to the emotion category information among the individual doll animated expression data included in the updated doll personality data as the doll animated expression information.

[0108] Step 108: Acquire motivational feedback information associated with the doll animation expression information.

[0109] In some embodiments, the execution entity may obtain incentive feedback information associated with the animated expression information of the doll. In practice, the execution entity may obtain incentive feedback information associated with the animated expression information of the doll from a preset database. The incentive feedback information may be preset information corresponding to the animated expression information of the doll (e.g., information indicating advertising content, virtual items, etc.).

[0110] Step 109: display the animated expression and motivational feedback information corresponding to the doll's animated expression information on a display screen of the doll.

[0111] In some embodiments, the execution entity may display the animated expression corresponding to the doll's animated expression information and the motivational feedback information on a display screen of the doll.

[0112] The above-mentioned embodiments of the present disclosure have the following beneficial effects: The method for displaying animated expressions on dolls in some embodiments of the present disclosure improves the user experience. Specifically, the poor user experience is caused by directly predicting the current user mood based on collected audio and displaying animated expressions that match the current user mood without performing background noise removal on the collected audio. Background sounds (such as television sound, traffic noise, and other people's conversations) can mix with the user's voice, causing distortion of the voice signal. This in turn leads to low accuracy in predicting the current user's mood. The animated expressions displayed by the doll may not match the user's actual mood, resulting in a poor user experience. Based on this, the method for displaying animated expressions on dolls in some embodiments of the present disclosure first obtains updated doll personality data. This allows the updated doll personality data used to generate doll animated expression information to be obtained. Then, a microphone built into the doll collects a sequence of ambient audio frames. Voice detection is performed on the collected ambient audio frame sequence to determine the timestamp corresponding to the audio frame at the start of the voice. This allows the timestamp of the voice start to be determined. Next, a microphone is used to capture a sequence of speech and audio frames within a preset time period after the timestamp, while a camera built into the doll captures an image of the doll's forward-looking scene. Consequently, after determining the speech start time, the speech and audio frames within the preset time period after the timestamp and the doll's forward-looking scene image are simultaneously captured. Next, background sound removal is performed on the speech and audio frames to obtain a target speech and audio frame sequence. This removes the background sound contained in the speech and audio frame sequence. Next, user mood information is generated based on the target speech and audio frame sequence and the forward-looking scene image. This can be combined with the background-removed target speech and audio frame sequence and the doll's forward-looking scene image to generate user mood information. The doll's forward-looking scene image may capture the user's facial expressions, and combining the target speech and audio frame sequence with the doll's forward-looking scene image to generate user mood information may further improve the accuracy of the user mood information. Next, based on the user mood information and the updated doll personality data, animated expression information for the doll is generated. Thus, the doll's personality, i.e., updated doll personality data, and the user's current mood, i.e., user mood information, can be combined to generate doll animated expression information representing the doll's animated expressions. Next, motivational feedback information associated with the doll animated expression information is obtained. Finally, the animated expression corresponding to the doll animated expression information and the motivational feedback information are displayed on the doll's display screen. Thus, the motivational feedback information can be displayed simultaneously with the animated expression, enhancing the doll's interactivity. Furthermore, because the background sound contained in the speech audio frame sequence is removed before generating the user mood information, the interference of the background sound on the speech signal is reduced. Generating the user mood information based on the target speech audio frame sequence after background sound removal improves the accuracy of the user mood information, ensuring that the animated expression displayed on the doll's display screen is more consistent with the user's actual mood, thereby improving the user experience.

[0113] Further references Figure 2 As an implementation of the methods shown in the figures, the present disclosure provides some embodiments of a doll animation expression display device. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.

[0114] like Figure 2 As shown, some embodiments of the doll animation expression display device 200 include: a first acquisition unit 201, a first collection unit 202, a voice detection processing unit 203, a second collection unit 204, a background sound removal processing unit 205, a first generation unit 206, a second generation unit 207, a second acquisition unit 208 and a display unit 209. Among them, the first acquisition unit 201 is configured to acquire and update the doll personality data; the first collection unit 202 is configured to collect an environmental audio frame sequence through a microphone built into the doll; the voice detection processing unit 203 is configured to perform voice detection processing on the collected environmental audio frame sequence to determine the timestamp corresponding to the voice starting point audio frame; the second collection unit 204 is configured to collect a voice audio frame sequence within a preset time period after the above timestamp through the microphone, and at the same time collect the doll's front view scene image through the camera device built into the doll; the background sound removal processing unit 205 is configured to process the above voice frame sequence. The background sound is removed from the audio frame sequence to obtain a target speech audio frame sequence; the first generation unit 206 is configured to generate user mood information based on the above target speech audio frame sequence and the above front-view scene image; the second generation unit 207 is configured to generate doll animation expression information based on the above user mood information and the above updated doll personality data; the second acquisition unit 208 is configured to obtain incentive feedback information associated with the above doll animation expression information; the display unit 209 is configured to display the animated expression corresponding to the doll animation expression information and the above incentive feedback information on the display screen of the above doll.

[0115] It is understood that the units described in the device 200 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 200 and the units included therein, and will not be repeated here.

[0116] Reference below Figure 3 , which shows a structural diagram of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0117] like Figure 3As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0118] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0119] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the functions defined in the methods of some embodiments of the present disclosure are performed.

[0120] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0121] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0122] The computer-readable medium may be included in the electronic device, or may exist independently and not be incorporated into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain updated doll personality data; collect an ambient audio frame sequence via a microphone built into the doll; perform voice detection processing on the collected ambient audio frame sequence to determine a timestamp corresponding to a voice starting point audio frame; collect a voice audio frame sequence within a preset time period after the timestamp via the microphone, while simultaneously collecting a front-view scene image of the doll via a camera built into the doll; perform background sound removal processing on the voice audio frame sequence to obtain a target voice audio frame sequence; generate user mood information based on the target voice audio frame sequence and the front-view scene image; generate doll animated expression information based on the user mood information and the updated doll personality data; obtain motivational feedback information associated with the doll animated expression information; and display the animated expression corresponding to the doll animated expression information and the motivational feedback information on the doll's display screen.

[0123] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0125] The units described in some embodiments of the present disclosure may be implemented in software or hardware. The units described may also be provided in a processor. For example, they may be described as follows: a processor including a first acquisition unit, a first acquisition unit, a voice detection processing unit, a second acquisition unit, a background sound removal processing unit, a first generation unit, a second generation unit, a second acquisition unit, and a display unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the first acquisition unit may also be described as a "unit for acquiring and updating doll personality data."

[0126] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0127] The above descriptions are merely some preferred embodiments of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of technical features, but should also encompass other technical solutions formed by any combination of technical features or their equivalents without departing from the inventive concept. For example, a technical solution formed by replacing a feature with a technical feature having similar functions as disclosed in the embodiments of the present disclosure (but not limited to) can be formed.

Claims

1. A method for displaying an animated expression of a doll, comprising: Get and update the doll's personality data; The ambient audio frame sequence is collected through the microphone built into the doll; Performing speech detection processing on the collected ambient audio frame sequence to determine the timestamp corresponding to the audio frame of the speech starting point; collecting a sequence of voice audio frames within a preset time period after the timestamp through a microphone, and collecting a front-view scene image of the doll through a camera device built into the doll; Performing background sound removal processing on the speech audio frame sequence to obtain a target speech audio frame sequence; generating user mood information based on the target speech audio frame sequence and the forward-looking scene image; generating doll animation expression information based on the user mood information and the updated doll personality data; Acquiring motivational feedback information associated with the doll animation expression information; The animated expression corresponding to the doll's animated expression information and the motivational feedback information are displayed on a display screen of the doll.

2. The method according to claim 1, wherein The step of obtaining and updating the character data of the doll includes: The user information and acquisition request information input by the user through the display screen are sent to the doll character configuration server, so as to obtain updated doll character data corresponding to the acquisition request information from the doll character configuration server.

3. The method according to claim 1, wherein The performing speech detection processing on the collected ambient audio frame sequence to determine the timestamp corresponding to the speech starting point audio frame includes: For the first ambient audio frame in the ambient audio frame sequence, perform the following speech detection process: Perform voice detection processing on the ambient audio frame to obtain detection information; In response to determining that the detection information indicates that the ambient audio frame includes a speech signal, determining the ambient audio frame as a speech onset audio frame, and determining a timestamp corresponding to the ambient audio frame as a timestamp corresponding to the speech onset audio frame; In response to determining that the detection information indicates that the ambient audio frame does not contain a speech signal, performing the following updating steps: Deleting the first ambient audio frame from the ambient audio frame sequence to update the ambient audio frame sequence; The speech detection process is performed again based on the first ambient audio frame in the updated ambient audio frame sequence.

4. The method according to claim 3, wherein: The performing speech detection processing on the ambient audio frame to obtain detection information includes: Performing feature extraction processing on the ambient audio frame to obtain ambient audio feature information; Inputting the ambient audio feature information corresponding to the ambient audio frame into a pre-trained wake-up word detection model to obtain a wake-up word detection label; In response to determining that the wake-up word detection tag indicates that the ambient audio frame does not include the wake-up word, determining information indicating that the ambient audio frame does not include a voice signal as detection information; In response to determining that the wake-up word detection tag represents that the ambient audio frame includes the wake-up word, information representing that the ambient audio frame includes a voice signal is determined as detection information.

5. The method according to claim 3, wherein: The performing speech detection processing on the ambient audio frame to obtain detection information includes: Performing voiceprint feature extraction processing on the ambient audio frame to obtain voiceprint feature information corresponding to the ambient audio frame; Obtaining a pre-stored user audio feature information set; For each piece of user audio feature information in the user audio feature information set, determining a similarity between the user audio feature information and the voiceprint feature information as an audio feature similarity; Determining the determined audio feature similarities as an audio feature similarity set; In response to determining that the audio feature similarity set includes an audio feature similarity that meets a preset condition, determining information representing that the ambient audio frame includes a speech signal as detection information; In response to determining that no audio feature similarity that meets a preset condition exists in the audio feature similarity set, information representing that the ambient audio frame does not contain a speech signal is determined as detection information.

6. The method according to claim 4, wherein: The wake-up word detection model is trained through the following steps: Acquire a sample set, wherein the samples in the sample set include sample environment audio feature information and a sample target wake-up word detection label corresponding to the sample environment audio feature information; The following training steps are performed based on the sample set: Inputting sample environment audio feature information of at least one sample in the sample set into the initial neural network to obtain a sample predicted wake-up word detection label corresponding to each sample in the at least one sample; Comparing the sample predicted wake-up word detection label corresponding to each sample of the at least one sample with the corresponding sample target wake-up word detection label; Determine whether the initial neural network achieves the preset optimization goal based on the comparison results; In response to determining that the initial neural network achieves the optimization goal, using the initial neural network as a trained wake-up word detection model; In response to determining that the initial neural network does not achieve the optimization goal, the network parameters of the initial neural network are adjusted, and the sample set is composed of unused samples. The adjusted initial neural network is used as the initial neural network, and the training step is performed again.

7. The method according to claim 1, wherein The performing background sound removal processing on the speech audio frame sequence to obtain a target speech audio frame sequence includes: Determining a preset number of ambient audio frames in the ambient audio frame sequence that are located before the voice starting point audio frame as respective ambient noise audio frames to obtain an ambient noise audio frame sequence; For each ambient noise audio frame in the ambient noise audio frame sequence, perform the following processing: Converting the ambient noise audio frame to obtain frequency domain complex spectrum information; Taking a modulus of a complex spectrum corresponding to the frequency domain complex spectrum information to obtain amplitude spectrum information; Based on the obtained information of each amplitude spectrum, an amplitude spectrum matrix is ​​generated; generating noise mean information based on the amplitude spectrum matrix; For each speech audio frame in the speech audio frame sequence, perform the following processing: Converting the speech audio frame to obtain target frequency domain complex spectrum information; Taking the modulus of the complex spectrum corresponding to the target frequency domain complex spectrum information to obtain target amplitude spectrum information; generating background sound removed amplitude spectrum information based on the target amplitude spectrum information and the noise mean information; Sorting the obtained background sound amplitude spectrum information according to the order of the corresponding target speech audio frames in the target speech audio frame sequence to obtain a background sound amplitude spectrum information sequence; An audio signal reconstruction process is performed on each background-removed sound amplitude spectrum information in the background-removed sound amplitude spectrum information sequence to obtain a target speech audio frame sequence.

8. A doll animation expression display device, comprising: A first acquiring unit is configured to acquire and update the character data of the doll; A first acquisition unit is configured to acquire an ambient audio frame sequence through a microphone built into the doll; A speech detection processing unit is configured to perform speech detection processing on the collected ambient audio frame sequence to determine a timestamp corresponding to a speech starting point audio frame; a second acquisition unit configured to acquire a speech audio frame sequence within a preset time period after the timestamp through a microphone, and simultaneously acquire a front-view scene image of the doll through a camera device built into the doll; A background sound removal processing unit is configured to perform background sound removal processing on the speech audio frame sequence to obtain a target speech audio frame sequence; a first generating unit configured to generate user mood information based on the target speech audio frame sequence and the front-view scene image; a second generating unit configured to generate doll animation expression information based on the user mood information and the updated doll personality data; A second acquiring unit is configured to acquire motivational feedback information associated with the doll animation expression information; The display unit is configured to display the animated expression corresponding to the doll's animated expression information and the motivational feedback information on a display screen of the doll.

9. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • AI stuffed toy with expression expression ability

    CN121944537A