Doll expression replacement display method and device, electronic equipment and medium
Through audio scene classification and emotion recognition using lightweight models and multi-scale temporal pooling layers, the lag and energy consumption issues of smart doll expression mapping are resolved, real-time and accurate expression display is achieved, the device damage rate is reduced, and the user experience is improved.
Patent Information
- Application Number
- CN202510894926.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-10
AI Technical Summary
In the existing technology, smart dolls cannot be mapped to expressions in real time according to emotional changes when offline. The expression conversion is delayed and has low accuracy. In addition, the transmission of large amounts of audio data leads to high energy consumption of the device, high damage rate, and poor user experience.
A lightweight scene recognition model and emotion recognition model are used, combined with a convolutional neural network and a multi-scale temporal pooling layer for audio scene classification and emotion recognition, to optimize the device acquisition state switching and achieve real-time and accurate expression mapping.
It improves the real-time and accuracy of expression mapping, reduces device energy consumption and damage rate, and enhances user experience.
Smart Images

Figure CN120762800A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method, device, electronic device, and medium for changing and displaying expressions of dolls. Background Art
[0002] Smart dolls can be low-power devices that use voice recognition technology and environmental awareness to interactively respond to user commands and locally display the user's emotions through the doll's facial expressions. To change the doll's expressions, a common approach is to send audio data collected by the doll to the cloud, where a large language model is used to perform emotion recognition on the audio data to determine the user's emotions. The user's emotions are then mapped and matched against a library of preset doll expressions to determine the current doll's expression. Finally, the current doll expression is sent to the doll for expression switching and display.
[0003] However, it has been found in practice that when the above method is used to change the expression display of the doll, the following technical problems often occur: Since the doll needs to rely on cloud computing resources for emotion recognition and expression mapping, it is impossible to achieve real-time mapping of emotion transformation to expression display in an offline situation, and the doll expression conversion is only performed through audio emotion recognition. The interaction is too simple, and there are problems such as expression conversion lag and low accuracy of expression mapping display. At the same time, dolls are generally portable, low-power devices. The transmission of a large amount of audio data will cause the device to consume more energy and the expression display time is short, resulting in lower performance of the doll, increased device damage rate of the doll, and reduced user experience.
[0004] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure concept and therefore it may contain information that does not form the prior art that is already known in this country to a person of ordinary skill in the art. Summary of the Invention
[0005] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0006] Some embodiments of the present disclosure provide methods, devices, electronic devices, and media for displaying and changing expressions of dolls to solve one or more of the technical problems mentioned in the above background technology section.
[0007] In a first aspect, some embodiments of the present disclosure provide a method for changing and displaying expressions of a doll, comprising: obtaining current device acquisition status information and device energy consumption information of the doll; controlling the doll to acquire scene audio according to the current device acquisition status information to obtain scene environment audio; determining model occupancy information required for the deployment and execution of an audio scene classification and recognition model and a doll audio emotion recognition model; in response to determining that the model occupancy information is greater than or equal to the device resource information corresponding to the doll, performing model lightweight processing on the audio scene classification and recognition model and the doll audio emotion recognition model to obtain a lightweight scene recognition model and a lightweight emotion recognition model; utilizing the lightweight scene recognition model to deploy and execute the scene environment audio; The scene environment audio is subjected to scene classification and recognition to obtain the current device scene information set of the above-mentioned doll; in response to determining that the above-mentioned current device scene information set is different from the previous device scene information, the device acquisition state switching processing is performed on the above-mentioned doll according to the above-mentioned device energy consumption information to obtain acquisition state switching information; the above-mentioned lightweight emotion recognition model is used to perform emotion recognition on the above-mentioned scene environment audio to obtain an audio emotion information set; based on the above-mentioned current device scene information set and the above-mentioned audio emotion information set, the expression of the electronic doll of the above-mentioned doll is mapped to obtain the current doll expression information; based on the above-mentioned current doll expression information and the above-mentioned acquisition state switching information, the above-mentioned doll is controlled to change the expression display of the doll on the device display screen.
[0008] In a second aspect, some embodiments of the present disclosure provide a doll expression replacement display device, comprising: an acquisition unit configured to acquire current device collection status information and device energy consumption information of the doll; a first control unit configured to control the doll to perform scene audio collection according to the current device collection status information to obtain scene environment audio; a determination unit configured to determine model occupancy information required for the deployment and execution of the audio scene classification and recognition model and the doll audio emotion recognition model; a model lightweight processing unit configured to perform model lightweight processing on the audio scene classification and recognition model and the doll audio emotion recognition model in response to determining that the model occupancy information is greater than or equal to the device resource information corresponding to the doll, to obtain a lightweight scene recognition model and a lightweight emotion recognition model; the scene classification and recognition unit is configured to utilize the lightweight scene recognition model to perform model lightweight processing on the audio scene classification and recognition model and the doll audio emotion recognition model. , perform scene classification and recognition on the above-mentioned scene environment audio to obtain the current device scene information set of the above-mentioned doll; the device acquisition state switching unit is configured to, in response to determining that the above-mentioned current device scene information set is different from the previous device scene information, perform device acquisition state switching processing on the above-mentioned doll according to the above-mentioned device energy consumption information to obtain acquisition state switching information; the emotion recognition unit is configured to use the above-mentioned lightweight emotion recognition model to perform emotion recognition on the above-mentioned scene environment audio to obtain an audio emotion information set; the expression mapping unit is configured to perform expression mapping on the electronic doll expression of the above-mentioned doll according to the above-mentioned current device scene information set and the above-mentioned audio emotion information set to obtain current doll expression information; the control unit is configured to control the above-mentioned doll to change the expression display of the doll on the device display screen according to the above-mentioned current doll expression information and the above-mentioned acquisition state switching information.
[0009] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.
[0010] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method described in any implementation manner in the first aspect is implemented.
[0011] The above-mentioned embodiments of the present disclosure have the following beneficial effects: the doll expression replacement and display method of some embodiments of the present disclosure can improve the real-time and accuracy of expression mapping display, reduce the energy consumption and damage rate of the doll device, and enhance the user experience. Specifically, the reasons for the related expression conversion lag and low expression mapping display accuracy, as well as the low performance of the doll, increased doll device damage rate, and reduced user experience are: because the doll relies on cloud computing resources for emotion recognition and expression mapping, it is impossible to achieve real-time mapping from emotion transformation to expression expression in offline conditions. The doll expression conversion is only performed through audio emotion recognition, which is too simple to interact with, resulting in expression conversion lag and low expression mapping display accuracy. In addition, dolls are generally portable and low-power devices. The transmission of large amounts of audio data will result in high device energy consumption and short expression display time, resulting in low performance, increased doll device damage rate, and reduced user experience. Based on this, the doll expression replacement and display method of some embodiments of the present disclosure can first obtain the doll's current device collection status information and device energy consumption information. Here, the current device acquisition status information and device energy consumption information can be used to grasp the performance of the doll device in real time, which is convenient for improving the battery life of the doll in the future. Secondly, based on the above-mentioned current device acquisition status information, the doll is controlled to perform scene audio acquisition to obtain scene environment audio. Here, the scene environment audio can be used for subsequent scene classification and emotion recognition, as well as to control the acquisition frequency of the doll and reduce the energy consumption of the doll. Thirdly, the model occupancy resource information required for the deployment and execution of the audio scene classification and recognition model and the doll audio emotion recognition model is determined. Here, it can be used to determine whether the model deployment and execution is applicable to the doll, so as to achieve a balance between device battery life and the accuracy of the doll expression display on low-power devices. Then, in response to determining that the above-mentioned model occupancy resource information is greater than or equal to the device resource information corresponding to the above-mentioned doll, the audio scene classification and recognition model and the doll audio emotion recognition model are subjected to model lightweight processing to obtain a lightweight scene recognition model and a lightweight emotion recognition model. Here, since the doll is a low-power device, it needs to meet the requirements of portability, real-time performance, low power consumption, and long battery life. Lightweighting the model can reduce the load consumption of the doll device and improve the performance of the doll device. Subsequently, using the lightweight scene recognition model, the above-mentioned scene environment audio is subjected to scene classification and recognition to obtain the current device scene information set of the doll. Here, it can be applied to the requirements of scene classification and recognition in all scenarios, and improve the accuracy of scene classification and recognition in offline conditions. Then, in response to determining that the above-mentioned current device scene information is different from the previous device scene information, the device acquisition state switching process is performed on the above-mentioned doll device based on the above-mentioned device energy consumption information to obtain acquisition state switching information. Here, the device battery life of the doll can be optimized, the energy consumption of the doll device can be reduced, and a balance between energy consumption and responsiveness can be achieved.Then, the light-weight emotion recognition model is used to perform emotion recognition on the scene environmental audio to obtain an audio emotion information set. Here, emotion recognition in an offline case can be implemented, emotion recognition accuracy in a discrete case can be improved, energy consumption in transmission can be reduced in a cloud connection case, and real-time emotion recognition can be implemented. Then, according to the current device scene information set and the audio emotion information set, an electronic doll expression of the doll is mapped to obtain current doll expression information. Here, the comprehensiveness and accuracy of expression mapping can be improved, the naturalness and fluency of user-doll interaction can be improved, and user experience can be improved. Finally, according to the current doll expression information and the collection state switching information, the doll is controlled to display doll expression replacement on the device screen. Here, real-time interaction and doll expression mapping display are implemented. Thus, the doll expression replacement display method can improve the real-time and accuracy of expression mapping display, reduce device energy consumption and damage rate of the doll, and improve user experience. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above and other features, aspects and advantages of the present disclosure will become more apparent with reference to the following detailed description when taken in conjunction with the accompanying drawings. Throughout the drawings, similar or same reference numerals are used to denote similar or same elements. It is to be understood that the drawings are schematic, and elements and features are not necessarily to scale.
[0013] Figure 1 is a flowchart of some embodiments of a doll expression replacement display method according to the present disclosure;
[0014] Figure 2 is a structural schematic diagram of some embodiments of a doll expression replacement display device according to the present disclosure;
[0015] Figure 3 is a structural schematic diagram of an electronic device suitable for use to implement some embodiments of the present disclosure. DETAILED DESCRIPTION
[0016] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0017] It should also be noted that, for the sake of brevity, only the parts of the drawings that are relevant to the present invention are shown. The embodiments and features in the present disclosure can be combined with each other as long as there is no conflict.
[0018] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0019] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0020] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0021] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0022] Figure 1 The flowchart 100 of some embodiments of the method for changing and displaying a doll expression according to the present disclosure is shown. The method for changing and displaying a doll expression includes the following steps:
[0023] Step 101: Obtain the current device collection status information and device energy consumption information of the doll.
[0024] In some embodiments, the execution entity (e.g., an electronic device) of the above-mentioned method for changing and displaying expressions of a doll can obtain the doll's current device collection state information and device energy consumption information via a wired or wireless connection. The current device collection state information may include the doll's device state information and the frequency of collecting audio data. The device state information may include: sleep state information, wake state information, and active state information. The frequency of collecting audio data for the sleep state information, wake state information, and active state information increases in sequence. For example, the sleep state information may include audio data collection every 10 seconds. The wake state information may include audio data collection every 2 seconds. The active state information may include audio data collection every 0.2 seconds. The device energy consumption information may include information on the power resources consumed by the doll to collect audio data and maintain the display of the doll's expressions. For example, the device energy consumption information may include a maximum value of 40 milliwatts for the device in the sleep state and a maximum value of 200 milliwatts for the device in the active state. The doll can be a portable, low-power device with a doll-like appearance and equipped with an electronic display, sensor acquisition equipment, a microphone, and a chip with a lightweight deep neural network model. For example, the doll can be a trendy electronic personalized bag.
[0025] Step 102: Control the doll to collect scene audio according to the current device collection status information to obtain scene environment audio.
[0026] In some embodiments, the execution entity may control the doll to collect scene audio based on the current device collection status information, thereby obtaining scene ambient audio. The scene ambient audio may include audio data collected within a preset area where the user is carrying the doll, including scene sounds and user interaction sounds, at a collection frequency determined by the current device collection status information. The preset area may be an area with a collection radius of 5 meters centered on the doll's location.
[0027] Step 103: Determine the model resource occupancy information required for the deployment and execution of the audio scene classification and recognition model and the puppet audio emotion recognition model.
[0028] In some embodiments, the execution entity may determine the model resource occupancy information required for the deployment and execution of the audio scene classification and recognition model and the puppet audio emotion recognition model. The audio scene classification and recognition model may be a deep neural network model that performs scene recognition and sound event recognition on the input scene environment audio to output information about the user's scene. For example, the audio scene classification and recognition model may be a YAMNet (Yet Another Music Recognition Network) model. The puppet audio emotion recognition model may be a deep neural network model that performs audio emotion recognition on the input scene environment audio to output recognized and detected emotion information. For example, the puppet audio emotion recognition model may be a lightweight version of a serial convolutional neural network model and a deep neural network model with a self-attention mechanism. The model resource occupancy information may include information on the memory, CPU (Central Processing Unit) resources, GPU (Graphics Processing Unit) resources, and other resources required to run the audio scene classification and recognition model and the puppet audio emotion recognition model on the collected scene environment audio.
[0029] Step 104 , in response to determining that the model resource occupancy information is greater than or equal to the device resource information corresponding to the doll, the audio scene classification recognition model and the doll audio emotion recognition model are lightweighted to obtain a lightweight scene recognition model and a lightweight emotion recognition model.
[0030] In some embodiments, the execution entity may, in response to determining that the resource information occupied by the model is greater than or equal to the device resource information corresponding to the doll, perform model lightweight processing on the audio scene classification and recognition model and the doll audio emotion recognition model to obtain a lightweight scene recognition model and a lightweight emotion recognition model. The device resource information corresponding to the doll may be information on the maximum resources that the doll can run. The lightweight scene recognition model may be a model that reduces the number of parameters and computational complexity of the audio scene classification and recognition model while ensuring the accuracy and performance of the model so that it can run on a low-power doll device. The lightweight emotion recognition model may be a model that reduces the number of parameters and computational complexity of the audio emotion recognition model while ensuring the accuracy and performance of the model so that it can run on a low-power doll device. The lightweight processing may be a lightweight processing performed by one or more combinations of methods such as pruning, distillation, neural network search (NAS), quantization, and low-rank decomposition.
[0031] Step 105 : Use a lightweight scene recognition model to perform scene classification and recognition on the scene environment audio to obtain a current device scene information set of the doll.
[0032] In some embodiments, the execution entity may utilize a lightweight scene recognition model to perform scene classification and recognition on the scene environment audio to obtain a current device scene information set of the doll. The current device scene information in the current device scene information set may be information about a scene where the user is located, as identified by the lightweight scene recognition model through recognition of the scene environment audio, and where the scene recognition probability value is greater than or equal to a preset scene recognition probability threshold. The preset scene recognition probability threshold may be a pre-set critical value for determining whether it is reasonable. The preset scene recognition probability threshold may be 0.8.
[0033] In the process of adopting technical solutions to solve the above-mentioned technical problem 1, the following technical problem 2 often arises: How to improve the recognition accuracy of sound events and event scenes in the doll's audio scene without increasing the doll's device energy consumption on a low-power doll device, so as to further reduce the low-power doll's energy consumption, extend the doll's device life, and improve the interactive experience with the user. In response to the above-mentioned technical problem 2, the conventional solution is generally to use the BoAW (Bag-of-Audio-Words) algorithm and TFL (Temporal Feature Learning, unsupervised temporal feature learning method) and a lightweight scene recognition model to perform scene classification and recognition on the above-mentioned scene environment audio to obtain the current device scene information set of the above-mentioned doll. However, the above conventional solutions still have the following problems: since the primitive representations of BoAW all obtain single audio features in a single-scale construction method, and TFL adopts a shallow-layer element-by-element optimization strategy to capture the temporal relationship between single-scale primitives for temporal modeling, it does not consider the temporal constraints between the input spectrum segments, making it unable to effectively capture the temporal relationship between the segments, resulting in less semantic information of the extracted temporal information and sound event scene information, and poor scene recognition accuracy, which in turn leads to high energy consumption and poor battery life for low-power dolls, reducing the user interaction experience. The inventors took into account the shortcomings of the above conventional solutions and combined with the advantages / technical status of the audio scene recognition technology owned by the inventor's company, we decided to adopt the following solution:
[0034] In some optional implementations of some embodiments, the lightweight scene recognition model includes: an audio segment feature extraction model, multiple local stacked temporal pooling layers, a global multi-scale temporal pooling layer, and an audio scene classification model. The audio segment feature extraction model may be a neural network model that extracts temporal information from an input scene environment audio segment set to output audio temporal information. The scene environment audio segments in the scene environment audio segment set may be audio segments in which the scene environment audio is segmented with a duration of 1 second and the overlap of adjacent audio segments is 50%. The audio segment feature extraction model may be a model including a first convolution extraction network, a first maximum pooling layer, a second convolution extraction network, a second maximum pooling layer, a third convolution extraction network, a third maximum pooling layer, and three fully connected networks in a series relationship. The first convolution extraction network may include a convolution layer with a convolution kernel of 3*3, 1 input channel, and 32 output channels, a BN (Batch Normalization) layer, an activation function layer with an LReLU (Leaky Rectified Linear Unit) activation function, a convolution layer with a convolution kernel of 3*3, 32 input channels, and 32 output channels, a BN layer, and an activation function layer with an LReLU activation function. The second and third convolution extraction networks may have the same network structure as the first convolution extraction network, but the input and output channels are (32, 64), (64, 64), (64, 128), and (128, 128), respectively. The first and second maximum pooling layers may be maximum pooling layers with a pooling kernel of 2*2, respectively. The third maximum pooling layer may be a maximum pooling layer with a pooling kernel of 6*2. The above three fully connected layers can be respectively a first fully connected network layer including a fully connected layer with input and output channels of 128 and 512, a BN layer, an LReLU activation function and a Dropout layer, a second fully connected network layer with input and output channels of 512 and 512, a BN layer, an LReLU activation function and a Dropout layer, and a third fully connected layer with input and output channels of 512 and 10, a BN layer and an LReLU activation function.
[0035] The local stacked temporal pooling layer in the above-mentioned multiple local stacked temporal pooling layers can be a local temporal pooling layer that extracts feature representations of multiple audio spectrum segment feature maps within the coding sliding window to obtain temporal information of new sub-feature maps arranged in chronological order, so that the subsequent stacked temporal pooling layer uses the newly formed sub-feature maps as primitive representation sequences to extract temporal information and semantic information between primitives. The above-mentioned local temporal pooling layer can be a regularized support vector regression machine that characterizes the correspondence between the input audio spectrum segment feature map sequence and the temporal sorting order through a learnable linear function, and describes the potential structure of the input audio spectrum segment feature map sequence through the generated hyperplane. The above-mentioned multiple local stacked temporal pooling layers can be a temporal pooling layer that uses a kernel mapping function to map the audio spectrum segment feature atlas output by the audio segment-level time-frequency domain feature extraction model one by one to obtain an expanded feature map; then, the expanded feature map is segmented using a sliding window and then regularized SVR (Support Vector Regression) encoding is performed to obtain a temporal encoded audio feature map set, which serves as the basic representation of the next layer of local stacked temporal pooling layer. The above-mentioned segmentation can be the number of audio spectrum segment feature maps included in the input audio spectrum segment feature map set minus the length of the sliding window, divided by the step size of the local stacked temporal pooling layer, rounded down, and added to 1. The global multi-scale temporal pooling layer can be a temporal pooling layer that globally encodes the local segment temporal semantic feature maps of multiple scales output by the above-mentioned multiple local stacked temporal pooling layers to output the semantic information of the audio feature map. The audio scene classification model can be a model that performs classification processing on the audio feature map set output by the above-mentioned global multi-scale temporal pooling layer. For example, the audio scene classification model may be a support vector machine.
[0036] Optionally, the lightweight scene recognition model is used to perform scene classification and recognition on the scene environment audio to obtain the current device scene information set of the doll, which may include the following steps:
[0037] The first step is to construct a spectrum for the above-mentioned scene environment audio to obtain an audio scene Mel spectrum set. The audio scene Mel spectrum in the above-mentioned audio scene Mel spectrum set can be a logarithmic Mel spectrum graph. In practice, the above-mentioned execution subject can first perform a short-time Fourier transform on the above-mentioned scene environment audio to obtain an audio frame spectrum graph set. The audio frame spectrum graph in the above-mentioned audio frame spectrum graph set can represent a graph structure representation of the change of audio frequency over time. Secondly, the above-mentioned audio frame spectrum graph set is mapped to obtain a Mel spectrum graph set. The Mel spectrum graph in the above-mentioned Mel spectrum graph set can be a spectrum graph obtained by mapping the linear frequency of the audio frame spectrum graph to the Mel frequency and then performing spectral filtering. Finally, the above-mentioned Mel spectrum graph set is logarithmically transformed to obtain an audio scene Mel spectrum set.
[0038] In the second step, the audio scene Mel spectrum set is segmented in the time domain to obtain an audio scene spectrum fragment set. The audio scene spectrum fragments in the audio scene spectrum fragment set can be spectrum fragments with a preset time length and a spectrum overlap of 50% with the adjacent audio scene Mel spectrum. The preset time length can be a pre-set time length. For example, the preset time length can be 1 second. It should be noted that the 1 second of time domain segmentation is obtained by counting a large amount of data. Using long spectrum fragments as input may result in too low resolution of the audio fragment level feature sequence in the time domain, which in turn results in less temporal information and semantic information being extracted during subsequent feature extraction, thereby reducing the model performance of the lightweight scene recognition model.
[0039] In the third step, the audio scene spectrum segment set is input into the audio segment-level time-frequency domain feature extraction model to obtain an audio spectrum segment feature graph set. The audio spectrum segment feature graphs in the audio spectrum segment feature graph set can represent the temporal information of the audio scene spectrum segment.
[0040] In the fourth step, the audio spectrum segment feature atlas is input into the multiple local stacked temporal pooling layers to obtain a local segment temporal semantic feature atlas. The local segment temporal semantic feature maps in the local segment temporal semantic feature atlas can represent the temporal and semantic information of the audio scene spectrum segment.
[0041] In a fifth step, nonlinear feature mapping is performed on the local segment temporal semantic feature atlas to obtain a local segment nonlinear mapping feature atlas. The local segment nonlinear mapping feature maps in the local segment nonlinear mapping feature atlas can represent nonlinear feature information on the audio temporal sequence. The nonlinear feature mapping can be performed using a kernel function.
[0042] In the sixth step, the local segment nonlinear mapping feature atlas is input into the global multi-scale temporal pooling layer to obtain the audio segment temporal semantic feature atlas. The local segment nonlinear mapping feature maps in the local segment nonlinear mapping feature atlas can represent the temporal relationship and semantic information between multi-scale primitives. The multiple local stacked temporal pooling processes and the global multi-scale pooling process can effectively capture the temporal relationship and non-stationary local variation information between multi-scale primitives to obtain audio semantic information containing multi-scale temporal information.
[0043] In the seventh step, the temporal semantic feature atlas of the audio clips is input into the audio scene classification model to obtain the current device scene information set of the doll.
[0044] The above technical problems and related contents are one of the invention points of the embodiments of the present disclosure, which solves the second technical problem mentioned in the background art. Since the primitive representation of BoAW is obtained by constructing a single scale audio single feature, and TFL adopts a shallow element-by-element optimization strategy to model the timing by capturing the timing relationship between single scale primitives, it does not consider the timing constraints between input spectral segments, which makes it difficult to effectively capture the timing relationship between segments, resulting in less semantic information of the extracted timing information and sound event scene information, lower scene recognition accuracy, and higher energy consumption of the low-power doll, poorer battery life, and reduced user interaction experience. If the above factors are solved, the energy consumption of the low-power doll can be reduced, the battery life of the doll can be improved, and the user interaction experience can be improved. In order to achieve this effect, the present disclosure first extracts audio timing information through an audio segment feature extraction model composed of a convolutional neural network model, which can learn more local context content in multiple time-frequency segments, thereby better coping with the complexity of scenes and events. Then, through multiple local stacked timing pooling layers and global multi-scale timing pooling layers, the multi-scale inter-primitive timing relationship is captured based on the pyramid timing pooling, and the global timing encoding operation is stacked into multiple local timing encoding operations, which can capture the inter-primitive timing relationship in multiple scales, thereby obtaining more expressive sample-level semantic features. The timing modeling method can effectively capture the timing relationship between multiple scale primitives, avoid capturing the timing relationship between primitives in a single scale, and analyze the multi-level semantic information of non-stationary signals in audio. Then, the audio scene classification model is used for audio scene recognition, and the high-quality audio segment timing semantic feature map set extracted from the multi-timing features and semantic features can further improve the accuracy and efficiency of audio scene recognition, thereby further improving the accuracy of scene recognition on low-energy devices. Finally, the current device scene information set and the subsequent emotional information are recognized, and the doll is dynamically mapped and transformed in expression, which can further improve the accuracy of scene recognition while improving the energy consumption of the doll device, balancing the accuracy of expression mapping of device energy consumption, improving the performance of the doll device, and improving the interaction experience of user and doll interaction.
[0045] In step 106, in response to determining that the current device scene information set is different from the last device scene information, the device collection state switching processing is performed on the doll according to the device energy consumption information, and the collection state switching information is obtained.
[0046] In some embodiments, the execution entity may, in response to determining that the current device scene information set is different from the previous device scene information, perform device acquisition state switching processing on the doll based on the device energy consumption information to obtain acquisition state switching information. The previous device scene information may be the scene information of the last time the doll's expression changed. The acquisition state switching information may be information obtained by changing the device acquisition frequency and state of the current doll. The device energy consumption information may be information about the current doll's power usage time and usage. As an example, the execution entity may, in response to determining that the current device scene information is different from the previous device scene information, perform device acquisition state switching processing on the doll based on the device energy consumption information using a doll state switching rule engine to obtain acquisition state switching information. The doll state switching rule engine may be a rule engine constructed using a pre-set business logic conversion rule table. The business logic switching rule table may be a table formed by multiple business logic switching rule information. For example, the business logic switching rule information can be rule information that switches the device status information to wake-up status information without changing the audio collection frequency when the device energy consumption information is less than 0.3 and the current device scene information is different from the previous device scene information. It can also be rule information that switches the device status information to wake-up status information and increases the audio collection frequency when the device energy consumption information is greater than or equal to 0.6 and the current device scene information is different from the previous device scene information.
[0047] While employing technical solutions to address the aforementioned technical problem (1), the following technical problem often arises: how to balance the energy consumption information of each doll's device and the accuracy of the doll's expression mapping conversion, so as to maintain or improve the accuracy of the doll's expression conversion mapping while ensuring the doll's battery life. A conventional solution to this technical problem (2) generally involves employing the traditional ant lion algorithm to perform device acquisition state switching on the doll based on the device energy consumption information, thereby obtaining acquisition state switching information. However, this conventional solution still suffers from the following problems: Because the traditional ant lion algorithm uses random initialization, which results in significant randomness, and because the updates to the ant colony are determined by elite ant lions and a roulette wheel selection algorithm during the ant colony update process, the algorithm's optimization efficiency decreases as the number of cycles increases, leading to trapping in local optimal solutions and resulting in low accuracy in device acquisition state switching. Increasing the acquisition frequency when it should be reduced increases the doll's device energy consumption and reduces the doll's battery life. Reducing or maintaining the acquisition frequency when it should be increased reduces the accuracy of the doll's expression conversion, increases the doll's damage rate, and reduces the accuracy of the expression mapping display, thereby diminishing the user's interactive experience. The inventors considered the shortcomings of conventional solutions and combined the advantages and current status of the equipment acquisition state switching technology owned by the inventors' company. We decided to adopt the following solution:
[0048] In some optional implementations of some embodiments, in response to determining that the current device scene information set is different from the previous device scene information, performing device collection state switching processing on the doll according to the device energy consumption information to obtain the collection state switching information may include the following steps:
[0049] The first step is to discretize the various energy consumption indicator information included in the energy consumption information of the above-mentioned device to obtain a discretized energy consumption indicator information set. Among them, the discretized energy consumption indicator information in the above-mentioned discretized energy consumption indicator information set can be information that characterizes the performance indicator information of the doll in a discretized form. The above-mentioned discretized energy consumption indicator information set may include but is not limited to at least one of the following: device memory load, power consumption rate, audio acquisition rate, CPU. The above-mentioned discretized energy consumption indicator information set can serve as a searchable discrete space for subsequent ant colonies and ant lion colonies. It should be noted that the above-mentioned discretization processing can divide the continuous interval into a finite number of discrete points and convert the infinite space into a finite set, which can significantly reduce the search complexity, and the collection state of the doll itself is discretized, which is more in line with the collection state transition scenario of the doll and more timely for subsequent heuristic optimization algorithms.
[0050] The second step is to generate a state switching fitness function. This state switching fitness function can represent a multi-objective fitness function that measures device energy consumption information and the accuracy of the doll's expression mapping conversion. This state switching fitness function can be a fitness function that minimizes energy loss and maximizes expression mapping conversion accuracy, introducing a heuristic weighted balance.
[0051] The third step is to generate an initialized ant colony and an initialized ant lion colony based on the above-mentioned discretized energy consumption indicator information set. Among them, each individual ant in the above-mentioned initialized ant colony and each individual ant lion in the initialized ant lion colony can be a switching scheme for the equipment acquisition state switching process. Each of the above-mentioned ants explores the solution space through random walks, and the ant lion guides the ants to explore the direction of the solution space through the trap mechanism, and jointly promotes the search process to explore the global optimal solution, so as to finally obtain the optimal solution through multiple rounds of iterations. As an example, the above-mentioned execution subject can use the tent chaos mapping algorithm to generate an initialized ant colony and an initialized ant lion colony based on the above-mentioned discretized energy consumption indicator information set. Among them, the above-mentioned initialized ant colony can be the ant colony initial position information set obtained by initializing the position of the ant colony. The above-mentioned initialized ant lion colony can be the ant lion initial position information set obtained by initializing the position of the above-mentioned ant lion colony.
[0052] Step 4: Based on the initial ant colony and the initial ant lion colony, perform the following determination steps:
[0053] Sub-step 1: Input the initialized ant colony and the initialized antlion colony into the aforementioned state switching fitness function to obtain an ant fitness function value set and an antlion fitness function value set. The ant fitness function values in the ant fitness function value set and the antlion fitness function values in the antlion fitness function value set can both represent the degree of balance between device energy consumption information and expression conversion accuracy of the solution corresponding to the ants.
[0054] Sub-step 2: Select the antlion fitness function value with the largest value from the antlion fitness function value set as the target antlion fitness function value, and determine the antlion corresponding to the target antlion fitness function value as the target antlion.
[0055] Sub-step 3: Determine the ant lion corresponding to each ant in the initialized ant colony to obtain the target associated ant lion group. In practice, the execution entity can use a random walk algorithm of the roulette wheel selection algorithm to determine the ant lion corresponding to each ant in the ant colony to obtain the associated ant lion group.
[0056] Sub-step 4: Performing a Lévy flight differential position update on the initialized ant colony based on the target-associated ant lion colony and the target ant lion to obtain an updated ant colony. The position of each ant colony in the updated ant colony may be a position information where the ant's random walk range is gradually reduced due to the random walk of the target ant lion and the target-associated ant lion.
[0057] As an example, the execution subject may, in the first step, perform the following update steps for each initialized ant individual in the initialized ant colony: first, perform position normalization processing on the initialized ant individual to obtain a position normalized ant. Secondly, determine the target associated ant lion of the initialized ant individual. Thirdly, perform Levy flight update processing on the initialized ant individual to obtain the initial ant individual after Levy update. Among them, the random step length in the Levy flight update may be the ratio of the first random number that obeys the standard normal distribution to the first power of the Levy factor of the modulus of the first random number that obeys the standard normal distribution. The Levy factor can characterize the heavy-tailed characteristic of determining the step length, and the value range is (0, 2). Subsequently, determine the dot product of the initial ant individual after Levy update and the target associated ant lion as the ant individual after position update. In the second step, three position-updated ant individuals are randomly selected from the obtained position-updated ant colony. The vector difference between the second and third random position-updated ant individuals is multiplied by a mutation factor and then added to the vector difference between the first and second random position-updated ant individuals, until every position-updated ant individual in the entire position-updated ant colony has mutated, thereby obtaining a differentially mutated ant colony. The mutation factor can be a random number in the range [0, 1]. In the third step, a binomial crossover operation is performed on the differentially mutated ant colony and the position-updated ant colony to obtain a crossover mutated ant colony. The binomial crossover operation can be performed such that a randomly selected random number uniformly distributed between (0, 1) is less than or equal to a crossover factor in the range [0, 1], and the random number is equal to [1, the vector dimension of the initial ant individual]. The crossover mutated ant individual is a position-updated ant; otherwise, the crossover mutated ant individual is a differentially mutated ant. In the fourth step, the selection operation in the differential evolution algorithm is performed on the crossover mutated ant colony to obtain a differentially selected updated ant colony, which serves as the updated ant colony.
[0058] Sub-step 5: Substitute the updated ant colony into the above-mentioned state switching fitness function to obtain the updated fitness function value set of the ant colony.
[0059] Sub-step 6: Compare the updated fitness function value of each ant colony in the ant colony's updated fitness function value set with the ant lion fitness function value corresponding to the ant colony's updated fitness function value in the ant lion fitness function value set, to obtain a comparison result set. Sub-step 7: Update the updated ant colony and the initial ant lion colony based on the comparison result set to obtain a target updated ant colony and an updated ant lion colony. The target updated ant colony can be the position information set of the remaining ant colony after some ants are captured by ant lions and the remaining ant lions randomly walk. The updated ant lions can be the position information of all ant lions randomly walking after an ant lion captures an ant and uses the ant's position as the updated ant lion's position.
[0060] As an example, the execution entity may first, in response to determining at least one comparison result in the comparison result set representing an ant colony's updated fitness function value greater than the corresponding ant lion's fitness function value, determine the position information of each ant corresponding to the updated fitness function value of each ant colony corresponding to the at least one comparison result as the position information of each corresponding ant lion. Then, the updated ant colony is removed from the at least one updated ant corresponding to the at least one comparison result, and the target updated ant colony is determined. Finally, the ant lions corresponding to the position information of each ant lion and at least one initialized ant lion whose updated fitness function value representing the ant colony is less than or equal to the corresponding ant lion fitness function value are determined as the updated ant lion colony.
[0061] Sub-step 8: Determine the number of times the above determination step has been performed.
[0062] In sub-step 9, in response to determining that the number of executions is greater than or equal to a preset execution threshold, the target update ant colony and the update ant lion colony are determined to be energy consumption weight coefficients corresponding to each energy consumption indicator information in each energy consumption indicator information, thereby obtaining an energy consumption weight coefficient set. The preset execution threshold may be a pre-set maximum number of iterations of the determination step. For example, the preset execution threshold may be 100.
[0063] In step 5, in response to determining that the number of executions is less than a preset execution threshold, the target update ant colony and the update ant lion colony are determined to be the initialization ant colony and the initialization ant lion colony, and the sum of the number of executions and the preset threshold is determined as the number of executions, thereby performing the above determination step again. The preset threshold may be a pre-set threshold. For example, the preset threshold may be 1.
[0064] The sixth step is to perform corresponding weighted summation on the energy consumption weight coefficient set and the above-mentioned various energy consumption index information to obtain the device target state information.
[0065] The seventh step is to perform device acquisition state switching processing on the doll according to the target state information to obtain acquisition state switching information.
[0066] The technical scheme and related content thereof serve as one of the invention points of the embodiments of the present disclosure, and solve the second technical problem mentioned in the background. Because the traditional ant lion algorithm adopts random initialization, there is great randomness, and in the ant colony updating process, the updating of the ant colony by the ant lion is determined by the elite ant lion and the roulette selection algorithm, which easily causes the optimization efficiency of the algorithm to be low with the increase of the number of cycles, the algorithm to fall into a local optimal solution, the accuracy of the device acquisition state switching to be low, the device energy consumption of the doll to be increased and the endurance of the doll to be reduced in the case of reducing the acquisition frequency, the accuracy of the doll expression transformation to be reduced in the case of increasing or keeping unchanged the acquisition frequency, the damage rate of the doll to be increased and the accuracy of the expression mapping transformation display to be reduced, and the user interactive experience to be reduced. The factors that cause the damage rate of the doll to be increased and the accuracy of the expression mapping transformation display to be reduced, and the user interactive experience to be reduced are often as follows: because the traditional ant lion algorithm adopts random initialization, there is great randomness, and in the ant colony updating process, the updating of the ant colony by the ant lion is determined by the elite ant lion and the roulette selection algorithm, which easily causes the optimization efficiency of the algorithm to be low with the increase of the number of cycles, the algorithm to fall into a local optimal solution, the accuracy of the device acquisition state switching to be low, the device energy consumption of the doll to be increased and the endurance of the doll to be reduced in the case of reducing the acquisition frequency, the accuracy of the doll expression transformation to be reduced in the case of increasing or keeping unchanged the acquisition frequency. If the above factors are solved, the damage rate of the doll can be reduced, the accuracy of the expression mapping transformation display can be improved, and the user interactive experience can be improved. In order to achieve this effect, the present disclosure first discretizes each energy consumption index information to generate a state switching fitness function, so as to facilitate subsequent determination of the weight coefficients of each energy consumption index information by using the improved ant lion algorithm, so as to more quantitatively balance the energy consumption of the device acquisition state switching and the accuracy of the expression transformation. Then, the ant colony and the ant lion colony are initialized by using the tent chaotic mapping algorithm, which can reduce the randomness of random initialization, enhance the traversal uniformity of the initial weight coefficients, improve the excellence of the ant lion colony and the optimization efficiency of the algorithm. Subsequently, the initialized ant colony is subjected to Levey flight differential position updating according to the target associated ant lion colony and the target ant lion, to obtain an updated ant colony. The long-term short-distance wandering of the Levey flight updating ensures the exploration of the nearby area and enhances the diversity of the population, and the occasional long-distance flight can expand the search range, which is helpful to improve the global search capability. The mutation, crossover and selection in the differential evolution algorithm can improve the optimization progress. Finally, the target ant lion and the associated ant lion are subjected to weighted updating, which can enhance the optimization efficiency of the algorithm at the early stage of the algorithm cycle, accelerate the convergence speed of the algorithm at the later stage of the cycle iteration, avoid the algorithm from falling into a local optimal solution to a certain extent, improve the weight values of each energy consumption index, reduce the damage rate of the doll, improve the accuracy of the expression mapping transformation display, and improve the user interactive experience.
[0067] In step 107, the light-weight emotion recognition model is used to perform emotion recognition on the scene environmental audio to obtain an audio emotion information set.
[0068] In some embodiments, the execution subject can use the light-weight emotion recognition model to perform emotion recognition on the scene environmental audio to obtain an audio emotion information set. The audio emotion information in the audio emotion information set can be information about the emotional tendency expressed by the recognized scene environmental audio. For example, the audio emotion information can be sad emotion information about the crying sound of a child and the high-pitched sound of a pet in the scene environmental audio.
[0069] In some optional implementations of some embodiments, the use of the light-weight emotion recognition model to perform emotion recognition on the scene environmental audio to obtain an audio emotion information set can include the following steps:
[0070] First, the scene environmental audio is input to the deep convolutional separable extraction network included in the light-weight emotion recognition model to obtain an audio channel spectrogram feature map. The light-weight emotion recognition model further includes a plurality of deep separable correction convolution networks, a multi-scale residual fusion network, and an emotion classification and recognition network. The deep convolutional separable extraction network can be a deep neural network model for extracting the emotional information of the input scene environmental audio through a convolutional neural network. The deep convolutional separable extraction network can be a convolutional neural network for extracting the semantic information of the input scene environmental audio through a series of deep separable convolution layers with a convolution kernel of 3*3, a normalization layer, and a ReLU (Rectified Linear Unit) activation function layer. The deep convolutional separable extraction network can ensure the independence of the features between the channels of the scene environmental audio during feature extraction, can reduce the redundant information of the audio channel spectrogram feature map in the feature space, and can to some extent avoid the problem of insufficient feature expression caused by the reduction of the number of parameters.
[0071] Second, the audio channel spectrogram feature map is equally divided in the channel to obtain a first audio channel equal division spectrogram feature map and a second audio channel equal division spectrogram feature map.
[0072] In the third step, the first audio channel averaged spectral feature map and the second audio channel averaged spectral feature map are input into the multiple depthwise separable corrected convolutional networks to obtain the first audio channel spectral corrected feature map and the second audio channel spectral corrected feature map. The depthwise separable corrected convolutional network in the multiple depthwise separable corrected convolutional networks can be a neural network consisting of a deep self-corrected convolutional network, a normalization layer, and a ReLU activation function connected in series. The deep self-corrected convolutional network can be a convolutional neural network with two branches executed in parallel. A branch network of the above-mentioned deep self-correction convolutional network can be inputting the averaged spectrogram feature map of the first audio channel into the average pooling layer for downsampling processing to expand the receptive field to obtain an audio pooling feature map; then, the audio pooling feature map is input into a depthwise separable convolution layer with a convolution kernel of 3*3 to obtain an audio convolution feature map; finally, the above-mentioned audio convolution feature map is upsampled by the same multiple as the average layer to obtain an upsampled feature map, which is added pixel by pixel to the first audio channel spectrogram and then input into the Sigmoid activation function to obtain a first branch spectrogram feature map; then, the above-mentioned first audio channel spectrogram feature map is input into a depthwise separable convolution layer with a convolution kernel of 3*3 for feature extraction to obtain a second branch spectrogram feature map, which is performed in parallel with the above-mentioned steps; finally, the first branch spectrogram feature map and the second branch spectrogram feature map are multiplied pixel by pixel and input into a depthwise separable convolution layer with a convolution kernel of 3*3 to obtain a deep neural network of the first audio channel spectrogram correction feature map. Another branch of the deep self-correcting convolutional network can be a deep neural network that inputs the second audio channel's averaged spectrogram feature map into a depthwise separable convolutional layer with a 3*3 convolution kernel for feature extraction, thereby outputting a spectrogram-corrected feature map for the second audio channel. Each of the multiple depthwise separable correcting convolutional networks is a series of stacked deep neural networks.
[0073] In a fourth step, the first audio channel spectrogram-corrected feature map, the second audio channel spectrogram-corrected feature map, and the audio channel spectrogram feature map are input into the multi-scale residual fusion network to obtain an audio multi-scale feature atlas. The multi-scale residual fusion network may be a deep neural network that first performs channel cascade concatenation on the first audio channel spectrogram-corrected feature map and the second audio channel spectrogram-corrected feature map, and then performs feature concatenation with the audio channel spectrogram feature map via residual connection to obtain a cascaded audio feature map; then, downsampling the cascaded audio feature map via average pooling to obtain the output of the audio multi-scale feature atlas.
[0074] In the fifth step, the multi-scale audio feature map set is input into the emotion classification and recognition network to obtain an audio emotion information set. The emotion classification and recognition network can be a deep neural network that inputs the input multi-scale audio feature map into a fully connected layer and then into a support vector machine for emotion classification.
[0075] It should be noted that the multiple deep separable rectified convolutional networks in the above-mentioned lightweight emotion recognition model adaptively construct the dependency relationship between input and output through self-correction operations, thereby allowing each spatial position to adaptively encode contextual information of distant areas, breaking the tradition of performing convolution in small areas. At the same time, the above-mentioned multi-scale residual fusion network can effectively utilize semantic information at different levels by extracting and fusing convolutional features in different scale spaces, thereby avoiding interference from irrelevant areas in the entire global information.
[0076] Step 108 : performing expression mapping on the electronic doll expression of the doll according to the current device scene information set and the audio emotion information set to obtain current doll expression information.
[0077] In some embodiments, the execution entity may perform expression mapping on the electronic doll expression of the doll based on the current device scene information set and the audio emotion information set to obtain current doll expression information. The current doll expression information may be an expression that matches the emotion in the current scene obtained by transforming the doll expression using the current device scene information set and the audio emotion information set and then rendering it in real time. As an example, the execution entity may utilize reinforcement learning to perform expression mapping on the doll expression of the doll based on the current device scene information set and the audio emotion information set to obtain current doll expression information.
[0078] In some optional implementations of some embodiments, performing expression mapping on the electronic doll expression of the doll based on the current device scene information set and the audio emotion information set to obtain current doll expression information may include the following steps:
[0079] The first step is to construct a feature hierarchy tree for the current device scene information set and the audio emotion information set, respectively, to obtain a scene feature hierarchy tree and an audio feature hierarchy tree. The scene feature hierarchy tree can be a tree structure that organizes the current device scene information set into a multi-level structure based on semantic relevance, with increasingly detailed subdivisions as the tree descends. The audio feature hierarchy tree can be a tree structure that organizes the audio emotion information set into a multi-level structure based on semantic relevance. The root node in the scene feature hierarchy tree can be the most abstract scene information strongly related to the business scenario. For example, the root node of the scene feature hierarchy tree can be a scene category. Each intermediate node in the intermediate node set of the scene feature hierarchy tree can be a semantic node that subdivides the root node into interpretable semantic groups. For example, the intermediate node can be, but is not limited to, at least one of the following: environmental scene or human activity scene. The bottom-level node in the scene feature hierarchy tree can be the specific scene information corresponding to the current device scene information set. The bottom-level node can be, but is not limited to, at least one of the following: dog barking, pet noise, laughter, speech, or footsteps.
[0080] In the second step, the current device scene information set and the audio emotion information set are subjected to label standardization processing according to the above-mentioned scene feature hierarchy tree and the above-mentioned audio feature hierarchy tree to obtain a standardized scene information set and a standardized emotion information set. Among them, the standardized scene information in the above-mentioned standardized scene information set can be information obtained by splicing the above-mentioned current device scene information set into a unified string using a separator, or it can be information obtained by querying the above-mentioned scene feature hierarchy tree and performing an upward merge operation if there is fuzzy classification in the above-mentioned current device scene information set. The standardized emotion information in the above-mentioned standardized emotion information set can be information obtained by converting the model emotion prediction score corresponding to the above-mentioned audio emotion information into a numerical value in the range of [0, 1], or it can be emotion information obtained after label standardization processing through emotion coverage priority rule information. For example, the above-mentioned audio emotion information set includes both laughter and short crying. Through the emotion coverage priority rule information, laughter and crying are covered and summarized as laughter. The above-mentioned emotion coverage priority rule information can be the audio emotion information corresponding to the largest model emotion prediction score covered by multiple model emotion prediction scores whose score difference in the model emotion prediction scores is above the preset prediction score difference threshold, or it can be the emotion information obtained by weighted summation of multiple model emotion prediction scores that are not above the preset prediction score difference threshold.
[0081] The third step is to use the puppet expression mapping engine to perform expression mapping on the standardized scene information set and the standardized emotion information set to obtain expression mapping information. The expression mapping information may be obtained by transforming the previous puppet expression information after performing emotion-scene fusion mapping on the standardized scene information set and the standardized emotion information set. The puppet expression mapping engine may be a multi-level decision engine, with the bottom layer being a mapping engine that utilizes a preset expression mapping rule information set and performs parameter fusion mapping of the standardized scene information set and the standardized emotion information set, while the middle layer utilizes context-aware dynamic optimization and high-level energy-adaptive output scheduling. In practice, the execution entity may determine that the model scene prediction score for the transformed scene information set is 0.78 for applause and the model emotion prediction score for the standardized emotion information set is 0.82 for laughter. The bottom layer may determine that the expression is happiness by searching the preset expression mapping rule information set, and determine that the product of 0.82 and 0.78 is 0.64, which serves as the intensity of happiness, to obtain the current initial expression information. Then, the above-mentioned middle layer, based on the expression similarity between the previous doll expression and the current initial expression information, responds to determining that the expression similarity is less than or equal to the preset expression similarity threshold, smoothly transitions the previous doll expression to the current initial expression information. In response to determining that the expression similarity is greater than the preset expression similarity threshold, the previous doll expression is transitioned to the current initial expression information through transition expression information. Among them, the above-mentioned preset expression similarity threshold can be a critical value of a pre-set expression transition method. For example, the above-mentioned preset expression similarity threshold can be 0.6. The above-mentioned transition expression information can be a surprised expression information. Finally, the way in which the above-mentioned current initial expression information is output is determined through the above-mentioned acquisition state switching information. For example, the above-mentioned acquisition state switching information is active state information, and a happy electronic expression (raised eyebrows + raised corners of the mouth, and the degree of both is 0.64) is fully displayed. The above-mentioned acquisition state switching information is sleeping state information, and the above-mentioned doll's happy expression is simplified and output.
[0082] In a fourth step, in response to determining that the expression mapping information representation is not successfully matched, fuzzy mapping matching is performed on the standardized scene information set and the standardized emotion information set to obtain expression fuzzy mapping matching information. The expression fuzzy mapping matching information may be expression information obtained by upward generalization and merging of the scene feature hierarchical tree and the audio feature hierarchical tree to the upper-level nodes for fuzzy mapping matching, if the expression is not accurately matched in the bottom-level node set of the scene feature hierarchical tree and the audio feature hierarchical tree.
[0083] The fifth step, in response to determining that the above-mentioned expression fuzzy mapping matching information representation is not successfully matched and the above-mentioned doll communication is normal, the above-mentioned standardized scene information set and the standardized emotion information set are dynamically fitted and predicted to obtain the expression prediction confidence. Among them, the expression prediction confidence in the above-mentioned expression prediction confidence set can be the probability value of the expression classification output by the deep neural network model. In practice, the above-mentioned execution entity can first perform word embedding processing and then splicing processing on the above-mentioned standardized scene information set and the standardized emotion information set to obtain a scene emotion splicing feature vector. Then, the above-mentioned scene emotion splicing feature vector is input into the lightweight expression mapping model to obtain the expression prediction confidence set. Among them, the above-mentioned lightweight expression mapping model can be an MLP (Multilayer Perceptron) located in the cloud, which performs expression mapping feedback on the input standardized scene information set and the standardized emotion information set.
[0084] In the sixth step, the expression information corresponding to the expression prediction confidence level and the historical expression mapping information set are fused at the decision-making level to obtain the target doll expression information. The decision-making level fusion can be a fusion process in which, if the expression prediction confidence level is greater than or equal to a preset confidence threshold, the doll expression corresponding to the expression prediction confidence level is determined as the target doll expression information; otherwise, the previous doll expression information is determined as the target doll expression information. The target doll expression information can be the expression information fused with the historical doll expression. The historical expression mapping information set can be the expression information of the doll that appeared before the current information. The historical expression mapping information set can verify the rationality of the previous doll expression to the target doll expression information based on prior knowledge. It should be noted that the expression mapping process simultaneously considers the doll's state information, the smooth transition of expression, and the communication level between the doll and the cloud. This ensures that the doll's expression changes are both responsive and natural and coherent. Furthermore, it can also ensure the accuracy and timeliness of expression mapping in a progressive manner while ensuring the doll's energy consumption.
[0085] In step 7, in response to detecting the user interaction information, dynamically adjust the target doll's expression information to obtain adjusted doll expression information as the current doll's expression information. The user interaction information may be voice information exchanged between the doll's owner and the doll. In practice, in response to detecting the user interaction information, dynamically weighted fusion is performed on the target doll's expression information and the expressions included in the user interaction information to obtain the current doll's expression information. The dynamic weighted fusion may be performed by dynamically fusion of the expression intensity included in the user interaction information and the target doll's expression information.
[0086] Step 109 : Control the doll to change its expression on the device display screen according to the current doll expression information and the acquisition state switching information.
[0087] In some embodiments, the execution entity may control the doll to change its expression on the device display screen based on the current doll expression information and the acquisition status switching information. The doll expression change display may be based on the doll's device energy consumption information and acquisition status switching information. For example, the device energy consumption information and acquisition status switching information may indicate an active state, with a battery level greater than or equal to 50%, and the expression information is fully rendered and displayed.
[0088] As an example, the execution subject may render the current doll expression information to obtain rendered expression information, and then control the doll to update and display the doll expression on the device display screen.
[0089] Optionally, after step 109, in response to determining that the current device scene information set is different from the previous device scene information, the execution subject performs device collection state switching processing on the doll according to the device energy consumption information, and after obtaining the collection state switching information, the method may further include the following steps:
[0090] The first step is to perform emotion recognition on the scene environment audio to obtain scene audio emotion information. The scene audio emotion information can be information obtained by recognizing emotional semantic information in the scene environment audio.
[0091] The second step is to perform doll expression mapping on the doll's screen expression based on the scene audio emotion information to obtain the current doll scene expression information. The specific implementation of the doll expression mapping process can refer to the implementation steps of steps 1 to 7 above, and the doll expression mapping process in steps 1 to 7 is performed with the weight of the current device scene information set to 0. The current doll scene expression information can be the expression information determined solely by mapping the scene audio emotion information.
[0092] The third step is to control the doll to change its expression on the device display screen according to the current doll scene expression information.
[0093] Further references Figure 2 As an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a doll expression replacement display device. These device embodiments are similar to Figure 1 Corresponding to the method embodiments shown, the doll expression changing and displaying device can be specifically applied to various electronic devices.
[0094] like Figure 2As shown, a puppet expression change and display device 200 includes: an acquisition unit 201, a first control unit 202, a determination unit 203, a model lightweight processing unit 204, a scene classification and recognition unit 205, a device acquisition state switching unit 206, an emotion recognition unit 207, an expression mapping unit 208, and a second control unit 209. The acquisition unit 201 is configured to obtain the puppet's current device acquisition state information and device energy consumption information. The first control unit 202 is configured to control the puppet to perform scene audio acquisition based on the current device acquisition state information to obtain scene environment audio. The determination unit 203 is configured to determine the model resource usage information required for deploying and executing the audio scene classification and recognition model and the puppet audio emotion recognition model. The model lightweight processing unit 204 is configured to, in response to determining that the model resource usage information is greater than or equal to the device resource information corresponding to the puppet, perform model lightweight processing on the audio scene classification and recognition model and the puppet audio emotion recognition model to obtain a lightweight scene recognition model and a lightweight emotion recognition model. The scene classification and recognition unit 205 is configured to: use a lightweight scene recognition model to perform scene classification and recognition on the above-mentioned scene environment audio to obtain the current device scene information set of the above-mentioned doll. The device acquisition state switching unit 206 is configured to: in response to determining that the above-mentioned current device scene information set is different from the previous device scene information, according to the above-mentioned device energy consumption information, perform device acquisition state switching processing on the above-mentioned doll to obtain acquisition state switching information. The emotion recognition unit 207 is configured to: use the above-mentioned lightweight emotion recognition model to perform emotion recognition on the above-mentioned scene environment audio to obtain an audio emotion information set. The expression mapping unit 208 is configured to: perform expression mapping on the electronic doll expression of the above-mentioned doll according to the above-mentioned current device scene information set and the above-mentioned audio emotion information set to obtain current doll expression information. The second control unit 209 is configured to: control the above-mentioned doll to change the expression display of the doll on the device display screen according to the above-mentioned current doll expression information and the above-mentioned acquisition state switching information.
[0095] It is understandable that the units described in the doll expression changing display device 200 are similar to those described in the reference Figure 1 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the doll expression changing display device 200 and the units included therein, and will not be repeated here.
[0096] Reference below Figure 3 , which shows a structural schematic diagram of an electronic device (eg, an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0097] like Figure 3 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. Various programs and data required for the operation of the electronic device 300 are also stored in the RAM 303. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0098] Typically, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Figure 3 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 3 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0099] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.
[0100] Note that the computer-readable medium or media used to provide the computer program sequence to the computer system can be embedded in a computer program product, which comprises all the respective features, which are provided with the computer program sequence, and which are enumerated above. It is understood that the computer-readable medium or media described herein are included in the computer program product, or are a component of the computer program product. In some embodiments of the disclosure, the computer-readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In some embodiments of the disclosure, a computer-readable storage medium can be any tangible medium that contains, or stores a program for use by or in connection with an instruction execution system, apparatus, or device. In some embodiments of the disclosure, a computer-readable signal medium can include a computer-readable storage medium in baseband or propagated as a carrier wave in a propagated data signal, which contains a computer-readable program code. Such a propagated signal can take a wide variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0101] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0102] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist independently without being assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: obtains the current device acquisition status information and device energy consumption information of the doll; controls the above-mentioned doll to perform scene audio acquisition according to the above-mentioned current device acquisition status information to obtain scene environment audio; determines the model occupancy resource information required for the deployment and execution of the audio scene classification and recognition model and the doll audio emotion recognition model; in response to determining that the above-mentioned model occupancy resource information is greater than or equal to the device resource information corresponding to the above-mentioned doll, performs model lightweight processing on the above-mentioned audio scene classification and recognition model and the above-mentioned doll audio emotion recognition model to obtain a lightweight scene recognition model and a lightweight emotion recognition model; utilizes the lightweight scene recognition The scene environment audio is subjected to scene classification and recognition by a recognition model, and the current device scene information set of the doll is obtained; in response to determining that the current device scene information set is different from the previous device scene information, the device acquisition state switching process is performed on the doll according to the device energy consumption information to obtain acquisition state switching information; the scene environment audio is subjected to emotion recognition by using the lightweight emotion recognition model to obtain an audio emotion information set; the electronic doll expression of the doll is mapped according to the current device scene information set and the audio emotion information set to obtain the current doll expression information; the doll is controlled to change its expression display on the device display screen according to the current doll expression information and the acquisition state switching information.
[0103] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0105] The units described in some embodiments of the present disclosure may be implemented by software or by hardware. The described units may also be provided in a processor, for example, may be described as: a processor including an acquisition unit, a first control unit, a determination unit, a model lightweight processing unit, a scene classification and recognition unit, a device acquisition state switching unit, an emotion recognition unit, an expression mapping unit, and a second control unit. The names of these units do not, in some cases, constitute a limitation on the units themselves. For example, the acquisition unit may also be described as a "unit for acquiring the current device acquisition state information and device energy consumption information of the doll."
[0106] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0107] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for changing and displaying a doll's expression, comprising: Get the doll's current device collection status information and device energy consumption information; According to the current device collection state information, the doll is controlled to collect scene audio to obtain scene environment audio; Determine the model resource usage information required for the deployment and execution of the audio scene classification and recognition model and the puppet audio emotion recognition model; In response to determining that the resource occupancy information of the model is greater than or equal to the device resource information corresponding to the doll, performing model lightweight processing on the audio scene classification and recognition model and the doll audio emotion recognition model to obtain a lightweight scene recognition model and a lightweight emotion recognition model; Using a lightweight scene recognition model, performing scene classification and recognition on the scene environment audio to obtain a current device scene information set of the doll; In response to determining that the current device scene information set is different from the previous device scene information set, performing device acquisition state switching processing on the doll according to the device energy consumption information to obtain acquisition state switching information; Using the lightweight emotion recognition model, emotion recognition is performed on the scene environment audio to obtain an audio emotion information set; Performing expression mapping on the electronic doll expression of the doll according to the current device scene information set and the audio emotion information set to obtain current doll expression information; According to the current doll expression information and the acquisition state switching information, the doll is controlled to change the expression display on the device display screen.
2. The method according to claim 1, wherein In response to determining that the current device scene information set is different from the previous device scene information, after performing device acquisition state switching processing on the doll according to the device energy consumption information to obtain acquisition state switching information, the method further includes: Performing emotion recognition on the scene environment audio to obtain scene audio emotion information; Performing doll expression mapping processing on the screen expression of the doll according to the scene audio emotion information to obtain current doll scene expression information; According to the current doll scene expression information, the doll is controlled to change the expression display on the device display screen.
3. The method according to claim 1, wherein The method of using the lightweight emotion recognition model to perform emotion recognition on the scene environment audio to obtain an audio emotion information set includes: Inputting the scene environment audio into the deep convolutional separable extraction network included in the lightweight emotion recognition model to obtain an audio channel spectrogram feature map, wherein the lightweight emotion recognition model further includes: multiple deep separable rectified convolutional networks, a multi-scale residual fusion network and an emotion classification and recognition network; Performing equal channel division on the audio channel spectral feature map to obtain a first audio channel equalized spectral feature map and a second audio channel equalized spectral feature map; Inputting the first audio channel averaged spectral feature map and the second audio channel averaged spectral feature map into the multiple depthwise separable rectified convolutional networks to obtain the first audio channel spectral rectified feature map and the second audio channel spectral rectified feature map; Inputting the first audio channel spectrogram correction feature map, the second audio channel spectrogram correction feature map, and the audio channel spectrogram feature map into the multi-scale residual fusion network to obtain an audio multi-scale feature map set; The audio multi-scale feature atlas is input into the emotion classification and recognition network to obtain an audio emotion information set.
4. The method according to claim 1, wherein The step of performing expression mapping on the electronic doll expression of the doll according to the current device scene information set and the audio emotion information set to obtain current doll expression information includes: Constructing feature hierarchical trees for the current device scene information set and the audio emotion information set, respectively, to obtain a scene feature hierarchical tree and an audio feature hierarchical tree; According to the scene feature hierarchical tree and the audio feature hierarchical tree, label normalization processing is performed on the current device scene information set and the audio emotion information set to obtain a normalized scene information set and a normalized emotion information set; Using a doll expression mapping engine, performing expression mapping processing on the standardized scene information set and the standardized emotion information set to obtain expression mapping information; In response to determining that the expression mapping information representation is not successfully matched, performing fuzzy mapping matching processing on the standardized scene information set and the standardized emotion information set to obtain expression fuzzy mapping matching information; In response to determining that the expression fuzzy mapping matching information representation is not successfully matched and the doll communication is normal, performing dynamic fitting prediction on the standardized scene information set and the standardized emotion information set to obtain expression prediction confidence; The expression information corresponding to the expression prediction confidence and the historical expression mapping information set are fused at the decision layer to obtain the target doll expression information; In response to detecting the user interaction information, the target doll expression information is dynamically adjusted to obtain the adjusted doll expression information as the current doll expression information.
5. A doll expression changing display device, comprising: an acquisition unit configured to acquire current device collection status information and device energy consumption information of the doll; A first control unit is configured to control the doll to collect scene audio according to the current device collection state information to obtain scene environment audio; a determination unit configured to determine model occupancy resource information required for deployment and execution of an audio scene classification and recognition model and a puppet audio emotion recognition model; A model lightweight processing unit is configured to, in response to determining that the resource occupied by the model is greater than or equal to the device resource information corresponding to the doll, perform model lightweight processing on the audio scene classification and recognition model and the doll audio emotion recognition model to obtain a lightweight scene recognition model and a lightweight emotion recognition model; A scene classification and recognition unit is configured to perform scene classification and recognition on the scene environment audio using a lightweight scene recognition model to obtain a current device scene information set of the doll; a device acquisition state switching unit configured to, in response to determining that the current device scene information set is different from the previous device scene information, perform device acquisition state switching processing on the doll according to the device energy consumption information to obtain acquisition state switching information; an emotion recognition unit configured to perform emotion recognition on the scene environment audio using the lightweight emotion recognition model to obtain an audio emotion information set; an expression mapping unit configured to perform expression mapping on the electronic doll expression of the doll according to the current device scene information set and the audio emotion information set to obtain current doll expression information; The second control unit is configured to control the doll to change its expression on the device display screen according to the current doll expression information and the acquisition state switching information.
6. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.
7. A computer-readable medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.