Method for detecting crying appeals of infants based on multimodal information

By combining multimodal information fusion of infant heart rate and crying sound signals, using a multi-layer dynamic learning graph convolution neural network, the accuracy and individualization of baby crying appeal information recognition are solved, and more accurate recognition of infant emotional needs is achieved.

CN115240714BActive Publication Date: 2025-07-22ZHENGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210864860.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-21
Publication Date
2025-07-22
Estimated Expiration
2042-07-21

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the demand information of infants crying, especially during individualization and growth.

Method used

The multimodal information fusion method is adopted, combining the baby's heart rate signal and cry sound signal, and the baby's crying appeal information is detected through a multi-layer dynamic learning graph convolution neural network, and the residual neural network and graph learning module are used to extract features, and the self-attention layer is used to classify.

Benefits of technology

Personalized and staged information detection of infant crying appeals has been achieved, which improves the accuracy of identification and the ability to adapt to infant growth, and provides individualized detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115240714B_ABST
    Figure CN115240714B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for detecting the crying appeals of infants based on multi-modal information, and its steps are as follows: First, collect the infant's heart rate signal and crying sound signal, and extract their features to obtain the heart rate signal and speech signal graph node feature vectors; Secondly, use the graph learning module to process the heart rate signal and speech signal graph node feature vectors respectively to obtain the heart rate dynamic learning graph and the speech dynamic learning graph; Then input the heart rate signal dynamic learning graph and the speech signal dynamic learning graph into a multi-layer feature fusion graph convolutional neural network for discrimination, and output the classification result of the infant's crying appeals; Finally, the infant guardian judges and corrects the feedback on the classification result of the infant's crying appeals according to the actual situation. The present invention can combine the heart rate signal and speech information when the infant cries, use the recent crying information of the infant to generate a dynamic learning graph, and realize more accurate identification of the crying appeal information of the infant with individual independence and growth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of biomedical engineering detection, and particularly to a method for detecting infant crying appeals based on multi-modal information. Background Art

[0002] Crying is the most direct way for infants to express their emotional needs. However, the information expressed by an infant's cry is relatively complex and ambiguous, such as hunger, pain, sleepiness, being restricted, and fear. For young parents or even very experienced nannies, it is a very difficult and urgent task to distinguish the emotional appeals expressed by an infant's cry. Therefore, detecting and identifying the appeal information of an infant's cry is of great significance for improving the happiness of parents and the quality of life of infants.

[0003] Newborn infants have instability in managing facial expressions. Identifying emotions through facial expression pictures and then detecting and identifying infant appeals is restricted. Both the changes in an infant's heart rate and the characteristics of crying sounds can, to a certain extent, accurately reflect the infant's current specific emotions. The change of a person's emotional state can directly affect the cardiovascular system. Therefore, heart rate variability (HRV), as a physiological signal, can truly reflect a person's emotional characteristics and is not controlled by subjective consciousness. A convolutional neural network can be used to perform emotional recognition on the effective features of HRV. However, the correlation between a single HRV signal and emotional characteristics is not sufficient to accurately identify the appeal information of an infant. Infant speech is the most convenient information that can reflect the emotional expression characteristics of an infant. Traditional speech emotion recognition has been able to identify a person's emotional information through speech signals, but it has poor robustness in identifying infant crying appeal information. Some research has used infant spectrograms as the feature vectors of a convolutional neural network to classify the cries of infants in three states: pain, hunger, and sleepiness, and achieved good results. However, infants have individual differences. Directly applying the trained model to the recognition of the cry of another infant results in reduced accuracy and poor robustness. Moreover, as infants grow, their language expression ability is also constantly changing, and the characteristics of crying speech will change accordingly. The recognition accuracy of a fixed deep learning model has a large difference in different stages. In addition, generally, the environment where infants are located is noisy, and accurately identifying the characteristics of infant crying signals also poses a challenge. In summary, the existing methods cannot accurately detect and identify infant crying appeal information. Summary of the Invention

[0004] Aiming at the problem that the existing information monitoring system cannot provide relatively accurate infant crying appeal information for infant guardians, the present invention proposes a method for detecting infant crying appeal information based on multi-modal information. The heart rate signal and the crying signal are combined and a multi-layer dynamic learning graph convolutional neural network is used to accurately detect and identify personalized and staged infant crying appeal information, and after the recognition, a reminder is sent to parents or other guardians through a terminal device.

[0005] The technical solution of the present invention is realized as follows:

[0006] A method for detecting the crying appeals of infants based on multimodal information, the steps are as follows:

[0007] Step 1: Use an optical module to collect the infant's heart rate signal, and use an acoustic module to collect the crying sound signal;

[0008] Step 2: Preprocess the collected infant heart rate signal and crying sound signal respectively, and use a residual neural network to extract features from the preprocessed infant heart rate signal and crying sound signal respectively to obtain the heart rate signal graph node feature vector and the speech signal graph node feature vector;

[0009] Step 3: Set the number of infant cries N as the dynamic learning graph size, take each infant cry as a node data, and input the previous N - 1 cries and the Nth cry with the most recent time; use the graph learning module to process the heart rate signal graph node feature vector and the speech signal graph node feature vector respectively to obtain the heart rate dynamic learning graph and the speech dynamic learning graph;

[0010] Step 4: Input the heart rate signal dynamic learning graph and the speech signal dynamic learning graph into a multi - layer feature fusion graph convolutional neural network for discrimination, and output the classification result of the infant crying appeal;

[0011] Step 5: The infant guardian judges the classification result of the infant crying appeal according to the actual situation. If the output information is correct, the detection work of the infant crying appeal information ends this time, and the classification result of this time is returned and stored; if the output information is incorrect, the guardian gives the correct classification result and feeds it back for storage.

[0012] Preferably, the method for preprocessing the collected infant heart rate signal is: use an amplifier, a rectifier, and a Butterworth filter to amplify and filter the collected infant heart rate signal in sequence to complete the preprocessing of the infant heart rate signal.

[0013] Preferably, the method for preprocessing the collected crying sound signal is: use a data - driven adaptive signal analysis method to perform frequency - scale analysis on the collected crying sound signal, remove the low - frequency trend component and the high - frequency noise component to obtain the infant timbre frequency band signal, use the wavelet empirical mode analysis method to perform adaptive signal decomposition on the infant timbre frequency band signal to obtain different frequency band components, and superimpose the normalized different frequency band components to generate the infant timbre signal to achieve the preprocessing of the crying sound signal.

[0014] Preferably, the method of using the graph learning module to process the heart rate signal graph node feature vector and the speech signal graph node feature vector respectively to obtain the heart rate dynamic learning graph and the speech dynamic learning graph is:

[0015] The node feature vectors of the heart rate signal diagram and the node feature vectors of the voice signal diagram are respectively dimensionally reduced to obtain the input X = (x1, x2, …, x n ) ∈ R n×p and the input Y = (y1, y2, …, y n ) ∈ R n×p ;

[0016] The heart rate dynamic learning diagram and the voice dynamic learning diagram are respectively expressed as:

[0017]

[0018]

[0019] Among them, represents the relationship between the heart rate signal nodes x i and x j , represents the relationship between the voice signal nodes y i and y j , ReLU(·) is an activation function that ensures the non-negativity of S ij , |x i P - x j P| and |y i P - y j P| represent the Euclidean distance between two nodes in the P-dimensional space; represents the optimization parameter for scaling the Euclidean distance; and i, j = 1, 2, …, n, and i ≠ j; the structures of the heart rate dynamic learning diagram and the voice dynamic learning diagram are optimized through the following loss function:

[0020]

[0021]

[0022] Among them, represents the structure loss function of the heart rate dynamic learning diagram, represents the structure loss function of the voice dynamic learning diagram, is the regularization term, Q k is the actual data label, is the predicted value of the model.

[0023] Preferably, the method of inputting the heart rate signal dynamic learning diagram and the voice signal dynamic learning diagram into the multi-layer feature fusion graph convolutional neural network for discrimination is:

[0024] The heart rate signal dynamic learning graph and the voice signal dynamic learning graph are respectively input into the multi-layer information aggregation module to obtain the heart rate mapping feature and the voice mapping feature;

[0025] After connecting the heart rate mapping feature and the voice mapping feature, they are input into the self-attention layer for processing, and the softmax activation function is used for prediction to obtain the classification prediction result, and through the semi-supervised loss function Optimize and output the predicted classification result.

[0026] Preferably, the structure of the multi-layer information aggregation module is: information aggregation layer I, information aggregation layer II, graph convolution layer I, graph convolution layer II, graph convolution layer III, graph convolution layer IV, and information splicing layer; the input end of graph convolution layer I is used to receive the input signal, the output end of graph convolution layer I is respectively connected to the input end of information aggregation layer I and the input end of graph convolution layer II, the output end of graph convolution layer II is respectively connected to the input end of information aggregation layer I and the input end of graph convolution layer III, the output end of information aggregation layer I is connected to the input end of information aggregation layer II, the output end of graph convolution layer III is respectively connected to the input end of information aggregation layer II and the input end of graph convolution layer IV, the output end of graph convolution layer IV is respectively connected to the input end of information aggregation layer II and the input end of the information splicing layer, the output end of information aggregation layer II is connected to the input end of the information splicing layer, and the output end of the information splicing layer is used to output the mapping feature.

[0027] Preferably, the semi-supervised loss function is:

[0028]

[0029] where Q i′j′ represents the data label of the corresponding node, and Z ∈ R n×4 represents the prediction result of the model for n nodes.

[0030] Preferably, the total loss function of the detection model is expressed as:

[0031]

[0032] where λ1 and λ2 are both trade-off parameters.

[0033] Compared with the prior art, the beneficial effects produced by the present invention are:

[0034] 1) The present invention fuses the signal features of two different modalities of heart rate and sound through a multi-layer feature fusion module, and classifies the crying appeal information of infants into four types.

[0035] 2) The present invention uses a single baby cry as a graph node, extracts the heart rate and sound characteristics during the baby cry as the graph node feature information respectively, and obtains a dynamic edge through the graph learning block to generate a heart rate or sound dynamic learning graph.

[0036] 3) The present invention allows the user to set the size of the dynamic learning graph. The larger the graph, the more accurate the classification but the longer the time-consuming; the smaller the graph, the faster the classification but the lower the accuracy.

[0037] 4) When initially using this method, except for the baby cry node at that time, other nodes are stored general cry examples without individual characteristics; through multiple uses by the user, a new dynamic learning graph generated by replacing the original cry example with the most recent baby cry example stored each time has both individual characteristics and stage growth characteristics. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 is a flowchart of the present invention.

[0040] Figure 2 is a schematic diagram of the multi-layer feature fusion module architecture of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0042] As Figure 1 shown, the embodiments of the present invention provide a method for detecting baby cry demands based on multi-modal information, and the steps are as follows:

[0043] Step 1: When the baby shows a crying phenomenon, use the light module on the baby bracelet to collect the baby's heart rate signal, and use the sound module on the baby bracelet to collect the crying sound signal;

[0044] Step 2: Preprocess the collected infant heart rate signals and crying sound signals respectively, and use the Residual Neural Network (ResNet-50) to extract features from the preprocessed infant heart rate signals and crying sound signals respectively, obtaining the heart rate signal graph node feature vector and the speech signal graph node feature vector;

[0045] The method for preprocessing the collected infant heart rate signals is as follows: use an amplifier, a rectifier, and a Butterworth filter to amplify and filter the collected infant heart rate signals in sequence to complete the preprocessing of the infant heart rate signals.

[0046] The method for preprocessing the collected crying sound signals is as follows: use a data-driven adaptive signal analysis method to perform frequency scale analysis on the collected crying sound signals, remove the low-frequency trend component and the high-frequency noise component to obtain the infant timbre frequency band signal, use the wavelet empirical mode analysis method to perform adaptive signal decomposition on the infant timbre frequency band signal to obtain different frequency band components, and superimpose the normalized different frequency band components to generate the infant timbre signal to achieve the preprocessing of the crying sound signals.

[0047] Step 3: Set the number of infant cries N as the dynamic learning graph size, take each infant cry as a node data, and input the previous N - 1 cries and the Nth cry with the closest time; use the graph learning module to process the heart rate signal graph node feature vector and the speech signal graph node feature vector respectively to obtain the heart rate dynamic learning graph and the speech dynamic learning graph;

[0048] Perform dimensionality reduction processing on the heart rate signal graph node feature vector and the speech signal graph node feature vector respectively to obtain the input X=(x1,x2,…,x n )∈R n×p of the heart rate dynamic learning graph and the input Y=(y1,y2,…,y n )∈R n×p ;

[0049] The heart rate dynamic learning graph and the speech dynamic learning graph are respectively expressed as:

[0050]

[0051]

[0052] Among them, represents the relationship between the heart rate signal node x i and x j , represents the relationship between the speech signal node y i and y j , ReLU(·) is an activation function that ensures the non-negativity of S ij non-negative, |xi P - x j P| and |y i P - y j P represents the Euclidean distance between two nodes in the P - dimensional space, and both are 1×P - dimensional vectors; represents the optimization parameter for scaling the Euclidean distance, which is a P×1 - dimensional vector; and i, j = 1, 2, …, n, and i ≠ j; optimize the structures of the heart rate dynamic learning graph and the speech dynamic learning graph through the following loss function:

[0053]

[0054]

[0055] Among them, represents the loss function of the heart rate dynamic learning graph structure, represents the loss function of the speech dynamic learning graph structure, is the regularization term, used to improve the generalization ability of the model and avoid overfitting; Q k is the actual data label, is the predicted value of the model.

[0056] Step 4: Input the heart rate signal dynamic learning graph and the speech signal dynamic learning graph into the multi - layer feature fusion graph convolutional neural network for discrimination, and output the classification result of the baby crying appeal;

[0057] Input the heart rate signal dynamic learning graph and the speech signal dynamic learning graph into the multi - layer information aggregation module respectively to obtain the heart rate mapping feature and the speech mapping feature;

[0058] As Figure 2 shown, the structure of the multi - layer information aggregation module is: information aggregation layer I, information aggregation layer II, graph convolutional layer I, graph convolutional layer II, graph convolutional layer III, graph convolutional layer IV, and information splicing layer; the input end of graph convolutional layer I is used to receive the input signal, the output end of graph convolutional layer I is respectively connected to the input end of information aggregation layer I and the input end of graph convolutional layer II, the output end of graph convolutional layer II is respectively connected to the input end of information aggregation layer I and the input end of graph convolutional layer III, the output end of information aggregation layer I is connected to the input end of information aggregation layer II, the output end of graph convolutional layer III is respectively connected to the input end of information aggregation layer II and the input end of graph convolutional layer IV, the output end of graph convolutional layer IV is respectively connected to the input end of information aggregation layer II and the input end of the information splicing layer, the output end of information aggregation layer II is connected to the input end of the information splicing layer, and the output end of the information splicing layer is used to output the mapping feature.

[0059] After connecting the heart rate mapping feature and the voice mapping feature, the combined feature is input into the self-attention layer for processing, and the softmax activation function is used for prediction to obtain the classification prediction result, which is optimized and output through the semi-supervised loss function Optimize and output the predicted classification result

[0060] The semi-supervised loss function is as follows

[0061]

[0062] where Q i′j′ represents the data label of the corresponding node, and Z ∈ R n×4 represents the prediction results of the model for n nodes

[0063] The total loss function of the detection model is expressed as

[0064]

[0065] where λ1 and λ2 are both trade-off parameters

[0066] Step Five: The baby's guardian judges the classification result of the baby's crying appeal according to the actual situation. If the output information is correct, the detection work of the baby's crying appeal information for this time ends, and the classification result for this time is returned and stored; if the output information is incorrect, the guardian gives the correct classification result and feeds it back for storage to improve the system performance

[0067] The specific operation steps are as follows

[0068] 1) When the baby cries Figure 1 the sound module on the baby bracelet first receives the baby's crying voice, and through the wireless transmission device, it prompts the guardian on the guardian's terminal device that the baby is crying

[0069] 2) The light module on the baby bracelet collects the baby's heart rate signal and voice signal, wirelessly transmits the collected signals to the processor end, and at the processor end, the signals are denoised and specific detection and effective information extraction processing are respectively performed on the heart rate signal and the voice signal

[0070] 3) The processor maps the preprocessed feature information to the feature information of the crying graph node for this time, calls the dynamic learning graph generated by the recent 50 crying nodes, and through Figure 2 the multi-modal information is fused through the multi-layer feature fusion block shown and four-class recognition of hunger, pain, sleepiness, and fear is performed, and the relevant result of "the baby may need comfort and company when crying" is prompted on the guardian's terminal device

[0071] 4) The guardian picks up the baby for comfort according to the prompt of the terminal device of the detection system, but finds that the crying state of the baby has not been relieved, indicating that the detection of the baby's crying appeal information by the system this time is incorrect;

[0072] 5) The guardian gradually stabilizes and calms down the baby's mood by feeding, and selects the correct appeal for the baby's crying this time, "The baby is crying because it is hungry", through the terminal device, and the background stores this crying and the result.

[0073] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for detecting the appeals of infant crying based on multi-modal information, characterized in that, The steps are as follows: Step 1: Use an optical module to collect the baby's heart rate signal and use an acoustic module to collect the crying sound signal; Step 2: Preprocess the collected baby's heart rate signal and crying sound signal respectively, and use a residual neural network to extract features from the preprocessed baby's heart rate signal and crying sound signal respectively to obtain the heart rate signal graph node feature vector and the speech signal graph node feature vector; Step 3: Set the number of baby cries N as the dynamic learning graph size, with each baby cry as a node data, and input the previous N - 1 cries and the Nth cry with the most recent time; Use the graph learning module to process the heart rate signal graph node feature vector and the speech signal graph node feature vector respectively to obtain the heart rate dynamic learning graph and the speech dynamic learning graph; The implementation method is: The node feature vectors of the heart rate signal graph and the node feature vectors of the voice signal graph are respectively subjected to dimensionality reduction processing to obtain the input X=(x1, x2, …, x n ) ∈ R n×p of the heart rate dynamic learning graph and the input Y=(y1, y2, …, y n ) ∈ R n ×p ; The heart rate dynamic learning graph and the speech dynamic learning graph are respectively represented as: Among them, represents the relationship between the heart rate signal node x i and x j relationship; represents the voice signal node y i and y j relationship, ReLU(·) is an activation function that ensures the non-negativity of S ij |x i P - x j P| and |y i P - y j P| represent the Euclidean distance between two nodes in the P-dimensional space; represents the optimization parameter for scaling the Euclidean distance; and i, j = 1, 2, …, n, and i ≠ j; the structures of the heart rate dynamic learning graph and the voice dynamic learning graph are optimized through the following loss function: Among them, represents the loss function of the heart rate dynamic learning graph structure, represents the loss function of the voice dynamic learning graph structure, is the regularization term, Q k is the actual data label, is the predicted value of the model; Step 4: Input the heart rate signal dynamic learning graph and the speech signal dynamic learning graph into a multi-layer feature fusion graph convolutional neural network for discrimination, and output the classification result of the baby's crying appeal; Step 5: The baby's guardian judges the classification result of the baby's crying appeal according to the actual situation. If the output information is correct, the detection work of the baby's crying appeal information ends this time, and the classification result of this time is returned and stored; if the output information is incorrect, the guardian gives the correct classification result and feeds it back for storage.

2. The method for detecting the crying appeal of an infant based on multimodal information according to claim 1, wherein The method for preprocessing the collected baby's heart rate signal is: use an amplifier, a rectifier, and a Butterworth filter to amplify and filter the collected baby's heart rate signal in sequence to complete the preprocessing of the baby's heart rate signal.

3. The method for detecting the crying appeal of an infant based on multimodal information according to claim 1, wherein The method for preprocessing the collected crying sound signal is: use a data-driven adaptive signal analysis method to perform frequency scale analysis on the collected crying sound signal, remove the low-frequency trend component and the high-frequency noise component to obtain the baby's timbre frequency band signal, use the wavelet empirical mode analysis method to perform adaptive signal decomposition on the baby's timbre frequency band signal to obtain different frequency band components, and superimpose the normalized different frequency band components to generate the baby's timbre signal to realize the preprocessing of the crying sound signal.

4. The method for detecting the crying appeals of infants based on multimodal information according to claim 1, wherein The method of inputting the heart rate signal dynamic learning graph and the speech signal dynamic learning graph into a multi-layer feature fusion graph convolutional neural network for discrimination is: Input the heart rate signal dynamic learning graph and the speech signal dynamic learning graph into a multi-layer information aggregation module respectively to obtain the heart rate mapping feature and the speech mapping feature; After concatenating the heart rate mapping features and the voice mapping features, the concatenated features are input into the self-attention layer for processing, and the softmax activation function is used for prediction to obtain the classification prediction result, and through the semi-supervised loss function Optimize and output the predicted classification result.

5. The method for detecting the crying appeal of an infant based on multimodal information according to claim 4, wherein The structure of the multi-layer information aggregation module is as follows: Information Aggregation Layer I, Information Aggregation Layer II, Graph Convolution Layer I, Graph Convolution Layer II, Graph Convolution Layer III, Graph Convolution Layer IV, and Information Concatenation Layer; the input end of Graph Convolution Layer I is used to receive the input signal, and the output end of Graph Convolution Layer I is respectively connected to the input end of Information Aggregation Layer I and the input end of Graph Convolution Layer II. The output end of Graph Convolution Layer II is respectively connected to the input end of Information Aggregation Layer I and the input end of Graph Convolution Layer III. The output end of Information Aggregation Layer I is connected to the input end of Information Aggregation Layer II. The output end of Graph Convolution Layer III is respectively connected to the input end of Information Aggregation Layer II and the input end of Graph Convolution Layer IV. The output end of Graph Convolution Layer IV is respectively connected to the input end of Information Aggregation Layer II and the input end of the Information Concatenation Layer. The output end of Information Aggregation Layer II is connected to the input end of the Information Concatenation Layer. The output end of the Information Concatenation Layer is used to output the mapped features.

6. The method for detecting the crying appeals of an infant based on multimodal information according to claim 4, wherein The semi-supervised loss function is as follows: Among them, Q i′j′ represents the data label of the corresponding node, and Z ∈ R n×4 represents the prediction results of the model for n nodes.

7. The method for detecting the crying appeal of an infant based on multimodal information according to claim 6, characterized in that, Total loss function of the detection model It is expressed as: Among them, both λ1 and λ2 are trade-off parameters.

Citation Information

Patent Citations

  • A multimodal speech emotion recognition method based on enhanced residual neural network

    CN109460737A

  • Method and system for providing monitoring service for kids

    KR1020170021216A