Infant crying identification method and device, terminal equipment and storage medium

By combining audio and video data in a multimodal fusion method, sound, facial, and motion features are extracted and fused, solving the accuracy problem of infant crying recognition in multi-infant environments and achieving higher recognition accuracy and reliable early warning.

CN121034355AInactive Publication Date: 2025-11-28SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511546347.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2025-11-28
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing methods for recognizing infant cries cannot accurately identify the cries of specific infants in complex audio environments, especially in environments with multiple infants, leading to a decrease in recognition accuracy.

Method used

A multimodal fusion method is adopted to acquire audio and video data, extract audio features, facial features and motion features, fuse them into a joint feature vector for infant crying classification, and use an infant crying classifier for recognition.

Benefits of technology

It improves the accuracy of infant crying recognition in complex audio environments and provides more reliable early warning information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121034355A_ABST
    Figure CN121034355A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biomedical engineering. The invention discloses an infant crying recognition method and device, terminal equipment and a storage medium, which can improve the accuracy of infant crying recognition in an environment with complex audio. The method comprises the following steps: acquiring audio data and video data of a first object, wherein the audio data comprises sound of the first object and sound of at least one second object; performing sound feature extraction processing on the audio data to obtain a target sound feature; performing facial feature extraction processing on the video data to obtain a target facial feature of the first object, and performing motion feature extraction processing on the video data to obtain a target motion feature of the first object; and fusing the target sound feature, the target facial feature and the target motion feature into a joint feature vector of the first object, and inputting the joint feature vector into an infant crying classifier for infant crying classification processing to obtain a crying recognition result of the first object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical engineering technology. More specifically, this application relates to a method, apparatus, terminal device, and storage medium for recognizing infant crying. Background Technology

[0002] Existing methods for recognizing infant cries primarily employ a single audio analysis pattern. More specifically, they utilize deep learning networks such as convolutional neural networks, recurrent neural networks, or Transformer models to directly learn higher-level audio feature representations from audio data, and then use these feature representations for recognition to obtain the infant's crying result. However, when the audio data contains the voices of multiple infants, existing methods fail to distinguish which voice belongs to the infant being detected, thus failing to accurately identify whether the detected infant is crying. For example, in a hospital's neonatal intensive care unit (NICU), multiple infants are typically in the same unit simultaneously. Multiple infants may cry at the same time. Audio data collected in this complex audio environment usually includes the cries of both the infants being monitored and those not being monitored. Therefore, when using only audio data to identify whether an infant is crying, existing methods for infant cry recognition cannot distinguish which sound in the audio data belongs to the monitored infant, thus failing to accurately identify whether the monitored infant is crying. This reduces the accuracy of infant cry recognition. Therefore, existing technologies need improvement. Summary of the Invention

[0003] The purpose of this application is to provide a method, apparatus, terminal device, and storage medium for recognizing infant cries, which can improve the accuracy of infant cry recognition in complex audio environments. This application is mainly achieved through the following technical solutions: A first aspect of this application provides a method for recognizing infant crying, including: Acquire audio data and video data of a first object, wherein the audio data includes the sound of the first object and the sound of at least one second object; The audio data is subjected to sound feature extraction processing to obtain the target sound features; Facial feature extraction processing is performed on the video data to obtain the target facial features of the first object, and motion feature extraction processing is performed on the video data to obtain the target motion features of the first object. The target sound features, target facial features, and target motion features are fused into a joint feature vector of the first object, and the joint feature vector is input into an infant crying classifier for infant crying classification processing to obtain the crying recognition result of the first object.

[0004] According to one embodiment of this application, the step of performing sound feature extraction processing on the audio data to obtain target sound features includes: The audio data is segmented and windowed to obtain multi-frame windowed signals; Perform sound feature extraction processing on each windowed signal frame to obtain the first sound feature of each windowed signal frame; Perform a Fast Fourier Transform on each windowed signal to obtain the spectral signal of each windowed signal; The spectral signal of each windowed signal is processed to extract sound features, thereby obtaining the second sound feature of each windowed signal. The target sound feature is formed by combining the first and second sound features of all windowed signals.

[0005] According to one embodiment of this application, the step of performing facial feature extraction processing on the video data to obtain the target facial features of the first object includes: A facial landmark detection model is used to extract facial feature points from any frame of the video data to obtain a target facial feature point set. Multiple feature point sets for facial regions are selected from the target facial feature point set, and the aspect ratio of each feature point set for facial regions is calculated to obtain the target aspect ratio of each facial region. The target aspect ratios of all facial regions are used as the target facial features of the first object.

[0006] According to one embodiment of this application, the step of selecting multiple feature point sets of facial regions from the target facial feature point set, calculating the aspect ratio of each feature point set of facial regions to obtain the target aspect ratio of each facial region, and using the target aspect ratios of all facial regions as the target facial features of the first object includes: The feature point set of the left eye is selected from the target facial feature point set, and the aspect ratio of the feature point set of the left eye is calculated to obtain the target aspect ratio of the left eye. The feature point set of the right eye is selected from the target facial feature point set, and the aspect ratio of the feature point set of the right eye is calculated to obtain the target aspect ratio of the right eye. The feature point set of the mouth is selected from the target facial feature point set, and the aspect ratio of the feature point set of the mouth is calculated to obtain the target aspect ratio of the mouth. The aspect ratio of the left eye, the aspect ratio of the right eye, and the aspect ratio of the mouth are used as the target facial features of the first object.

[0007] According to one embodiment of this application, the step of performing motion feature extraction processing on the video data to obtain the target motion features of the first object includes: The displacement of each pixel in the video data between consecutive frames is calculated using the optical flow method to obtain the displacement amplitude of each pixel; A preset model is used to perform human body part region recognition processing on each frame of the video data to obtain multiple target human body part regions in each frame of the image; The average displacement amplitude of all pixels in each target human body part region of each frame image is calculated to obtain the average displacement amplitude of each target human body part region of each frame image. The average displacement amplitude of all target human body parts is calculated and processed according to preset rules to obtain the target motion characteristics of the first object.

[0008] According to one embodiment of this application, the step of using a preset model to perform human body part region recognition processing on each frame of the video data to obtain multiple target human body part regions in each frame includes: An image segmentation model is used to segment human body parts in each frame of the video data to obtain target masks for multiple human body parts in each frame. Based on the target mask of multiple human body parts in each frame of the image, human body part region recognition processing is performed on each frame of the image to obtain multiple target human body part regions in each frame of the image.

[0009] According to one embodiment of this application, the step of using a preset model to perform human body part region recognition processing on each frame of the video data to obtain multiple target human body part regions in each frame includes: The human key point detection method is used to detect head key points and limb key points in each frame of the video data to obtain the target key point set of each frame. Based on the set of target key points in each frame of the image, human body part region recognition processing is performed on each frame of the image to obtain multiple target human body part regions in each frame of the image.

[0010] A second aspect of this application provides a device for recognizing infant crying, comprising: The acquisition module is used to acquire audio data and video data of a first object, wherein the audio data includes the sound of the first object and the sound of at least one second object; The sound feature extraction module is used to perform sound feature extraction processing on the audio data to obtain target sound features; The facial feature and motion feature extraction module is used to perform facial feature extraction processing on the video data to obtain the target facial features of the first object, and to perform motion feature extraction processing on the video data to obtain the target motion features of the first object. The classification module is used to fuse the target sound features, the target facial features, and the target motion features into a joint feature vector of the first object, and input the joint feature vector into an infant crying classifier for infant crying classification processing to obtain the crying recognition result of the first object.

[0011] A third aspect of this application provides a terminal device, including a processor and a memory, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to execute the steps of the infant crying recognition method provided in the first aspect of this application.

[0012] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program that causes a computer to perform the steps of the infant crying recognition method provided in the first aspect of this application.

[0013] The beneficial effects of the embodiments of this application include: This application embodiment employs a multimodal fusion approach to achieve the recognition of infant crying. Specifically, this application embodiment acquires audio data and video data of a first object, the audio data including the voice of the first object and the voice of at least one second object; performs sound feature extraction processing on the audio data to obtain target sound features; performs facial feature extraction processing on the video data to obtain target facial features of the first object, and performs motion feature extraction processing on the video data to obtain target motion features of the first object; fuses the target sound features, target facial features, and target motion features into a joint feature vector of the first object, and inputs the joint feature vector into an infant crying classifier for infant crying classification processing to obtain the crying recognition result of the first object. Compared with existing technologies that only use single-modal recognition methods based on audio data, this application embodiment, based on the audio data, adds visual information (i.e., video data) of the first object from the perspective of increasing the features of the first object. Therefore, this application embodiment can accurately identify whether the first object (which can be understood as the detected infant) is crying without needing to distinguish which sound in the audio data is the first object's (i.e., the infant being detected) voice, thereby improving the accuracy of infant crying recognition in complex audio environments. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 The flowcharts for some embodiments of the infant crying recognition method of this application are shown below; Figure 2 A flowchart of the infant crying recognition method of this application in some other embodiments; Figure 3 A schematic diagram of the infant crying recognition device of this application in some embodiments; Figure 4 This is a schematic block diagram of the terminal device of this application in some embodiments. Detailed Implementation

[0016] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0017] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0018] The terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0019] The terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are expressly listed, but may include other steps or units that are not expressly listed or that are inherent to such process, method, product, or apparatus.

[0020] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The term "and / or" as used in this application includes any and all combinations of one or more of the associated listed items.

[0021] The specific embodiments of this application will be further described below with reference to the accompanying drawings.

[0022] refer to Figure 1 The diagram shown is a flowchart of a method for recognizing infant crying provided in the first aspect of an embodiment of this application. Figure 1 The method for recognizing infant crying includes: S1. Acquire audio data and video data of a first object, wherein the audio data includes the sound of the first object and the sound of at least one second object.

[0023] The first target group is infants and toddlers, which is a general term for babies and toddlers, usually referring to young children aged 0-3 years.

[0024] The second object is also an infant or toddler, which is an infant or toddler in the same indoor space as the first object, excluding the first object.

[0025] The sound of the first object can be the sound of the first object crying, laughing, or speaking. The sound of the second object can be the sound of the second object crying, laughing, or speaking.

[0026] The number of the second object can be one or more.

[0027] The audio data is obtained by recording equipment. The recording equipment can be a microphone or a smart voice recorder. The recording equipment can be configured by someone skilled in the art according to actual needs.

[0028] The video data was obtained by a camera. Each frame of the video data contains the entire first subject.

[0029] S2. Perform sound feature extraction processing on the audio data to obtain the target sound features. Step S2 can also refer to... Figure 2 The "Target Sound Features" step in the process.

[0030] Further, step S2 includes: performing frame segmentation and windowing processing on the audio data to obtain multi-frame windowed signals; performing sound feature extraction processing on each frame windowed signal to obtain a first sound feature of each frame windowed signal; performing fast Fourier transform processing on each frame windowed signal to obtain a spectral signal of each frame windowed signal; performing sound feature extraction processing on the spectral signal of each frame windowed signal to obtain a second sound feature of each frame windowed signal; and combining the first and second sound features of all windowed signals to form the target sound feature.

[0031] The audio data is an audio signal.

[0032] Furthermore, the steps of performing frame segmentation and windowing processing on the audio data to obtain multi-frame windowed signals include: setting a sliding window and a step size; performing frame segmentation processing on the audio data based on the sliding window and the step size to obtain multi-frame sub-signals; and performing windowing processing on each frame sub-signal using a window function to obtain multi-frame windowed signals.

[0033] The sliding window is 16 milliseconds, and the step size is 8 milliseconds. In other embodiments, the specific values ​​of the sliding window and the step size can be set by those skilled in the art according to actual needs.

[0034] The window function is a Hamming window. In other embodiments, the window function may also be a Hamming window, and the specific setting can be determined by those skilled in the art according to actual needs.

[0035] Furthermore, the step of performing sound feature extraction processing on each windowed signal frame to obtain the first sound feature of each windowed signal frame includes: performing short-time energy calculation processing on each windowed signal frame to obtain the target short-time energy of each windowed signal frame; performing zero-crossing rate calculation processing on each windowed signal frame to obtain the target zero-crossing rate of each windowed signal frame; and combining the target short-time energy and the target zero-crossing rate of each windowed signal frame to form the first sound feature of each windowed signal frame.

[0036] Further, the step of extracting sound features from the spectral signal of each windowed signal to obtain the second sound feature of each windowed signal includes: calculating the centroid of the spectral signal of each windowed signal to obtain the target centroid of the spectral signal of each windowed signal; calculating the spectral spread feature of the spectral signal of each windowed signal to obtain the target spread feature of the spectral signal of each windowed signal; calculating the Mel-frequency cepstral coefficients of the spectral signal of each windowed signal to obtain the target Mel-frequency cepstral coefficients of the spectral signal of each windowed signal; calculating the chromaticity feature of the spectral signal of each windowed signal to obtain the target chromaticity feature of the spectral signal of each windowed signal; and combining the target centroid, target spread feature, target Mel-frequency cepstral coefficients, and target chromaticity feature of the spectral signal of each windowed signal to form the second sound feature of each windowed signal.

[0037] The chromaticity feature can be understood as the timbre feature.

[0038] Further, the step of combining the first and second acoustic features of all windowed signals to form the target acoustic feature includes: in all windowed signals, sequentially calculating the first-order difference features of the target short-time energy of the first acoustic features of two adjacent windowed frames in time sequence to obtain multiple first-order difference features of the target; in the multiple first-order difference features of the target, sequentially calculating the second-order difference features of two adjacent first-order difference features in time sequence to obtain multiple first-order difference features of the target; in all windowed signals, sequentially calculating the first-order difference features of the target zero-crossing rate of the first acoustic features of two adjacent windowed frames in time sequence to obtain multiple second-order difference features of the target; in the multiple second-order difference features of the target, sequentially calculating the second-order difference features of the target zero-crossing rate of the first acoustic features of two adjacent windowed frames in time sequence to obtain multiple second-order difference features of the target; in the multiple second-order difference features of the target, sequentially calculating the second-order difference features of the target zero-crossing rate of the first acoustic features of the target; in the multiple second-order difference features of the target, sequentially calculating the second-order difference features of the target zero-crossing rate of the target; in the multiple windowed signals ... zero-crossing rate of the target; in the multiple windowed signals, sequentially calculating the second-order difference features of the target zero-crossing rate of the target zero-crossing rate of the target zero-crossing rate of the target zero-crossing rate of the target zero-crossing rate of the target zero-crossing rate of the target zero-crossing The second-order difference features of two adjacent second-target first-order difference features are calculated sequentially to obtain multiple second-target second-order difference features. In all windowed signals, the first-order difference features of the target centroids of the second sound features of two adjacent windowed frames are calculated sequentially in time order to obtain multiple third-target first-order difference features. Among these multiple third-target first-order difference features, the second-order difference features of two adjacent third-target first-order difference features are calculated sequentially in time order to obtain multiple third-target second-order difference features. In all windowed signals, the first-order difference features of the target diffusion features of the second sound features of two adjacent windowed frames are calculated sequentially in time order to obtain multiple fourth-target first-order difference features. Among these multiple fourth-target first-order difference features, the second-order difference features of the target diffusion features of the second sound features are calculated sequentially in time order to obtain multiple fourth-target first-order difference features. The second-order difference features of two adjacent fourth target first-order difference features are used to obtain multiple fourth target second-order difference features. In all windowed signals, the first-order difference features of the target Mel-frequency cepstral coefficients of the second acoustic features of two adjacent windowed frames are calculated sequentially in time to obtain multiple fifth target first-order difference features. In these multiple fifth target first-order difference features, the second-order difference features of two adjacent fifth target first-order difference features are calculated sequentially in time to obtain multiple fifth target second-order difference features. In all windowed signals, the first-order difference features of the target chromaticity features of the second acoustic features of two adjacent windowed frames are calculated sequentially in time to obtain multiple sixth target first-order difference features. In these multiple sixth target first-order difference features, the second-order difference features of the target chromaticity features of the second acoustic features of two adjacent windowed frames are calculated sequentially in time to obtain multiple sixth target first-order difference features. Calculate the second-order difference features of two adjacent first-order difference features of the sixth targets to obtain multiple second-order difference features of the sixth targets; calculate the statistics of the multiple first-order difference features, the multiple second-order difference features, the multiple third-order difference features, the multiple fourth-order difference features, the multiple fifth-order difference features, and the multiple sixth-order difference features to obtain a first statistic; calculate the statistics of the multiple first-order difference features, the multiple second-order difference features, the multiple third-order difference features, the multiple fourth-order difference features, the multiple fifth-order difference features, and the multiple sixth-order difference features to obtain a second statistic;The first statistic and the second statistic are combined to form the target sound feature.

[0039] The statistics are the mean, variance, and / or median.

[0040] S3. Perform facial feature extraction processing on the video data to obtain the target facial features of the first object, and perform motion feature extraction processing on the video data to obtain the target motion features of the first object. Step S3 can also be referenced... Figure 2 The steps are "Target Facial Features" and "Target Motion Features".

[0041] Further, the step of performing facial feature extraction processing on the video data to obtain the target facial features of the first object includes: using a facial key point detection model to perform facial feature point extraction processing on any frame of the video data to obtain a target facial feature point set; selecting multiple feature point sets of facial parts from the target facial feature point set, and performing aspect ratio calculation processing on the feature point set of each facial part to obtain the target aspect ratio of each facial part, and using the target aspect ratio of all facial parts as the target facial features of the first object.

[0042] The facial landmark detection model is HRNet-R90JT. HRNet stands for High-Resolution Network, meaning a high-resolution network that maintains high spatial accuracy of feature maps throughout the facial landmark detection process. R90 indicates that a ±90° rotation data augmentation strategy was used during training, significantly improving the model's robustness to large head rotations. JT stands for Jointly Trained, meaning the model was jointly trained on a 300-W (300-W refers to a dataset containing facial images of 300 different individuals) adult dataset and an InfAnFace (Infant Annotated Faces) dataset, enabling it to accurately detect facial landmarks in both adults and infants simultaneously. In other embodiments, the facial landmark detection model can also be ASM (Active Shape Models) or DAN (Deep Alignment Network).

[0043] The target facial feature point set consists of 68 feature points on the face of the first object. Each of the 68 feature points is a coordinate point.

[0044] Further, the step of selecting multiple feature point sets for facial regions from the target facial feature point set, and performing aspect ratio calculation on the feature point set for each facial region to obtain the target aspect ratio for each facial region, and using the target aspect ratios of all facial regions as the target facial features of the first object includes: selecting the feature point set for the left eye from the target facial feature point set, and performing aspect ratio calculation on the feature point set for the left eye to obtain the target aspect ratio for the left eye; selecting the feature point set for the right eye from the target facial feature point set, and performing aspect ratio calculation on the feature point set for the right eye to obtain the target aspect ratio for the right eye; selecting the feature point set for the mouth from the target facial feature point set, and performing aspect ratio calculation on the feature point set for the mouth to obtain the target aspect ratio for the mouth; and using the target aspect ratios of the left eye, the right eye, and the mouth as the target facial features of the first object.

[0045] Furthermore, the feature point set of the left eye is the 37th to 42nd feature points in the target facial feature point set.

[0046] Furthermore, the formula for calculating the aspect ratio of the feature point set of the left eye to obtain the target aspect ratio of the left eye is as follows: ; in, It is the aspect ratio of the target in the left eye; It is the coordinate of the first feature point in the feature point set of the left eye, and also the coordinate of the 37th feature point in the feature point set of the target face; It is the coordinate of the second feature point in the feature point set of the left eye, and also the coordinate of the 38th feature point in the feature point set of the target face; It is the coordinate of the third feature point in the feature point set of the left eye, and also the coordinate of the 39th feature point in the feature point set of the target face; It is the coordinate of the fourth feature point in the feature point set of the left eye, and also the coordinate of the 40th feature point in the feature point set of the target face; It is the coordinate of the fifth feature point in the feature point set of the left eye, and also the coordinate of the 41st feature point in the feature point set of the target face; It is the coordinate of the sixth feature point in the feature point set of the left eye, and also the coordinate of the 42nd feature point in the feature point set of the target face. The full English name is Eye Aspect Ratio, which is the aspect ratio of the eyes. The aspect ratio can be used to estimate the degree of opening or closing of the eyes of the first subject.

[0047] Furthermore, the feature point set of the right eye is the 43rd to 48th feature points in the target facial feature point set.

[0048] Furthermore, the formula for calculating the aspect ratio of the right eye by performing aspect ratio calculation on the feature point set of the right eye is as follows: ; in, It is the aspect ratio of the target in the right eye; It is the coordinate of the first feature point in the feature point set of the right eye, and also the coordinate of the 43rd feature point in the feature point set of the target face; It is the coordinate of the second feature point in the feature point set of the right eye, and also the coordinate of the 44th feature point in the feature point set of the target face; It is the coordinate of the third feature point in the feature point set of the right eye, and also the coordinate of the 45th feature point in the feature point set of the target face; It is the coordinate of the fourth feature point in the feature point set of the right eye, and also the coordinate of the 46th feature point in the feature point set of the target face; It is the coordinate of the fifth feature point in the feature point set of the right eye, and also the coordinate of the 47th feature point in the feature point set of the target face; It is the coordinate of the sixth feature point in the feature point set of the right eye, and also the coordinate of the 48th feature point in the feature point set of the target face.

[0049] Furthermore, the feature point set of the mouth is the 49th, 51st, 53rd, 55th, 57th and 59th feature points in the target facial feature point set.

[0050] Furthermore, the formula for calculating the aspect ratio of the mouth by performing aspect ratio calculation on the feature point set of the mouth is as follows: ; in, The target aspect ratio of the mouth; It is the coordinate of the first feature point in the feature point set of the mouth, and also the coordinate of the 49th feature point in the feature point set of the target face; It is the coordinate of the second feature point in the feature point set of the mouth, and also the coordinate of the 51st feature point in the feature point set of the target face; It is the coordinate of the third feature point in the feature point set of the mouth, and also the coordinate of the 53rd feature point in the feature point set of the target face; It is the coordinate of the fourth feature point in the feature point set of the mouth, and also the coordinate of the 55th feature point in the feature point set of the target face; It is the coordinate of the fifth feature point in the feature point set of the mouth, and also the coordinate of the 57th feature point in the feature point set of the target face; It is the coordinate of the sixth feature point in the feature point set of the mouth, and also the coordinate of the 59th feature point in the feature point set of the target face. The full English term is Mouth Aspect Ratio, which is used to accurately assess the degree of mouth opening of the first subject.

[0051] Further, the step of performing motion feature extraction processing on the video data to obtain the target motion features of the first object includes: calculating the displacement of each pixel in the video data between consecutive frames using optical flow to obtain the displacement amplitude of each pixel; performing human body part region recognition processing on each frame of the video data using a preset model to obtain multiple target human body part regions in each frame; calculating the average displacement amplitude of all pixels in each target human body part region of each frame to obtain the average displacement amplitude of each target human body part region in each frame; and calculating the average displacement amplitude of all target human body part regions according to preset rules to obtain the target motion features of the first object.

[0052] Furthermore, the calculation formula for obtaining the displacement amplitude of each pixel in the video data by using optical flow to calculate the displacement of each pixel between consecutive frames can be as follows: ; in, It is the first in the video data The displacement amplitude of each pixel; It is the first in the video data The displacement of each pixel in the X direction; It is the first in the video data The displacement of each pixel in the Y direction; It is the first in the video data The coordinates of each pixel on the X-axis; It is the first in the video data The coordinates of each pixel on the Y-axis.

[0053] In other embodiments, the step of calculating the displacement of each pixel in the video data between consecutive frames using optical flow method to obtain the displacement amplitude of each pixel can also be implemented by existing technology.

[0054] Furthermore, the step of performing human body part region recognition processing on each frame of the video data using a preset model to obtain multiple target human body part regions in each frame includes: performing human body part segmentation processing on each frame of the video data using an image segmentation model to obtain target masks for multiple human body parts in each frame; and performing human body part region recognition processing on each frame based on the target masks for multiple human body parts in each frame to obtain multiple target human body part regions in each frame.

[0055] The image segmentation model is a Mask2Former (Masked-attention Mask Transformer) model. In other embodiments, the image segmentation model can also be an FCN (Fully Convolutional Networks) or a PSPNet (Pyramid Scene Parsing Network), and the image segmentation model can also be set by those skilled in the art according to actual needs.

[0056] The multiple human body parts in each frame of the image include the head and limbs. In other embodiments, the multiple human body parts in each frame of the image may also include body parts, which can be specifically set by those skilled in the art according to actual needs.

[0057] Furthermore, the step of performing human body part region recognition processing on each frame image based on the target mask of multiple human body parts in each frame image to obtain multiple target human body part regions in each frame image includes: taking the target mask of the head in each frame image in the corresponding pixel area of ​​each frame image as the target human body part region corresponding to the head in each frame image; taking the target mask of the limbs in each frame image in the corresponding pixel area of ​​each frame image as the target human body part region corresponding to the limbs in each frame image.

[0058] Furthermore, the formula for calculating the average displacement amplitude of all pixels within each target human body region of each frame image is as follows: ; in, It is the first of each frame of the image The average displacement amplitude of each target human body part region; It is the first of each frame of the image The total number of valid pixels within a target human body part region; It is the first of each frame of the image All valid pixels within the target human body part area.

[0059] Further, the step of calculating the average displacement amplitude of all target human body parts regions according to preset rules to obtain the target motion features of the first object includes: calculating the average displacement amplitude of the target human body parts regions corresponding to the head in all images to obtain a first average value; calculating the average displacement amplitude of the target human body parts regions corresponding to the limbs in all images to obtain a second average value; and summing the first average value and the second average value to obtain the target motion features of the first object.

[0060] In other embodiments, the step of calculating the average displacement amplitude of all target human body parts according to preset rules to obtain the target motion characteristics of the first object can also be implemented in other ways, and can be set by those skilled in the art according to actual needs.

[0061] S4. The target sound features, the target facial features, and the target motion features are fused into a joint feature vector of the first object, and the joint feature vector is input into an infant crying classifier for infant crying classification processing to obtain the crying recognition result of the first object.

[0062] The infant crying classifier can be a pre-trained support vector machine, multilayer perceptron, or decision tree classifier. In other embodiments, the infant crying classifier can also be other classifiers, which can be set by those skilled in the art according to actual needs.

[0063] The term "crying" in this article can be understood as "the sound of crying".

[0064] Through the above-described embodiments, this application adds visual information (i.e., video data) of the first object, based on the audio data, from the perspective of increasing the features of the first object. Therefore, this application can accurately identify whether the first object (which can be understood as the detected infant) is crying without needing to distinguish which sound in the audio data belongs to the first object. This improves the accuracy of infant crying recognition in complex audio environments. This application can provide medical personnel with more reliable early warning information.

[0065] In some implementations, before performing framing and windowing processing on the audio data to obtain multi-frame windowed signals, the step of extracting sound features from the audio data to obtain target sound features further includes: normalizing the audio data. The normalized audio data is then used for framing and windowing processing to obtain the multi-frame windowed signals.

[0066] In some implementations, the step of using a preset model to perform human body part region recognition processing on each frame of the video data to obtain multiple target human body part regions in each frame includes: using a human key point detection method to perform head key point and limb key point detection processing on each frame of the video data to obtain a set of target key points for each frame; and performing human body part region recognition processing on each frame based on the set of target key points for each frame to obtain multiple target human body part regions in each frame.

[0067] The human keypoint detection method can be HRNet (High-Resolution Network). In other embodiments, the human keypoint detection method can be configured by those skilled in the art according to actual needs.

[0068] Furthermore, the step of performing human body part region recognition processing on each frame image based on the target key point set of each frame image to obtain multiple target human body part regions of each frame image includes: taking the area enclosed by all head key points in the target key point set of each frame image as the target human body part region corresponding to the head in each frame image; and taking the area enclosed by all limb key points in the target key point set of each frame image as the target human body part region corresponding to the limb in each frame image.

[0069] In other embodiments, the step of performing human body part region recognition processing on each frame image based on the target key point set of each frame image to obtain multiple target human body part regions of each frame image includes: using all head key points in the target key point set of each frame image to form the target human body part region corresponding to the head of each frame image; and using all limb key points in the target key point set of each frame image to form the target human body part region corresponding to the limb of each frame image.

[0070] In some embodiments, the infant crying recognition method further includes a training method for the infant crying classifier, the training method for the infant crying classifier including: A training dataset and a set of real labels are obtained, wherein each training data point in the training dataset corresponds one-to-one with one of the real labels in the set of real labels, and each training data point is formed by fusing voice training features, facial training features, and motion training features. The target training data is input into the original classifier for infant crying classification to obtain a predicted value. The target training data is any training data point in the training dataset. A loss function is calculated based on the predicted value and the corresponding real label of the target training data. The parameters of the original classifier are adjusted based on the loss function to obtain the infant crying classifier.

[0071] The data format of the voice training features is the same as that of the target voice features; the data format of the facial training features is the same as that of the target facial features; and the data format of the motion training features is the same as that of the target motion features.

[0072] The original classifier can be an untrained support vector machine, multilayer perceptron, or decision tree classifier. In other embodiments, the original classifier can also be other classifiers, which can be set by those skilled in the art according to actual needs.

[0073] The formula for calculating the loss function is as follows: ; in, It is the loss function; It is the length of the training dataset, that is, the total number of all training data; It is the true label corresponding to the target training data, that is, the first label in the training dataset. The true label corresponding to each training data point; It is the predicted value, that is, the predicted value corresponding to the target training data.

[0074] refer to Figure 3 The diagram shown is a schematic block diagram of a device for recognizing an infant's crying, provided in the second aspect of an embodiment of this application. Figure 3 The infant crying recognition device 100 includes: The acquisition module 101 is used to acquire audio data and video data of a first object, wherein the audio data includes the sound of the first object and the sound of at least one second object; The sound feature extraction module 102 is used to perform sound feature extraction processing on the audio data to obtain target sound features; The facial feature and motion feature extraction module 103 is used to perform facial feature extraction processing on the video data to obtain the target facial features of the first object, and to perform motion feature extraction processing on the video data to obtain the target motion features of the first object. The classification module 104 is used to fuse the target sound features, the target facial features and the target motion features into a joint feature vector of the first object, and input the joint feature vector into an infant crying classifier for infant crying classification processing to obtain the crying recognition result of the first object.

[0075] A third aspect of this application provides a terminal device, the schematic diagram of which is as follows: Figure 4As shown. The terminal device includes a processor, memory, network interface, display screen, and temperature sensor connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface of the terminal device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for recognizing infant cries. The display screen can be a liquid crystal display (LCD) or an e-ink display screen, and the temperature sensor is pre-installed inside the terminal device to detect the operating temperature of the internal components.

[0076] Those skilled in the art will understand that Figure 4 The schematic diagram shown is only a partial structural diagram related to the present invention and does not constitute a limitation on the terminal device to which the present invention is applied. The specific terminal device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0077] In some embodiments, this application provides a terminal device, which includes a processor and a memory for storing computer programs. The processor is used to call and run the computer programs stored in the memory to perform the steps of the infant crying recognition method provided in the first aspect of this application.

[0078] A fourth aspect of this application provides a computer-readable storage medium for storing a computer program that causes a computer to perform the steps of the infant crying recognition method provided in the first aspect of this application.

[0079] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0080] The technical features of the above embodiments can be combined without changing the basic principles of this application. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0081] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the patent protection scope of this application should be determined by the appended claims.

Claims

1. A method for recognizing infant crying, characterized in that, include: Acquire audio data and video data of a first object, wherein the audio data includes the sound of the first object and the sound of at least one second object; The audio data is subjected to sound feature extraction processing to obtain the target sound features; Facial feature extraction processing is performed on the video data to obtain the target facial features of the first object, and motion feature extraction processing is performed on the video data to obtain the target motion features of the first object. The target sound features, target facial features, and target motion features are fused into a joint feature vector of the first object, and the joint feature vector is input into an infant crying classifier for infant crying classification processing to obtain the crying recognition result of the first object.

2. The method for recognizing infant crying according to claim 1, characterized in that, The steps for extracting sound features from the audio data to obtain target sound features include: The audio data is segmented and windowed to obtain multi-frame windowed signals; Perform sound feature extraction processing on each windowed signal frame to obtain the first sound feature of each windowed signal frame; Perform a Fast Fourier Transform on each windowed signal to obtain the spectral signal of each windowed signal; The spectral signal of each windowed signal is processed to extract sound features, thereby obtaining the second sound feature of each windowed signal. The target sound feature is formed by combining the first and second sound features of all windowed signals.

3. The method for recognizing infant crying according to claim 1, characterized in that, The steps of performing facial feature extraction processing on the video data to obtain the target facial features of the first object include: A facial landmark detection model is used to extract facial feature points from any frame of the video data to obtain a target facial feature point set. Multiple feature point sets for facial regions are selected from the target facial feature point set, and the aspect ratio of each feature point set for facial regions is calculated to obtain the target aspect ratio of each facial region. The target aspect ratios of all facial regions are used as the target facial features of the first object.

4. The method for recognizing infant crying according to claim 3, characterized in that, The steps of selecting multiple feature point sets for facial regions from the target facial feature point set, calculating the aspect ratio of each feature point set for each facial region to obtain the target aspect ratio of each facial region, and using the target aspect ratios of all facial regions as the target facial features of the first object include: The feature point set of the left eye is selected from the target facial feature point set, and the aspect ratio of the feature point set of the left eye is calculated to obtain the target aspect ratio of the left eye. The feature point set of the right eye is selected from the target facial feature point set, and the aspect ratio of the feature point set of the right eye is calculated to obtain the target aspect ratio of the right eye. The feature point set of the mouth is selected from the target facial feature point set, and the aspect ratio of the feature point set of the mouth is calculated to obtain the target aspect ratio of the mouth. The aspect ratio of the left eye, the aspect ratio of the right eye, and the aspect ratio of the mouth are used as the target facial features of the first object.

5. The method for recognizing infant crying according to claim 1, characterized in that, The steps of performing motion feature extraction processing on the video data to obtain the target motion features of the first object include: The displacement of each pixel in the video data between consecutive frames is calculated using the optical flow method to obtain the displacement amplitude of each pixel; A preset model is used to perform human body part region recognition processing on each frame of the video data to obtain multiple target human body part regions in each frame of the image; The average displacement amplitude of all pixels in each target human body part region of each frame image is calculated to obtain the average displacement amplitude of each target human body part region of each frame image. The average displacement amplitude of all target human body parts is calculated and processed according to preset rules to obtain the target motion characteristics of the first object.

6. The method for recognizing infant crying according to claim 5, characterized in that, The steps of using a preset model to perform human body part region recognition processing on each frame of the video data to obtain multiple target human body part regions in each frame include: An image segmentation model is used to segment human body parts in each frame of the video data to obtain target masks for multiple human body parts in each frame. Based on the target mask of multiple human body parts in each frame of the image, human body part region recognition processing is performed on each frame of the image to obtain multiple target human body part regions in each frame of the image.

7. The method for recognizing infant crying according to claim 5, characterized in that, The steps of using a preset model to perform human body part region recognition processing on each frame of the video data to obtain multiple target human body part regions in each frame include: The human key point detection method is used to detect head key points and limb key points in each frame of the video data to obtain the target key point set of each frame. Based on the set of target key points in each frame of the image, human body part region recognition processing is performed on each frame of the image to obtain multiple target human body part regions in each frame of the image.

8. A device for recognizing infant crying, characterized in that, include: The acquisition module is used to acquire audio data and video data of a first object, wherein the audio data includes the sound of the first object and the sound of at least one second object; The sound feature extraction module is used to perform sound feature extraction processing on the audio data to obtain target sound features; The facial feature and motion feature extraction module is used to perform facial feature extraction processing on the video data to obtain the target facial features of the first object, and to perform motion feature extraction processing on the video data to obtain the target motion features of the first object. The classification module is used to fuse the target sound features, the target facial features, and the target motion features into a joint feature vector of the first object, and input the joint feature vector into an infant crying classifier for infant crying classification processing to obtain the crying recognition result of the first object.

9. A terminal device, characterized in that, include: A processor and a memory, the memory for storing a computer program, the processor for calling and running the computer program stored in the memory to perform the steps of the infant crying recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program that causes a computer to perform the steps of the infant crying recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Infant incubator system and method based on artificial intelligence

    CN113762085A

  • Infant crying detection method and device based on audio and video fusion

    CN114582355A

  • Truck overload detection method based on sound Mel frequency characteristics

    CN117332293A

  • Voiceprint matching method

    CN119763585A

  • Video-based Moire reflection identification method and device, terminal equipment and storage medium

    CN120823550A