Computer-implemented method for providing data for automatic baby cry determination

The method enhances baby cry determination by using high-quality data and personalized parameters, improving the accuracy of automated baby cry identification through acoustic monitoring and transfer learning.

JP7734904B2Active Publication Date: 2025-09-08ZOUNDREAM AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023502740
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-13
Filing Date
2021-07-13
Publication Date
2025-09-08
Estimated Expiration
2041-07-13

AI Technical Summary

Technical Problem

Existing automated methods for identifying the reason why a baby is crying are not sufficiently reliable, and personalization techniques often fail to improve the accuracy of these determinations.

Method used

A computer-implemented method for automatic baby cry determination involves acoustic monitoring, selecting high-quality cry data, establishing personal baby data, and using personalized parameters for accurate cry assessment, which can include transfer learning from a general baby cry detection model to enhance accuracy.

Benefits of technology

The method improves the reliability and accuracy of baby cry determination by utilizing high-quality data and personalized parameters, allowing for more precise identification of the reason for crying, even in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007734904000001
    Figure 0007734904000001
  • Figure 0007734904000002
    Figure 0007734904000002
  • Figure 0007734904000003
    Figure 0007734904000003
Patent Text Reader

Abstract

A computer-implemented method for providing data for automatic baby cry determination is proposed, comprising the steps of acoustically monitoring a baby and providing a corresponding stream of audio data, detecting a cry in the stream of audio data, selecting cry-related data from the audio data in response to detecting the cry, determining personal baby data for personalized cry determination, preparing a determination stage for determination according to the personal baby data, and supplying the cry-related data to the cry determination stage prepared according to the personal baby data. Further, an automatic baby cry determination arrangement is proposed.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the sound of crying babies. [Background technology]

[0002] Newborns literally cry for help whenever they experience discomfort due to more or less serious causes, for example, they are hungry, have difficulty breathing, are tired, want their diaper changed, are in some form of pain, etc. Parents not only need to notice that their baby is crying, but also find the current reason why their baby is crying based on their own experience, their often limited understanding of the baby's signals, and finally, their own intuition.

[0003] This can cause stress for parents for two simple reasons: on the one hand, they need to be listened to immediately whenever their baby cries, and on the other hand, they need to identify the reason, which is a particular problem for first-time parents of a newborn, although more experienced parents often understand that their baby's crying is an indication that they need attention.

[0004] It has been proposed to place an audio transmitter near the cradle to transmit the audio to a receiver near the parent - this solves the first problem, but the second problem of identifying why the baby is crying remains with a simple transmitter / receiver combination. In view of this, several proposals have been made to identify why the baby is crying in an automated manner. For example, smartphones are used as both transmitters and receivers, and a baby cry detection app is installed on one of the smartphones to help identify why the baby is crying. Thus, even when the appropriate hardware is provided, the problem of identifying why the baby is crying remains, since an appropriate app is needed to identify why the baby is crying.

[0005] There are already several proposals in the scientific literature for such identification methods.

[0006] The paper "Harnessing Infant Cry for Swift, Cost-Effective Diagnosis of Perinatal Asphyxia in Low-Resource Settings" by Charles C. Onu proposes that perinatal asphyxia, one of the top three causes of infant mortality in developing countries, can be recognized by a pattern recognition system that models patterns in the cries of known suffocating infants and normal infants. It proposes that cries are sampled, and each cry sample passes through several signal processing stages, at the end of which a feature vector representing the coefficients of the MEL frequency cepstrum is extracted. The recognition process then involves the steps of speech sampling, feature extraction, mean normalization, cross-validation, and training with testing. It is ensured that the feature vectors used all have the same length and sampling rate.

[0007] In the paper "Ubenwa: Cry-based Diagnosis of Birth Asphyxia" by Charles Udeogu, Eyeni Ndiomu, Urbain Kengni, Doina Precup, Guilherme M. Sant'anna, Edward Ali-kor, and Peace Opar, published at the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, the authors propose that cry input samples be segmented, preprocessed, features extracted, a multi-segment classification determined, and then a judgment made regarding the reason for crying.

[0008] The paper "Neural Transfer Learning for Cry-Based Diagnosis of Perinatal Asphyxia" by Charles C. Onu, Jonathan Lebenso, William L. Hamilton, and Doina Precup describes significant changes in the cry patterns of newborns affected by suffocation. The authors hypothesize that model parameters learned from adult speech can serve as a better initial setting (rather than random) for training models on infant speech. The authors also note that the physiological coupling between crying and breathing has long been recognized, and that crying presupposes the function of respiratory muscles. Additionally, it is stated that cry production and breathing are both regulated by the same brain region. The authors propose models and evaluate their robustness in different noise situations, such as children playing, barking dogs, and sirens. The authors also evaluate each model's response to varying lengths of speech data, stating that real-world diagnostic systems need to work with as much data as possible.

[0009] In the paper "Time-frequency analysis in infant cry classification using quadratic time frequency distributions" by J. Saraswathy, M. Hariharan, Wan Khairunizama, J. Sarojini, N. Thiyagar, Y. Sazali, and Shafriza Nisha, published in Biocybernetics and Biomedical Engineering 38 (2018) 634-645, the authors suggest that research on infant crying may provide automated tools for differentiating infant conditions, such as organic disorders, feeding management, sleep management, maternal health, and sensorimotor integration status. The authors mention parameters such as pitch information, noise concentration, spectral energy features, harmonic analysis-based attributes, linear prediction cepstral coefficients, and MEL-frequency cepstral coefficients. They also state that representation of infant cry signals can use time-frequency-based techniques, namely, wavelet packet transform, short-time Fourier transform (STFT), and empirical mode decomposition (EMD). The authors also state that in joint TF analysis, the time and frequency domain representations of a signal can be combined into TF spectral energy content, leading to clearer insights into the properties of multi-component signals. It is suggested that TF spectral energy content can be used to derive salient features that can characterize different patterns of cry signals, highlighting the importance of TF analysis-based methods in classification and detection using multi-component signals, especially for efficiently distinguishing between different cry utterances.

[0010] In the paper "Monitoring Infant's Emotional Cry in Domestic Environments using the Capsule Network Architecture" by M.A. Tugtekin Turan and Engin Erzin, published at Interspeech 2018, September 2-6, 2018, Hyderabad, the authors propose taking spectrogram representations from short segments of audio signals representing a baby's cry as input to a specific deep learning topology. To achieve accurate performance, the authors apply a high-pass FIR filter to remove speech sounds and other low-frequency noise from the signal. The authors argue that baby-cry audio does not have completely continuous characteristics, and accordingly, impulse-like sequences of different sizes or durations are segmented before a voice activity detection algorithm is applied.

[0011] In the paper "A Hybrid System for Automatic Infant Cry Recognition II" by Carlos Alberto Reyes-Garci'a, Sandra E. Barajas, Esteban Tlelo-Cuautle and Orion Fausto Reyes-Galaviz, the authors propose using genetic algorithms and also suggest that automatic infant cry recognition is very similar to the automatic speech recognition process.

[0012] In a review entitled "Acoustic Analysis of Baby Cry" by Rodney Petrus Balandong R, Department of Biomedical Engineering Faculty of Engineering University of Malaya, May 2013, it is stated that there are several approaches to obtaining cry samples.

[0013] In "A review: Survey on automatic infant cry analysis and classification" by Saraswathy Jeyaraman, Hariharan Muthusamy, Wan Khairunizam, Sarojini Jeyaraman, Thiyagar Nadarajaw, Sazali Yaacob5 & Shafriza Nisha, Health and Technology https: / / doi.org / 10.1007 / s12553-018-0243-5, the authors state that the automatic infant cry classification process is a pattern recognition problem similar to automatic speech recognition. The authors report that because silent intervals typically convey little information but increase computational cost, removal or segmentation is a known preprocessing technique in infant cry classification analysis. The authors also mention different crying types such as spontaneous crying during diaper changes, before meals, during soothing, during pediatric evaluation and in pathological conditions such as vena cava thrombosis, meningitis, peritonitis, respiratory arrest, lingual frenulum, IUGR-microcephaly, tetralogy of Fallot, hyperbilirubinemia, gastroschisis, IUGR-respiratory arrest, bovine protein allergy, cardiac complex etc.

[0014] According to the paper "Infant Cries Identification by using Codebook as Feature Matching, and MFCC as Feature Extraction" by MDRenanti et al., published in Journal of Theoretical and Applied Information Technology, I-ESS 1817-31 95, it would be inconvenient if silence were only cut out from the audio data stream at the beginning and end of the audio signal.

[0015] In "Audio Pattern Recognition of Baby Crying Sound Events," published in the Journal of the Audio Engineering Society, Vol. 63, No. 5, May 2015, by Stavros Ntalampiras, a methodology is proposed for distinguishing between five different states: (a) hungry, (b) uncomfortable (needing change), (c) needing to burp, (d) in pain, and (e) needing to sleep. The periodic nature of the involved audio signal is described as burden. The authors consider several groups of acoustic parameters, including perceptual linear prediction parameters, Mel-frequency cepstral coefficients, perceptual wavelet packets, Teager Energy Operator (TEO)-based features, and time-modulation features. Multiple methods, including support vector machines and multilayer perception, for distinguishing between crying sounds are discussed.

[0016] The paper "Automated Baby Cry Classification on a Hospital-acquired Baby Cry Database" by Rodica Ileana Tuduce, Mircea Sorin Rus, Horia Cucu, and Corneliu Burileanu proposes that a baby cry recognition system that can distinguish between different types of baby cries would help parents distinguish between the needs of their particular baby and, at the same time, help parents learn to make such distinctions on their own. The authors test several classifiers but find that most perform worse on actual recordings of baby cries than on cries extracted from carefully selected samples.

[0017] In a paper titled "Infant Cry Analysis and Detection" presented at the 2012 IEEE 27th Convention of Electrical and Electronics Engineers in Israel, Rami Cohen and Yizhar Lavner proposed an algorithm that includes three main stages: a voice activity detector stage, a classification stage, and a post-processing stage for validating the classification stage to reduce negative errors. The algorithm is based on three classification levels at different time scales: the frame level, where each frame (tens of milliseconds) is classified as either "cry" or "no cry" based on its spectral features; sections of several hundred milliseconds; and segments of several seconds, where the final decision is made according to the number of "cry" sections they contain. The multiple time scale analysis and classification levels are said to be aimed at providing a classifier with a very high detection rate while maintaining a low false positive rate. The authors believe that performance evaluation using recordings of infant cries and other natural sounds, such as car engines, car horns, and speech, demonstrates both a high detection rate and robustness in the presence of noise.

[0018] The paper, "An Investigation into Classification of Infant Cries using Modified Signal Processing Methods," by Shubham Asthana, Naman Varma, and Vinay Kumar Mittal, suggests that a baby's cry is a combination of vocalizations, contractile silences, coughs, choking, and interruptions.

[0019] Methods and devices have also been proposed in the patent literature.

[0020] CN103530979A shows a remote baby cry alarm device for hospitals, which includes a baby cry detection module, an alarm planning module, an alarm receiving module and an alarm module, some components are connected by wires while other components are connected wirelessly.

[0021] CN104347066A shows "Infant crying sound recognition method and system based on deep neural network", which proposes to distinguish between pathological and non-pathological conditions by considering recorded crying sounds.

[0022] CN106653001A shows an infant cry recognition method and system. It is stated that the main problem is that only one crying reason is given. A method for recognizing the reason why an infant is crying is proposed, and in this context it is stated that a number of the following features can be extracted and analyzed: average cry duration, cry duration variance, average cry energy, cry energy variance, pitch frequency, mean pitch frequency, maximum pitch frequency, minimum pitch frequency, pitch frequency dynamic range, pitch mean rate of change of frequency, first formant frequency, mean rate of change of first formant frequency, mean first formant frequency, maximum first formant frequency, minimum first formant frequency, first resonant peak frequency dynamic range, second formant frequency, mean rate of change of second formant frequency, mean second formant frequency, second formant frequency maximum, second formant frequency minimum, second resonant peak frequency dynamic range, Mel frequency cepstral parameters, and inverse Mel frequency cepstral parameters. With regard to the pre-processing step, it is proposed that noise reduction be performed on the crying signal to suppress background noise, and that an auto-detection algorithm be used to remove particularly noisy data fragments, thereby improving the signal-to-noise ratio of the crying signal to be extracted into subsequent features. It will be understood that features extracted in accordance with CN106653001A and the methods by which they are extracted may also be used in the context of the present invention. Accordingly, the cited documents are incorporated herein in their entirety by reference.

[0023] CN106653059A discloses a method and system for automatically recognizing infant cries. To identify the reason for a baby's crying, it is suggested that the baby's age at the time of crying and the duration of the cry can help determine the probability of a pathological reason for the cry. Explicit reference to the last feeding time is made regarding the crying time interval. It is also mentioned that performing image analysis of a video capturing the baby's face while recording the baby's crying audio can be beneficial. It is noted that amateur recordings under non-laboratory conditions may reduce the accuracy of the judgment, providing an inaccurate reason for the cry or misleading an inexperienced parent. Implementing the known method as an app on a smartphone is explicitly mentioned.

[0024] CN107591162A presents a pattern matching-based cry recognition method and intelligent care system. It is mentioned that young parents spend more and more time outside the home, but hiring a babysitter is expensive, and therefore, a crying baby may not be dealt with in a timely manner. Given a smartphone, an infant care function is proposed to solve this problem.

[0025] From GB2234840A we see an automatic baby cry detector that automatically generates a sound when it detects that a baby is crying. The sound continues long enough to ensure that the baby is soothed and put to sleep. The cry detector is then muted long enough to ensure that genuine cries of distress are not ignored by the parent.

[0026] US 2008 / 000 3550A1 proposes teaching new parents the meaning of particular cries by storing the sounds of an infant in a playable audio format. The storage medium may be a DVD.

[0027] KR 2008 003 5549A shows a system for notifying a baby's crying to a mobile phone, which automatically calls the mother's mobile phone when a cry is detected.

[0028] From KR 2010 000 466A is known a paediatric diagnostic device which allows early diagnosis of childhood pneumonia and childhood pneumonia through the crying of a child.

[0029] From KR 2011 0113359A a method and device for detecting a baby's cry using frequency and sequential patterns is known.

[0030] A method and system for analyzing a digital voice audio signal associated with a baby's cry is also known from US 2013 / 031 7815 A1, which proposes to determine the special needs of the baby by inputting time-frequency characteristics determined by processing the digital voice signal with a pre-trained artificial neural network.

[0031] From US 2014 / 004 4269 A1 we know an intelligent ambient sound monitoring system. The system is proposed to monitor the ambient sound environment and compare it with predefined sounds, e.g., in terms of frequency signature, amplitude, and duration, to detect important or significant background sounds, such as alarms, car horns, directed voice communications, crying babies, doorbells, telephones, etc. The system is stated to be useful for people who listen to music using headphones to block out ambient sounds.

[0032] US 2019 / 180772 A1 proposes that the audio capture device can store audio data for long or short periods of time and that the audio capture device can wirelessly transmit the audio. It also suggests that a mobile device, such as a smartphone, can be used to record and display crying sounds, and that the accuracy of automatic detection decreases somewhat in unfavorable environments (e.g., noisy environments). It states that the system is more resilient to faults by displaying multiple reasons for crying on the device screen. It states that the classifier can be implemented using a deep neural network. It also proposes performing segmentation and identifying the source of information for each segment. Furthermore, this document considers the relationship between the age of crying and typical duration. It also proposes that the audio stream segmentation process involves a machine learning algorithm to automatically analyze the audio data dataset into labeled time segments that, for example, distinguish the baby to be detected from other children, environmental noise, or silence. However, any such personalization is proposed only for crying detection. It is further noted that vocalization, crying and constant signal / growth sleep sound models can be created for multiple age groups, for example groups each containing babies aged two months apart.

[0033] A method and system for detecting audio events for a smartphone device is known from US 2016 / 036 4963 A1. When an electronic device acquires audio data, it is proposed to divide the audio data into a plurality of audio components, each of which is associated with a respective frequency in a frequency band and includes a series of time windows. The electronic device is then proposed to extract feature vectors from these audio components and classify the extracted feature vectors. In this way, the smartphone device will be able to distinguish between different audio events.

[0034] US 2017 / 017 8667 A1 discloses a technique for robust cry detection using temporal characteristics of acoustic features. It proposes dividing audio data into frames, then determining an acoustic feature vector for each frame, and determining parameters based on each acoustic feature that changes over time corresponding to the frame. Whether the audio matches a predefined audio is then determined based on the parameters. The use of a baby monitor and identifying baby cries is mentioned. It is stated that generating a small number of parameters from a dataset can be an important aspect of using machine learning techniques such as neural networks, and is therefore useful for identifying desired audio. It is stated that the known audio identification device can be embodied in a computer, smartphone, laptop, camera device, consumer electronics, or other device.

[0035] CN 107657963A discloses a cry identification and recognition method suitable for collecting different cry samples and corresponding crying reasons according to different infants to identify the reason for the infant's crying and provide a comparison for better cry recognition. It is noted that babies' cries generally have higher volume and energy than pure background noise. It is also described that a cry database for storing at least one cry sample can be provided, and that additional cry samples can be stored in the database after the cause is identified during use of the device for identifying the cause of the cry. It is also proposed to store additional cry information in the database if the cause of the cry cannot be determined based on the sound sample database.

[0036] CN 107886953A discloses an infant cry translation system based on facial expression and speech recognition. It proposes that a cry microprocessor is used to continuously train and optimize sample feature data in a sample cry database through a learning memory and feedback self-checking function. It proposes that an intensity greater than a threshold is considered to determine whether an audio segment corresponds to a baby's cry.

[0037] CN 109243493A shows a crying baby emotion recognition method based on improved long-term and short-term memory networks. In this context, the long-term and short-term memory networks need to be trained.

[0038] CN 110085216A discloses a method and apparatus for detecting baby cries. This document describes drawbacks in detection techniques for detecting baby cries, including support vector machine learning algorithms, which have low separation accuracy between baby cries and other sounds, and that speech detection is not sufficiently accurate. It proposes performing feature extraction of perceptual linear predictive coefficients to obtain speech features corresponding to speech data within training samples. At least two speech types are provided, and an acoustic model of baby cries is proposed that takes into account the posterior probability that each frame corresponds to a particular speech type.

[0039] From CN 1564 2458A we learn a baby cry detection method that relies on comparison with several stored samples.

[0040] As can be seen, there are multiple ways to identify why a baby is crying, and similarly, multiple different situations can be distinguished, and therefore the above cited documents are incorporated herein in their entirety with respect to methods for identifying cries, in particular with respect to machine learning methods, and also with respect to the different reasons why a baby is crying that can be identified by analyzing the cry.

[0041] However, although much research has been done in the past to identify the reason why a baby is crying from the baby's own cry, and although it has been suggested that several different conditions can be distinguished, the results obtained with practical devices still need to be improved. In this regard, it should be noted that certain conditions are known to have a large influence on the cry characteristics, and therefore different babies will cry differently under similar circumstances.

[0042] In this regard, a master's thesis by Dror Lederman entitled "Automatic Classification of Infant's Cry" examines the physiological function of newborns in relation to the acoustic signature of their cries, comparing histograms for the stationary cry of full-term versus preterm newborns. Other comparisons include, among others, the cries of infants exposed to cocaine in utero versus those not exposed, and the cries of infants with disorders such as metabolic or chromosomal disorders. The authors note that when working with cry signals, the accuracy of automatic segmentation is less critical than with speech / word segmentation, where inaccurate segmentation can result in the loss of important information. The authors also note that age has been found to be a critical parameter in the analysis of cry signals, and that cry features, including fundamental frequency and formants, have been found to change significantly as infants grow, especially during the first few months of life.

[0043] KR20030077489A emphasizes that infants develop rapidly and that cry characteristics, such as race and gender, can categorize different groups of infants. It states that mass-produced machines cannot analyze the individual characteristics of crying infants. It proposes using a local Internet terminal to acquire audio data from crying babies and utilizing an Internet server for analyzing the audio data. It notes that the data can be stored for future use in infant crying research. It also proposes a service method for providing an instantaneous condition analysis service, and details of the infant population can be stored in a database. However, while a determination of why a baby is crying can be based on a large database, it is inconvenient that a connection to a server must be provided, and as a result, characterization of the cry is not possible without the connection.

[0044] From KR 2005 0023812A we learn of a system for analyzing infant cries using a wireless Internet connection. It is proposed to provide a server management system for managing a wireless Internet service system, which in turn provides a wireless Internet terminal infant voice application for a wireless Internet terminal. It is stated that a personalized voice database can be configured and that the information required for the infant voice device application can be changed so that the user always receives an accurate analysis of the cry in accordance with the latest research. However, there is no mention of how the database can best be expanded, nor is there any mention of how changes to the infant voice device application can be achieved particularly efficiently.

[0045] Another device for analyzing infant cries is known from KR 2012 0107382A. It states that if a baby's cry audio frequency distribution information is recognized a minimum number of times within a predetermined period, the cry audio frequency distribution information can be statistically processed to adjust and optimize the cry audio for the specific baby at the location where the device is located. It is proposed that an adult user of the device can verbally state the reason for the baby's crying, and that this utterance is recognized. Therefore, if it is confirmed that the user's utterance is recognized during or within a certain period after the baby is crying, the utterance content can be processed to associate it with a service function related to the crying baby. Such utterances could be "38.5°" or "The diaper is not wet."

[0046] CN 109658953A discloses a baby cry recognition method and device. It states that a cloud server can be provided to which the audio feature vector and collected audio data segments can be sent. When the device is connected to the server, the cloud server can send the latest version of the identification model to the device, and the device can compare and, if the identification model is not the latest version, send its own identification model to the cloud server. Furthermore, if a network connection to the cloud server is unavailable, the audio feature vector can be identified by a locally stored neural network model. Summary of the Invention [Problem to be solved by the invention]

[0047] Accordingly, in the past, it has been proposed to automatically identify the reason why a baby is crying. However, even though it has been proposed in the past that personalization can help identify the reason why a baby is crying, the determinations proposed by automated methods are often not considered to be sufficiently reliable. In view of this, it would be useful to be able to improve automatic cry determination techniques.

[0048] The object of the present invention is to provide novelty for industrial applications.

[0049] This object is achieved by the subject matter claimed in the independent claims. Some preferred embodiments are set out in the dependent claims. [Means for solving the problem]

[0050] According to a first concept, a computer-implemented method for providing data for automatic baby cry determination is proposed, comprising the steps of acoustically monitoring a baby to provide a corresponding stream of audio data, automatically detecting a cry in the stream of audio data, automatically selecting cry data from the audio data in response to detecting the cry, determining parameters from the selected cry data that enable cry determination, establishing personal baby data for personalized cry determination, preparing a determination stage for determination according to the personal baby data, and supplying parameters to the cry determination stage prepared according to the personal baby data.

[0051] The inventors of the present invention have realized that personalized baby cry assessment requires high-quality cry data to be used in the assessment. If the data provided for automatic baby cry assessment is of insufficient quality, the effect of personalization will not be as fully achieved as it would otherwise be, and the quality of the assessment, as estimated, for example, by the percentage of correct assessments, will not improve or will not improve significantly above non-personalized assessments. In contrast, if the quality of the data is sufficiently high, personalization will typically not only be more reliable. Furthermore, personalization usually only needs to be affected at a much later stage; in particular, it is often possible to use the same set of parameters for all babies despite the personalized assessment. This simplifies the assessment.

[0052] Nevertheless, once the correct voice input data has been selected, it may be possible to determine different sets of parameters depending on the personal baby data established, even though very good results can be obtained using the same parameter determination step for all babies.

[0053] While the personal baby data can be established in a variety of different ways, it will be apparent that requesting personal baby data in a personalized manner from a parent or other caregiver prior to a determination is the most preferred method and the easiest to implement. It will also be appreciated that requesting corresponding input from a parent or other caregiver is only necessary during initial setup of the device used to carry out the method and for later updating of some of the input. While establishing personal baby data by requesting input from a parent and / or caregiver is believed to be the most reliable and simple method, it is also possible to identify at least some of the data by cry analysis; for example, a single cry or multiple cries from the same baby can be evaluated, resulting in a personalization of the baby's most likely age, weight, height, or gender.

[0054] High-quality cry data is ensured by acoustically monitoring the baby and automatically selecting relevant cry data from the stream of audio data. The selected data can be separated from the audio data stream, i.e., they can be extracted or marked as cry parts or potential cry parts; for example, if the baby starts crying due to a prolonged loud noise in the environment, and it is not entirely clear whether the audio data belongs to a cry, the corresponding data can be marked as a "potential cry part." Such marking can be different from marking in which a higher confidence is given that the audio data belongs to a cry.

[0055] In this regard, it will be appreciated that baby monitoring is typically performed continuously, preferably such that audio is recorded from the baby's vicinity over an extended period of time. This has various advantages over situations where, for example, a parent only triggers the collection of audio data when they notice that the baby is crying. Monitoring a baby over an extended period of time allows for the use of audio data that includes both crying and non-crying periods. This, in turn, simplifies the consideration of typical background behavior. It should be appreciated that acoustic background characteristics vary in terms of audio level, spectral distribution of noise, and the duration and occurrence of significant background noise due to, for example, dogs barking, car horns, slamming doors, crying of an older child, etc. A clear understanding of such background behavior can aid in the selection of data from the audio stream as crying data, and thus improve the quality of the data provided for personalization decisions.

[0056] For example, if an air conditioning system produces noise in a particular frequency band, in some implementations of baby cry determination, such frequency band should be ignored when determining parameters describing a baby cry. Continuous monitoring of a baby can make it possible to notice the presence of noise in a particular frequency band by viewing audio data obtained during periods when the baby is not crying. Accordingly, in each implementation of baby cry determination that relies on removing noise in a particular frequency band, it is understood that the corresponding frequency band should be ignored and corresponding information can be added to the selected audio data. This is preferable to simply removing the noise-affected frequencies because the remaining frequency bands are known to be generally associated with a baby's cry but are not considered for the particular case. Furthermore, this does not mean that a particular audio data stream needs to undergo (computationally intensive) band filtering; it is sufficient to feed the corresponding information into the parameter determination stage; therefore, for example, rather than assigning a value representing the spectral intensity in each band, such value may be presented as "not available" (N / A). It will be appreciated that if a frequency band is ignored, a different algorithm for cry detection, e.g., using different filter parameters, may be required. It will be appreciated that the cry data can alternatively and / or additionally be selected by selecting frames from a period after the baby begins to cry. However, in some embodiments, it may be preferable not to remove certain frequency bands that are particularly susceptible to noise. In embodiments where the amount of data is not significantly reduced, it will be appreciated that the adverse effects of noise are less pronounced, since more complete information related to the baby's cry is fed into the convolutional neural network, rather than specific parameters, such as formant-related parameters, pitch frequency-related parameters, first performance maxima, etc. (compare in particular the parameters listed below). At the same time, it will be appreciated that the quality of the determination is improved.It will be appreciated that processing more complete information, such as with a convolutional neural network, requires greater computational effort, but this additional computational effort is at least partially compensated for by eliminating optional filtering steps and / or allowing the same processing to be achieved regardless of the specific noise characteristics present. Thus, from a general perspective, the inventors have realized that in analyzing a baby's cry, it is more useful to feed more complete information to an artificial intelligence-based decision stage rather than expending significant computational effort to reduce the complexity of the data fed to the decision stage. However, what can be done is to ensure that such more complete information actually determined is suitable for determining a baby's cry, which can be easily assumed if patterns typically found in crying are identified and separated from the audio data. Accordingly, in a preferred embodiment, selecting the correct audio input for audio evaluation involves identifying audio-related patterns in the data stream and preferably separating such audio-related data from non-audio-related data.

[0057] Although babies typically cry for long periods of time, there are also short periods during which no loud crying is recorded, for example, because the baby needs to breathe. Information associated with these short periods does not need to be completely discarded. In particular, when determining parameters for further evaluation, these short periods are preferably not removed from the audio data stream, as they may also contain useful information. It should be understood that in some cases, the length of time during which no very loud sounds are recorded after the baby begins to cry may provide important clues in determining why the baby is crying. Therefore, for such personalized evaluations, it may be useful to at least include an indication of the length of the audio data. In other cases, it may be useful to at least determine the time, or time tag, at which a crying-related pattern occurs within the audio data.

[0058] However, if cry parameters are determined, it is even more preferable that they be determined from a longer, uninterrupted period, since thus clues can be obtained from the repeated onset of crying, even if the baby is not particularly fussy during the repeated onset of crying.

[0059] When longer, uninterrupted periods of crying are considered, the cry can be isolated by cutting out prolonged pre-cry and / or post-cry noises. It will also be appreciated that babies who are not receiving adequate care can cry for very long periods of time. Therefore, it will be appreciated that the reason why a baby is crying is preferably determined even if the crying is still ongoing. In such cases, if the parent or caregiver does not respond immediately, the determination may be repeated, and if the determination obtained during such a prolonged period of crying varies, a best-hypothesis determination may be made among the different determinations. It will be appreciated that the determination may change over time because the reason why a baby is crying gradually changes over time, for example, because a baby who was previously in pain gradually becomes tired.

[0060] Accordingly, when providing or obtaining data for automatic baby cry determination, the selection, extraction, or identification of cry data involves identifying the time period to be analyzed and / or the frequency bands or frequencies to be analyzed (or not to be analyzed). With regard to omitting frequency bands, it should be explicitly mentioned that bandpass filtering to avoid Nyquist aliasing is not considered "omitting" frequencies. Rather, when omitting frequencies needs to be mentioned, it is understood that the omitted frequencies are lower than the sampling frequency and that the omitting typically affects the digital data. As a result, frequencies can be omitted by ignoring certain frequency bands that are above the lowest processable frequency and below the highest processable frequency. However, while specific filtering is not essential, particularly with respect to embodiments in which cry patterns can be detected based on a spectrogram-like representation and audio processing can be minimized accordingly, it should be noted that it may be preferable to normalize the audio level, for example, in such a way that the normalized maximum audio level is the same for each window. It should be noted that once a cry pattern is identified and isolated within a window, it is possible that the maximum audio level occurring within the window and used as a reference for audio level normalization does not constitute part of the cry pattern. This would be the case, for example, if a door slammed very loudly, causing a baby to cry. Therefore, it is possible to use a cry pattern with a normalized audio level for subsequent cry translation, or to re-normalize the cry pattern if it is preferable to use a cry determination for a different period. It is also noted that in some implementations of cry translation, particularly those using convolutional or neural networks for translation, it is preferable to have cry patterns of standardized length. Therefore, it is possible to add data representing silence, for example, by extending the audio data with corresponding periods of silence, or by adding regions representing silence to the spectrogram, for example, by making them completely black.It should be noted that using both normalized and standardized lengths in cry pattern translation is particularly preferable when using machine learning models, such as those based on convolutional or neural networks, to which spectrogram-like representations are provided as input. It should also be noted that although reference is made many times throughout the description and claims to spectrogram-like representations of audio data, which are segmented windows and / or cry patterns, it is not necessary to use linear spectrogram-like representations of audio data; rather, non-linear spectrogram-like representations of audio data, in particular mel-spectrogram-like and / or lock-spectrogram-like representations, may be used. These non-linear spectrogram-like representations of audio data can be used both for cry pattern translation and for cry pattern identification and separation.

[0061] As mentioned above, personalization does not need to be implemented in a way that follows the preceding steps of detecting cries, selecting cry data, or determining parameters from the selected cry data. This is advantageous because the computational complexity and / or configuration effort of personalization is minimized, and if a personalized determination is not possible for a baby with specific personal data, such as gender, age, height, weight, medical prerequisites, etc., because the peer group is still too small, at least a non-personalized determination can be achieved that is not impaired by insufficient specific data. Note that "similar" peer groups can also be selected, and / or the number of peer groups can be small until the database is sufficiently grown. With regard to personalization, such personalization can be implemented as privatization, which uses separate and different parameters for every individual baby, or as clustering, which determines peer groups or clusters of babies with very similar crying patterns. It will be appreciated that while privatization is possible by specifically training and modeling on baby cries obtained from just one particular baby, privatization can also be achieved by slightly adapting the filter parameters so that they are better suited to a particular baby, for example, by first determining a more general model, for example, based on cries from a peer group or cluster of babies with very similar cry patterns and / or very similar personal data (such as weight, age, height, and gender). This is known as transfer learning, and it should be appreciated that the particular method proposed in this application for providing data for baby cry determination is particularly useful in personalized baby cry determination via transfer learning.

[0062] The above is based on the paper "Neural Transfer Learning for Cry-based Diagnosis of Perinatal Asphyxia" by Charles C. Onu, Jonathan Lebenso, William L. Hamilton, and Doina Precup, which proposes that model parameters learned from adult speech can serve as a better initial setting (rather than random) for training a model on infant speech, and the applicant is unaware of any attempt to personalize baby cry detection by transfer learning from a more general baby cry detection model, specifically in a way where the initial model on which transfer learning is based is not obtained by clustering of database entries, specifically in a fine-grained clustering that distinguishes between different clusters of database entries, e.g., more than 6, 8, 10, 15, 20, for each determined reason for the baby's cry.

[0063] It will also be appreciated that the method of the present invention facilitates the generation of a database of cries from different babies, allowing for clustering with finer distinctions than previously known, such as grouping babies with weight intervals of 500g, 400g, 300g, 200g, or 100g or less; height intervals of 5cm, 4cm, 3cm, 2cm, or 1cm or less; and age intervals of 8 weeks, 6 weeks, 4 weeks, and 2 weeks. Needless to say, any intervals in between may be selected. It will be appreciated that intervals greater than the maximum values ​​indicated for weight, height, or age may result in rather inaccurate personalization and thus may not maximize the use of the high-quality cry data obtainable by the present invention, while the lower limits indicated for weight and height reflect the imprecision of measurements typically observed privately at home, making finer personalization less useful. When clustering is based entirely or partially on each individual baby's personal data, additional parameters, such as the baby's current temperature in full 0.1°C, 0.2°C, or 0.3°C increments, or known medical conditions, may also be taken into account.

[0064] The audio data is sampled at a sampling frequency of, for example, 4 kHz, 8 kHz, 10 kHz, or 16 kHz. The sampling frequency is typically determined based on the frequency content of the baby's cry, the frequency response of the microphone used in baby monitoring, and / or the available computing power and / or bandwidth for uploading the audio data to the cloud and / or server used in automatic baby cry detection. The bandwidth may be updated, for example, to reflect the bandwidth available for uploading the audio data to the cloud. However, while relevant crying information can be found in the frequency range above 8 kHz, recording these frequencies is often difficult, both in terms of the microphones used and their directivity. This becomes even more important at higher frequencies due to unfavorable polar patterns of microphone sensitivity, even when the microphones are sufficiently sensitive at higher frequencies. Therefore, without limiting the present invention, for many users, sampling frequencies up to 8-10 kHz will produce results indistinguishable from those obtained at higher frequencies. The audio signal from the microphone is pre-conditioned, such as by amplification, low-pass and / or band-pass filtering, and then digitized. For further processing of the audio data and / or for transmitting the audio data to a server, cloud, etc., it is preferable to define frames that include a number of samples, in particular a fixed number of samples, such as 64 samples, 128 samples or 256 samples. Although it is not necessary to use a fixed number of frames, or even to use fixed frames at all, in the following reference will often be made to frames, as using frames reduces computational complexity.

[0065] With respect to the parameters determined from the cry data, one or more of the following parameters may be determined: Average cry energy during the current crying event, a sliding average of cry energy over a specific number of consecutive frames, specifically within 2, 4, 8, 16 or 32 frames, and / or over a specific period of time such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; cry duration variance between interruptions during an event; Cry energy distribution over 2, 4, 8, 16 or 32 frames, and / or over a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the current pitch frequency; specifically the pitch frequency averaged over 2, 4, 8, 16 or 32 frames, and / or over a specific time period such as 1, 2, 5, 10, 15 or 30 seconds; the maximum pitch frequency during a cry event and / or over 2, 4, 8, 16 or 32 frames of cry data, and / or over a specific time period such as 1, 2, 5, 10, 15 or 30 seconds; Changes in sliding maximum pitch frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; a minimum pitch frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds;a sliding minimum pitch frequency change during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; the dynamic range of pitch frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; the mean rate of change of pitch frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; the first formant frequency in 2, 4, 8, 16 or 32 frames of a cry event or cry data, and / or during a specific period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds (note that in the context of the present invention, the term formant may relate to the spectral shaping that occurs as a result of the human vocal tract, and reference to formants may refer to peaks in the spectrum, i.e., maxima, and / or to partial harmonics that are amplified by resonance); the mean rate of change of the first formant frequency averaged over 2, 4, 8, 16, or 32 frames of cry data and / or over specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the sliding mean percentage change of the average first formant frequency sliding across 2, 4, 8, 16, or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the mean value of the first formant frequency, averaging over 2, 4, 8, 16 or 32 frames of cry data and / or over specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the maximum value of the first formant frequency in 2, 4, 8, 16 or 32 frames of cry data and / or during a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; the minimum value of the first formant frequency in 2, 4, 8, 16 or 32 frames of cry data and / or during a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; a first resonant peak frequency dynamic range during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the second formant frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the mean rate of change of the second formant frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; second formant frequency average during the cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; The second formant frequency maximum during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the second formant frequency minimum during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; a second resonance during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during a specific period of time such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; Peak frequency dynamic range during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; Mel frequency cepstrum parameters (note that the cepstrum is the result of a series of mathematical operations: a - transform the signal from the time domain to the frequency domain, b - log of the spectral amplitude, c - transform to the quefrency domain, where the last independent variable, the quefrency, actually has a time scale), the parameters are determined for the entire cry event and / or during 2, 4, 8, 16 or 32 frames of cry data and / or during specific periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds etc; and / or Inverted Mel frequency cepstral parameters.

[0066] It should be understood that although the parameters or some of the parameters listed above can be determined for each cry in order to feed pre-computed parameters into the neural network, this is not necessarily necessary. In particular, it is possible to feed the machine learning model a representation of the recorded crying audio containing all relevant information, in which case the machine itself "evaluates" which parameters are actually relevant. An example of such a representation would be the mel spectrogram of the audio.

[0067] It should be noted that where reference is made above to specific periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc., reference may also be made to any other fixed period up to the explicitly mentioned respective time, such as 7 seconds or 28 seconds. This also applies to the other parameters mentioned below. However, it will be understood that the respective times and / or frame numbers are advantageous, with short periods up to 5 seconds being suitable for determining typical parameters in cry data that are well suited for interpreting the cry. A medium length of up to 15 seconds may be useful if some of the information is buried in ambient noise, while a longer period of up to 30 seconds may be useful for identifying the primary reason for the baby's crying if several such reasons coexist, such as the baby being in pain and simultaneously hungry and sleepy.

[0068] It will be understood that not all of the above-mentioned parameters are necessary for the assessment, even for personalized assessment. Conversely, a highly successful assessment can rely on only two or three of the above-mentioned parameters. This is especially true because some of the parameters involve somewhat redundant information, such as the average voice level during the entire cry event, or a sliding average of voice level over 2, 4, 8, 16, 32, or 64 frames. When techniques such as convolutional neural networks are used in personalized assessment, different listed parameters may be optimal for different groups of babies with similar crying patterns due to the same age, ethnicity, gender, weight, baby height, etc. Nevertheless, even then, a set of parameters that are typically common to different babies can be found, so that the overall number of parameters that need to be determined can be somewhat reduced, yet still provide a useful set of parameters for personalized assessment. The same applies to parameters when some frequency bands are unusable due to noise. This can help to keep the computational and configuration effort associated with personalization to a minimum, particularly when one or more relevant descriptors for personalization refer to weight, baby height or age, and using a set of parameters that is most suitable for a range of weights, a range of baby heights and / or a range of ages can still help to obtain very good assessment results when an update of the personalization assessment phase is due or has passed the deadline, or when a non-personalized assessment needs to be performed, for example because a cloud server that is normally addressed for personalization assessment of crying data is currently unavailable.

[0069] It should be emphasized that in the above reference is made to the use of sliding parameters. Techniques such as the use of sliding parameters or cross-correlation techniques are particularly preferred, as in this way the impact of determining the exact onset of a cry in sufficient detail can be reduced.

[0070] However, as mentioned above, it is not absolutely necessary to determine low-dimensional parameters from the audio data for subsequent determination (note that reference to low-dimensional parameters is made, e.g., by using a limited number of parameters; the dimensionality is clearly reduced compared to the situation where audio data of, e.g., 500 ms chunks recorded at 16 bits and 48 kHz is determined). Another possibility rather than calculating low-dimensional parameters is to identify in advance patterns that can normally be associated with a baby's cry, separate them, and then provide the separated parameters to the subsequent determination stage. A particularly easy possibility to implement is to generate spectrograms for chunks of a given length and then search for patterns in those spectrograms that are normally associated with a baby's cry. The advantage is that generating spectrograms is simple and computationally inexpensive, and searching for patterns in those spectrograms that are normally associated with a baby's cry can therefore be easily achieved by common image analysis techniques. Therefore, the necessary steps are particularly easy to implement. In this context, it is preferable to identify and isolate cry patterns having a length of at least 0.25 seconds, and preferably the minimum identified and isolated cry pattern is even longer, specifically at least 0.3 or 0.4 seconds. In practical embodiments, cry patterns having a minimum length of 0.4 seconds can be reliably identified despite any background noise observed in a typical field installation using a spectrogram-like representation of the monitored audio data, particularly by implementing an object localization method to search for cry patterns having a duration of at least 0.4 seconds within a sound-level-normalized Melspectrum-like representation of a 5-second window of the monitored audio data. At the same time, the length of the patterns can also be limited; for example, identified and isolated cry patterns with a length longer than 2 seconds can be clipped or completely excluded, while cry patterns longer than 2 seconds can be accepted. However, each cry is often segmented into multiple distinct cry patterns, each shorter than 2 seconds.It can be understood that using a standardized length of a crying pattern for subsequent translation implies that most of the identified and separated crying pattern needs to be elongated, for example, by a black shoulder on one side, for subsequent translation, which may impair the accuracy of the crying determination. Therefore, the use of crying patterns with a length of 4 seconds or less, preferably 3 seconds or less, and especially 2 seconds or less, is highly preferred. Using shorter crying patterns also helps with privacy concerns, especially if only short crying patterns or their representations are uploaded to the cloud, rather than a continuous audio stream.

[0071] Even when normalizing the lengths of the cry patterns for translation by adding periods of "silence" or corresponding spectrogram patterns, the original lengths are still fed into the convolutional neural network or machine learning model used in determining baby cries. For example, in embodiments that first calculate a probability for each cry pattern to represent a specific reason for a baby's crying, and then determine a general reason for a baby's crying from the collected probabilities, it would be useful to also include the length of each pattern considered.

[0072] In a preferred embodiment, it is proposed that the baby is continuously acoustically monitored and that pre-cry audio data is at least temporarily stored until it is determined that the subsequent audio data is not related to a cry. In this regard, it should be noted that this is applicable both to situations in which low-dimensional parameters are determined from audio data for subsequent baby cry determination, and to situations in which a simple search for patterns that can be associated with baby cries and their separation is performed. For example, if a typical pattern to be searched for in baby cry audio data has a length of, for example, 2 seconds or less, a number of frames corresponding to a 5-second period can be grouped together and processed in such a way that the pattern found therein is particularly simple, for example, by determining a spectrogram. Obviously, this means that data acquired for more than 5 seconds must first be stored for such subsequent analysis. It should also be noted that, since a cry pattern should never be expected to begin and end only within a given period, such as 5 seconds, the stride used should be low so that there is overlap between periods. The step size should preferably be such that each pattern searched for is completely within at least one of the successive windows. With a short step size, this is easy to ensure. While this may result in a situation where some patterns are identified within two subsequent periods, it will be appreciated that it is easy to discard patterns that appear twice; more specifically, timestamp techniques can be easily used to discard such crying patterns that are found twice due to overlapping periods. When implementing crying pattern detection and / or crying pattern separation by an object localization algorithm based on a spectrogram-like representation of audio data, particularly a male spectrogram-like representation of audio data, it should be noted that situations may arise where a pattern is not completely contained within the window into which the underlying audio data was segmented prior to obtaining the spectrogram-like representation. This is easy to notice, since the identified pattern extends to the window boundary.Even if the same cry pattern may be separated somewhat differently within each of a series of overlapping windows, a cry pattern that is incomplete within one window will be found again in another window where the cry pattern is subsequently completed. Therefore, subsequent steps of translation will yield only incomplete cry patterns, typically with low accuracy, so it is preferable to ignore any patterns that extend to the window boundaries. Even if an incomplete cry pattern cannot be identified within a subsequent window, this will rarely have any adverse effect on the cry translation, since in typical situations, multiple cry patterns will be identified and observed so that omitting an incomplete pattern does not result in any significant disadvantage.

[0073] Storing pre-cry data significantly reduces the overall computational effort, as cry detection can be well separated from cry determination without reducing the accuracy of the cry determination. In this regard, it should be noted that cry identification can be simplified and / or performed in a multi-step process. Typically, a baby's cry is significantly louder than any background noise. Preferably, such an increase in audio level is used as the first step of cry detection. Even if cry detection is performed in the cloud, implementing such a first step locally may be preferable, as it significantly reduces the amount of data transferred to the cloud and reduces the overall energy consumed in cry determination. It will be appreciated that implementing the first step of cry detection locally, for example, using a simple comparator, is very easy to implement given that audio level detection is very easy. Therefore, it is not necessary to have a particularly powerful local processor or microcontroller.

[0074] Accordingly, the first important criterion is the absolute audio level of the sample or frame. Rather than using the absolute audio level, the increase in audio level over a given fairly short period of time can also be used, so that adaptation to acceptable ambient noise is automatically achieved. It should be understood that the background / crying distinction can rely on artificial intelligence / neural network filtering techniques, and that in such cases different filters can and preferably are used than those used in the actual crying determination.

[0075] As mentioned above, the corresponding test of the audio level can be performed with very little computational effort, since it only requires a comparison of the current audio level, which is a binary value of the current audio data, against a predetermined or learned threshold. However, background noises such as a dog barking or a door slamming can also result in significantly high audio levels. Therefore, once a significantly high audio level has been detected by comparing the current audio level to a threshold or by detecting a sudden increase in volume, it should still be possible to determine whether the sudden high audio level is associated with a sudden loud background noise or with a baby crying. For this reason, recording audio data before a cry is useful, as it allows for evaluation of audio recorded immediately before that data if the threshold is exceeded. Storing such pre-cry audio data for subsequent evaluation requires significantly less energy than continuously checking for several conditions that, in combination with a sufficiently high probability, indicate that the baby is crying. It should be understood that pre-cry audio data does not need to be stored for an especially long period of time, and therefore a small amount of memory is usually sufficient. With this small amount of memory, new data can be written cyclically over the oldest data. It will also be appreciated that clues indicating that a baby is crying can also be derived from non-acoustic data, for example, from video surveillance of the baby showing movements or showing that the baby's facial expressions are typical for a crying baby.

[0076] One preferred possibility for setting the threshold is to continuously measure the noise level for successive segments of the data stream (e.g., samples or frames), for example, by determining the average value for the frames. Note that this is possible even using an analog implementation. Also, instead of using the average during each segment, and taking into account that the audio level is likely to fluctuate within each segment, the minimum of these fluctuating audio levels can be determined as the background level. This background level can be considered alone or from multiple background levels per segment, and a new, overall background level, such as a sliding average background level, can be determined. The threshold that needs to be exceeded to consider the baby crying can then be determined by taking into account each background level(s), for example, considering only samples that are at least x dB higher than the previous background level, where X is, for example, 6, 12, or 18. However, it is understood that the threshold that needs to be exceeded (or "x" in the example) can be a function of the overall audio level, since it is not reasonable to consider the baby crying even louder if the background level is particularly high, and therefore the baby should not be expected to cry particularly loudly due to background noise. Therefore, if the surrounding environment has a larger background noise, X will typically be smaller. In this context, it is understood that the actual audio level of a baby's cry depends on both the distance from the microphone to the baby and the baby itself. However, the recording microphone can typically be placed at a distance of 1-2 m from the baby. Furthermore, despite some variability, the overall audio level of a baby's cry can be considered to be within a certain useful range, specifically given the audio level resolution achievable even with a conventional, inexpensive digital-to-analog converter. When the background noise is extremely loud, it may be prudent to continuously search for the baby's cry within the audio data stream. This is reasonable, since if audio level cannot be used as the primary clue, other parameters, such as frequency content / spectrograms and the like, should be determined.It should be noted that while the above-described first stage of baby cry identification is preferred, other possibilities may exist, for example using a fixed threshold, and using a threshold that is determined initially or periodically taking into account sound levels during sound events that are positively associated with a particular reason for the baby crying, etc.

[0077] In a preferred embodiment, it is also proposed that a baby's cry, specifically an occurrence of a baby's cry in a continuous acoustic monitoring stream, is detected based on at least one, and preferably at least two, particularly at least three, of the following: a current sound level above a threshold; a current sound level above the average background noise by a given margin; a current sound level in one or more frequency bands above a threshold; a current sound level in one or more frequency bands or at one or more frequencies above the corresponding average background noise by a given margin; a temporal pattern of the sound; and a model including acoustic features not only from the time domain but also from the frequency domain. In other words, the temporal and / or spectral patterns of the sound stream can be established, and a decision can be made taking into account each pattern(s). It should be noted that the sound data can be processed in such a way that a decision related to the pattern can be made using conventional image analysis techniques. If the cry is detected relative to the average background noise, the background noise can be averaged over, for example, the previous 5, 10, 20, 30 seconds, or 1 minute.

[0078] It will be appreciated that multiple conditions can be established that must be met in common to consider a baby crying. For example, a loud noise can be considered a cry only if the spectral distribution of the sound energy corresponds to the typical spectral distribution of the sound energy of a baby crying, and if it is long enough. Because the amount of computation varies for different conditions, it is reasonable to have a multi-step / multi-stage cry discrimination method, with at least a computationally intensive discrimination step running continuously, and remaining and / or additional discrimination steps running only if the continuously running discrimination step indicates that a sound pattern requiring more detailed analysis has been found. In this way, it will be appreciated that energy consumption can be relatively low, which is particularly advantageous when the method is performed entirely, or at least for the initial part, on a battery-powered device. It will also be appreciated that any cry discrimination step can be performed sufficiently fast, even on processing devices that are considered slow in application, such as DSPs, FPGAs, microcontrollers, microprocessors, etc. Accordingly, despite the multi-step cry discrimination approach, latency will be negligible. In other words, this does not result in any noticeable or significant delay in the cry determination. In any event, since there is no direct communication established such as in video or telephone communication between adults, the typical delay will be very small and easily tolerable.

[0079] Taking this into consideration, in a preferred embodiment it is proposed that an identification step or steps requiring a low computational load are performed continuously and the remaining and / or additional identification steps are performed only if the continuously performing identification step(s) indicate that an audio pattern requiring a more detailed analysis has been found, wherein baby cries, in particular occurrences of baby cries in the continuous acoustic monitoring stream are detected based on at least one of the current audio level above a threshold, the current audio level above the average background noise by a given limit, the current audio level in one or more frequency bands above a threshold, the current audio level in one or more frequency bands above the corresponding average background noise by a given limit, the temporal pattern of the audio, the temporal and / or spectral pattern of the audio, and preferably to determine whether a baby cry is present in the audio data stream by multi-step / multi-stage cry identification.

[0080] It will be appreciated that if a baby cry or a more accurate representation of a baby cry or audio that is said to be a baby cry has been detected, several possibilities exist: First, it is possible to evaluate each of several frames or specified periods following the detection of a cry to determine whether the baby is still crying; this can be done, for example, independently of the current audio level; thus, data from periods in which the baby needs to breathe is also analyzed, since breathing sounds can also give important clues as to why the baby is crying. (Note that the number of samples within the period or window in which the search for a cry pattern is performed may differ from the number of samples grouped together within a frame, for example, to obtain a file that can be easily transferred to a cloud server; for example, for file transfer, a certain number of samples should be grouped together in a way that allows for simple error correction, but this number is typically small, as is the number of samples grouped together to form a window.)

[0081] In embodiments where a small number of parameters are extracted from the audio data for determining and / or detecting a cry within an audio stream, a counter may be used that counts up to the minimum number of frames that need to be acquired and analyzed after a cry event or purported cry event is detected. Also, when a period within an audio stream is analyzed to determine whether a pattern has been found, such a period will typically consist of several frames, and therefore a counter should also be provided.

[0082] It is noted that it is preferable to monitor whether the baby is (still) crying, so that loud noises should also be searched for in the further incoming audio data, in parallel with the calculation of parameters suitable for a personalized determination of the reason why the baby is crying and / or in parallel with the identification and isolation of patterns typically associated with a baby crying. Note that with regard to detecting whether crying is continuing, it may not be necessary to require that the average audio level of the recorded frames be significantly louder than the audio level of previous frames, but that the audio level does not fall below a given minimum value. A hysteresis-like behavior can therefore be implemented once the alleged crying audio has been further analyzed.

[0083] If this is done, the counter can be reset whenever it is confirmed that the baby is still crying. This approach ensures that the end phase of the cry is fully recorded if the parent and / or caregiver does not respond to the baby's cry before the crying stops. This can be useful for determining whether a baby who has quieted down should be left alone afterwards. However, it would also be possible to consider only frames in which the baby was confirmed to have been crying during the sample recording. A particularly useful way to achieve this is to search the audio stream for patterns that typically correspond to a baby's cry. Such a search can be achieved, for example, by obtaining a two-dimensional representation of an audio excerpt or period, e.g., by determining a spectrogram showing the frequency content of an audio period (e.g., a 5-second audio period) over time. Searching the audio stream for this pattern that typically corresponds to a baby's cry can then be achieved by training an artificial intelligence model with patterns known to be associated with baby cries in such a way that the model's output is the portion of the spectrogram that best corresponds to a baby's cry. For example, a portion of the audio stream having a likelihood of association with a baby's cry greater than 50%, 60%, 70%, 75%, 80%, 90%, or 95% can be selected. Also, it is not necessary to isolate an actual portion of the spectrogram (or sonograph, voiceprint, or voicegram) or a spectrogram-like two-dimensional representation of the audio stream; rather, specifying the start time of the cry pattern—and, if cry patterns of different lengths are considered in the determination—identifying the time of the cry pattern is sufficient. This is particularly useful if a similar determination of the identified pattern is then achieved in the determination stage, since it may be desirable to use a different resolution for pattern identification and separation than the resolution used for pattern translation. Specifically, both frequency resolution and time resolution may be low for pattern translation. A reduction in computational effort may result if a significant proportion of the audio stream is discarded as not related to a baby's cry.On the other hand, if a significant percentage of the audio stream is selected for translation, it may be more useful from a computational point of view to search for cry patterns using the same temporal and / or frequency resolution as that used later for pattern translation. Note, however, that the windows or periods over which the search for cry-related patterns is accomplished should have a significant overlap, which means that cry-related patterns, especially for short cry-related patterns, may be found both at the end of a preceding period and at the beginning of a following period. This should be taken into consideration when deciding whether it is advantageous to accomplish the search for cry patterns at a different temporal and / or frequency resolution than that used for pattern translation.

[0084] In any case, it is preferable to use timestamps so that the length of any interruptions in the cry, for example because the baby is gasping or taking a breath, can be ascertained so that it can be used to improve interpretation or determination of the cry pattern.

[0085] It may be advantageous to sample and / or analyze the general background, particularly the background where an average sound level is observed, so that short, pulse-like loud sounds, such as a door slamming or a dog barking, do not adversely affect the background analysis. The background analysis may help to establish the most useful parameters for cry determination and may also help to establish a threshold for the first cry detection stage. It should also be understood that certain parameters that may be optimal for assessing the reason a baby is crying in a very quiet environment may not be adequately measured in a real environment due to background noise or because the monitoring microphone would need to be placed too far from the baby. In such cases, otherwise obtainable relevant information may be buried under the noise, and other parameters for determining the reason a baby is crying should be selected.

[0086] From this it can be seen that information related to the acoustic background can be very helpful in establishing optimal parameters and / or for cry discrimination, especially for personalized determination, and should provide very high quality data. It is understood that when a variety of different devices for implementing the method, such as different smartphones, can be used, taking into account variations due to microphone placement and / or microphone characteristics, the selection of parameters should take into account the "stability" of the parameters.

[0087] This can be easily understood for parameters such as the overall audio level of the cry, which varies with the distance between the baby and the monitoring microphone. However, other factors also play a role, such as whether the baby is in a crib, whether the curtains in the room are closed (i.e., whether higher frequencies are more absorbed), what the polar pattern of the microphone sensitivity looks like (e.g., cardioid, hypercardioid, supercardioid, subcardioid, or unipolar), and how it is oriented relative to the baby. It should be understood that this can affect not only a few low-dimensional parameters derived from the audio stream, but can also adversely affect the identification and isolation of crying patterns within the audio stream using spectrograms or spectrogram-like transforms. Therefore, when training an AI model, it is highly preferable to rely on data obtained using more than one setup / microphone, but to use a training set that includes audio samples acquired with a variety of different devices. In particular, it is possible to simultaneously record the same audio with multiple devices to establish a training set. This is also helpful because the multiple devices will not be perfectly synchronized, so that audio patterns associated with the exact same baby's cry will have different onset times within their respective 5-second periods. Furthermore, it will be appreciated that the recorded audio may vary from device to device, even when placed in the exact same location to record the exact same audio. There are many reasons for these variations, such as variations in microphone sensitivity, microphone response, amplifier response used to condition the analog signal before digitization, etc. It should be appreciated that it is even possible to synthesize a training set using multiple devices recording playbacks of several previously recorded baby cries, preferably previously recorded using high-quality microphone placement.

[0088] In particular, when a neural network filter needs to be established for personalized judgment of crying based on only a few parameters, it may be useful to also consider the behavior of the acoustic background. Therefore, it may be advantageous to upload some non-crying background sound patterns to a server so that typical background patterns, especially neural network filter / neural network filter parameters, can also be taken into account in the evaluation and determination. Here, it is generally referred to as a neural network filter or neural network parameters in the present application, but it is understood that reference may also be made to classification, classification model, and the like, and this is not considered to be a difference in the techniques and methods implemented in the terms used to describe such techniques.

[0089] It is particularly preferred if the cry data from which the parameters enabling the cry determination are determined include audio data from the onset of a cry event, in particular audio data from the first 2 seconds of the cry, preferably from the first 1 second of the cry, and particularly preferably from the first 500 ms of the cry. This is easily possible if the determination that a baby is crying is made in an automated manner, and taking into account the onset of a cry event can be useful in the determination because the crying itself increases discomfort, for example, because the baby experiences additional stress due to the fact that they have to wait too long for a response, and / or because the crying itself, if continued for a long period of time, can exhaust the baby. Also, if parameter changes, such as changes in the first frequency of formants, are taken into account as shown above, initial changes can contain particularly valuable information. It should be noted that if the cry pattern is separated from the audio stream period using convolutional or neural networks or other artificial intelligence methods, this is often done using very simple circuitry, such as analog or digital comparators, following an initial audio level determination. In such cases, it is also preferable to assess the actual onset of loud noise, and therefore the audio data should preferably be stored so that loud noise detectable by the comparator is near the end of the window within which the search for the correct parameters is performed, for example a 4, 5, 6, 7, 8, 9 or 10 second window.

[0090] It should be noted that in some cases, it may be advantageous to frequently change the method for determining the baby's cry, for example, by frequently changing the filter coefficients in the neural network filter used to determine the baby's cry, taking into account the baby's growth and development. This may be advantageous, for example, because, very soon after birth, the vocal characteristics of a newborn baby change rapidly, and rapidly changing medical conditions such as high fever may have a strong impact on how the data should be determined, so for useful personalization, the filter coefficients of the neural network filter should also be changed frequently. Depending on the exact implementation of personalization, implementing a local execution of the personalization determination step may not be feasible, as it may require a significant amount of memory for different filter coefficients and / or because currently appropriate filter coefficients need to be identified and downloaded; therefore, the determination is preferably performed from time to time on a centralized server and / or in the cloud.

[0091] Accordingly, a preferred embodiment also proposes implementing all or part of, for example, at least one, cry detection step in audio obtained from a locally acoustically monitored baby and uploading the data to a (cloud) server configuration used in centralized automatic baby cry determination, particularly uploading data for determining the baby's cry to the cloud. Even a simple first step of cry detection, such as a threshold comparator, helps reduce the data stream that needs to be uploaded, saving energy and bandwidth. Note that such a simple first local step of cry detection is therefore preferable even when the actual determination of the baby's cry, or a major part of it, is performed in a cloud server. It is also noted that uploading the audio to a distributed remote cloud server is not necessary; for example, there may be cases where a device located near the baby has very limited computing power and the device used to notify the parent that the baby is crying is a smartphone, which nowadays has at least the computing power to accomplish the determination. In such cases, rather than uploading the audio to a remote cloud server, the audio from the device located near the baby can also be uploaded to the smartphone for further determination, thus mitigating any privacy concerns parents may have.

[0092] While it is clearly preferable to locally determine at least some probability of whether a cry event is currently being recorded in order to save the bandwidth otherwise required to continuously upload audio data to the cry identification stage, it has already been mentioned that improved results can be obtained by taking into account at least some of the baby's characteristics, specifically those most suitable for personalization, such as gender, age, weight, and height, and, even without full personalization, taking into account other personal data such as current body temperature, current medical condition, and the time since the baby was fed. Because age, weight, and height change only slowly, personalization determination can also be implemented locally, especially for situations where uploading audio data is impaired. Even in such situations, it will be appreciated that establishing a connection between the device recording the baby's audio and a (cloud) server is highly preferable so that the neural network filter for personalized baby cry determination can be frequently updated. Furthermore, during the connection period, locally collected data can be uploaded to the server, and new filters or executable instructions for determining why the baby is crying can be downloaded, taking into account the recorded audio data.

[0093] It will be appreciated that in order to improve personalized determination of baby cries, it is preferable to upload data associated with the acoustic monitoring of baby cries and / or parameters associated with selected cry data that enable cry determination to a cloud and / or centralized server. While it may be preferable not to upload any data if locally available computing power is sufficient to achieve cry identification, at least in cases of insufficient bandwidth, it is highly preferable to upload as many samples as possible related to the cry pattern, since the database of audio samples will preferably grow when uploading a large number of patterns with tags from feedback that confirm or disagree with previous determinations, and the model can be retrained using the expanded database of tagged samples.

[0094] Therefore, to improve the personalized decision, at least some of the cries and / or parameters derived from the cries can be stored on a server together with the respective personal baby data, allowing the personalized decision to take the information stored on the server into account. For example, it is understood that it may be sufficient to send the device ID if a subscription to a new filter is being obtained for a particular device, and if the parent or caregiver initially indicates the baby's age and further details. However, since other parameters such as the baby's weight and height should also be updated frequently, the parent and / or caregiver is preferably requested to enter the corresponding information periodically. It is understood that entering such information can be done, inter alia, using a separate device, such as a smartphone running an appropriate app, and / or by allowing the user to enter the corresponding information into the device through speech, using either local or centralized speech recognition.

[0095] As can be seen from the above, a preferred embodiment is proposed to include a step of downloading information from a centralized server that enables local personalized baby cry determination. It is assumed that, as babies grow and age, after a period of time, personalized filters will no longer provide the most favorable results. Therefore, it is possible and useful to limit the use of local personalized baby cry determination to a specific period of time. Once such period has elapsed, a warning can be issued that personalization is no longer reliable, and / or a standard, non-personalized filter can be used, and / or a message can be issued to the user requesting a filter update instead of indicating the reason for the baby's cry. In certain cases where a parent or caregiver has subscribed to periodic filter updates and such a filter has not been updated for an extended period of time, for example because the connection to the centralized server is impaired and / or blocked, a warning can be generated shortly before the use of the personalized filter is stopped entirely and / or before determination is achieved exclusively in a non-personalized manner. It is also clear that personalized determination can be achieved within a cloud server.

[0096] As mentioned above, in a preferred embodiment, it is proposed that audio data acquired before the onset of a cry is (also) used to determine the acoustic background and / or to determine additional parameters for baby cry determination. Regarding the determination of additional parameters for baby cry determination, situations may arise in which the exact onset of a cry cannot be determined with a sufficiently high probability, for example, due to the co-occurrence of a loud acoustic background. However, it should be noted that by relying on spectrograms and the like for pattern identification, cry-related audio data can be separated significantly more reliably than simply relying on a few parameters derived from the audio stream. It can thus be seen that spectrogram-based identification and separation of crying durations or crying patterns is significantly more robust against acoustic disturbances, which significantly aids in obtaining better cry translation results.

[0097] It is understood that the exact onset of a cry may nevertheless not be possible to determine with a sufficiently high probability, and that evaluation of additional (possibly preceding) frames may be helpful in making the determination. This may preferably be done by evaluating a sliding parameter and / or by cross-correlation techniques, or by analyzing a period, e.g., one, two, or three standardized periods used for identifying and isolating crying patterns in the audio stream preceding a loud noise above a threshold. Also, if a baby's cry is detected following a loud noise, it is likely that the baby needs to be comforted, and accordingly, such events may be useful in making the determination, even if they are not considered background patterns that need to be removed or filtered out from the audio data.

[0098] With regard to the minimum probability at which an onset of crying is considered to have been detected without requiring evaluation of audio data acquired prior to the onset of crying, in typical cases of multi-stage crying detection, such a likelihood or probability can be determined and crying is considered to have been detected if such probability is, for example, higher than 70%, 80%, 90%, 95% or 99%, it should first be understood that the exact threshold at which the probability is considered sufficiently high depends, inter alia, on the pattern of background noise and / or the quality of the multi-stage crying detection.

[0099] Given the current standard already achieved by the applicant, the probability that a cry is detected for the first time in a frame, and thus that a cry onset has been detected, is easily above 99%. However, lower thresholds such as 97%, 95%, 90%, or 80% can be set. Note that even when a very high probability of a cry onset being accurately determined is achieved, it is still possible to feed previous frames into the baby cry determination stage, along with frames recorded after the (likely) onset of a cry. This can be particularly useful when techniques such as cross-correlation are used in the determination. Even the number of frames preceding the supposed onset of a cry to be fed into the cry determination can be determined in terms of probability, for example, by determining the number of preceding frames by the formula (100 - probability (in %)) x A, 0.5 or 1 or 1.5, or any number in between; obviously, the number of preceding frames obtained from these formulas is rounded up to the next higher integer.

[0100] It will be appreciated that the aforementioned method of providing data for evaluation is particularly useful in field environments outside of a speech or acoustics laboratory. In the field, appropriate data parameters are particularly important to improve the accuracy of cry determination. For example, in a typical laboratory setup, audio has a clean, low-noise background, allowing for clear recordings of cries. In contrast, in a typical field environment, background noise is significantly higher, the volume of cries varies widely, and recordings are less "clear" with respect to high-frequency components due, for example, to less-than-optimal microphone positioning. These differences typically result in significantly lower accuracy in the field than in a laboratory environment. However, by using cross-correlation techniques and / or sliding averages, the accuracy of identifying the onset of crying and / or identifying the crying itself in the field becomes fully comparable to the accuracy obtained in a laboratory environment, despite the presence of significant noise. Furthermore, with regard to the final accuracy of the translation, it should be noted that significantly better results can be obtained by identifying and separating crying patterns using AI methods, such as image processing involving spectrogram-like transformations of audio periods.

[0101] Nevertheless, it is understood that the absolute accuracy obtainable and determined both in the laboratory and in the field may still depend on, for example, the sample used, the quality of the actual determination, as represented, for example, by a neural network filter, the length of the recording, or the definition and mathematical determination of the "accuracy" measure. Therefore, accuracy determined by different methods is not easily comparable. Typically, accuracy is defined such that a method results in greater than 90% accuracy in the laboratory.

[0102] Nevertheless, using the same method, the overall accuracy in the field may drop from ⇒90% in the lab to no less than 80% in the field if the determination stage is provided with appropriate data that allows it to consider, for example, sliding averages and / or cross-correlations, or searching for patterns in the spectrogram-like representation of the audio period. Note that the spectrogram-like representation may be a standard spectrum, or may differ in that the frequency resolution is different with respect to the audio spectrum, and / or may have different dynamic ranges for different frequencies.

[0103] From the foregoing, it will be understood that the parameters are preferably fed into the cry determination stage in a manner that allows for the determination of the cry using neural networks, convolutional neural networks and / or other known artificial intelligence techniques. It will be seen that typically such techniques rely on multiple so-called "layers" and that personalization can be achieved if at least one such layer is personalized; however, it is possible to personalize the cry determination by personalizing two or more layers, for example by selecting different filter parameters for one or more convolutional layers depending on personal data related to the baby. If multiple layers are personalized, each layer could be fully personalized, for example by selecting different filter parameters depending on gender, age and weight; whereas it would also be possible to select different filter parameters depending only on gender in a first personalization layer, different filter parameters depending only on weight in a second personalization layer, and different filter parameters depending only on age in a third personalization layer (the first, second and third personalization layers do not necessarily process data in that sequence). It will be clear to those skilled in the art that when multiple layers are personalized, it may be possible to, for example, personalize a first layer with, for example, one, two, three or more personal parameters, and then personalize a second layer with two, three, four or more personal parameters that may or may not overlap with the personal parameters used to personalize the first layer, making them fully or partially personalized.

[0104] The computer-implemented multi-stage crying determination method may have one or more stages for crying detection and crying pattern separation, for example based on a (mel) spectrogram-like representation of an overlapping sequence of windows segmenting the original audio data, based on an object localization method searching for crying patterns of at least 0.4 seconds and preferably less than 3 seconds in length, and further preferably using a two-dimensional representation of the crying patterns, for example with a different time and / or frequency resolution than the representation used for crying pattern detection and separation, and / or by setting the maximum audio level within each audio pattern to a specific value and / or by interpolating short crying patterns to a desired length, so that each crying pattern considered after separation can be classified into several different classes (or baby groups). It is possible to implement each stage separately from the others, for example, in such a way that the first stage runs in proximity to the baby, the second stage runs in the cloud, and the third stage runs on the parent's or caregiver's smartphone. In such a case, it is easy to see that personalizing one stage, preferably the final collector stage and / or the stage that assigns probabilities to each cry pattern, may be sufficient to obtain an overall personalized assessment of the baby's cry. However, it is not necessary to separate the different stages, and it would be entirely possible to implement a computer-implemented multi-stage cry assessment method in such a way that the different stages and interface steps, such as renormalization of duration and audio level, are performed in such a way that the user is unaware of the different stages but sees the overall assessment as a single process. This can be achieved by concatenating the stages for execution in the same place, for example by uploading the entire audio stream into the cloud and receiving only the cry determination from the cloud, nevertheless this can still be considered a multi-stage cry determination.It is expressly stated that providing cry pattern detection and separation, cry pattern probability determination of a sequence of cry patterns and personalized automated collective determination of the probabilities obtained from the probability determination is considered inventive, particularly when object localization is used for cry pattern detection and separation.

[0105] It will also be appreciated that in a preferred embodiment, parameters and / or data streams of the recorded audio are uploaded along with the baby data information. While uploading the entire audio recording is preferred in situations where an existing baby cry database is being augmented, it will be appreciated that uploading only extracted parameters may otherwise be preferable since less data needs to be transmitted, allowing for faster response, especially when data transmission bandwidth is low. In this context, it should be appreciated that one of the key steps in artificial intelligence evaluation of data is dimensionality reduction. For example, if a chunk of audio data is considered to contain 64 frames of 128 consecutive 16-bit samples, the initial space is (64 * 128 * 16 =) 131072 dimensions. To address this, parameters such as those listed above, e.g., mean voice level, first formant frequency shift, etc., can be determined (alternatively, a spectrogram-like representation of the source spectrogram can be used). Now, as can be seen from the above, if it is desired to determine a baby's cry based on a few parameters rather than a two-dimensional spectrogram-like representation, there are many different parameters that can be used to describe and determine a baby's cry, and this number of different parameters is typically further reduced by selecting only the most appropriate parameters.

[0106] In personalized determination of a baby's cry, patterns found in audio data from other babies are identified and a set of parameters that best describe these patterns is searched for. When a spectrogram-like representation-based technique is used, such identification of cry patterns in the spectrogram-like representation can be based on an image-like analysis that compares the spectrogram-like representation of the cry to be determined with cry patterns associated with known cries. Known cry patterns can be selected from a large database so that only cries from babies of similar age, sex, weight, medical condition, etc. are grouped together. In such cases, the filters used to compare the patterns have distinct personalization because the cries of each similar (peer) group will be somewhat different from the cries of babies belonging to different peer groups, even if the babies cry for the same reason. However, if the database is not large enough to establish large, diverse peer groups, it may be possible to first determine the cry patterns identified and separated in the original audio stream recorded by one or more layers using the exact same filter parameters for all babies. Since a cry includes multiple cry patterns, this first step results in multiple possible reasons why the baby cries, with a reason or likelihood of the reason being assigned to each cry pattern. Such multiple possible reasons or likelihoods for the baby's cry can then be determined in a further step, which then personalizes this step. It should be noted that by processing the data in this way, the number of personalized filter coefficients for personalization is significantly lower than when personalizing the first step of assigning a reasonable likelihood of a reason for the cry to each single cry pattern.

[0107] It should be noted that when different (“peer”) groups of babies are established for personalization, situations may arise in which different sets of parameters may be optimal for each different group of babies. It is desirable to reduce the amount of computation for determining the parameters and thus select a small set of parameters sufficient for personalization. However, if only a limited number of parameters, or even worse, only a reduced set of parameters, are sent to the server instead of the complete audio stream or instead of at least a portion of the audio stream associated with the cry parameters or a spectrogram-like representation thereof, the identification of new patterns may be impaired. Accordingly, it is preferable to send the complete audio data—or the fully extracted / separated cry data, rather than just the parameters extracted therefrom—in order to at least identify new patterns from the additional data.

[0108] In a preferred embodiment, the computer-implemented method of the present invention comprises uploading information related to parameters and / or recorded audio data streams and / or parts thereof and / or such fragments such as identified and isolated crying patterns along with baby data information, in particular related to at least one, preferably at least two, three or four of the following: age, sex, height, weight, ethnicity, only child / twins / triplets, current medical condition, known medical prerequisites, in particular known current illnesses and / or fevers, parent and / or caregiver languages.

[0109] It will be understood that information such as date of birth, only child / twins / triplets, etc. does not need to be transmitted every time data is transmitted from the local device to the server or cloud. However, since part of the baby data information is necessary, at least information that allows identifying the local device and can be associated with the corresponding necessary baby data, such as the ID of the locally used device, may be transmitted; in such a case, the actual baby data may be transmitted alone before the personalization decision and stored and retrieved in the cloud or on the server according to the transmitted information, such as the ID of the locally used device. In this context, it will be understood that it is sufficient for parents to enter the respective baby data when registering their baby or device using an app, website form, or the like.

[0110] Furthermore, it is preferable if (feedback) information related to the accuracy of one or more previous determinations is uploaded to the server. This can help recalibrate the filters (or classifications) used in the machine learning model and / or remove previous errors. Again, information related to the accuracy of previous determinations can be uploaded when a determination of the current cry event is not necessary. It is preferable to transmit information related to the accuracy of one or more previous determinations together with data dependent on the cry event, such as a cry event ID, a device ID + time tag, or the like; this can also be a combination of additional information, particularly if the determination is determined to be poor, such as the actual automated determination, feedback of the parent or caregiver's determination, and the corresponding baby cry parameters and / or raw audio data; instead of the raw data, timestamps of previously determined cry events, preferably those for which data has already been transmitted and remains stored on a centralized server such as a cloud server, can also be transmitted. Whether it is preferable to retransmit audio data or parameters determined taking audio data into account, or to retransmit only the ID or time tag, depends, among other things, on the storage space on the server. Simply statistical data may also be sent indicating how often the decisions were correct overall, or how often a particular decision, such as "baby needs to be comforted" or "baby needs to be burped", was correct or incorrect. The use of statistical information about the decisions allows different decision algorithms / filters or decision results to be provided to different users, albeit related to babies in the same peer group, using different filters and / or algorithms, and then the different decisions can be evaluated in a statistical way. This is particularly useful if the group of users is large enough.

[0111] It will be appreciated that different channels and / or different times may be used to transmit different types of data.

[0112] The data provided for personalized determination preferably enables the determination to distinguish at least one of the following conditions: "baby is tired," "baby is hungry," "baby is uncomfortable and needs attention," "baby needs to be burped," and "baby is in pain." If parameters are transmitted, such parameters are preferably selected and provided to enable the distinguishing of at least two, particularly at least three, and especially all, of the different conditions. Once a sufficiently large database is available, it can be assumed that certain medical conditions, such as "baby has reflux," "baby has bloating," or "baby has middle ear inflammation," or more detailed reasons for discomfort, such as "baby is too hot," "baby is too cold," or "baby is bored," can also be identified. In this context, it should be understood that the method of providing data for personalized determination proposed in the present invention is also very useful for expanding the existing database of baby cries and thus for improving baby cry determination. Therefore, by properly implementing the present invention, the database of cries can be expanded to enable highly sophisticated personalization of the determination in a short period of time.

[0113] Although the above method can be implemented using a wide variety of devices and / or systems, protection is particularly sought for an automatic baby cry determination arrangement including a microphone for continuously acoustically monitoring the baby, a digital conversion stage for converting the monitored audio stream into a stream of digital data, a memory stage for storing personal baby data information, a communication stage for transmitting the data to a centralized server arrangement, and indication means for indicating the result of the determination, such as a loudspeaker arrangement for acoustically indicating the result of the determination, a display and / or an interface to a display, wherein a cry identification stage is provided for identifying episodes of crying within the digital data, and the communication stage is adapted for transmitting data related to the cry to the centralized server arrangement for determination taking into account the personal baby data information, and for receiving data related to the personalized determination of the baby's cry from the centralized server arrangement. The cry determination arrangement is located in an apparatus separate from the apparatus including the microphone, and a display loudspeaker arrangement for acoustically or visually indicating the result of the determination.

[0114] It will be appreciated that one or more stages, in particular the cry identification stage for identifying episodes of crying in a stream of digital data, can be implemented by a combination of hardware and software, and a personalized filter for local determination of automatically identified crying can be received from a centralized server configuration, or the results of the determination can be received where some or all of the parameters obtained taking into account the audio data are sent to a centralized server or cloud for determination.

[0115] In a preferred embodiment, it is proposed that the automatic baby cry determination arrangement comprises a feedback arrangement for obtaining feedback information related to the accuracy of one or more previous determinations, and that the communication (or I / O) stage is adapted for transmitting the feedback information to a centralized server arrangement. In a preferred embodiment, the feedback arrangement is integrated into the device used to acoustically monitor the baby.

[0116] Furthermore, in a preferred embodiment, the automatic baby cry determination arrangement comprises a local determination stage adapted to determine the baby's cry taking into account data received from the centralized server arrangement related to personalized determination of the baby's cry. The local determination may be an auxiliary determination stage that allows determination when audio data cannot be transmitted to the centralized server arrangement, or it may be the primary or only determination stage from which all determinations presented to the parent and / or caregiver are generated.

[0117] It has been mentioned above that the personalization determination of baby cry data depends on factors such as the baby's age, height and weight, which change significantly as the baby grows, resulting in the fact that the personalization may become outdated. A corresponding check should be made to prevent the personalization determination from being attempted using an outdated filter. Accordingly, if the automatic baby cry determination configuration includes a timer and an evaluation stage that evaluates the current age of the personal baby cry determination information and / or the validity of the age or (filter / algorithm) data received from the centralized server or cloud configuration prior to determining the baby cry, it is preferable that the baby cry determination configuration that is adapted to output the baby cry determination depends on that evaluation.

[0118] The invention will now be described, by way of example, with reference to the drawings in which: [Brief explanation of the drawings]

[0119] [Figure 1a] 1 illustrates a series of steps in determining a baby cry, with some of these steps implementing one embodiment of the present invention. [Figure 1b] This shows part of the crying detection / data preprocessing. [Figure 2] Show multiple symbols that can be used to indicate your baby's current needs. [Figure 3]It shows 3D spectrograms extracted from several audio recordings representing crying babies exhibiting different needs - time increases along the X axis, frequency increases along the Y axis, and intensity increases along the Z axis. The units are arbitrary but are the same for all parts. [Figure 4a] A comparison of spectrograms is shown to visualize the intensity variations for multiple frequencies over time for different cries, and in more detail: Figure 4a is associated with cries from different hungry babies; Figure 4b is associated with different cries from the same hungry baby; Figure 4c is associated with cries from different babies in pain; Figure 4d is associated with different cries from the same baby in pain; Figure 4e is associated with cries from different babies needing to be burped; and Figure 4f is associated with different cries from the same baby needing to be burped. [Figure 4b] A comparison of spectrograms is shown to visualize the intensity variations for multiple frequencies over time for different cries, and in more detail: Figure 4a is associated with cries from different hungry babies; Figure 4b is associated with different cries from the same hungry baby; Figure 4c is associated with cries from different babies in pain; Figure 4d is associated with different cries from the same baby in pain; Figure 4e is associated with cries from different babies needing to be burped; and Figure 4f is associated with different cries from the same baby needing to be burped. [Figure 4c] A comparison of spectrograms is shown to visualize the intensity variations for multiple frequencies over time for different cries, and in more detail: Figure 4a is associated with cries from different hungry babies; Figure 4b is associated with different cries from the same hungry baby; Figure 4c is associated with cries from different babies in pain; Figure 4d is associated with different cries from the same baby in pain; Figure 4e is associated with cries from different babies needing to be burped; and Figure 4f is associated with different cries from the same baby needing to be burped. [Figure 4d]A comparison of spectrograms is shown to visualize the intensity variations for multiple frequencies over time for different cries, and in more detail: Figure 4a is associated with cries from different hungry babies; Figure 4b is associated with different cries from the same hungry baby; Figure 4c is associated with cries from different babies in pain; Figure 4d is associated with different cries from the same baby in pain; Figure 4e is associated with cries from different babies needing to be burped; and Figure 4f is associated with different cries from the same baby needing to be burped. [Figure 4e] A comparison of spectrograms is shown to visualize the intensity variations for multiple frequencies over time for different cries, and in more detail: Figure 4a is associated with cries from different hungry babies; Figure 4b is associated with different cries from the same hungry baby; Figure 4c is associated with cries from different babies in pain; Figure 4d is associated with different cries from the same baby in pain; Figure 4e is associated with cries from different babies needing to be burped; and Figure 4f is associated with different cries from the same baby needing to be burped. [Figure 4f] A comparison of spectrograms is shown to visualize the intensity variations for multiple frequencies over time for different cries, and in more detail: Figure 4a is associated with cries from different hungry babies; Figure 4b is associated with different cries from the same hungry baby; Figure 4c is associated with cries from different babies in pain; Figure 4d is associated with different cries from the same baby in pain; Figure 4e is associated with cries from different babies needing to be burped; and Figure 4f is associated with different cries from the same baby needing to be burped. [Figure 5a]The clustering of different cries is shown in Figure 5a, where Figure 5b shows the entire cluster of cries, Figure 5c shows the "sleepy" cries within the entire cluster, Figure 5d shows the "need to burp" cries within the entire cluster, Figure 5e shows the "discomfort" cries within the entire cluster, and Figure 5f shows the "pain" cries within the entire cluster. (The separation of the clusters in the 2d graph is not perfect when additional differentiation parameters are considered, but it can be seen that even in the 2d graph shown, clusters are beginning to emerge.) [Figure 5b] The clustering of different cries is shown in Figure 5a, where Figure 5b shows the entire cluster of cries, Figure 5c shows the "sleepy" cries within the entire cluster, Figure 5d shows the "need to burp" cries within the entire cluster, Figure 5e shows the "discomfort" cries within the entire cluster, and Figure 5f shows the "pain" cries within the entire cluster. (The separation of the clusters in the 2d graph is not perfect when additional differentiation parameters are considered, but it can be seen that even in the 2d graph shown, clusters are beginning to emerge.) [Figure 5c] The clustering of different cries is shown in Figure 5a, where Figure 5b shows the entire cluster of cries, Figure 5c shows the "sleepy" cries within the entire cluster, Figure 5d shows the "need to burp" cries within the entire cluster, Figure 5e shows the "discomfort" cries within the entire cluster, and Figure 5f shows the "pain" cries within the entire cluster. (The separation of the clusters in the 2d graph is not perfect when additional differentiation parameters are considered, but it can be seen that even in the 2d graph shown, clusters are beginning to emerge.) [Figure 5d]The clustering of different cries is shown in Figure 5a, where Figure 5b shows the entire cluster of cries, Figure 5c shows the "sleepy" cries within the entire cluster, Figure 5d shows the "need to burp" cries within the entire cluster, Figure 5e shows the "discomfort" cries within the entire cluster, and Figure 5f shows the "pain" cries within the entire cluster. (The separation of the clusters in the 2d graph is not perfect when additional differentiation parameters are considered, but it can be seen that even in the 2d graph shown, clusters are beginning to emerge.) [Figure 5e] The clustering of different cries is shown in Figure 5a, where Figure 5b shows the entire cluster of cries, Figure 5c shows the "sleepy" cries within the entire cluster, Figure 5d shows the "need to burp" cries within the entire cluster, Figure 5e shows the "discomfort" cries within the entire cluster, and Figure 5f shows the "pain" cries within the entire cluster. (The separation of the clusters in the 2d graph is not perfect when additional differentiation parameters are considered, but it can be seen that even in the 2d graph shown, clusters are beginning to emerge.) [Figure 5f] The clustering of different cries is shown in Figure 5a, where Figure 5b shows the entire cluster of cries, Figure 5c shows the "sleepy" cries within the entire cluster, Figure 5d shows the "need to burp" cries within the entire cluster, Figure 5e shows the "discomfort" cries within the entire cluster, and Figure 5f shows the "pain" cries within the entire cluster. (The separation of the clusters in the 2d graph is not perfect when additional differentiation parameters are considered, but it can be seen that even in the 2d graph shown, clusters are beginning to emerge.) [Figure 6] 3D representation of the T-SNE dimensionality reduced mel spectrogram from two different perspectives. [Figure 7] The K-Means clustering is shown with the centroids for each cluster of five different labels depicted as white crosses, and the division into different cells. DETAILED DESCRIPTION OF THE INVENTION

[0120] FIG. 1 illustrates steps useful in baby cry determination performed by a computer-implemented method for providing data for automatic baby cry determination, the computer-implemented method for providing data for automatic baby cry determination including the steps of acoustically monitoring a baby to provide a corresponding stream of audio data, detecting a cry or a portion of a cry within the stream of audio data, selecting data from the audio data in response to detecting the cry or portion thereof, determining personal baby data for personalized cry determination, preparing a determination stage for determination according to the personal baby data, processing the selected data for cry determination, and providing the processed information to the cry determination stage prepared for personalized determination according to the personal baby data.

[0121] In this regard, Figure 1 suggests that for baby crying detection, first an appropriate sound processing or pre-processing device is activated and placed close enough to the baby to be monitored and switched on.

[0122] In a preferred embodiment, audio processing is accomplished near the baby, and then, as long as a connection to the centralized server is available, the pre-processed audio data is uploaded to the centralized server, which may be a cloud server, together with the personal baby data information. In such a preferred embodiment, the audio pre-processing device is part of an automatic baby cry determination arrangement (not shown), which includes a microphone for continuously acoustically monitoring the baby, a digital conversion stage for converting the monitored audio stream into a stream of digital data, a memory stage for storing the personal baby data information, and a communication stage for transmitting data to the centralized server arrangement, wherein a cry identification stage is provided for identifying onsets of crying within the stream of digital data, and the communication stage is adapted to receive data related to personalized determination of the baby's cry from the centralized server arrangement.

[0123] It is understood that a typical smartphone can be used as a preprocessing device, since it essentially includes a battery, a microphone and appropriate microphone signal conversion circuitry, a processing unit, and wireless I / O connections. When the preprocessing device is implemented using a smartphone, appropriate apps can be installed to implement the functions and processing steps, so that all necessary preprocessing (and, where applicable, both preprocessing and determination) can be performed on the smartphone. However, because not every parent and / or caregiver has an extra smartphone, and some applications, e.g., hospital applications and pediatric stations, require significantly more preprocessing, it is preferable to integrate the necessary hardware into a standalone package or other baby monitoring device, such as a video camera for baby monitoring or a sensor configuration to monitor whether the baby is breathing. It should be noted that a non-smartphone device can be used near the baby, from which audio data can be transferred via short-range communication, such as Bluetooth and / or Wi-Fi, to the parent's or caregiver's smartphone, where additional (pre)processing can be accomplished, so that the preprocessed audio-related data can be uploaded to a centralized server. In such an arrangement, audio data need only be transferred to the smartphone device if particularly loud noises are detected near the baby; nevertheless, many parents will wish to receive a continuous stream of audio from their baby, in which case it would clearly be possible to transmit the continuous audio stream from the baby to the parent's or caregiver's smartphone, laptop, tablet, etc., and accomplish there any processing of the audio stream necessary for determining whether the baby is crying, including detection of portions of the audio stream that are associated with crying or can be assumed to be associated with crying with some non-zero probability.

[0124] A preferred integrated standalone device (not shown) is now described. The standalone device includes a power supply, microphone, processing unit, memory, wireless I / O connections and input / output means, and preferably a timer. It will be appreciated that such a device can be constructed in such a way that it boots particularly quickly, so that there is no significant delay between switching on and actual operation.

[0125] The power source can be a battery, for example a rechargeable battery, or it can be a power source that is plugged into a power outlet.

[0126] The microphone can be any microphone sensitive in the 150 Hz to 3000 Hz range, with a wider range being preferred, for example, from 100 or even 80 Hz as a lower limit to 3500, preferably 4000 Hz as an upper limit. It will be appreciated that modern microphones readily record frequencies in this range, but that differences in spectral sensitivity can nevertheless adversely affect the baby's cry detection if certain frequencies are suppressed or overemphasized. This is particularly problematic when smartphones are permitted as standalone devices, since different smartphones from various manufacturers may have widely varying spectral sensitivities. However, the problem is less pronounced, and in that case, better results are expected using one or a few models of standalone device, since the same microphone model can be used. It is even possible to calibrate the microphone and install calibration data on the device so that any recorded sound can be corrected for the actual (spectral) sensitivity of a given device. Nevertheless, it should be noted that variations in spectral sensitivity can also be caused by variations in the environment, for example, because more or less absorbing material is placed around the baby, resulting in higher or lower absorption, especially at higher frequencies. Consequently, the overall sensitivity of the microphone should be such that, when placed at a distance of approximately 0.25 m to 1.5 m, a very loud baby cry should result in a digital signal close to, but not exceeding, the maximum digital signal strength. In preferred embodiments, the sensitivity of the device is set either manually or automatically. The microphone's polar pattern is such that the orientation of the device does not significantly affect the overall sensitivity and / or spectral sensitivity; therefore, a unipolar pattern is preferred. The microphone signal is amplified, preferably bandpass filtered, and converted to a digital audio signal. It is understood that the sample frequency of the analog-to-digital conversion is sufficiently high to avoid aliasing problems according to Nyquist theory.Accordingly, if the microphone is sensitive up to 4000 Hz as an upper limit, a sample frequency of 8 kHz is considered as a minimum. Also, if an 8 kHz sampling frequency is used, it is useful to clip the analog signal at 4 kHz using an appropriate analog (band-pass or low-pass) filter. In a typical implementation, the analog-to-digital conversion produces an output signal of at least 12 bits, and preferably 14 bits. Since there is usually inevitably some background noise, a higher dynamic resolution typically does not improve the decision.

[0127] The I / O connection forms part of the communication stage for transmitting information to a nearby parent or caregiver and for transmitting data to a centralized server configuration. Different connections can be selected for communicating with the parent or caregiver, on the one hand, and with the centralized server, on the other. For example, short-range wireless protocols such as Bluetooth, Bluetooth LE, or ZigBee can be used to transmit information to the caregiver, while wide-area wireless protocols such as G4, G5, GSM, UMTS, or WiFi communication with an Internet access point can be used to communicate with the centralized server. In this context, it is understood that only a limited amount of information needs to be transmitted from the device near the baby to the parent or caregiver. For example, a periodic device heartbeat signal can be transmitted indicating that the device is operating correctly or another status, such as "low battery." Furthermore, if the baby is crying, a crying indicator can be transmitted independently of the actual determination of crying, which should be indicated once it becomes available. A skilled person will understand that this can be done by transmitting very few bits, and accordingly, both bandwidth and energy consumption can be quite low. Nevertheless, in a preferred embodiment, the parent may have the possibility to decide whether or not the transmission of any audio is desirable. In some cases, the parent may wish to perform permanent acoustic monitoring of the baby.

[0128] However, it should be understood that one mode of operation simply notifies the parent or caregiver that the baby is crying, so that the caregiver can then move to the device where the actual verdict is displayed, and therefore it is not even necessary to transmit the actual verdict to the parent or caregiver. In contrast, when transmitting data to a centralized server, typical cry data from the current cry to be rated and / or data collected from multiple cries should be transmitted along with personalized baby information. Because soothing a baby is often difficult, even if the reason for the baby's crying is known, such cry data may be collected over an extended period of time, such as several minutes, and can be expected to result in a significant amount of data being transmitted. Therefore, having a broadband connection to the server is useful. It is not absolutely necessary to transmit large amounts of data to a caregiver parent who is in a room away from the baby, and therefore does not require a broadband connection to be used, but it is understood that it is not necessary to use low-energy protocols such as Bluetooth LE or ZigBee. Rather, (broadband) I / O such as Wi-Fi can also be used for communication with the parent or caregiver.

[0129] The input / output means of the preferred standalone device serve, on the one hand, to input personalized baby data information into the device, such as the baby's age, weight, height, current status, and / or current or persistent medical conditions. The input means can be implemented using the aforementioned I / O connections when used in connection with a smartphone, laptop, tablet, PC, or the like, where data is input and sent to the standalone device for storage. However, an even more preferred method for inputting personalized baby data information may be using a microphone and additional speech recognition; if this method for inputting personalized baby data information is selected, a button or the like may be provided so that entry into the personalized baby data information input mode can be requested by pressing the button. Note that it is not necessary to have a speech recognition stage implemented on the device itself, but parent or caregiver speech data can be uploaded to the cloud for speech processing there, sending back personalization information and / or information that can be more easily determined than from speech, e.g., a text file. It will be appreciated that speech recognition related to personalization information can be implemented using services already available within the web.

[0130] It may even be possible to use the integrated speaker to guide the user by requesting specific input information, and preferably to confirm the input as understood through the speaker using a machine-synthesized voice. This is particularly useful for personalized baby data that must be updated regularly, such as information related to the baby's weight or high fever, since using the microphone allows the parent or caregiver to easily and quickly update the personalized information. If personalized baby information is desired to be entered using a different device and a wireless or wired connection such as USB—which may be used to provide power anyway—preferably, the only input means provided would still be a confirm / reject button to confirm or reject the assessment, thus providing feedback on the quality of the assessment. Also, in some cases, such as the use of standalone devices and pediatric stations, it would be preferable if medical conditions could be entered as personalized information. This is advantageous in pediatric stations, where cries from babies with abnormal medical conditions are more abundant, allowing the database to grow rapidly.

[0131] Regarding the need to provide feedback, while it is not necessary to enable feedback for every single device for every single cry, it is still highly preferable to do so because providing feedback helps to expand the database of available samples and thus improve the assessment. Furthermore, it is understood that if appropriate feedback is provided, a large number of “tagged” samples will be available. It is understood that techniques such as neural network filters are used for assessment and / or to detect cries in more or less noisy backgrounds, and samples are required to train a model to determine the appropriate filter. Here, if feedback is provided, the database of available samples tagged with feedback on whether the parent or caregiver's automated assessment was correct or not can be significantly larger than otherwise, and can grow rapidly, especially once a sufficient number of devices are deployed. Furthermore, a sufficiently large database with samples from multiple babies of different ages, genders, heights, weights, etc., allows for more personalized assessments to be provided. A larger database may also be useful in identifying recency, and it is highly desirable to solicit feedback from the parent or caregiver accordingly and provide that feedback to a centralized server, preferably in a manner that allows the feedback to be combined with the personalized information and audio data to which the cry is associated. However, in some cases, it is not necessary to transmit or re-transmit the entire audio data if the audio data has been previously determined and remains stored on the server until feedback is received.

[0132] With regard to personalized assessment, the current understanding of baby cries is that for newborns up to a certain age, such as 4-6 months, there are no significant differences between babies from different countries, ethnicities, or "races." Rather, applicants' current understanding is that differences in crying are due to physiological differences between small and large babies, newborns and older infants, and even medical conditions that have significant impact on babies. It is understood that it may be possible to more clearly distinguish between different cries and / or between the multiple reasons why a baby cries. It is understood from the foregoing description of the prior art that certain medical conditions can change the way a baby cries, and therefore, important medical clues can be obtained from analysis of audio data. It is further understood that providing new samples to the database can also provide data for automatic cry assessment, and that depending on how the samples are prepared to expand the sample database, a computer-implemented method in accordance with the present invention may be configured.

[0133] The memory of a preferred standalone device is used to store executable instructions for a processing device, such as a microcontroller, CPU, DSP, and / or FPGA in the standalone box. Personalization information is then stored within the standalone device, including a device ID for uploading to the server, audio data / feedback data, and filter data for local crying identification and personalized local crying determination. Additionally, the memory allows for buffering of the most recent audio data, so that once a cry is detected in the audio data, audio data immediately prior to the cry is also available, e.g., a preceding period between 20 seconds and 0.5 seconds; preferably, a period of at least 5 seconds, particularly at least 10 seconds, and especially at least 15 seconds, if a 5-second window is used to search for crying patterns in the audio stream. The length of the data immediately prior to the cry can be determined taking into account the cost and availability of appropriate buffer memory, as well as the expected noise level. If a particularly noisy environment is expected or tolerated, it may be useful to also store samples of background / ambient noise, e.g., to identify particularly noisy or particularly quiet frequency bands. It will be understood that different types of memory may be used for the specific different purposes shown, such as ROM, e.g., EEPROM memory, RAM memory, flash memory, etc. Furthermore, it will be understood that the size of the required memory can be easily estimated taking into account the intended use and the period allowed between two transmissions of audio data samples to the centralized server, as well as the type of data that should be stored locally, at least for a while. It should be noted that this could just be feedback data, parameters derived from the cry data, the original (raw) audio data of all cries identified since the last upload, samples of background noise at different levels, e.g., particularly loud non-cry background noise, or non-cry background noise with frequently observed frequency components, the latter implying a local statistical analysis of the background behavior.

[0134] The size of the memory provided also depends on the length of the baby's cry being considered. As indicated above, the determination can be achieved locally and / or centrally using an AI / CNN filter that is trained before the determination to distinguish the cry from background noise. Such filtering can be highly accurate, given the typical length of a baby's cry, provided, of course, that the entire cry to be determined is made available during the determination stage. Often, of course, a determination is required before the baby stops crying. Accordingly, the fragment of a baby's cry typically evaluated should be taken into account when determining the size of the memory required to store the cry for later uploading and / or buffering data. It should be noted that in a preferred embodiment, which has resulted in a very high correct detection rate in practical implementations, cry patterns with a length of less than 5 seconds, specifically between 1.5 and 4.5 seconds, and particularly about 2, 3, or 4 seconds, are isolated from the audio stream for cry determination. Since a baby is likely to cry for an extended period of time unless a parent or caregiver continues to soothe the baby, it is preferable to store at least 10, preferably at least 20, and especially at least 30 isolated crying patterns of the lengths indicated above, as several such isolated crying patterns are preferably analyzed. Note, for example, that even with a sampling rate corresponding to CD quality, only 0.5 MB to 8 MB should typically be required to implement very useful memory.

[0135] As mentioned above, the standalone box has some kind of data processing capability, e.g., a microcontroller, CPU, DSP, and / or FPGA, as well as memory for storing instructions and / or configurations for each of these devices. These instructions preferably include, among other things, instructions for performing at least the cry detection locally. This allows for selecting data related to the cry, thus significantly reducing the amount of data that needs to be transmitted to the centralized server compared to when the overall determination is performed solely on the centralized server. It should be noted that even a first stage determination based solely on audio intensity to detect particularly loud noises already results in a very significant reduction in transmitted data, since each baby has long periods when the baby is not crying.

[0136] The processing power required for local determination can be readily estimated. In this regard, it should be noted that while the processing device should preferably be able to achieve local cry detection, full personalization to the degree possible on a centralized server may not be necessary or possible on a local device, given memory and processing constraints, e.g., with regard to frequency of updates. It will be understood, however, that some personalization is possible locally.

[0137] Therefore, a preferred processing device is typically arranged such that the automatic baby cry determination arrangement includes a local determination stage, which is adapted to determine a baby cry taking into account filter or determination instruction data received from the centralized server arrangement and relating to personalized determination of the baby cry.

[0138] Accordingly, it will be appreciated that local determination need not be effective for each and every cry, but may be limited to cases where the centralized server is known to be inaccessible or accessible only at particularly low data transmission rates, such conditions being easily determined by the I / O stage of the standalone box as described.

[0139] The local decision stage may function as an auxiliary decision stage without personalization, but is also typically personalized, albeit to a lesser extent than is possible on the server, the degree of personalization depending, for example, on the availability of filters and / or the processing power available locally. However, given that in a typical application only fairly low resolution audio data needs to be analyzed and processed, the processing power available locally is typically sufficient to allow even processing steps such as Fourier filtering, cross-correlation and the like without undue strain on the processing device.

[0140] Thus, if a sufficient broadband connection is not available to upload data to a server, in preferred embodiments, local determination of crying can be achieved, provided that appropriate determination steps are implemented on the device. By analyzing audio data locally, parents do not need to have permanent internet access, which can be advantageous when traveling or if parents are irrationally concerned that Wi-Fi radiation may harm their baby. In any case, in preferred embodiments, it is possible and preferable to activate wireless transmission only once crying has been detected. This reduces battery consumption and addresses the concerns of parents who fear electromagnetic radiation sources in close proximity to their baby.

[0141] The standalone box also includes a timer. The timer can be a conventional clock, but it is understood that it will typically need to count for at least several days to determine whether a personalization decision is still valid, or whether a previous personalization related to the baby's high fever should still be considered valid given the time that has passed since the high fever event. Also, the time since the last personalization can be measured, and a warning can be issued if the data is out of date, for example, since a healthy baby should, under normal conditions, have gained more than 10% in weight since the last personalization data entry.

[0142] Accordingly, the local standalone box may include a timer and an evaluation stage for evaluating, prior to the baby cry determination, the current age and / or age of the personal baby data information and / or the validity and / or relevance of the data received from the centralized server configuration for the personalized baby cry determination, and the baby cry determination configuration is adapted to output the baby cry determination depending on the evaluation.

[0143] The timer is also useful for extrapolating personalized data. For example, the device can be initialized upon first use, and the time since initialization upon first use can be determined for each voice sample. It is not absolutely necessary, but this is particularly preferred, that the baby's age be entered upon initialization, since cries can be determined in a non-personalized manner until the parent has had time to enter all the information. As a result, the baby's age at the time of recording a particular voice sample can be determined and entered into the database. The baby's age can be easily calculated later based on this information. However, further information, such as height, can be extrapolated. For example, if the initial height is entered along with the baby's age and gender, it can then be determined whether the baby was of average height or above or below a certain given percentile of babies of the same age and gender peer group.

[0144] Filter updates can be offered as a service that the user has to pay for, for example, by subscription. Once the subscription expires, the user can use the device either with the final filter, a generic filter, or simply as a baby phone without any cry detection functionality. Since the subscription is limited to a specific time, once the device is sold, which is typically the case after the baby has grown significantly, the subscription period typically ends and a new subscription needs to be paid. This can also be tolerated or enabled, for example, by sending a corresponding reset code.

[0145] The device also includes output means for outputting the result of the baby's cry determination. The output means may be a screen or an LED-illuminated symbol as shown in FIG. 2, which is particularly useful when no additional different reasons are indicated other than those for which an LED is provided. The output means may additionally or alternatively be or include a speaker and / or I / O for communicating with a user's smartphone so that an immediate determination can be provided at a remote location, for example, on the smartphone screen if the parent or caregiver is in a different room away from the baby.

[0146] Using the previously described apparatus, crying can be determined in the following manner, and to this end, data can be provided for automatic personalized baby crying determination using the following method, which, as can be understood, will be a computer-implemented method.

[0147] First, the local device, which is placed near the baby's cradle, is switched on and started up. Once the device boots, a check is made to determine whether the centralized server is reachable with sufficient bandwidth. Any data that needs to be uploaded to the centralized server, such as previously collected cry data along with the local determination and / or feedback on the previous determination, is uploaded to the centralized server. Communication with a remote station near the parent, such as a smartphone, is then established by sending appropriate data via the I / O communication interface. A check is made to determine whether the current subscription to the personalized determination is still valid or should be updated. If the current subscription is no longer valid, a warning message is sent to the remote station. If the current subscription is still valid, audio sampling begins, and a message is sent to the remote station indicating that the local device is now "listening to your baby."

[0148] Audio sampling is then achieved such that the microphone is set to active mode, and appropriate amplification of the electrical signal from the microphone is set so that during loud audio events, the signal is well above the electrical noise level of the device without overloading the black signal. The electrical input signal is then filtered with a cutoff of 4 kHz. The filtered and amplified electrical analog signal is then converted to a digital signal with a sampling rate of 8 kHz and a dynamic resolution of 14 bits.

[0149] The digital signal undergoes automatic multi-stage cry detection. In the embodiment described herein, to detect crying, samples in the digital audio stream are first grouped into frames of 128 samples each, with the frames therefore having a length of 16 ms. The frames are written to a frame ring buffer, which in the described embodiment stores 1024 frames, cyclically storing the most recent frame in the memory location where the current oldest frame was previously stored. However, it should be noted that rather than using frames of 128 samples each, the number of samples in a frame can vary. Using frames with fewer than 128 samples allows for more precise slicing or cutting out of irrelevant data, while frames with more samples are more manageable by low-power CPUs. Note that 1024 frames of 128 samples acquired at an 8 kHz sampling rate correspond to 1024 * 128 samples * (1 / 8000) seconds = 16.38 seconds. As a result, for a 5-second window within which crying patterns are searched, it is possible to search for crying patterns or onsets within three windows preceding loud noises.

[0150] Then, as an initial step in crying detection, for every frame, the root mean square of the digital values ​​of the 128 samples is determined to provide an estimate of the current average frame speech level. The estimate of the current average frame speech level is also stored. From the current average frame speech level, a threshold is determined that the speech level of a new frame must exceed to satisfy the first criterion that crying may have been detected. Note that the threshold can be determined in an adaptive manner and need not be constant regardless of the current average frame speech level.

[0151] If it is detected that the average frame audio level does not exceed the average frame audio level of the preceding frame by an amount corresponding to the threshold, it is determined that crying has not been detected and the next frame is to be analyzed.

[0152] It will be appreciated that the first cry detection stage can be implemented in different ways, for example using the average of the preceding average frame speech levels or the minimum of the average frame speech levels within a number of preceding frames, for example 4, 8 or 16 preceding frames, where the minimum value should be chosen as it is a good estimate of at least how noisy the environment is. Accordingly, the cry detection first stage removes speech below an adaptive threshold in the preferred embodiment described herein.

[0153] If the average frame audio level of a new frame is detected to exceed the average frame audio level of the preceding frame by at least an amount corresponding to a threshold, a first criterion indicating that crying may have been detected is met, and the 1024 frames in the frame buffer are saved to another memory location. In one embodiment, the alleged crying audio-related data remains protected until an additional cry detection stage determines that no audio is present despite the satisfaction of the first criterion, or until the data can otherwise be selected as related to crying. This allows for later consideration of seemingly unrelated audio data that actually already contains audio data related to crying episodes. In another embodiment, the alleged crying audio-related information undergoes a pattern identification and separation step. In this regard, overlapping windows, e.g., five-second long windows with three-second strides, can be defined. It is possible to perform the crying pattern analysis non-locally, e.g., on a central cloud server, in which case the other memory location would be on the cloud server. In other cases, particularly if the locally provided processing power is sufficiently high, it may be possible to search for cry patterns locally without even storing the frames stored in a separate memory location, given that the supposed cry pattern can be directly detected within the previously stored 1,023 frames, assuming the processing is fast enough. As a result, only the isolated cry pattern would need to be transferred to a cloud server or further local processing stages. It will be appreciated that cry pattern identification and separation can be affected by first defining a window having cry patterns of multiple lengths to be detected, defining a spectrogram-like representation of the defined window, and searching for the cry pattern within the spectrogram-like representation using artificial intelligence techniques, particularly convolutional neural networks. It will be appreciated that for such a search for cry patterns, it is not absolutely necessary to use a model trained in a personalized manner, since baby patterns have very similar characteristics for a wide variety of babies.This is useful in local searches for cry patterns, since it is not absolutely necessary to search for cry patterns in a personalized way. It has already been mentioned above that the searched cry patterns preferably have a minimum length, and such a minimum length of 0.25 seconds, or 0.3 seconds, or preferably 0.4 seconds, is advantageous in that, even if personalization of cry detection and cry separation is not implemented, the distinction between cry patterns and non-crying periods is significantly more reliable for longer patterns considered. In other words, by searching for longer cry patterns, personalization of cry detection becomes unnecessary, especially when known object localization algorithms are essentially used for detecting cry patterns within spectrogram-like representations of audio data.

[0154] However, it is not absolutely necessary to implement the second stage of cry detection as cry pattern detection and separation—accordingly, a more rigorous analysis can be achieved by means other than a search for cry patterns to determine whether a sudden increase in sound intensity is actually due to a baby crying, as determined by the average frame sound level of the new frame exceeding the average frame sound level of the preceding frame by at least an amount corresponding to a threshold. For such a more rigorous analysis without cry pattern identification and separation, it should be understood that babies cry for long periods of time, and therefore only sound data associated with sustained high sound intensity should be considered anyway. Accordingly, subsequent frames that exceed a certain intensity, for example, because their root-mean-square frame intensity as defined above exceeds the current minimum noise by an adaptive threshold level, are copied to the cry detection frame buffer for further analysis. On the other hand, if cry pattern detection and separation is implemented using an object localization method taking into account a spectrogram-like representation of the sound data, it is not absolutely necessary to initiate such cry detection and separation only when the sound level exceeds a given threshold. The cry pattern detection and isolation method can be performed continuously based on object localization in a spectrogram-like representation. Thus, the decision whether to implement a proactive check for particularly loud sound levels can be made based on, for example, energy consumption considerations when operating a local device from a battery, or the available data bandwidth for uploading sound data to a central device such as a cloud server when implementing the cry pattern detection and isolation in the cloud.

[0155] If it is found that too few subsequent frames need to be copied to the cry detection frame buffer for further analysis within a certain period of time, it can be safely assumed that the baby is not crying, and the data stored in the buffer can be deleted. Otherwise, an additional test is performed. Consequently, a further cry detection stage is now implemented by counting the frames that are entered into the buffer within a given period of time. While this is the preferred method of rejecting noise as audio that could be crying, in some implementations such additional rejection is not used.

[0156] It is emphasized that because only frames with a sufficiently high average audio level are copied to the buffer, the buffer does not represent a complete sequence of frames, as one or more intermediate frames may have a low average audio level and therefore not be copied to the cry detection frame buffer. It should be understood that this approach differs from a typical laboratory setup, where audio without background noise is available for analysis, and as a result, no frames need to be omitted to account for estimated background environmental noise levels. Note that in embodiments where the crying pattern is identified and isolated as part of the cry detection itself, obviously portions of the originally recorded audio stream will also be excluded.

[0157] However, in some embodiments and implementations, frames are not filtered out after the onset of a potential cry. In implementations that do not filter out frames, cross-correlation / sliding average techniques can be easily used. An advantage of removing frames with low audio levels is that a smaller overall amount of data is handled and analyzed, while one advantage of not filtering out frames is that higher accuracy / precision can be achieved, especially when sliding / cross-correlation techniques are used. In this context, it is understood that once a cry is established to be present in the audio data, it is usually useful to analyze the complete sequence to determine the reason for the cry; therefore, the complete sequence including all frames following the first frame that satisfies the first cry detection stage criteria should be stored anyway. Obviously, when identifying and isolating a cry pattern from a window defined within the originally recorded audio stream, the complete sequence including all frames is also determined, and it can be noted that identifying and isolating a cry pattern as part of the cry detection provides particularly good results in determining the cry.

[0158] It will also be appreciated that providing a separate buffer for the complete sequence may not be necessary if the circular buffer containing all pre-cry data is large enough, in which case the complete sequence may simply be stored in the circular frame ring buffer.

[0159] It should also be emphasized that if the number of frames above a given - and, if applicable, adaptive - threshold is not used to identify and reject short events, the buffer may still be closed, for example because a number of consecutive frames are identified as irrelevant, for example due to low audio levels. Obviously, in that case the buffer will not be completely filled.

[0160] A neural network filter trained on previously identified baby cries is then used to establish whether particularly loud frames constitute part of a baby's cry. Note that this neural network filter may be different from the neural network filter used in the cry determination (or "translation"). As previously mentioned, one way to use a neural network filter is to define a spectrogram-like representation of multiple frames and then search for cry patterns within this representation. As a result, with respect to the cry detection neural network filter, it is understood that, under certain circumstances, certain frequency ranges of the audio energy and spectrogram may provide important clues that differ both between babies and between environments, but the reason the baby is crying is not important, and personalization is not absolutely necessary to improve the accuracy of cry detection. Nevertheless, given that personalization of cry detection and / or cry determination is quite computationally intensive when artificial intelligence techniques such as convolutional neural networks are used, it may be advantageous to limit the necessary personalization to a minimum in order to still obtain desirable results.

[0161] For example, in some environments, frequency bands where babies cry particularly loudly may also experience stronger background noise, making them less suitable for cry detection. Unfortunately, background noise patterns may change even faster than the baby's personalized cry determination, for example, due to frequent changes in the baby's monitored location, and background noise changes due to windows being opened or closed depending on weather conditions, etc. Accordingly, in the most preferred embodiment, the cry detection itself is not personalized. Nevertheless, in practical implementations, cry detection with an accuracy of over 99% in the field can be easily achieved using a properly trained neural network filter as a stage following audio level determination. One advantage of a spectrogram-like representation of an audio window containing extracted crying audio is its robustness to noise; in other words, it is noted that, regardless of noise, the cry pattern is very reliably identified within the window representation using such a technique.

[0162] However, in some cases it may be preferable not to isolate the cry patterns found in the spectrogram-like representation, and in such cases, for cry detection using a neural network filter, the original complete audio data, e.g., each frame in the buffer, can be directly input into the appropriate neural network filter or parameters. Above, several parameters are disclosed that can be extracted from the cry data. Similar parameters can be determined for crying detection, such as the average alleged crying audio of frames in the buffer being considered crying, a sliding average of the alleged crying audio energy over a certain consecutive number and / or frames, in particular 2, 4, 8, 16 or 32 frames in the buffer, and / or over a certain period of time such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the alleged crying audio duration variance of frames in the buffer; the alleged crying audio energy variance over 2, 4, 8, 16 or 32 frames, in particular 2, 4, 8, 16 or 32 frames, and / or over a certain period of time such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the current pitch frequency; the pitch frequency averaged over 2, 4, 8, 16 or 32 frames, and / or over a specific time period such as 1, 2, 5, 10, 15 or 30 seconds; the maximum pitch frequency during the alleged crying audio event and / or over 2, 4, 8, 16 or 32 frames of the alleged crying audio data in the buffer, and / or over a specific time period such as 1, 2, 5, 10, 15 or 30 seconds; a change in sliding maximum pitch frequency during an allegedly crying audio event and / or during 2, 4, 8, 16 or 32 frames of allegedly crying audio data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; a minimum pitch frequency during an audio event attributed to a cry and / or during 2, 4, 8, 16 or 32 frames of audio data attributed to a cry according to buffered frames, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; a sliding minimum pitch frequency change during an audio event attributed to a cry and / or during 2, 4, 8, 16 or 32 frames of audio data attributed to a cry, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; the dynamic range of pitch frequency during the alleged crying audio event and / or during 2, 4, 8, 16 or 32 frames of the alleged crying audio data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the average rate of change of pitch frequency during the alleged crying audio event and / or during 2, 4, 8, 16 or 32 frames of the alleged crying audio data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; Also, assuming the frames represent crying data, formant-related parameters may be determined, such as: a first formant frequency in 2, 4, 8, 16 or 32 frames of an allegedly crying audio event or audio data of the allegedly crying audio, and / or during a particular time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; the average rate of change of the first formant frequency averaged over 2, 4, 8, 16, or 32 frames of audio data purported to be crying, and / or over specified time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the average percentage change of the sliding average first formant frequency over 2, 4, 8, 16, or 32 frames of audio data identified as crying, and / or over specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the mean value of the first formant frequency averaged over 2, 4, 8, 16, or 32 frames of the alleged crying audio data and / or over specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, or 30 seconds; the maximum value of the first formant frequency in 2, 4, 8, 16, or 32 frames of audio data identified as crying, and / or within a specific time period, such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, or 30 seconds; the minimum value of the first formant frequency in 2, 4, 8, 16, or 32 frames of audio data identified as crying, and / or within a specific time period, such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, or 30 seconds; a first resonant peak frequency dynamic range during the purported crying audio event and / or during 2, 4, 8, 16 or 32 frames of the purported crying audio data, and / or during a specific time period, such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; the second formant frequency during the alleged crying audio event and / or during 2, 4, 8, 16 or 32 frames of the alleged crying audio data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the mean rate of change of the second formant frequency during the alleged crying audio event and / or during 2, 4, 8, 16 or 32 frames of the alleged crying audio data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the second formant frequency average during the alleged crying audio event and / or during 2, 4, 8, 16 or 32 frames of the alleged crying audio data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the second formant frequency maximum during the audio event identified as crying and / or during 2, 4, 8, 16 or 32 frames of audio data identified as crying, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the second formant frequency minimum during the identified crying audio event and / or during 2, 4, 8, 16 or 32 frames of identified crying audio data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; a second resonance during the purported crying audio event and / or during 2, 4, 8, 16 or 32 frames of purported crying audio data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the peak frequency dynamic range during the alleged cry audio event and / or during 2, 4, 8, 16 or 32 frames of the alleged cry audio data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; and, again assuming the baby is actually crying, it is possible to determine Mel frequency cepstral parameters, which are determined for the entire alleged cry audio event and / or during 2, 4, 8, 16 or 32 frames of the alleged cry audio data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; and / or inverted Mel frequency cepstral parameters, etc.

[0163] While it is possible to use such parameters for cry detection, it will be readily apparent that, given the large variance noted above for the background, accuracy will frequently not be improved by using more precise, computationally more intensive parameters. Accordingly, the determination of whether a baby is currently crying may be based on parameters selected to remain computationally low while still providing highly accurate cry detection. This in turn allows for cry detection to be achieved locally, even if multi-stage cry detection is used, with at least one stage using neural network filter techniques.

[0164] Accordingly, from the foregoing, it can be seen that as a step in the cry detection stage, the contents in the buffer are further analyzed. Typically, a number of n frames are stored in the buffer and an output signal is generated indicating whether a frame in the buffer is associated with a baby crying; alternatively, it is possible to first identify and isolate crying patterns and then store these crying patterns for further determination of the reason the baby is crying and / or send them to a centralized server and / or process them immediately.

[0165] In a most typical implementation that relies on parameter evaluation rather than cry pattern discrimination, a determination is made as to whether a frame in a buffer is associated with a baby's cry, indicating the probability that the audio data is associated with a cry and / or the degree of reliability of the determination. While such a determination can be made for each frame in the buffer, handling is significantly easier if the determination is made on a buffer-by-buffer basis. In this regard, multiple buffers can be analyzed in sequence, for example, because a first buffer is completely filled and / or a previous buffer has been closed since the temporal distance between two frames exceeding the threshold value became too large. Once several buffers have been analyzed, a final output can be determined, which is therefore a function of the results calculated for each single bar for each single frame. This can be done by averaging the probabilities that each buffer is associated with a cry of the same weight, or by the probability that each buffer is associated with a cry in a manner that takes into account the reliability of each probability. In a preferred embodiment, the number of buffers analyzed during cry detection is set to two or three. In a preferred embodiment, to determine that a cry has been detected, it is required that the probability of the audio data in at least one buffer should exceed 75%, and that the (linear) average of the probabilities over the N=3 buffers is greater than 50%. If all criteria are met, the audio data is determined to belong to a baby's cry. However, if the criteria are not met, the corresponding analysis is repeated for subsequent buffers until no frames are detected that should be considered as cry candidates, taking into account the initial first threshold criterion for a long period of time. In other words, if the output of the last stage of the cry detection analysis is negative for buffer (n, n+1, n+2), the analysis is repeated for buffer (n+1, n+2, n+3) as long as some frames show particularly loud audio.

[0166] It is understood that if no crying is detected in the final stage, the data in the cry detection frame buffers is erased unless it is determined that the background audio data in one or more cry detection frame buffers should be uploaded to a server for training a cry detection neural network filter and / or for identifying typical or highly significant background patterns.

[0167] Such a decision to upload non-crying audio data to the server may be random, or buffered non-crying audio data may be marked for upload because the average probability is very close to the probability determined to indicate crying, or because the average probability is extremely low. In this context, it is understood that the purpose of periodically uploading non-crying data to the server is to identify new background behavior patterns and improve the neural network cry detection filter. Even if non-crying data is to be uploaded, it is understood that data protection regulations are to be adhered to. Specifically, it may be possible to allow uploading only after the person owning the local device agrees to uploading specific non-crying patterns. Also, speech detection may be achieved to prevent the uploading of speech-related audio data, and it may be possible to upload non-crying audio without reference to a specific device.

[0168] It was mentioned above that if no cry is detected in the final stage, the data in the cry detection frame buffer is cleared. In a similar approach, if the number of cry patterns in a given period is very small, i.e., if the baby is only crying for a very short period, such as 5 or 6 seconds, it would be possible to discard the cry pattern even if it has been clearly identified as having a very high likelihood of being associated with a cry. This would be especially true if the final decision relies on a large number of cry patterns, e.g., 5-10 cry patterns.

[0169] Otherwise, if it is determined that a cry has been detected, the parent or caregiver is first notified by sending a corresponding message to the remote station. This is useful because the parent or caregiver needs time to reach the local station, during which time a determination of the reason for the baby's crying can often be made. In addition, or as a message that the baby is crying, audio data of the cry, preferably including pre-onset audio data stored in a circular buffer, can be sent to the remote station for audio playback. It should be understood that sending pre-cry audio can be useful in attracting the parent's attention because the audio played at the remote station is more similar to what the caregiver would hear while near the baby. Furthermore, it can be useful for acoustically assessing the baby's surroundings when the baby begins to cry, which may be useful if the crying is caused by an external influence, such as a sibling or pet entering the room. However, it should be understood that some embodiments require parents and caregivers not to be notified, even if they are in another room, because they remain within hearing range of the baby.

[0170] Then, after sending the message to the remote station, it is determined whether the cry can be determined by a centralized server or whether the determination must be performed locally. To this end, a data transmission file is prepared containing all relevant frames of the crying audio data stream and / or all crying patterns up to the preparation of the file. In embodiments that do not rely on local cry pattern identification and separation, this file may include not only frames that exceed a given threshold after the first frame exceeds the threshold (and thus includes more frames than were buffered for simple cry detection), but also the complete sequence of frames recorded since the first frame exceeded the threshold. Furthermore, the file may include frames from the frame buffer before the cyclical cry that was locked after the detection of the first extremely high audio level. In embodiments where local cry pattern identification and separation is achieved, only the isolated crying patterns are transmitted, and only the first cry identification stage, such as threshold comparison, is achieved, but the cry pattern identification and separation is achieved on a central device. In embodiments, the file uploaded to the central device preferably includes the entire contents of the ring buffer. Personalized baby data, such as weight, gender, age, and medical condition, are then added to the file. This can be done in a coded manner, for example, by including a previously assigned ID. In this case, corresponding information stored in a central database on a centralized server can be retrieved corresponding to the specific ID assigned to the device. This is preferable because data considered confidential by many users does not need to be transmitted very frequently, thus avoiding confidentiality issues. In a preferred embodiment, the device negotiates with the server to obtain permission to upload data using a token system. Negotiating with the server to obtain permission to upload data reduces the load from unnecessary data arriving on the server. Furthermore, incoming data from a specific user can be stored in a predetermined location uniquely assigned to that specific user, thus improving confidentiality.It will be appreciated that if the overall data rate sent to the server is not a concern, then a first local simple determination need not even be achieved, such as by use of a comparator.

[0171] It is understood that the exact content of the data file and / or the exact structure of the data file may vary. It is also possible, though less preferable, to omit pre-cry data—or cry patterns that may be identified in the audio representation even if the baby is no longer particularly fussy—and, if bandwidth is particularly an issue, voids may be left in the sequence of frames, for example, excluding frames that are very close to the minimum background level determined before the onset of a cry. However, this is clearly significantly less preferable, since the results typically obtained with such a method are less accurate. In particular, the possibility of using cross-correlation techniques may be impaired. It should also be emphasized that, with regard to ongoing crying, after the first number of frames have been sent to the server so that a determination can begin, additional data may be collected and sent to the server to refine the determination of the ongoing cry.

[0172] A preferred method of cry translation or cry determination will not be described here, and it is understood that there are multiple ways to determine a baby's cry. It is also understood that the method of providing data for determination proposed by the present application is useful for all, or at least a variety of, such different methods for determining a baby's cry. Nevertheless, by describing an exemplary embodiment of cry determination, it will become clearer how the method of providing data for determination is best implemented.

[0173] To understand cry translation, it should be understood that a baby's cry exhibits distinct characteristics relative to the baby's current specific needs. This can be seen, for example, in the 3D spectrograms shown in Figures 3a-e, which clearly demonstrate the differences between the spectrograms of cries recorded from the exact same baby while they had different needs. What is clearly visible in Figure 3 is that the 3D spectrograms are clearly different. Note that a spectrogram shows a three-dimensional plot of energy content (z-axis) over time for multiple different frequencies (x-axis and y-axis, respectively). Essentially, the same information for different cries is given in Figures 4a and 4b.

[0174] The patterns shown are typical for each reason, so significant differences in the cry could, in principle, make it possible to distinguish one cry from another, or determine why a baby is crying given the audio, as compared in Figures 4a-4f. However, it will be appreciated that patterns will look different not only for different cries from the exact same baby, but also for different babies crying for the same reason, and that the differences will depend, among other things, on the baby's age, weight, height, etc. Nevertheless, significant differences between different types of cries can still be identified, particularly when isolating specific parameters from the cry and / or using machine learning algorithms, as compared in Figures 4a-4f.

[0175] It should then be understood that often there is no single reason for a baby's crying, for example the baby may be tired and need to be burped; the baby may be hungry and need to be comforted, etc. This is reflected in the cries and, accordingly, in the respective crying patterns, so that any given crying pattern may simultaneously determine multiple patterns if the baby is crying, and the appropriate determination preferably takes that into account.

[0176] Regarding the first method for automatically identifying differences between different cry patterns, it may be useful to describe each cry using sufficient parameters or to provide specific parameters derived from the cries to the model. Using sufficient parameters, it is possible to define groups (or "clouds") of cries, with each cloud containing cries for a different reason. This is illustrated in Figures 5a-5e and 6. Note that the patterns shown in Figure 5 are from a raw dataset of cries augmented by an unsupervised deep learning technique called self-organizing maps. In Figure 6, each cry is represented by a dot in a multidimensional parameter space. Different types of dots represent different reasons why a baby cries, and Figure 6 clearly shows that it is possible to group different types of dots, and therefore different reasons why a baby cries. Even if the actual cry determination is based on identified and separated cry patterns in the original audio data stream, learning about new ones and / or peer grouping may rely on self-organizing maps based on specific parameters. As a result, once such peer groups are established, previously obtained cry patterns for each peer group can be used to train a respective model personalized for each peer group. Another way to improve personalization is to determine, for each cry pattern, the probability that this cry pattern is associated with each of a plurality of different reason classes, and from the sequence of cry patterns, an overall determination of why the baby is crying can then be provided by considering the sequence of probabilities for each of the different classes. A simplification of personalization with very good results is to personalize only the determination of this sequence of probabilities. In this way, computationally intensive personalization can be reduced to a minimum while still providing excellent personalized results. It will be appreciated that this simplifies personalization because the assignment of a probability to each cry pattern can be achieved in a non-personalized manner.

[0177] The goal of cry translation, or "cry classification," is to classify audio data identified as belonging to a baby's cry into one of several distinct classes. Figure 3 shows cries for five different baby crying reasons: hungry, uncomfortable, needing to burp, in pain, and sleepy. These five distinct reasons can be used as the classes into which each baby's cry is classified. However, while these classes are very useful and easy for young parents to distinguish, they should not be interpreted as limiting the probability of cry classification. Rather, fewer classes could be implemented, for example, combining "uncomfortable" and "in pain," or more classes could be used, such as classes describing breathing patterns associated with "coughing," "hiccups," and "sneezing." It is also possible to use additional classes not associated with crying at all, such as "silent" or "undefined." The use of such additional classes, intentionally unrelated to crying, helps eliminate potential false positives, since even with a good cry detection stage, audio data incorrectly identified as crying may be forwarded to the cry classification stage. Providing one or more non-crying classes to the filter helps reduce the number of false positives.

[0178] It will be appreciated that this is often the situation encountered in data mining and data analysis, and that as a result, artificial intelligence techniques and in particular neural network techniques such as CNN (Convolutional Neural Network) techniques can be applied to distinguish between different cries, provided that suitable training data can be provided and sufficient parameters can be found.

[0179] Accordingly, some data needs to be provided to either the local or centralized crying determination stage. In both cases, similar techniques can be used, such as artificial intelligence / neural network filtering techniques. It will also be understood that in both cases, the audio data can be determined in a personalized manner, but the degree of personalization and / or the amount of computation available can differ for the local and centralized cases, respectively. Specifically, in the centralized case, the available processing power is often greater, and in some cases significantly greater, than in the case of local determination on a local device. Therefore, since the calculation parameters such as those listed above require at least some computational effort, the number of parameters used as input to the neural network filter in the determination will be even greater.

[0180] As a result, neural network filters used in a centralized versus server configuration may be more complex than filters that can be implemented locally on local stations with less processing power. It is also understood that filter coefficients for local determination are, in the most typical case, determined on the centralized server and transferred from the centralized server to the local device. (Note that with respect to a centralized server, this does not exclude the possibility that the “server” may be spatially distributed, as in the case of a “cloud server,” since this server may be used by multiple users who have voice data transferred to the server from their local devices.) Because personalized filters are typically determined on the server more frequently and then downloaded to the local device, and because in some cases only partial personalization is possible on the local device, filter updates / personalization are better possible when a centralized server is used. For example, if a baby has a fever or a fever within a certain range, the baby's cry may change. Because fevers can occur frequently and naturally, each time the personalized filter coefficients are updated, the corresponding set of filter coefficients needs to be stored locally on the local device. Since similar conditions such as fever need to be taken into account, the memory size required to store a wide variety of different filter coefficients may be very large, and the amount of data that may need to be transferred from the server to the local device to update the filter coefficients for each different condition that may be distinguished on the server will often be too large. Therefore, local determination is often doomed to be less accurate than centralized determination, given the technical difficulties.

[0181] Nevertheless, even for local voice determination, the voice data needs to be provided to the determination stage, and depending on whether the voice data is determined locally or not - the determination can be considered to be personalized if it is done locally, given that a specific filter for a specific baby has been downloaded or retrieved by a push service, and with regard to uploading, if personalization data related to the baby associated with the ID has previously been stored on the server, the voice data can be linked to the ID, and it is understood that the complete personal information can be transmitted, and the decision on whether frequent transmission of personal information or transmission of the ID associated with the personal information stored on the server can be used can be made taking into account data protection regulations and the desire to protect privacy in the best possible way.

[0182] Despite challenges such as available computing power that may affect the exact way in which the audio data is determined in the audio determination stage, it is believed that this example suffices to illustrate what can be done if sufficient computing power is available, for example after uploading the audio data to a centralized server. From this, it is also easy to deduce how local decisions can be affected. For example, if cross-correlation techniques are too computationally intensive, rather than determining the decision by calculating the best correspondence when the input signal is shifted sample by sample, it would be possible to only consider the results obtained when the input signal is shifted frame by frame, or when the input signal is shifted over two frames, thus reducing the computational load by a factor of two.

[0183] In one embodiment, it is assumed that the number of frames initially transferred to the centralized server is sufficient to implement and perform the cross-correlation step, and that once the initially transferred cry data has been determined, the parent can reach the local device and confirm any determinations made initially. Note that confirmation of any determination does not need to be immediate; on the other hand, in many cases, the parent will be convinced to confirm or reject the determination once the baby has stopped crying in response to the parent's actions. Also, even if the determination is immediately considered correct, the parent should care for the baby before evaluating the determination. Therefore, the determination can generally be made later, for example, using a smartphone running an appropriate app. Nevertheless, in one embodiment, the parent may have the possibility to immediately input a determination or feedback into the device. In this embodiment, if the parent has not yet determined the determination, additional data can be uploaded, and the cry determination can be achieved as if a larger file had been transmitted initially. The only difference compared to when data is transmitted repeatedly rather than as a large file is that the first determination can be made and transmitted using the first portion of the data. As a result, if more data is received, the determination may be correct or confirmed; if the initial determination does not change with analysis of more data, the user may not even know that additional data is being determined; specifically, it can be expected that providing more data may increase the probability that the determination is correct, unless a probability that the determination is correct is indicated. If the determination changes over time, the user can be clearly advised that the determination is changing, lest the user think that they might refer to the initial determination as a glitch. It will be understood that transferring data to the cry determination stage may continue until the user confirms the cry determination and / or until the baby has stopped crying. Nevertheless, in light of the foregoing, it will be sufficient for this application to describe when the determination is achieved, considering only the first file transferred to the centralized server.

[0184] From the foregoing, it can be seen that analyzing a long sequence of frames is both possible and useful. Reference was made above to a first buffer containing, for example, 1024 frames. Multiple such buffers can be packed into a single file, which is then analyzed to determine why the baby is crying. While detecting that the baby is crying as early as possible is useful because the corresponding information should be communicated to the parent or caregiver as soon as possible, it should be understood that the determination may take somewhat longer, regardless of the reaction time required by the parent or caregiver. Therefore, using a large number of frames for cry determination typically does not constitute a significant problem. Accordingly, if cry detection is preferred to function based on three or fewer buffers, the cry determination can be achieved significantly earlier with a larger number of buffers, such as 5, 6, 7, 8, or 16 buffers (each buffer holding, for example, 1024 frames). Nevertheless, it is preferable that the first determination be made available to the user within less than 15 seconds, preferably 10 seconds or less, after the cry detection. Otherwise, the user may consider the local device unresponsive.

[0185] From the foregoing it will also be appreciated that it is preferable to use multiple buffers for cry detection, and that the local device should preferably have sufficient memory to store at least 16 buffers, with more preferably needing to be stored on the local device for later loading in the event that the detective cries cannot be analyzed on the central server.

[0186] Once enough data has been collected and uploaded to the centralized server for determination in the cry determination stage, pre-processing can begin. During pre-processing, a set of filter parameters corresponding to the personalization information is determined, for example, by referring to a filter parameter set database, and a neural network filter is configured according to this set of filter parameters.

[0187] The audio data itself is then fed to the personalized neural network filter, for example frame by frame, or parameters describing the audio data are determined and the determined parameters describing the audio data are input to the personalized neural network filter.

[0188] As mentioned above, parameters that may describe the audio data may be input to the neural network filter and may include: average cry energy during the current crying event, a sliding average of cry energy over a certain number of consecutive frames, specifically within 2, 4, 8, 16, or 32 frames, and / or over a certain period of time, such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; cry duration variance between breaks in an event; specifically within 2, 4, 8, 16, or 32 frames, and / or over a certain period of time, such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the cry energy variance over a period of time, such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the current pitch frequency; specifically, the average pitch frequency over 2, 4, 8, 16 or 32 frames, and / or over a period of time, such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; during a cry event and / or over 2, 4, 8, 16 or 32 frames of cry data, and / or over 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc. The maximum pitch frequency during a specific time period, such as 5 seconds, 30 seconds, etc.; the sliding maximum pitch frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods, such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the sliding maximum pitch frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods, such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc. the minimum pitch frequency; the sliding minimum pitch frequency change during a cry event and / or during 2, 4, 8, 16, or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; the dynamic range of pitch frequency during a cry event and / or during 2, 4, 8, 16, or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds;mean rate of change of pitch in frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; first formant frequency during a cry event or during 2, 4, 8, 16 or 32 frames of cry data, and / or during specific time periods such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; cry data the average percentage change of the first formant frequency averaged over 2, 4, 8, 16 or 32 frames of cry data;the average percentage change of the first formant frequency averaged over 2, 4, 8, 16 or 32 frames of cry data;the average percentage change of the first formant frequency averaged over 2, 4, 8, 16 or 32 frames of cry data;the average percentage change of the first formant frequency averaged over 2, 4, 8, 16 or 32 frames of cry data; the mean value of the first formant frequency over 2, 4, 8, 16 or 32 frames of cry data and / or over a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; the maximum value of the first formant frequency over 2, 4, 8, 16 or 32 frames of cry data and / or over a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds; the maximum value of the first formant frequency over 2, 4, 8, 16 or 32 frames of cry data and / or over a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds a first formant frequency minimum during a specified time period, such as 30 seconds; a first resonant peak frequency dynamic range during a cry event and / or during 2, 4, 8, 16, or 32 frames of cry data, and / or during a specified time period, such as 1, 2, 5, 10, 15, or 30 seconds; a second formant frequency during a cry event and / or during 2, 4, 8, 16, or 32 frames of cry data, and / or during a specified time period, such as 1, 2, 5, 10, 15, or 30 seconds;the mean rate of change of the second formant frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data and / or during a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the mean second formant frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data and / or during a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the maximum second formant frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data and / or during a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc.; the maximum second formant frequency during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data and / or during a specific time period such as 1 second, 2 seconds, 5 seconds, 10 seconds, 15 seconds, 30 seconds, etc. a second formant frequency minimum during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during a specific time period such as 1, 2, 5, 10, 15, 30 seconds; a second resonance during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during a specific time period such as 1, 2, 5, 10, 15, 30 seconds; a peak frequency dynamic range during a cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during a specific time period such as 1, 2, 5, 10, 15, 30 seconds; Mel frequency cepstral parameters, parameters determined for the entire cry event and / or during 2, 4, 8, 16 or 32 frames of cry data, and / or during a specific time period such as 1, 2, 5, 10, 15, 30 seconds; and / or inverted Mel frequency cepstral parameters.

[0189] As noted above, where references in the list of parameters are made to specific times, they may also refer to any other fixed period up to the respective time explicitly mentioned, and the reader is reminded that reasons why certain lengths are convenient are given above.

[0190] It has also been repeatedly emphasized above that advantages can be gained if patterns represented by the audio data are analyzed cross-correlatedly by considering different potential expressions of crying, which are typical for different reasons, and compared to the patterns represented by the neural network filter; considering different potential expressions of crying can be easily done if sliding averages or other sliding parameters are determined for each frame or sample, and highly corresponding sequences of parameters are used as inputs to the neural network filter, respectively. However, it will be appreciated that such techniques are more computationally intensive.

[0191] In another embodiment, if a cry pattern has been isolated during the cry detection stage and / or the time at which the cry pattern occurs has been determined, rather than calculating specific parameters for the frames associated with that time, a spectrogram-like representation of the audio recorded during the period in which the cry pattern was observed can be prepared, and if the same temporal and frequency resolution can be used for cry identification and separation, the previously isolated cry pattern can be used, eliminating the need to prepare an additional spectrogram-like representation. The spectrogram-like representation can then be fed into a convolutional neural network, which outputs the likelihood that each cry pattern fed into the convolutional neural network belongs to one of the predefined classes. In a particularly accurate, but more computationally intensive, embodiment, each convolutional neural network filter parameter is personalized, i.e., a different set of filter parameters is used for each defined peer group. In another implementation, it would be possible to use the same set of filter parameters for all peer groups, determine the probability that each isolated cry pattern belongs to each of the different reasons why a baby cries, and then evaluate the combined reasons obtained from each cry pattern in a personalized manner. It should be appreciated that by personalizing only one or two or three final layers, both the amount of computation required to train the model and the number of parameters used in the decision and that need to be stored can be reduced.

[0192] For clarity, it is emphasized that although a neural network filter can be used in the final stage of cry detection, and a neural network filter implementation can also be used for cry determination, the neural network filter used in the final stage of cry detection is different from the filter used in cry determination, as are the overall inputs to the respective neural network filters for cry detection and cry determination, respectively. Filters used for cry translation are likely to be more complex, e.g., using more layers and / or more inputs in the convolutional neural network, e.g., using higher resolution, and / or having personal parameters and / or additional inputs to more layers.

[0193] It should be understood that once a cry is detected, it is not absolutely necessary to determine the cry. If the parent is confident that they understand why the baby is crying without help from an electronic device, it is sometimes sufficient to notify the parent that the baby is crying. Therefore, in such cases, there is no need to translate the cry, thus, for example, saving energy. Accordingly, in some embodiments, it may even be possible to trigger automatic cry determination only if the parent explicitly requires assistance. In yet another configuration, it may be possible to not allow cry translation at all and to use only cry detection to improve the responsiveness of the baby monitor. In such a configuration, transmission may be reflected only in response to the detection of a cry, and / or audio may be transmitted in such a way that it is extra loud at the receiver, for example, by changing the gain of the digital signal in response to the detection of a cry.

[0194] In the cry detection stage, the probability that the audio data input to the neural network filter belongs to one of several predefined classes, such as "hungry," "sleepy," or "needs to burp," is determined. Accordingly, several probabilities are obtained, resulting in n sets of probabilities, with the components of the n sets representing the probability that the baby was crying for reasons related to each class. Going from frame to frame or buffer to buffer, the n sets of components obtained each time are different. Therefore, an overall detection must be calculated from the n sets of sequences.

[0195] There are various probabilities for determining the overall decision. For example, the average of each component of the N sets can be calculated, and the component with the largest average value, and therefore the highest overall probability, is selected as the decision. This average can be a linear average, a root-mean-square average, or the like. In a preferred, simple embodiment, a linear average is calculated. Also, taking into account that cross-correlation techniques can yield very high probabilities for very good matches, the maximum values ​​for each component across all sets can be compared; since the pattern matching achievable with cross-correlation techniques is not perfect due to sampling and noise, it may be even preferable to consider the maximum value of a sliding average, for example, averaging each component of 2, 3, 4, or 5 consecutive n sets. If the maximum value for a given component does not exceed a certain threshold for all frames considered, some components can be completely excluded from consideration; in this way, components that never yield a satisfactory match will not cause an erroneous decision.

[0196] Another possibility would be to construct an N×M matrix from M consecutive N-tuples acquired for M consecutive frames, and then feed this M over M matrix to a further neural network filter for final decision (or use a similar technique by implementing a corresponding layer of a convolutional neural network). Although reference has been made to a neural network filter, it should be understood that a wide variety of data processing techniques are contemplated, particularly for neural network filter implementations, particularly with regard to the number and size of filter layers, and therefore details of such filters will not be described herein. It will of course be understood that, in general, fewer layers and / or less complex layers are typically implemented on a local device.

[0197] In a particularly preferred embodiment, the cry detection comprises an (optionally) very simple (optionally: first stage) that detects only recorded sound levels that are above a threshold, and a (second) stage of the cry detection comprises identifying and isolating (possibly crying) patterns within overlapping time windows that completely cover the sound stream (or optionally both the time immediately before and after the excessive sound levels), where the crying patterns are identified by searching for patterns known to correspond to baby cries within a spectrogram-like representation of the time windows, and matching these to the identified and isolated crying patterns. The method further comprises subjecting the baby to a turn classification, establishing probabilities for reasons why the baby is crying taking into account the crying patterns, in particular taking into account spectrogram-like representations of the isolated crying patterns, and then determining a specific reason why the baby is crying taking into account the probabilities obtained for the sequence of isolated crying patterns, wherein at least one of the pattern classification and / or identification taking into account the probabilities of specific reasons why the baby is crying is performed by personalized determination according to personal baby data, in particular personalizing the step of determining a specific reason why the baby is crying taking into account the probabilities obtained for each crying pattern.

[0198] Regardless of the final decision on how to implement the neural network filter and / or algorithm that selects a decision given the output of a given input, it will be understood that the reliability of the decision strongly depends on the quality of the data available for analysis. It will be understood that techniques such as cross-correlation and / or sliding parameters are particularly useful for more accurate analysis and, in particular, personalizing the decision, and that providing data suitable for such techniques is extremely important for enabling improved decisions. In this regard, it will be mentioned that a practical implementation of an embodiment in which the first cry detection stage simply compares the current sound level with the average preceding background sound level, and subjects the sound data recorded just before and after a loud noise above a comparable tort threshold to cry pattern identification and separation based on a spectrogram-like representation has been found to produce particularly good results in cry determination and to be particularly robust, for example, regardless of the precise placement of the sound recording device relative to the sound source, the specific sound recording device and microphone used, and / or any typical background noise present when recording a baby's cry.

[0199] Once a decision has been obtained, a corresponding output needs to be generated. To this end, the result of the decision is fed into an output stage (or "output manager") which generates a corresponding output signal, which can be an audible signal, a visual signal, for example a pattern displayed on a monitor or a flashing LED.

[0200] It will be appreciated that showing the output can be achieved by a specific output stage adapted to improve the user's experience. For example, the output manager may suppress output of the initial decision if an initial decision related to only audio data from one or two buffers is forwarded to the output stage, or may suppress output of the initial decision if not much time has passed since notifying the user about the onset of crying. Not showing the initial decision immediately helps to avoid confusing the user with decisions that change over time. Also, if the user sets preferences, for example indicating that they want the two most likely reasons shown along with their respective probabilities, the output manager may prepare such output on demand.

[0201] If the determination step indicates that the cry detection may have triggered due to a false positive, corresponding information may also be shown to the user and / or the corresponding request to move the baby may be cancelled. Also, if a first determination is initially made based on a large number of frames and / or crying patterns, especially with a sufficiently high probability, situations may arise in which a second determination different from the first may also be justified, for example because the reason for the baby's crying has actually changed. However, again in this case it may be preferable to prevent the determination from changing for a certain period of time, such as 2 or 3 minutes, to avoid confusing the user.

[0202] Once the baby has stopped crying after a period of time, for example, 30 seconds, 1 minute, 2 minutes, 3 minutes, 4 minutes, or 5 minutes after the crying has stopped, the display of the corresponding flashing LED or generation of the audible signal may cease and a standard message such as "listening to baby cry" or "tracking audio stream" may be generated instead. Indicating to the user why the baby was crying may be useful in some cases, often if the baby is particularly exhausted, because the reason for the previous discomfort is still valid, but the baby has fallen asleep. Displaying the previous determination for a while may therefore be helpful to the parent or caregiver.

[0203] However, in a typical situation, it is expected that a parent or caregiver will tend to the baby while the baby is still crying. They will then typically attempt to soothe the baby, for example, by feeding the hungry baby, by helping the baby burp, or by soothing the baby until it falls asleep. Feedback can be entered into the local device depending on the displayed determination and the success or failure of their attempts to soothe the baby in light of the determination. This is particularly useful if the feedback is transmitted to a centralized server, preferably in such a way that the feedback can be related to previously determined audio data. It will be appreciated that such feedback can be useful for improving the cry database, and in particular for providing users with tagged samples to improve the database. It should be appreciated that uploading the feedback, information associating the feedback with specific audio data and any previously derived decisions, and personalization information to the centralized server benefits both the operator of the centralized server, since the upload helps to expand the database, and the parent, since having multiple tagged audio cries, and preferably also the personalization information, helps to improve personalization, for example, by identifying peer groups of other babies with similar crying patterns. This helps to distinguish between groups of babies even if other parameters, such as gender, age, height, weight, etc., are identical, resulting in improved personalization. Such personalization, which takes actual cries into account, is also useful when parameters entered by parents, such as height or weight, are out of date or entered incorrectly.

[0204] It will also be appreciated that once a peer group of other babies has been identified, information obtained for such a peer group can be used for a particular baby known to have a crying pattern similar to that of the peer group. For example, if all of a given baby's cries are found to closely resemble those of a peer group of other babies with a particular rare disease, a corresponding alert can be issued to the parents. It will be appreciated that methods can be implemented to reward parents and other caregivers when they provide feedback; for example, if a subscription model is implemented, a refund can be made or a current subscription can be extended without additional payment. Therefore, in preferred embodiments, incentive generation means and / or incentive steps are provided for uploading feedback to a centralized server. It will be appreciated that if the connection to the centralized server is interrupted, the relevant data is stored in the local device until a connection is established and the data is transferred.

[0205] It should be understood that the feedback need not only relate to the accuracy of the translation, but the feedback can also be related to the accuracy of the cry detection. It should also be understood that there are multiple ways to implement the feedback, for example, using an app on a smartphone used as a remote station to the local device, pressing a button on the local device, or speaking into the microphone of the local device to confirm or reject the determination.

[0206] Depending on the size of the database, the neural network filter and therefore the decision may not be as clear initially as later decisions as more samples are collected and more cases can be distinguished, so that a general cry detection filter may be used, but as the database grows, the filter becomes increasingly clear.

[0207] Thus, once sufficient samples are collected from, for example, particularly heavy or large babies, which sound different from small, light babies, further differentiation is expected. Filter updating can be automated, for example, by adapting the filter once a week to an average, generic filter for a particular age / peer group. Current knowledge also suggests that while there are no significant differences in the cries of very young newborns, as an infant grows, it is expected that cries will become more differentiated periodically, for example, depending on the infant's associated country of birth or native language. Consequently, properly adapting the cry detection filter can enable longer device use and / or more accurate results for older infants, which is particularly useful when non-standard cries unknown to the parent, such as those associated with a particular illness, are analyzed.

[0208] In this context, it can be assumed that as a baby grows, it is likely to develop similarly to other babies, and that similar babies in its peer group of babies of the same age / same sex / same height will therefore experience similar development of their vocal tract. This assumption can be considered valid unless contradictory information is input by parents and / or the cries tagged by parental feedback do not differ from the corresponding cries of babies in the same peer group, so that in many cases a filter can be determined for each peer group. However, peer group determinations can also be made each time the filter is updated and / or after a certain period of time has passed.

[0209] The data uploaded to the centralized server may be entered into a database, and the audio samples in the database, tagged with user feedback, may be repeatedly used to retrain the neural network filters used in cry translation, and, as long as background noise is also sent to the database, to retrain the neural network filters used in cry detection. It is believed that such techniques are well known in the art for training neural network filters, retraining databases to take into account recency, etc. This makes it possible to provide adaptive filters even when only a limited amount of data has been uploaded for a particular child, for example because a parent does not want to send the data for privacy reasons.

[0210] It has been mentioned above that the method is applicable and the device can be used in a pediatric station. In a pediatric station, the local device may pick up sounds from more than one baby. The same applies, for example, to twin monitoring. In a setup where there is a risk that the local device may pick up sounds from more than one baby, several possibilities exist. First, it would be possible to connect multiple microphones to the local device, with each microphone placed very close to one of the babies. Then, a determination of which baby is crying can be made by considering the sound intensity received from each of these microphones. If multiple local devices are used, rather than using multiple microphones connected by wire cables to a single local device, the devices can exchange information about the sound intensity recorded by each local device, and the decision can be based on the exchanged information. It will be understood that this is useful even for identical twins. Another possibility would be to detect any crying and determine it using multiple different personalizations, each corresponding to one of the babies being monitored. For each personalization, the likelihood that the cry judgment is correct can be determined and the judgment with the highest likelihood can be issued. Another possibility is to show all possible judgments and let the caregiver decide which baby is crying and, accordingly, which judgment is relevant. This would be a preferred embodiment for a pediatric station. As a result, the device can easily be used for multiple babies simultaneously.

[0211] Thus, what is proposed above is, inter alia, and without limitation to application, a computer-implemented method for providing data for automatic baby cry determination, comprising the steps of acoustically monitoring a baby and providing a corresponding stream of audio data, detecting a cry in the stream of audio data, selecting cry-related data from the audio data in response to detecting the cry, determining parameters from the selected cry data to enable cry determination, determining personal baby data for personalized cry determination, preparing a determination stage for determination according to the personal baby data, and supplying the parameters to the cry determination stage prepared according to the personal baby data. It is noted that in certain embodiments the parameters may be a spectrogram-like representation of the audio data corresponding to the time at which a cry pattern is identified and / or the isolated cry pattern.

[0212] It is further disclosed that with respect to the proposed method, the baby can be continuously acoustically monitored and pre-cry audio data is at least temporarily stored, for example until subsequent audio data is known not to be associated with a cry because no crying pattern has been identified therein with a significantly high probability. The method, as proposed, also discloses that within the continuous acoustic monitoring stream, baby cries, and in particular onsets of baby cries, are detected based on at least one of: a current audio level above a threshold; a current audio level above average background noise by a given limit; a current audio level in one or more frequency bands above a threshold; a current audio level in one or more frequency bands above corresponding average background noise by a given limit; a temporal pattern of audio; a temporal and / or spectral pattern of audio levels that deviates from the temporal and / or spectral pattern of sudden loud non-cry noises; and non-acoustic cues, in particular a motion detector and / or a breathing detector, derived from video monitoring data of the baby.

[0213] It is also noted that the selected cry data from which parameters enabling cry determination are determined may include audio data from the onset of a cry event, in particular audio data from the first 2 seconds of the cry, preferably from the first second of the cry, and particularly preferably from the first 500 ms of the cry. This can be done by examining the cry pattern in a way that isolates a window that includes the time preceding the rise in audio level, if present. Also disclosed is a computer-implemented method that may additionally include the steps of locally detecting cries in audio obtained from an acoustically monitored baby and uploading the data to a server configuration for use in centralized automatic baby cry determination, in particular uploading selected data for determining baby cries in the cloud.

[0214] It is also disclosed that the method may then include a step of uploading data related to the acoustic monitoring of the crying baby and / or parameters related to the selected cry that enable the cry determination to the cloud, and / or storing at least a portion of the cries and / or parameters derived from the cries together with the personal baby data on a server and establishing the determination taking into account the information stored on the server.

[0215] It should be understood that the disclosed and proposed method may also include the step of downloading information from a centralized server that enables local personalized baby cry determination, in particular that enables local personalized baby cry determination for a limited period of time. It is noted that in the preferred computer-implemented method as disclosed and proposed, monitored audio data acquired prior to the onset of a cry is used to determine the acoustic background and / or to determine additional parameters for baby cry determination, in particular if the exact onset cannot be determined with a sufficiently high probability.

[0216] It is then also proposed that the parameters be supplied to the cry determination stage in a way that allows for the determination of the cry using neural networks and / or artificial intelligence techniques, particularly when the parameters supplied to the cry determination stage are obtained by transfer learning and / or by training a model on the cries of only one baby.

[0217] It is further disclosed that the method may also include uploading the parameters and / or the recorded audio data stream together with baby data information, particularly baby data information related to at least one of age, sex, height, weight, ethnicity, only child / twins / triplets, current medical condition, known medical prerequisites, particularly known current illness and / or fever, parent and / or caregiver language, and / or uploading baby data information related to the accuracy of one or more previous determinations.

[0218] It is stated that it is disclosed that the parameters determined from the selected cry data are selected to enable determination of at least one of the following conditions: "baby is tired," "baby is hungry," "baby needs to be soothed," "baby needs to be burped," and "baby is in pain."

[0219] An automatic baby cry determination arrangement is also disclosed, which in one embodiment includes a microphone for continuously acoustically monitoring the baby, a digital conversion stage for converting the monitored audio stream into a stream of digital data, a memory stage for storing personal baby data information, a communication stage for transmitting data to a centralized server arrangement, and a cry identification stage is provided for identifying onset of crying within the stream of digital data, wherein the communication stage is adapted to receive data related to personalized determination of the baby's cry from the centralized server arrangement.

[0220] It is further disclosed that the automatic baby cry determination arrangement may further include a feedback arrangement for obtaining feedback information related to the accuracy of one or more previous determinations, and the communication step is adapted for sending the feedback information to the centralized server arrangement.It is proposed that the automatic baby cry determination arrangement may include a local determination step, and the local determination step is adapted to determine the baby's cry taking into account data received from the centralized server arrangement related to personalized determination of the baby's cry.It is noted that the automatic baby cry determination arrangement may include a timer and an evaluation step for evaluating the current age of the personal baby data information and / or the age or validity of data received from the centralized server prior to determining the baby's cry and related to personalized determination of the baby's cry, and the baby cry determination arrangement is configured to output a baby cry determination depending on the evaluation.

[0221] Although certain features are expressly disclosed as combinable in view of the dependency of the claims as filed, it is not intended that the disclosure be limited to only the combinations disclosed in the claims as originally filed. For example, one embodiment of the method according to attached claim 2 may be preferred, in which the convolutional neural network for cry pattern detection and separation is not personalized, and the search for cry patterns is achieved using a convolutional neural network to identify cry patterns but is not based on a spectrogram-like representation of audio data. Also, for example, using a computer-implemented method as recited in claim 5, it would be possible to upload any audio-related data, such as the identified and separated cry pattern representations, along with baby data information related to at least one of age, gender, height, weight, ethnicity, whether only child / twins / triplets, current medical condition, known medical prerequisites, particularly known current illnesses and / or fevers, parent and / or caregiver language, and / or baby data information related to the accuracy of one or more previous determinations, even if the machine learning model for cry pattern detection and separation is personalized.

Claims

1. 1. A computer-implemented method for providing data for automated personalized baby cry determination, comprising: obtaining personal information related to the baby, including at least the baby's age, for personalized crying assessment; acoustically monitoring the baby in a background noise environment to provide a corresponding stream of audio data samples; detecting patterns associated with crying within the stream of audio data samples of said acoustic monitoring, taking into account at least temporal and / or spectral patterns of said audio; selecting a pattern associated with the detected crying for further determination; and the selected crying-related patterns; Together with the obtained personal information, providing the data to a determination stage for further determination, a predefined age range that includes the age of the baby, successively evaluating the patterns associated with each selected cry by comparing the pattern associated with each selected cry with patterns known to correspond to different classes of reasons for crying for babies in the predefined age range, to yield a plurality of probabilities that each of the patterns associated with each selected cry belongs to each of the different classes; establishing a sequence of such probabilities by said successive evaluations of patterns associated with cries; determining a reason for the baby's crying based on the sequence of the plurality of probabilities; performing a determination of a pattern associated with the selected cry based on the obtained personal information; A computer-implemented method comprising the steps of:

2. 2. The method of claim 1, wherein a sequence of audio data windows is established, a spectrogram-like representation is established for each window, a pattern associated with a cry is identified in the window, and data associated with the pattern associated with the cry is selected for further determination using windows that overlap in time.

3. 3. The method of claim 2, wherein searching for cry-associated patterns is accomplished using a convolutional neural network to identify the cry-associated patterns within the spectrogram-like representation of the audio data.

4. 4. The method of claim 3, comprising at least temporarily storing audio data in such a way that a temporal and / or spectral pattern can be established for the search for patterns associated with crying based on audio data acquired at least in part before the audio level of the baby's crying exceeds a threshold.

5. 5. A computer-implemented method according to any one of claims 1 to 4, wherein classes are used such that determination of at least one of the following conditions can be achieved: "baby is tired", "baby is hungry", "baby needs to be soothed", "baby needs to be burped", "baby is in pain".

6. uploading the voice-related data to a central device together with baby data information related to at least some of the following: age, sex, height, weight, ethnicity, single / twins / triplets, current medical condition, known medical pre-conditions, in particular known current illnesses and / or fevers, parent and / or caregiver language; and / or uploading to the central device the baby data information related to the accuracy of one or more previous determinations; The computer-implemented method of any one of claims 1 to 5, comprising:

7. Detecting a baby cry in the stream of audio data samples of said acoustic monitoring by considering the temporal and / or spectral patterns of said audio is achieved as part of a multi-step / multi-stage cry discrimination with a preceding detection step of considering whether an audio level above a threshold is observed, said step of considering whether an audio level above a threshold is observed: the current audio level above the threshold, the current voice level above the average background noise by a given limit, the current sound level in one or more frequency bands that are above a threshold; the current sound level in one or more frequency bands above the corresponding average background noise by a given limit; the temporal pattern of the audio; a temporal and / or spectral pattern of speech levels that deviates from the temporal and / or spectral pattern of a sudden loud non-crying noise; In particular, non-acoustic cues derived from video surveillance data of the baby; Motion and / or breathing detectors is based on at least one of and / or such comparison is accomplished locally, in particular where identifying patterns associated with crying in a spectrogram-like representation of said audio data using a convolutional neural network is accomplished on a data processing arrangement remote from said baby, in particular in a cloud server. A computer-implemented method according to any one of claims 1 to 6.

8. locally detecting whether sounds from the acoustically monitored baby are above said threshold; In response to detecting the sounds exceeding the threshold, uploading data to a server configuration for use in centralized automatic crying-related pattern detection. The computer-implemented method of claim 7, comprising:

9. An automated baby cry determination device suitable for carrying out the automated personalized baby cry determination method according to any one of claims 1 to 8, said device comprising: a microphone for continuously acoustically monitoring sounds from the baby to provide an audio stream; a digital converter for converting the audio stream into a stream of digital data; a memory for storing personal information relating to said baby; A communication module for sending data to a centralized server configuration a cry identification module for identifying patterns associated with crying within said stream of digital data; Equipped with the communication module is adapted to transmit digital data identified by the cry identification module as a pattern associated with a cry to the centralized server configuration that performs a determination of the digital data, and to receive data related to a personalized determination of the pattern associated with the cry from the centralized server configuration. Automatic baby cry detector.

10. 10. The automatic baby cry determination device of claim 9, further comprising a feedback arrangement for obtaining feedback information related to the accuracy of one or more previous determinations, wherein the communication module is adapted to transmit the feedback information to a centralized server arrangement.

11. 11. The automatic baby cry determination device according to claim 9 or 10, further comprising a local determination stage, wherein said local determination stage is adapted to determine the baby's cry taking into account data received from said centralized server arrangement related to personalized determination of the baby's cry.

12. A timer and An evaluation stage, Personal Baby Data Information current age and / or an evaluation step, prior to said determining the baby's cry, of evaluating the age or validity of data received from said centralized server arrangement and related to personalized determination of the baby's cry; Equipped with the baby cry determination device is adapted to output a baby cry determination in response to the evaluation. The automatic baby cry detection device according to any one of claims 9 to 11.

Citation Information

Patent Citations

  • Automatic identification method and system for the cause of infant crying

    CN106653059B

  • Analysis system for baby'S voice

    JP2002278582A

  • Cry monitor system for infant

    JP2002367049A

  • Program and method for supporting childcare

    JP2019095931A