A method, device and system for analysing audio data
The method and system analyze ultrasonic breathing patterns using CNN and RNN algorithms to improve the accuracy of identifying individuals, addressing the limitations of voice biometrics in existing systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- THE SEC OF STATE FOR DEFENCE IN HER BRITANNIC MAJESTYS GOVERNMENT OF THE UK OF GREAT BRITAIN & NORTHERN IRELAND
- Filing Date
- 2026-01-22
- Publication Date
- 2026-07-30
AI Technical Summary
Existing speaker verification systems rely on voice biometrics, which have not been widely applied to breath signatures, limiting their effectiveness in determining the presence and identity of individuals based on breathing patterns.
A method and system that analyze audio data, particularly in the ultrasonic frequency range, to identify the presence and condition of individuals by comparing a trained dataset of breathing information with sampled audio data using image classification algorithms, such as CNN and RNN, to determine the presence and identity of individuals.
Enhances the accuracy of determining the presence and identity of individuals by utilizing unique acoustic features in breathing patterns, providing conclusive results and reducing uncertainty through the use of ultrasonic frequency analysis.
Smart Images

Figure IB2026050587_30072026_PF_FP_ABST
Abstract
Description
[0001] Applicant file ref: D1899 / GBD
[0002] 1
[0003] A METHOD, DEVICE AND SYSTEM FOR ANALYSING AUDIO DATA
[0004] Technical Field of the Invention
[0005] The invention relates to a method for analysing audio data, specifically a method for analysing audio data to determine the presence of one or more persons. The invention extends to a related device and system for analysing audio data to determine the presence of one or more persons.
[0006] Background to the Invention
[0007] Breathing analysis is the study of sounds made from breathing and is relevant to the fields of medical diagnosis, therapy, criminal justice and intelligence.
[0008] It is generally known that breathing comprises unique acoustic features, which can be captured, recorded and digitally processed using available signal processing techniques. These unique acoustic features can include both anatomical and behavioural components, including pitch, loudness, rate and patterns.
[0009] Many electronic and online communication devices and business services rely on voice biometrics to provide speaker verification. Speaker verification generally requires the user to record a predefined word or phrase to be stored in an authentication database. In order to access the service or device subsequently, the user is requested to input their vocal password, which is then compared to the stored vocal signature for that user. If the input vocal sample matches the stored vocal signature, the user’s identity is authenticated and access to the device or service is permitted. Such a technique has not been widely applied to breath signatures.
[0010] It is an aim of the invention to provide an improved method, device and system for analysing audio data to determine the presence of one or more persons in an environment.
[0011] Summary of the Invention
[0012] According to a first aspect, the invention provides a method for analysing audio data, the method comprising the steps of:Applicant file ref: D1899 / GBD
[0013] 2
[0014] i. Providing a trained dataset of one or more individuals, the trained dataset including breathing information of one or more individuals collected at least in part in the ultrasonic frequency range;
[0015] ii. Sampling audio data, at least in part in the ultrasonic frequency range; iii. Processing the sampled audio data to generate an evaluation dataset; iv. Comparing the evaluation dataset with the trained dataset to determine whether one or more persons is present; and
[0016] v. Providing an output comprising results.
[0017] The one or more individuals in the trained dataset concerns one or more humans whose identity may be known or unknown. For each of the one or more individuals, the trained dataset includes breathing information collected in part in the ultrasonic frequency range. For example, the trained dataset may comprise breathing information of the one or more individuals recorded solely in the ultrasonic frequency range or breathing information recorded in the ultrasonic frequency range and at least one other frequency range, such as the audible frequency range and / or infrasonic frequency range.
[0018] In its broadest sense, the method according to the first aspect provides the advantageous effect of determining whether one or more persons is present in an environment. A person being a human whose identity may be known or unknown. The environment being the location in which audio data is sampled. The environment may be isolated (e.g. in a room with access control) or open (e.g. in a public space). The method according to the first aspect can therefore determine if one or more live humans in general is present in an environment.
[0019] The human ear can ordinarily hear sound at frequencies between 20 Hz and 20 kHz. This frequency range is referred to as the human audible range or audible sound. Sound having frequencies equal to and above 20 kHz is referred to as ultrasound and sound having frequencies below 20 Hz is referred to as infrasound. Sound in the ultrasonic range (>20 kHz) or in the infrasonic range (<20 Hz) is generally inaudible to humans. Audio data for the purposes of the invention is sound information that may fall within the audible, ultrasonic and / or infrasonic frequency range.Applicant file ref: D1899 / GBD
[0020] 3
[0021] The inventors have identified that breathing information sampled at least in part in the ultrasonic frequency range includes additional features that are unique to one or more persons. Accordingly, breathing information sampled at the ultrasonic frequency range can advantageously provide information that can reduce uncertainty when determining if one or more persons is present in an environment.
[0022] The trained dataset of the one or more individuals may include breathing information of one or more known individuals. A known individual refers to a human of known identity and from whom breathing information is stored as part of the trained dataset. Preferably, the trained dataset includes breathing information of one or more known individuals collected in part in the ultrasonic frequency range. The breathing information may include metadata corresponding to the one or more known individuals. The metadata may comprise personal information of the known individual. For example, the personal information may include basic identification information such as their name. Further personal information, such as address, age and photograph may also be included in the metadata.
[0023] In addition to the trained dataset including breathing information of one or more known individuals, at least one person of the one or more persons whose presence is determined by the method according to the first aspect may be a known individual. For example, the method may advantageously determine whether one or more known individuals is present through comparing the evaluation dataset with the trained dataset, where the trained dataset may comprise breathing information of one or more known individuals including respective metadata associated with the one or more known individuals. The first aspect of the invention may therefore provide a method of determining whether or not a known individual is present in an environment. Equally, the provided method according to the first aspect may determine whether or not multiple known individuals are present in an environment.
[0024] In some embodiments, the method may comprise the step of comparing the evaluation dataset with the trained dataset to determine the condition of one or more persons. The condition of the one or more persons may include whether the one or more persons is asleep, stressed, fatigued or sick. In such embodiments, the trained datasetApplicant file ref: D1899 / GBD
[0025] 4
[0026] may comprise breathing information of one or more individuals collected at least in part in the ultrasonic frequency range when the condition of the one or more individuals is known. For example, for each of the one or more individuals, the trained dataset may include multiple entries of breathing information collected at least in part in the ultrasonic frequency range under different conditions, e.g. when the one or more individuals is asleep, stressed, fatigued and / or sick. This additional breathing information allows for the comparison method step of the first aspect, i.e. where the evaluation dataset is compared with the trained dataset, to determine whether one or more persons is present and the respective condition of said one or more persons. The breathing information may include metadata comprising information about the condition of the one or more known individuals, e.g. whether the one or more known individuals is asleep, stressed, fatigued and / or sick.
[0027] In some embodiments, the breathing information may comprise inhalation information of one or more individuals collected at least in part in the ultrasonic frequency. The breathing information of the one or more individuals may comprise solely inhalation information or the breathing information may include inhalation and exhalation breathing information.
[0028] The inventors have identified that inhalation breathing information of an individual includes multi-frequency harmonics and more complex structures compared to exhalation breathing information, which in contrast is found to be more stable and steady with respect to time. The added complexity and frequency harmonics in the inhalation breathing signatures provide additional information and features for comparison between the evaluation dataset and the trained dataset. Accordingly, by providing a trained dataset comprising inhalation breathing information collected in part in the ultrasonic frequency range, as well as sampling audio data of one or more inhalation events at least in part in the ultrasonic frequency range, additional features can be used to compare the evaluation dataset and trained dataset. These additional features specific to inhalation can increase certainty in the method output results and thus lead to a more accurate method for determining the presence and / or identity of one or more persons.Applicant file ref: D1899 / GBD
[0029] 5
[0030] Preferably, the trained dataset and the evaluation dataset are of the same format in order to facilitate and increase accuracy of dataset comparison. In some embodiments, the trained dataset and / or evaluation dataset may be an image format. The image format may show sound frequencies, at least in part in the ultrasonic frequency range, from one or more individuals breathing. Preferably, the image format may be a visual representation showing frequencies over time of a breathing event. Suitable image formats may include a sonograph, voiceprint, waterfall plot or more preferably a spectrogram. A spectrogram is an image format well suited for comparisons between the trained dataset and evaluation dataset because it captures frequency changes across a breathing event. In some embodiments, the spectrogram may include breathing information from multiple breathing events. The spectrogram may for example include breathing information of two or more distinct and consecutive breathing events. A single breathing event may include inhalation and / or exhalation.
[0031] In some embodiments, the method step of comparing the trained dataset and evaluation dataset may be carried out using a computer algorithm, preferably an image classification algorithm. The image classification algorithm may for example be a Convolutional Neural Network (CNN) and / or a CNN and Recurrent Neural Network (RNN) based algorithm. Advantageously, image classification algorithms provide an indication of the likelihood that the evaluation dataset matches at least part of the trained dataset. Image classification algorithms output a level of confidence value which represents the likelihood that the evaluation dataset matches at least part of the trained dataset. The level of confidence value may be outputted as a percentage. For example, the level of confidence may be outputted as a percentage >70% to mean a high confidence that the evaluation dataset matches at least part of the trained dataset, a percentage <30% to mean a high confidence that the evaluation dataset does not match at least part of the trained dataset, a percentage of 50% to 70% to mean a low confidence that the evaluation dataset matches at least part of the trained dataset, and a percentage of 30% to 50% to mean a low confidence that the evaluation dataset does not match at least part of the trained dataset.
[0032] One or more threshold values may be associated with the level of confidence. A threshold value may determine whether the output results of the method include aApplicant file ref: D1899 / GBD
[0033] 6
[0034] conclusive positive message (e.g. ‘person detected’ or ‘person X detected [along with metadata on person X]’), or a conclusive negative message (e.g. ‘no person detected’ or ‘person X not detected’), or an inconclusive message indicating that a conclusion was not reached (e.g. ‘unknown’ or ‘undetermined’). For instance, the output results of the method may include a conclusive positive message when the level of confidence is equal to or above an upper threshold value of 60%, or preferably 70%, or preferably 75%, or preferably 80%, or preferably 85%, or preferably 90%, or preferably 95% or preferably 99%. Additionally or alternatively, the output results of the method may include a conclusive negative message when the level of confidence is equal to or below a lower threshold value of 40%, or preferably 30%, or preferably 25%, or preferably 20%, or preferably 15%, or preferably 10%, or preferably 5%, or preferably 1%. Additionally or alternatively, the output results of the method may include an inconclusive message indicating that a conclusion was not reached when the level of confidence is within a threshold range of 5% to 95%, or preferably 10% to 90%, or preferably 15% to 85%, or preferably 20% to 80%, or preferably 25% to 75%, or preferably 30% to 70%, or preferably 40% to 60%.
[0035] In some embodiments, the method step of comparing the evaluation dataset with the trained dataset may finish and output results generated when the level of confidence is equal to or above the upper threshold value, and / or when the level of confidence is equal to or below the lower threshold value.
[0036] The first aspect of the invention may further comprise the method step after step v. of:
[0037] updating the trained dataset to incorporate the evaluation dataset and corresponding output results.
[0038] For example, the evaluation dataset and output results may be fed back into the trained dataset to expand the trained dataset. The step of updating the trained dataset may occur if the earlier step of comparing the trained dataset and evaluation dataset yields a level of confidence value that falls below an update threshold value. This indicates a high level of confidence that the one or more persons whose breathing information is represented in the evaluation dataset is not included in the trained dataset. In such instances, and if the identity and personal information of the one or more persons areApplicant file ref: D1899 / GBD
[0039] 7
[0040] known from another method, metadata including personal information of the one or more persons may also be included in the output results and incorporated in the trained dataset along with the evaluation dataset during the updating step. The output results and evaluation dataset included in the update may then form part of the breathing information of the trained dataset. The update threshold value below which the trained dataset will be updated to incorporate the evaluation dataset and corresponding output results may be set to 30%, or preferably 25%, or preferably 20%, or preferably 15%, or preferably 10%, or preferably 5%, or preferably 1%. Additionally or alternatively, the step of updating the trained dataset may occur if the earlier step of comparing the trained dataset and evaluation dataset yields a level of confidence that falls within an update threshold range. The update threshold range may cover level of confidence values that fall below a threshold value indicative of a high confidence that a known person is present. For example, if the threshold value is set at 80% to indicate a high confidence that a known person is present, the update threshold range may be 70% to 80% or preferably 75% to 80%. In some embodiments, the step of updating the trained dataset may occur when the level of confidence value is within the update threshold range and the identity of the one or more persons is determined and authenticated through another known authentication technique, such as face / voice recognition, password, etc. By updating the trained dataset when a level of confidence lays within the update threshold range, the trained dataset more comprehensively captures the variations in breathing information associated with one or more known individuals irrespective of their condition and / or location.
[0041] The first aspect of the invention may further comprise the method step after step v. of:
[0042] transmitting the output results to a secondary unit configured to receive and process the output results to authenticate the one or more persons for access.
[0043] The secondary unit may be for example a mobile phone, a computer or a type of digital lock. In some embodiments the secondary unit may be a sub-component or subsystem of the device. The secondary unit may process the output results to determine whether to unlock and grant user access to the secondary unit. For instance, the output results may include information confirming or declining the presence of one or more persons with the permissions to access the secondary unit. The secondary unitApplicant file ref: D1899 / GBD
[0044] 8
[0045] may process such information in an authentication step to determine whether to unlock and thus providing access to the secondary unit. Alternatively, the secondary unit may be configured to process the output results to determine whether the one or more persons are authorised to access the secondary unit and then grant / decline access accordingly. In some embodiments, the output results may include information that one or more persons present are unknown or known unauthorized individuals. In such embodiments, the secondary unit may process the output results and sound an alarm and / or activate another safe-guarding action to alert the presence of an intruder and take protective measures to mitigate the risk of unauthorized access.
[0046] In some embodiments the output results are transmitted to the secondary unit periodically such that the secondary unit can continually authenticate the one or more persons. The frequency at which the output results are transmitted may be determined by security protocols of the secondary unit.
[0047] According to a second aspect, the invention provides a device for analysing audio data, the device comprising:
[0048] an acoustic sensor configured to have a measured frequency response at least in part in the ultrasonic range;
[0049] a means for collecting audio data from the acoustic sensor;
[0050] a storage means for storing a trained dataset of one or more individuals, wherein the trained dataset includes breathing information of one or more individuals collected in part in the ultrasonic frequency range;
[0051] a processing means configured to generate an evaluation dataset from the collected audio data and compare the evaluation dataset with the trained dataset to determine whether one or more persons is present; and
[0052] an output means for outputting results.
[0053] The device of the second aspect is a self-contained unit, i.e. a unit including components that are each integral to the device. This differs from a ‘System’ (covered later in the third aspect of the invention), where at last some of the components of the system are located remotely to the device. A self-contained device offers manyApplicant file ref: D1899 / GBD
[0054] 9
[0055] advantages including portability, simplicity and silent operation (since it does not require data transmission to or from the device).
[0056] The device is configured to perform the method of the first aspect of the invention, and thus embodies all of the advantages associated with the method according to the first aspect of the invention.
[0057] Sound propagates as a pressure wave to a receptor, such as an ear or other acoustic sensor, where it is converted into nerve or electrical impulses. For the purposes of electronic digital signal processing a sound wave is converted into a series of discrete samples overtime using an analogue-to-digital converter (ADC), after which it may be reconstructed. The inventors have found that the useful energy for human breathing falls within the 20 Hz - 100 kHz frequency range and that an audio sampling rate of 200 kHz is adequate. Preferably, the audio sampling rate is at least double the frequency range, i.e. the Nyquist Frequency.
[0058] The acoustic sensor may be any type of sensor, which is capable of detecting sound, and is capable of being configured to have a measured frequency response in the ultrasonic range. The acoustic sensor may be able to detect sound in the audible, ultrasonic and infrasonic frequency range. Preferably, the acoustic sensor may be able to detect sound in the frequency range 20 Hz to 100 kHz, and more preferably 20 kHz to 100 kHz. The acoustic sensor may be a microphone and preferably a Digital Microelectromechanical systems (MEMS) microphone. MEMS microphones have the advantage of providing precise and readily matched performance characteristics. Furthermore, they inherently have increased sensitivity at higher frequency and are well suited to operating in the ultrasonic range. In practice, the inventor has discovered that a MEMs microphone can be sampled much faster than its specification suggests and the capability to record ultrasonics has been verified using an anechoic chamber and full audio lab.
[0059] In some embodiments, the device is configured to provide a broadband frequency response, for example in the range from 20 Hz to 100 kHz, and the processing means is configured to analyse audio data across the broadband range of the system, so thatApplicant file ref: D1899 / GBD
[0060] 10
[0061] acoustic features in at least the ultrasonic range are analysed and contribute to the output. This can be achieved by using a single acoustic sensor having a broadband frequency response or a number of acoustic sensors configured to respond to different frequency ranges.
[0062] To allow for digital signal processing, the input sound wave is sampled to ensure it is provided in digital form for further analysis. Accordingly, the means for collecting audio data in the device may be an analogue to digital converter (ADC). The ADC may be a separate component within the device or may be an integral part of the acoustic sensor. For example, a MEMS microphone may combine an acoustic sensor with an ADC so that a sound wave input is sampled over time and collected as audio data.
[0063] The storage means may be a standard digital storage unit for storing data such as a hard disk drive, flash drive, SSD, etc. The storage means is configured to store the trained dataset including breathing information of one or more individuals collected in part in the ultrasonic frequency range, and may enable the writing of additional trained dataset thereon (specifically in the embodiments where the sampled audio data and output results are fed back into the trained dataset to expand the trained dataset).
[0064] The processing means may be a standard microcomputer programmed to perform the steps of generating the evaluation dataset based on the collected audio data (e.g. converting all or part of the audio data to the evaluation dataset), as well as comparing the evaluation dataset and trained dataset to determine whether one or more persons is present. The evaluation dataset may for example be of the format of an image, preferably a visual representation of the evaluation dataset such as a spectrogram, sonograph, waterfall plot or voiceprint, which displays frequency against time. The processing means may be configured to extract and analyse ultrasonic features in the collected audio data. The ultrasonic signals in breathing may be related to inhalation and exhalation from the lungs, trachea, nasal cavity and larynx.
[0065] The processing means may be configured to compare the evaluation dataset with the trained dataset through an image classification algorithm. The image classification algorithm may output results including a determination of the presence and / orApplicant file ref: D1899 / GBD
[0066] 11
[0067] identification of one or more persons and preferably a corresponding level of confidence value.
[0068] Additionally or alternatively, the processing means may be a standard microcomputer programmed to perform the steps of extracting a plurality of acoustic features in the ultrasonic range and comparing the extracted features with stored reference data. In some embodiments, the processing means is configured to operate across a broad band frequency range from 20 Hz to 100 kHz.
[0069] The output means outputs results from the processing means, specifically, the results generated by the comparisons between the evaluation dataset and trained dataset. The output results can be conveyed to a device operator. For example, the output results may include a visual representation of the conclusion from the comparison, e.g. a screen displaying a message indicative of whether one or more persons is present, and possibly identifying said one or more persons. The visual representation may also show a corresponding level of confidence value and / or personal information about the one or more individuals in the trained dataset for which strong similarities with the evaluation dataset have been detected. Alternatively and / or additionally, the output results may include an audible and / or haptic representation of the conclusion from the comparison.
[0070] In some embodiments, the output results may be transmitted or broadcasted as a signal to a secondary unit such as a phone, computer, digital lock, or a sub-component / sub-unit of the device etc. The signal comprising output results may confirm the presence and / or identity of one or more persons who are using the secondary unit for the purposes of user authentication. The signal comprising output results may be transmitted / broadcasted by the device of the second aspect in a wired or wireless manner to the secondary unit. The transmission / broadcast of the signal comprising output results may occur automatically and / or upon demand. Automatic transmission / broadcast of the signal comprising output results may happen periodically, such that the user of the secondary unit can be continually or near continually authenticated. The device may apply the method according to first aspect of the invention alongside, and be complementary to, other known authenticationApplicant file ref: D1899 / GBD
[0071] 12
[0072] techniques, for example, a password, face recognition, voice recognition, etc.). The method according to the first aspect of the invention being advantageous in that the one or more persons is only required to be in proximity to the acoustic sensor of the device to be authenticated.
[0073] The device may comprise an updating means configured to update the trained dataset to incorporate the evaluation dataset and corresponding output results. The updating means may write on to the storage means when the level of confidence is below an update threshold value or falls within an update threshold range.
[0074] The updating means may update the trained dataset while the device is operational and still actively collecting audio data. Alternatively, the updating means may update the trained dataset while the device is idle and in a state dedicated to allowing the trained dataset to be updated.
[0075] According to a third aspect, the invention provides a system for analysing audio data, wherein the system comprises:
[0076] a device and a network entity;
[0077] a network storage means, on the network entity, configured to store a trained dataset of one or more individuals, the trained dataset including breathing information of one or more individuals collected at least in part in the ultrasonic frequency range;
[0078] a device collecting means, on the device, for collecting audio data from an acoustic sensor, the acoustic sensor configured to have a measured frequency response at least in part the ultrasonic range;
[0079] a device transmitter, on the device, configured to transmit the audio data; a network receiver, on the network entity, configured to receive the audio data; and
[0080] a network processing means, on the network entity, configured to generate an evaluation dataset from the received audio data and compare the evaluation dataset with the trained dataset to determine whether one or more persons is present.
[0081] By providing a system including a network entity and device, at least some of the data storage and data processing demands can happen external to the device. The deviceApplicant file ref: D1899 / GBD
[0082] 13
[0083] may for example simply consist of an acoustic sensor, a data collecting means and a transmitter, while the bulkier and more power consuming components, such as the processing means and data storage means may be located on the network entity. Advantageously, the device of the system according to the third aspect of the invention may therefore be significantly smaller and lighter than an equivalent device according to the second aspect of the invention. Furthermore, the energy demands to power the device can be significantly reduced. The system accordingly to the third aspect may be particularly advantageous during covert operations since the device is easier to conceal.
[0084] The system is configured to perform the method of the first aspect of the invention, and thus embodies all of the advantages associated with the method according to the first aspect of the invention.
[0085] In some embodiments, the system may further comprise a device storage means, on the device, configured to store at least a subset of the trained dataset of one or more individuals, the trained dataset including breathing information of one or more individuals collected at least in part in the ultrasonic frequency range. In such embodiments, the network storage means, on the network entity, may store the entirety of, or another subset of, the trained dataset.
[0086] There are numerous other combinations for which the trained dataset may be stored on the network storage means and / or on the device storage means. For instance, in some embodiments, the network storage means may store the entirety of the trained dataset for data backup purposes, while the device storage means may store a subset of the trained dataset to enable simultaneous data processing on the device and network entity (and thus increase overall processing speed of the system). Alternatively, the network storage means and the device storage means may store different subsets of the trained dataset to minimise overall storage demand while still enabling simultaneous data processing on the device and network entity.
[0087] In order to transfer data from the network entity to the device, for example, from the network storage means to the device storage means, the system may further compriseApplicant file ref: D1899 / GBD
[0088] 14
[0089] a network transmitter, on the network entity, configured to transmit data, and a device receiver, on the device, configured to receive the data. In some embodiments, the network entity may comprise an output means configured to output the results. In some embodiments, the network transmitter may be configured to transmit the output results from the network processing means and the device receiver may be configured to receive the output results from the network transmitter. The output results may include a conclusion from the network processing means, e.g. a determination of whether one or more persons is present. The output results may comprise metadata of a known individual extracted from the trained dataset.
[0090] The system may further comprise a device output means, on the device, for outputting the output results in a visual, audible format and / or haptic format.
[0091] In some embodiments, the output results may be transmitted and / or broadcasted as a signal to a secondary unit such as a phone, computer, digital lock, or a sub-component / sub-unit of the system etc. The signal comprising output results may confirm the presence and / or identity of one or more persons who are using the secondary unit for the purposes of user authentication. The signal comprising output results may be transmitted / broadcasted by the system of the third aspect in a wired or wireless manner to the secondary unit. The signal comprising results may be transmitted / broadcasted by the device entity and / or network entity via, for example, the device transmitter and / or network transmitter. The transmission / broadcast of the signal comprising output results may occur automatically and / or upon demand. Automatic transmission / broadcast of the signal comprising output results may happen periodically, such that the user of the secondary unit can be continually authenticated. The system may apply the method according to first aspect of the invention alongside, and be complementary to, other known authentication techniques, for example, a password, face recognition, voice recognition, etc.).
[0092] In some embodiments, the system may further comprise a device processing means, on the device, configured to generate an evaluation dataset from a whole of, or a subset of, the audio data collected by the device collecting means. In such embodiments, the network processing means may be configured to generate anApplicant file ref: D1899 / GBD
[0093] 15
[0094] evaluation dataset from a whole of, or another subset of, the audio data collected by the device collecting means.
[0095] In some embodiments, the device processing means may be configured to compare a whole of, or a subset of, the evaluation dataset with the trained dataset stored on the device storage means. Additionally or alternatively, the network processing means may be configured to compare a whole of the evaluation dataset with the trained dataset stored on the network storage means, or the network processing means may be configured to compare a subset of the evaluation dataset with the trained dataset stored on the network storage means.
[0096] The amount of the trained dataset stored on the network storage means and / or device storage means (and thus the proportion of data processing performed by the network processing means vs. by the device processing means (if included)) may be determined by at least one of the following: the data transfer speed of the device transmitter and / or network receiver, the maximum capacity of the device storage means, and the difference in processing performance between the network processing means and device processing means, access to the network entity, etc.
[0097] In some embodiments, the network entity may comprise a network updater configured to update the trained dataset to incorporate the evaluation dataset and corresponding output results. After the network processing means has compared the trained dataset and evaluation dataset and produced output results, the network updater may transfer the evaluation dataset to the trained dataset found on the network storage means along with the corresponding output results. The network updater may only transfer the evaluation dataset to become part of the trained dataset if the level of confidence value associated with the comparison is below an update threshold value or within an update threshold range. In embodiments where at least some of the trained dataset is stored locally on the device storage means, the network updater may work in conjunction with the network transmitter and device receiver to transfer the evaluation dataset and output results to the device, specifically, transfer the evaluation dataset and output results to the trained dataset found on the device storage means.Applicant file ref: D1899 / GBD
[0098] 16
[0099] Any feature in one aspect of the invention may be applied to any other aspects of the invention, in any appropriate combination. In particular device aspects may be applied to method or use aspects and vice versa. The invention extends to a method, device or system substantially as herein described, with reference to the accompanying drawings and Examples. In all aspects, the invention may comprise, consist essentially of, or consist of any feature or combination of features.
[0100] Brief Description of the Drawings
[0101] The invention will now be described, purely by way of example, with reference to the accompanying drawings, in which;
[0102] Figure 1 is a diagram of a method for analysing audio data according to the first aspect of the invention.
[0103] Figure 2 is a schematic illustrating how data is generated, processed and transferred for a method according to the first aspect of the invention.
[0104] Figure 3 is a spectrogram of an inhalation and an exhalation breathing event;
[0105] Figure 4 is a schematic of a device according to a second aspect of the invention.
[0106] Figure 5 is a schematic of a system according to a third aspect of the invention.
[0107] Figure 6 is three spectrograms of an inward nasal breath by three different individuals, respectively.
[0108] Figure 7a is two spectrograms of an inward nasal breath by the same individual in an ‘At Rest condition’ and in an ‘Exerted condition’, respectively.
[0109] Figure 7b is two spectrograms of an inward nasal breath by the same individual in a ‘Normal condition’ and in an ‘Illness condition’, respectively.
[0110] The drawings are for illustrative purposes only and are not to scale.Applicant file ref: D1899 / GBD
[0111] 17
[0112] Detailed Description
[0113] Figure 1 is a diagram of a method 100 for analysing audio data according to the first aspect of the invention. The method 100 comprises six steps including a first step 102 of providing a trained dataset of one or more individuals, the trained dataset including breathing information of one or more individuals collected in part in the ultrasonic frequency range; a second step 104 of sampling audio data, at least in part in the ultrasonic frequency range; a third step 106 of processing the sampled audio data to generate an evaluation dataset; a fourth step 108 of comparing the evaluation dataset with the trained dataset to determine whether one or more persons is present; a fifth step 110 of providing an output comprising results; and a sixth step 112 of updating the trained dataset to incorporate the evaluation dataset and corresponding output results.
[0114] Figure 2 is a schematic illustrating how data is generated, processed and transferred for a method according to the first aspect of the invention. The direction of data transfer is illustrated through the use of arrows.
[0115] In a first method step, a trained dataset 202 of one or more individuals is provided. The trained dataset 202 may be acquired (e.g. purchased) or generated beforehand. In some embodiments, the trained dataset 202 is generated by: sampling audio data at least in part in the ultrasonic frequency range of one or more known individuals; and processing the sampled audio data to generate a trained dataset 202 comprising breathing information of the one or more individuals collected at least in part in the ultrasonic frequency range including associated metadata of personal information about the one or more known individuals. Additionally or alternatively, the trained dataset 202 may be generated by: sampling audio data at least in part in the ultrasonic frequency range of one or more unknown individuals; and processing the sampled audio data to generate a trained dataset 202 including breathing information of the one or more individuals collected at least in part in the ultrasonic frequency range.
[0116] The trained dataset includes a library of trained dataset entries 204a - 204e each comprising breathing information of one or more individuals collected at least in part in the ultrasonic frequency range. The breathing information in each of the trainedApplicant file ref: D1899 / GBD
[0117] 18
[0118] dataset entries 204a - 204e is described as a spectrogram, where each spectrogram was generated from previous breathing event(s) of the one or more individuals along with corresponding metadata outlining personal information of the one or more individuals. Example breathing information generated by three different individuals that could form part of the trained dataset is provided in Figure 6 and a description of how the breathing information differs between the three individual is provided towards the end of this section.
[0119] If the breathing information was generated from one or more individuals with a known condition (e.g. the one or more individuals was asleep, stressed, fatigued or sick), the metadata will also include personal information related to that condition. Examples of breathing information generated from a first individual with a first known condition (fatigue / exertion) and a second individual with a second known condition (illness) is provided in Figures 7a / 7b, respectively. This breathing information could form part of the trained dataset. In both these examples, a noticeable difference in the breathing information was observed between the individual in their baseline state and the same individual with a condition. A description of how the breathing information differs when an individual has a condition is provided at the end of this section.
[0120] In a second method step, audio data is sampled using an acoustic sensor combined with an analogue to digital converter to produce audio data 206. Sampling happens at least in part in the ultrasonic frequency range, and the audio data 206 is temporarily stored in its raw digital format.
[0121] In a third method step, the audio data 206 is processed by an evaluation dataset generator 208, which converts the sampled audio data 206 from its raw digital format to an evaluation dataset 210. The evaluation dataset 210 has a spectrogram format to facilitate comparison with the spectrograms of the trained dataset entries 204a -204e. Following audio data processing and format conversion, the evaluation dataset generator 208 outputs the evaluation dataset 210.
[0122] In a fourth method step, a classification algorithm 212 is used to compare the evaluation dataset 210 and the trained dataset 202. In particular, the classificationApplicant file ref: D1899 / GBD
[0123] 19
[0124] algorithm 212 is a Convolutional Neural Network (CNN). The CNN produces a model of the trained dataset 202 that has been trained to classify spectrograms of breathing information of one or more individuals collected at least in part in the ultrasonic frequency range. The model has been trained using a CNN in a known way. In more detail, the model is trained using known and labelled spectrograms of breathing information of one or more individuals collected at least in part in the ultrasonic frequency range. The labelled spectrograms pass through one or more layers of the CNN and the CNN extracts features and patterns of features from the labelled spectrograms. In this way, the model learns to associate features and patterns of features in the spectrographs with the corresponding label. The model is then tested and validated using known and understood unlabelled spectrograms. Once the model is trained and tested, it calculates a level of confidence between the evaluation dataset 210 and trained dataset 202 by comparing spectrogram (s) in the evaluation dataset 210 with the model of the trained dataset 202. For example, the model calculates a level of confidence >70% when it predicts a high confidence that the evaluation dataset matches at least part of the trained dataset, a level of confidence <30% when it predicts a high confidence that the evaluation dataset does not match at least part of the trained dataset, a level of confidence of 50% to 70% when it predicts a low confidence that the evaluation dataset matches at least part of the trained dataset, and a level of confidence of 30% to 50% when it predicts a low confidence that the evaluation dataset does not match at least part of the trained dataset. If the level of confidence equals or exceeds an upper threshold value of 85%, the classification algorithm will output results 214 that are conclusive. In this instance, the output results 214 include a positive conclusive message, e.g. ‘person X detected’, along with other personal information about person X. If the classification algorithm 212 calculates a level of confidence equal to or below a lower threshold value of 20%, the classification algorithm will output results 214 that are conclusive. In this instance, the output results 214 simply include a negative conclusive message, e.g. ‘no person detected’ or ‘person X not detected’. If the classification algorithm 212 calculates a level of confidence within a threshold range of 20% to 85%, the classification algorithm will output results 214 that are inconclusive. In this instance, the output results 214 simply include an inconclusive message, e.g. ‘unknown’ or ‘undetermined’.Applicant file ref: D1899 / GBD
[0125] 20
[0126] In some embodiments, the classification algorithm 212 includes a Recurrent Neural Network (RNN). In more detail, once the CNN has extracted the features and patterns of features from the labelled spectrograms, the RNN takes the features and patterns of features and processes them in sequence over time. Therefore, the RNN determines a temporal relationship between the features and patterns of features. The model learns to associate temporal relationships of features and patterns of features in the spectrographs with the corresponding label.
[0127] In a sixth method step, an updater 216 updates the trained dataset 202 to incorporate the evaluation dataset 210 and corresponding output results 214. The updater 216 will update the trained dataset 202 if comparisons between the evaluation dataset 210 and trained dataset 202, specifically comparisons between evaluation dataset 210 and the trained dataset entries 204a - 204e, yields a level of confidence that falls below an update threshold value of 10% or within the update threshold range of 75% to 85%. The former indicates that the one or more persons whose breathing generated the evaluation dataset 210 is not included in the trained dataset 202, and thus the updater 216 is programmed to expand the trained dataset 202. If personal information about the one or more persons is known, the output results 214 will further include metadata associated with the identity of the one or more persons for incorporation into the trained dataset 202.
[0128] Figure 3 is a spectrogram 300 of Frequency in Hertz vs. Time in Seconds generated from sampled audio data of an inhalation breathing event followed by an exhalation breathing event. The spectrogram 300 shows an inhalation region 302 comprising inhalation information and an exhalation region 304 comprising exhalation breathing information. The inhalation breathing information includes multi-frequency harmonics 306a - 306c at are not present in the exhalation region 304. Furthermore, the magnitude of inhalation breathing information is significantly greater at higher frequencies, i.e. above 50kHz, than the exhalation breathing information. Inhalation breathing information therefore provides additional features indicative to a specific person than compared to exhalation breathing information. These additional features improve the confidence in comparisons between spectrograms of the evaluationApplicant file ref: D1899 / GBD
[0129] 21
[0130] dataset to spectrograms of the trained dataset, and thus the confidence in the determination of the presence / identification of one or more persons.
[0131] Figure 4 is a schematic of a device 400 according to a second aspect of the invention. The direction of data transfer within the device 400 is illustrated through the use of arrows. The device 400 is a self-contained unit including electronic components integral to the device. The device 400 include a Digital MEMS microphone 402 that is configured to have a measured frequency response at least in part in the ultrasonic range and a sampling rate of 200 kHz, specifically, a sample rate of at least double the ultrasonic frequency range. The MEMS microphone 402 includes an ADC (not shown) to sample acoustic sound waves and generate audio data. The audio data is then collected and temporarily stored on computer memory, e.g. RAM 404 in its raw format. The device 400 also includes a non-volatile memory, e.g. SSD 406 with a trained dataset of one or more individuals stored thereon. The trained dataset includes breathing information of one or more individuals collected in part in the ultrasonic frequency range.
[0132] The device 400 further has a computer processor 408 comprising an evaluation dataset generation module 410 and a dataset comparison module 412. The evaluation dataset generation module 410 is configured to process audio data stored on RAM 404 to generate an evaluation dataset. Specifically, the evaluation dataset generation module 410 converts the audio data, which is stored in its raw sampled format on the RAM 404, into a spectrogram. The dataset comparison module 412 is configured to compare the evaluation dataset with the trained dataset and calculate a level of confidence between the spectrograms of the evaluation dataset and the spectrograms of the training dataset. The level of confidence represents the likelihood that the evaluation dataset matches at least part of the trained dataset. In particular, the dataset comparison module 412 reads the evaluation dataset generated by the dataset generation module 410 (i.e. the generated spectrogram) and the trained dataset stored on the storage means 406. The dataset comparison module 412 includes an image classification algorithm programmed to provide an indication of the likelihood that the spectrogram of the evaluation dataset matches the spectrograms of the trained dataset. The image classification algorithm calculates a level of confidence valueApplicant file ref: D1899 / GBD
[0133] 22
[0134] representing the likelihood that the spectrogram of the evaluation dataset matches at least part of the spectrogram (s) of the trained dataset. If the level of confidence equals or exceeds an upper threshold value of 85%, the image classification algorithm will output results comprising a positive conclusive message to an output means 414. If the level of confidence equals or falls below a lower threshold value of 20%, the image classification algorithm will output results comprising a negative conclusive message to an output means 414. If the level of confidence falls within a threshold range of 20% to 85%, the image classification algorithm will output results comprising an inconclusive message to an output means 414. As described above, the classification algorithm 212 may be a CNN and / or a CNN and RNN based algorithm.
[0135] The output means 414 is a display including a screen for providing a visual representation of the output results from the processor dataset comparison module 412 of the processor 408. The output means 414 can show a conclusive or inconclusive message related to whether one or more persons is present, and / or the identity of said one or more persons including personal information based on metadata associated with one or more individual in the trained dataset, and / or the level of confidence calculated by the image classification algorithm.
[0136] The output means 414 of the device 400 can also include a device transmitter (not shown) to transmit the output results as a signal to a secondary unit (also not shown), such as a mobile phone or computer, for the purpose of secondary unit user authentication. In such an embodiment, the secondary unit has a receiver for receiving the signal comprising output results. The device transmitter may automatically and periodically transmit the signal comprising output results such that the secondary unit can continually authenticate one or more persons in close proximity to the device 400 based on their breathing signature. The secondary unit may rely on the output results from the device 400 in conjunction with other authentication techniques, such as a password, face / voice recognition, etc., to increase the reliability of authentication with minimal additional effort from the user.
[0137] The device 400 also includes a trained dataset updater 416 that is configured to update the trained dataset to incorporate the evaluation dataset and corresponding outputApplicant file ref: D1899 / GBD
[0138] 23
[0139] results. The updater 416 therefore expands the trained dataset on the SSD storage 406. The updater 416 is activated when the level of confidence value between the evaluation dataset and trained dataset is equal to or below an update threshold value of 10% or an update threshold range of 75% to 85%. The former implies that the breathing information in the evaluation dataset is from one or more persons who have not previously be recorded and stored in the trained dataset. The update allows for metadata associated with the one or more persons (identified through other means) to be included as part of the output results. For example, personal information about the one or more persons whose breathing is sampled.
[0140] Figure 5 is a schematic of a system 500 according to a third aspect of the invention. The direction of data transfer across the system 500 is illustrated through the use of arrows. The system 500 comprises a device 510 and a network entity 560. The device 510 and network entity 560 are physically separate, but can communicate wirelessly in order to transfer data.
[0141] The network entity 560 comprises a network storage means 570 configured to store a trained dataset of one or more individuals, the trained dataset including breathing information of one or more individuals collected at least in part in the ultrasonic frequency range.
[0142] The device 510 includes an acoustic sensor 512, such as a MEMs microphone, and a means of collecting audio data 514, such as a RAM. The device 510 also has a transmitter 516 for transmitting audio data from the device 510 to the network entity 560. Wireless data transmission is illustrated in Figure 5 through the use of dashed arrows. Any mode of wireless data known in the art may be used to transfer the data including for example Wi-Fi, Bluetooth, Cellular, Satellite Communications, etc.
[0143] The network entity 560 includes a receiver 562 for receiving audio data transmitted from the device transmitter 516. The network receiver 562 transfers the audio data to a network processing means 564. The network processing means 564 comprises a network evaluation dataset generation module 566 and a network dataset comparison module 568.Applicant file ref: D1899 / GBD
[0144] 24
[0145] The network evaluation dataset generation module 566 is configured to process audio data received by the network receiver 562 to generate an evaluation dataset. Specifically, the evaluation dataset generation module 566 converts the audio data received by the network receiver 562 into a spectrogram.
[0146] The network dataset comparison module 568 is configured to compare the evaluation dataset with the trained dataset. In particular, the network dataset comparison module 568 reads (i) the evaluation dataset generated by the network dataset generation module 566 (specifically, the spectrogram generated from the sampled audio data) and (ii) the trained dataset stored on the network storage means 570. The network dataset comparison module 568 includes an image classification algorithm programmed to compare the spectrogram of the evaluation dataset with the trained dataset. As described above, the image classification algorithm may be a CNN and / or a CNN and RNN based algorithm. For each comparison, the image classification algorithm calculates a level of confidence value which quantifies the likelihood that the spectrogram of the evaluation dataset matches the trained dataset. If the level of confidence equals or exceeds an upper threshold value of 85%, the image classification algorithm will stop and the network dataset comparison module 568 will output results including a positive conclusive message to a network transmitter 572. If the level of confidence equals or falls below a lower threshold value of 20%, the image classification algorithm will stop and the network dataset comparison module 568 will output results including a negative conclusive message to a network transmitter 572. If the level of confidence falls within a threshold range of 20% to 85%, the image classification algorithm will stop and the network dataset comparison module 568 will output results including an inconclusive message to a network transmitter 572. The network transmitter 572 is configured to then transmit the output results to the device 510.
[0147] The output results include a message which summarises the outcome from the network dataset comparison module 568. The message can be conclusive positive (e.g. ‘Person X detected’), conclusive negative (e.g. ‘No person detected’) or inconclusive e.g. (‘Detection undetermined’). When a person is detected who matches a known individual in the trained dataset, the output results will also include metadataApplicant file ref: D1899 / GBD
[0148] 25
[0149] describing personal information about the known person such as their address, age and photograph. The output results will also include the level of confidence value to indicate certainty of the likelihood of the match between the person(s) and individual(s) stored in the trained dataset is (and thus provide an indicator of confidence in the prediction conclusion).
[0150] The device 510 further comprises a device receiver 518 configured to receive the output results transmitted by the network transmitter 572 and transfer the output results to a device output means 520. The device output means 520 provides a visual or audible representation of the output results. For example, in some embodiments, the device output means 520 is a speaker that provides an audible message describing the output results.
[0151] The network transmitter 572 can be further configured to transmit a signal comprising the output results to a secondary unit (not shown) such as a mobile phone or computer, for the purpose of secondary unit user authentication. In such an embodiment, the secondary unit has a receiver for receiving the signal comprising output results. The network transmitter 572 may automatically and periodically transmit the signal comprising output results such that the secondary unit can continually authenticate one or more persons in close proximity to the device 510 based on their breathing signature. The secondary unit may rely on the output results from the system 500 in conjunction with other authentication techniques, such as a password, face / voice recognition, multi-factor authentication, etc., to increase the reliability of authentication with minimal additional effort from the user.
[0152] The network entity 560 also has a trained dataset updater 574 that is configured to update the trained dataset to incorporate the evaluation dataset and corresponding output results. The network updater 574 expands the trained dataset on the network storage means 570. The network updater 574 is activated when a level of confidence between the evaluation dataset and trained dataset is equal to or below an update threshold value of 10% or an update threshold range of 75% to 85%. Calculating a level of confidence below this update threshold value indicates that the breathing information in the evaluation dataset is from one or more persons who have notApplicant file ref: D1899 / GBD
[0153] 26
[0154] previously be recorded and stored in the trained dataset. As well as incorporation the evaluation dataset, the network updater 574 also incorporates output results including known metadata in the trained dataset. The metadata can include personal information about the one or more persons whose breathing is sampled.
[0155] The network entity 560 includes a network output means 576. The network output means 576 provides a visual or audible or haptic representation of the output results. For example, the network output means 576 in the present embodiment is a screen, which display the output results and evaluation dataset in textual and graphical form. The network output means 576 allows an operator of the system 500 to see and understand the output results away from the device 510.
[0156] Supporting information for determining the presence and identity of one or more persons
[0157] Figure 6 shows three spectrograms 602, 604, 606 of an inward nasal breath by three different individuals, respectively: Individual X; Individual Y; and Individual Z. Even through visual inspection, it can be clearly observed that each spectrogram 602, 604, 606 includes unigue features and patterns that are specific to an individual. Therefore, were the three spectrograms included in a trained dataset, comparison between the trained dataset and an evaluation dataset including sampled audio data would allow for the presence and identity of one or more persons (specifically, individuals X, Y and / or Z) to be determined.
[0158] For instance, by visually analysing the three spectrograms 602, 604, 606, it can be observed that the spectrogram associated with Individual X 602 is roughly symmetrical around a central time 608; the spectrogram associated with Individual Y 604 has a slightly longer later-time tail 610 with a significantly lower peak freguency 612; and the spectrogram associated with Individual Z 606 has a very significant trailing tail 614 with more defined freguency bands visible 616. Such differences can be easily guantified for automated comparison purposes. Taking peak freguency as an example, the spectrogram for Individual Y 604 contains relatively less ultrasonic freguency components than the spectrogram for both Individual X 602 and Individual Z 606,Applicant file ref: D1899 / GBD
[0159] 27
[0160] peaking at ~32 kHz. The spectrogram for Individual X 602 maintains 64 kHz over a significant portion of the sample, while the spectrogram for Individual Z 606 peaks at 64 kHz and quickly decays with time.
[0161] Other noticeable differences between the spectrograms 602, 604, 606, which may be basis for comparison between a trained dataset and an evaluation dataset, include: the dominant frequency band(s) of the spectrogram, the time period for which a breathing event generates signal, and the levels of frequency “smoothness”. The term ‘smoothness’ or ‘smoothing’ is intended to mean a continuously strong signal (represented as white in the spectrograms) across several frequences with no obvious banding. For example, the spectrograms 602, 604, 606 show that Individual X generates strong signal (represented in white) in the ultrasonic frequency range for ~0.3s, while Individual Y and Z generate a strong signal in the ultrasonic frequency range for ~0.5s and ~1.5s, respectively. Moreover, both the spectrograms for Individual X 602 and Individual Y 604 have some smoothing over the frequency bands in different but specific frequency ranges, whereas the spectrogram for Individual Z 606 contains sharper, more well-defined frequency bands 616.
[0162] The use of an CNN and / or RNN to compare similar breathing information of the spectrograms 602, 604, 606 in Figure 6 allows for additional and more accurate comparisons to be made relative to the above examples which were purely based on visual inspection. Significantly expanding on the trained dataset to include breathing information from different individuals or multiple entries of the same individual taken at different times would result in a trained dataset with improved accuracy and utility to determine the presence of one or more persons.
[0163] Supporting information for determining the condition of one or more persons
[0164] Figure 7a shows two spectrograms 702, 704 of an inward nasal breath by an individual in an ‘At Rest condition’ and the same individual in an ‘Exerted condition’, respectively. In order to bring the individual into the Exerted condition, the Chester Step Test was used, whereby the individual repeatably stepped up and down on a 30cm step until their heart rate reached 75-80% of their estimated maximum heart rate. Contrastingly,Applicant file ref: D1899 / GBD
[0165] 28
[0166] and as the name suggests, the ‘At rest’ condition is the state of the individual when they have a resting heart rate.
[0167] Figure 7a shows that the breathing information for an individual in the ultrasonic frequency range differs depending on whether the individual is exerted or at rest. For example, the spectrogram for an individual in an exerted condition 704 has a higher peak frequency 706 (>64 kHz peak frequency in exerted condition compared to ~40 kHz peak frequency in the At Rest condition). Another difference includes the levels of frequency “smoothness”, e.g. the spectrogram for the At Rest condition 702 shows frequency bands between 8-32 kHz that are relatively discrete, with clear lines dividing them. On the other hand, the spectrogram for the Exerted condition 704 shows frequency bands that are smoothed over the same range, with less discrete banding.
[0168] While there are clear differences in the breathing information of a spectrogram for an individual in an exerted condition 704 and an at rest condition 702 at the ultrasonic frequency range, there is also some commonality indicative of the spectrograms representing the same individual. For example, the actual frequency band structure, i.e. the dominant bands and the rough structure of ultrasonic harmonics, stay the same.
[0169] In another example, Figure 7b shows two spectrograms 752, 754 of an inward nasal breath by an individual in a ‘Normal’ condition and the same individual in an ‘Illness condition’, respectively. The illness condition in this case is a respiratory condition (e.g. cold or flu). The individual represented in Figure 7b is different to the individual represented Figure 7a.
[0170] Figure 7b shows that the breathing information for an individual in the ultrasonic frequency range differs depending on whether the individual has an illness. For instance, the spectrogram for an individual with an illness 754 has a lower peak frequency 756 and less ultrasonic components: the peak frequency for the individual in a normal condition is >64kHz and is maintained for most of the sample time range, while the peak frequency 756 for the individual with an illness condition is lower and lasts for a significantly smaller portion of the time range. Another difference is that frequency banding 758 is sharper and more defined in the spectrogram for anApplicant file ref: D1899 / GBD
[0171] 29
[0172] individual with illness condition 754, i.e. the frequency bands are clear and less smooth compared to the spectrogram of the normal condition 752 which has a frequency response that is smooth in the major regions 760.
[0173] Similarly to the spectrograms in Figure 7a, the two spectrograms in Figure 7b also have some common features that are indicative of them representing the same individual. For instance, the general time frequency structure is maintained, e.g. both spectrograms show high energy concentration (indicated by an abundance of white on the spectrogram to show a high signal) across a broad frequency range. Additionally, the time-frequency structure (i.e. the general shape) of the spectrograms 752, 754 is very similar.
[0174] By including breathing information of an individual in a trained dataset, which includes said individual in different but known conditions, more accurate determination of the individual’s identity (regardless of their condition), as well as determination of their respective condition, can be achieved. For example, for each individual, the trained dataset can include multiple entries of breathing information collected at least in part in the ultrasonic frequency range under different conditions, e.g. when the one or more individuals is asleep, stressed, exerted and / or sick. The additional breathing information allows for the comparison method step of the first aspect, i.e. where the evaluation dataset is compared with the trained dataset, to determine whether one or more persons is present and the respective condition of said one or more persons. The breathing information can include metadata comprising information about the condition of the one or more known individuals, e.g. whether the one or more known individuals is asleep, stressed, fatigued and / or sick.
[0175] It will be understood that the present invention has been described above purely by way of example, and modification of detail can be made within the scope of the invention.
Claims
Applicant file ref: D1899 / GBD30CLAIMS1. A method for analysing audio data, the method comprising the steps of:i. providing a trained dataset of one or more individuals, the trained dataset including breathing information of one or more individuals collected at least in part in the ultrasonic frequency range;ii. sampling audio data, at least in part in the ultrasonic frequency range; iii. processing the sampled audio data to generate an evaluation dataset; iv. comparing the evaluation dataset with the trained dataset to determine whether one or more persons is present; andv. providing an output comprising results.
2. A method according to claim 1, wherein the trained dataset of one or more individuals includes breathing information of one or more known individuals.
3. A method according to claim 2, wherein at least one person of the one or more persons is a known individual.
4. A method according to any preceding claim, wherein the breathing information comprises inhalation information of one or more individuals collected in part in the ultrasonic frequency.
5. A method according to any preceding claim, wherein the trained dataset and the evaluation dataset format are in an image format.
6. A method according to claim 5, wherein the image format is a spectrogram.
7. A method according to claim 5 or 6, wherein step iv) comprises an image classification algorithm.
8. A method according to any of the preceding claims, wherein the method further comprises the step:Applicant file ref: D1899 / GBD31updating the trained dataset to incorporate the evaluation dataset and corresponding output results.
9. A method according to any of the preceding claims, wherein the method further comprises the step:transmitting the results to a secondary unit configured to receive and process the results to authenticate the one or more persons for access.
10. A method according to claim 9, wherein the results are transmitted to the secondary unit periodically such that the secondary unit can continually authenticate the one or more persons.
11. A method according to any of the preceding claims, wherein step (iv) further comprises comparing the evaluation dataset with the trained dataset to determine the condition of the one or more persons.
12. A device for analysing audio data, the device comprising:an acoustic sensor configured to have a measured frequency response at least in part in the ultrasonic range;a means for collecting audio data from the acoustic sensor;a storage means for storing a trained dataset of one or more individuals, wherein the trained dataset includes breathing information of one or more individuals collected in part in the ultrasonic frequency range;a processing means configured to generate an evaluation dataset from the collected audio data and compare the evaluation dataset with the trained dataset to determine whether one or more persons is present; andan output means for outputting results.
13. A device according to claim 12, wherein the trained dataset of one or more individuals includes breathing information of one or more known individuals.
14. A device according to claim 13, wherein at least one person of the one or more persons is a known individual.Applicant file ref: D1899 / GBD3215. A device according to any one of claims 12 to 14, wherein the processing means is further configured to determine the condition of the one or more persons.
16. A device according to any one of claims 12 to 15, wherein the breathing information comprises inhalation information of one or more individuals collected in part in the ultrasonic frequency.
17. A device according to any one of claims 12 to 16, wherein the trained dataset and the evaluation dataset comprises an image format.
18. A device according to claim 17, wherein the image format is a spectrogram.
19. A device according to claim 17 or 18, wherein the processing means is configured to compare the evaluation dataset with the trained dataset through an image classification algorithm.
20. A device according to any one of claims 12 to 19, wherein the acoustic sensor is configured to operate with a sampling rate of double the frequency of the audio data.
21. A device according to any one of claims 12 to 20, wherein the means for collecting audio data is an analogue to digital converter.
22. A device according to any one of claims 12 to 21, wherein the device further comprises:an updating means configured to update the trained dataset to incorporate the evaluation dataset and corresponding output results.
23. A system for analysing audio data, wherein the system comprises:a device and a network entity;a network storage means, on the network entity, configured to store a trained dataset of one or more individuals, the trained dataset including breathingApplicant file ref: D1899 / GBD33information of one or more individuals collected at least in part in the ultrasonic frequency range;a data collecting means, on the device, for collecting audio data from an acoustic sensor, the acoustic sensor configured to have a measured frequency response at least in part the ultrasonic range;a device transmitter, on the device, configured to transmit the audio data;a network receiver, on the network entity, configured to receive the audio data; anda network processing means, on the network entity, configured to generate an evaluation dataset from the received audio data and compare the evaluation dataset with the trained dataset to determine whether one or more persons is present.
24. A system according to claim 23, further comprising:a network transmitter, on the network entity, configured to transmit an output result from the network processing means; anda device receiver, on the device, configured to receive the output result.
25. A system according to claim 23 or 24, further comprising:a network updater, on the network entity, configured to update the trained dataset to incorporate the evaluation dataset and corresponding output results.