Acoustic monitoring system and method, corresponding computer program

The acoustic monitoring system addresses spatial, temporal, and spatiotemporal ambiguities by grouping sound extracts and generating consolidated alerts, reducing false positives and enhancing response efficiency.

FR3166995A1Pending Publication Date: 2026-04-03OSO-AI
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Acoustic monitoring systems in communal living spaces for vulnerable individuals face issues of spatial, temporal, and spatiotemporal ambiguities, leading to redundant and unnecessary alerts due to the propagation and repetition of sound events, causing confusion and inefficiency in response teams.

Method used

An acoustic monitoring system with deduplication means that groups sound extracts based on similarity criteria to generate consolidated alerts, reducing redundant alerts by eliminating the need for multiple sensors and using machine learning to suppress unnecessary alerts.

Benefits of technology

The system effectively reduces false positives by consolidating alerts, minimizing redundant responses and improving efficiency in alert handling without requiring complex processing, thus enhancing the reliability of acoustic monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

This acoustic monitoring system (18) includes an interface (40) for receiving a plurality of sound extracts (Si) provided by at least one acoustic sensor located in a monitored area (Z), and means (60) for analyzing sound extracts (Si) to generate alerts (Aj). It further includes deduplication means (62) designed to group at least some of the received sound extracts or alerts (Aj) generated by the analysis means (60) according to at least one sound similarity criterion, so as to create a reduced number of group(s) of sound extracts or alerts, and to determine a consolidated sound extract or a consolidated alert (CAk) for each group created. Figure for the abstract: Fig. 4
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Acoustic monitoring system and method, corresponding computer program

[0001] The present invention relates to an acoustic monitoring system, an acoustic monitoring method implemented by this system and a computer program comprising instructions for implementing such an acoustic monitoring method when executed by the system.

[0002] This type of monitoring system is designed to receive sound samples from at least one acoustic sensor located in an area to be monitored, for example, a communal living space for vulnerable and / or elderly people requiring close monitoring and intervention in the event of a detected problem or anomaly. In this case, several acoustic sensors are advantageously distributed throughout several private and common areas of the monitored zone. Generally, the in-depth analysis of the captured sounds for the detection of an anomaly, for triggering an alert, and for making a possible decision to intervene is not performed locally but remotely by a monitoring server equipped with substantial computing resources, possibly shared between several monitored zones.Each acoustic sensor is then integrated into a communicating or connected object, for example compatible with the IoT (Internet of Things) technology, according to which this object is locally equipped with a sensor, limited processing means and means of transmitting data from the sensor to the remote server via a local network or the Internet.

[0003] Acoustic monitoring, compared to other possible technologies such as video, wearable devices, or call buttons, has the advantage of not being overly intrusive, as well as being affordable and capable of triggering real-time alerts. Video monitoring has the advantage of providing detailed visual information that allows for a more precise analysis of a situation that can sometimes be complex to interpret acoustically, but it has the disadvantage of raising privacy concerns, high costs, and dependence on lighting conditions and field of view.Wearable devices offer the advantage of providing personalized information directly sensitive to the mobility of the wearers, but they have the disadvantage of being highly dependent on user compliance and regular maintenance, making them impractical for large-scale deployment. Call buttons, while intuitive and moderately priced, suffer from an inability to detect certain serious emergencies because they rely on interaction with the user. people likely to experience these emergency situations: for example, it has been observed that approximately 85% of frail and / or elderly people living in communities do not activate them in situations which would nevertheless justify the issuing of alerts, particularly in situations where these people are on the ground following a fall.

[0004] Acoustic monitoring thus strikes a balance between practicality and efficiency. Thanks to acoustic sensors, for example simple microphones, strategically placed in private living rooms and common corridors of a collective establishment for vulnerable and / or elderly people, it can detect key sound information, such as calls for help or abnormal noises, and alert caregivers in real time.

[0005] The invention therefore applies more particularly to an acoustic monitoring system, comprising: - an interface for receiving a plurality of sound extracts provided by at least one acoustic sensor located in an area to be monitored; and - means of analyzing sound extracts for the generation of alerts.

[0006] Increasingly, analysis methods include an artificial intelligence software module with at least one machine learning engine that receives input sound clips and provides output recognition by classification of these clips, with at least one class provided for generating alerts. In a medical application where monitoring is involved in establishing a diagnosis, an example of such a system is detailed in patent document EP 3 762 942 Bl. In another application for remote monitoring of care provided to vulnerable individuals where monitoring is involved in detecting potential mistreatment of these individuals, another example of such a system is detailed in patent document US 2022 / 0108704 AL

[0007] However, this acoustic approach is not without its problems, particularly because it is indiscriminate. The alerts must be reliable. While false negatives (i.e., the absence of an alert in a situation that warrants one) must be absolutely avoided for obvious safety reasons, false positives (i.e., the generation of an alert in a situation that does not warrant one) can lead to unnecessary or untimely interventions. The main difficulty lies in the ambiguity of the sound source or the sound excerpt itself, whether spatial or temporal, in certain situations other than those mentioned in the aforementioned documents EP 3 762 942 Blet US 2022 / 0108704 A1, for example, in a group living facility for vulnerable and / or elderly people living with relative independence, each having their own private living space and sharing common areas with the others.In terms of spatial ambiguity, the same sound excerpt, for example a cry for help, can repeat itself by spreading through space, first in the private place in which it is located. The sound can occur not only within the product itself, but also in a nearby private space and / or a nearby public space, potentially triggering numerous alerts and interventions, only one of which is ultimately relevant. In terms of temporal ambiguity, the same sound clip, for example, a disturbing scream, can be repeated over time, also potentially triggering numerous alerts and interventions, only one of which is relevant. In terms of spatiotemporal ambiguity, the same sound clip, for example, a scream, can be repeated in both time and space, particularly if it originates from a person moving from one location to another within the monitored area while emitting the same scream repeatedly, also potentially triggering numerous alerts and interventions, only one of which is relevant.

[0008] By way of non-limiting example in an equally non-limiting application context, according to a first spatial ambiguity scenario illustrated in [Fig.1], a first resident 10 present in a first private place 12 of a monitored area Z, for example a specialized collective residence for elderly or vulnerable people, emits a sound 14 such as a call for help.

[0009] The first private space 12 is equipped with a communicating device 16 in telecommunication with a remote monitoring server 18. The communicating device 16 itself comprises an acoustic sensor 20, for example a microphone, a local computer 22 coupled to the acoustic sensor 20, and a transceiver 24 coupled to the local computer 22. It is, for example, located on the ceiling of the first private space 12. The remote monitoring server 18 is located outside the area to be monitored Z. It may be shared by several communicating devices located respectively in several different private or common spaces within the specialized collective residence and housed in a dedicated monitoring room. In this case, it is advantageously connected to the various communicating devices in the residence by a local telecommunications and / or data transmission network.It can also be shared by several communicating objects in several different private or common areas of several residences. In this case, it is connected to the different communicating objects by a wide area telecommunications and / or data transmission network such as the Internet 26. In all cases, it is equipped with means to trigger an alarm or alert an intervention team 28 when an anomaly is detected.

[0010] The monitored area Z includes a second private space 30, also equipped with a second communicating object 32 similar to the communicating object 16, in which a second resident 34 is present. This second private space 30 is adjacent to the first private space 12. The monitored area Z further includes a third common space 36, for example a corridor serving, in particular, the private spaces 12 and 30, also equipped with a third communicating object 38 similar to the communicating objects 16 and 32. The monitored area Z may obviously include many other private or common areas equipped with communicating objects such as communicating objects 16, 32 and 38, since it can extend to the whole of a specialized collective residence as indicated above.

[0011] According to the first scenario, the distress call sound 14 emitted by the first resident 10 is detected by the acoustic sensor 20 of the first communicating device 16. The latter then transmits a first sound clip S1 to the remote monitoring server 18. However, this sound 14 propagates through the walls and is also detected, albeit attenuated, by the acoustic sensor of the second communicating device 32 located in the adjacent second private space 30. The second communicating device 32 then transmits a second sound clip S2 to the remote monitoring server 18. The same sound 14 is also detected, albeit attenuated, by the acoustic sensor of the third communicating device 38 located in the adjacent corridor 36. The third communicating device 38 then transmits a third sound clip S3 to the remote monitoring server 18.

[0012] Assuming that the first sound clip S1, as analyzed by the remote monitoring server 18, generates an alert, there is no reason why sound clips S2 and S3, which have the same characteristics, albeit attenuated, should not also generate alerts. Three redundant but independent alerts A1, A2, and A3 are then generated based on a single sound event. As is known, these three independent alerts can trigger independent interventions by three different members of the response team 28, one heading towards the first private area 12, and the other two heading respectively towards the second and third private and common areas 30 and 36. Of course, only the first alert A1 is justified, the other two, A2 and A3, being false positives. This therefore leads to some confusion and inefficiency.Furthermore, if the first resident 10 is known to the intervention team 28 to call for help inappropriately while in their private space 12, then the responder called to the first private space 12 may ultimately decide not to go. However, the other two responders will not be able to make such a decision given the different origins of the audio clips associated with alerts A2 and A3.

[0013] By way of non-limiting example also, according to a second scenario of temporal ambiguity illustrated in [Fig.2], the first resident 10 present in the first private place 12 of the area to be monitored Z emits the sound 14, but this time in the form of a succession of blows on the ground with his cane.

[0014] According to this second scenario, the sound 14 of knocks on the ground emitted by the first resident 10 at time T1 is perceived by the acoustic sensor 20 of the first communicating object 16. The latter then transmits a first sound extract SI to the remote monitoring server 18. But the sound 14, i.e. the same succession of knocks on the ground or a completely similar series of knocks, is repeated by the first resident 10 at time T2 while he is still present in the first private space 12. It is again perceived by the acoustic sensor 20 of the first communicating object 16. The latter then transmits a second sound extract S2 to the remote monitoring server 18. This sound 14, i.e. the same series of knocks on the ground or a completely similar series of knocks, is again repeated by the first resident 10 at time T3 while he is still present in the first private space 12. It is again perceived by the acoustic sensor 20 of the first communicating object 16. The latter then transmits a third sound extract S3 to the remote monitoring server 18.

[0015] Assuming that the first sound clip S1, as analyzed by the remote monitoring server 18, generates an alert, there is no reason why the subsequent sound clips S2 and S3, which have the same or very similar characteristics, should not also generate alerts. Three redundant but independent alerts A1, A2, and A3 are then generated based on a single, repeated sound event. As is known, these three independent alerts can trigger independent interventions by three different members of the response team 28, all three being successively called upon at times T1, T2, and T3 to proceed to the first private location 12. Of course, only the first alert A1 is justified; the other two, A2 and A3, are false positives. This again leads to confusion and inefficiency.

[0016] By way of non-limiting example also, according to a third spatiotemporal ambiguity scenario illustrated in [Fig.3], the first resident 10 present in the first private place 12 of the area to be monitored Z emits the sound 14, but this time in the form of an ominous cry.

[0017] According to this third scenario, the alarming scream 14 emitted by the first resident 10 at time T1 is detected by the acoustic sensor 20 of the first communicating object 16. The latter then transmits a first sound clip S1 to the remote monitoring server 18. However, this sound 14 propagates through the walls and is also detected, albeit attenuated, by the acoustic sensor of the second communicating object 32 located in the adjacent second private space 30. The second communicating object 32 then transmits a second sound clip S2 to the remote monitoring server 18. The same sound 14 is also detected, albeit attenuated, by the acoustic sensor of the third communicating object 38 located in the adjacent corridor 36. The third communicating object 38 then transmits a third sound clip S3 to the remote monitoring server 18.

[0018] Still according to this third scenario, the sound 14, i.e. the same disturbing cry or a very similar disturbing cry, is repeated by the first resident 10 at a time T2 when he has just left the first private place 12 and is in the corridor 36. It is detected by the acoustic sensor of the third communicating object 38. The latter then transmits a fourth sound clip S4 to the remote monitoring server 18. But this sound 14 propagates through the walls and is also detected, albeit attenuated, by the acoustic sensor 20 of the first communicating object 16 located in the first private space 12. The first communicating object 16 then transmits a fifth sound clip S5 to the remote monitoring server 18. The same sound 14 is also detected, albeit attenuated, by the acoustic sensor of the second communicating object 32 located in the second private space 30. The second communicating object 32 then transmits a sixth sound clip S6 to the remote monitoring server 18.

[0019] According to this third scenario, the sound 14, i.e., the same disturbing cry or a very similar disturbing cry, is repeated by the first resident 10 at time T3 while he has moved in the corridor 36 and is near the second private space 30. It is perceived by the acoustic sensor of the third communicating object 38. The latter then transmits a seventh sound clip S7 to the remote monitoring server 18. But this sound 14 propagates through the walls and is also perceived in a muted form by the acoustic sensor of the second communicating object 32 located in the second private space 30. The second communicating object 32 then transmits an eighth sound clip S8 to the remote monitoring server 18. The same sound 14 is also perceived in a muted form by the acoustic sensor 20 of the first communicating object 16 located in the first private space 12.The first communicating object 16 then transmits a ninth sound extract S9 to the remote monitoring server 18.

[0020] Assuming that the first sound extract SI, as analyzed by the remote monitoring server 18, generates an alert, there is no reason why the subsequent sound extracts S2 to S9, which have the same or very similar characteristics, should not also generate alerts. Nine redundant but independent alerts A1, A2, A3, A4, A5, A6, A7, A8, A9 are then generated based on a single, repeated sound event. As is known in itself, these nine independent alerts can trigger independent interventions by up to nine different responders from the intervention team 28, all being successively called at times T1, T2 and T3 to go to the first private place 12, the second private place 30 and the common corridor 36. Of course, only the first alert Al is justified, the other eight A2 to A9 being false positives.This, therefore, again creates confusion and inefficiency.

[0021] Solutions have been explored to try to overcome this kind of acoustic ambiguity. But they are generally complex, like the solution taught in the article by Chakraborty et al., entitled "Joint model-based recognition and localization of overlapped acoustic events using a set of distributed small "Microphone arrays," published on arXiv.org under the reference arXiv:1712.07065, on December 19, 2017. This method requires the installation of distributed arrays of small microphones to locate and identify a plurality of captured sound fragments, as well as a complex conjoint analysis model. Furthermore, it is relevant for identifying and locating a sound in a single room, such as a meeting room, but not for eliminating ambiguities in sound distributed across more complex monitoring areas.

[0022] It may therefore be desirable to provide an acoustic monitoring system which makes it possible to overcome at least some of the aforementioned problems and constraints.

[0023] An acoustic monitoring system comprising: is therefore proposed. - an interface for receiving a plurality of sound extracts provided by at least one acoustic sensor placed in an area to be monitored; - means for analyzing sound extracts for generating alerts; and - deduplication means designed to group at least some of the sound extracts received or the alerts generated by the analysis means according to at least one criterion of sound similarity, so as to constitute a reduced number of group(s) of sound extracts or alerts, and to determine a consolidated sound extract or a consolidated alert for each group constituted.

[0024] Thus, by simply grouping sounds by similarity, with these groupings potentially involving one or more acoustic sensors, and by determining a consolidated sound extract (to be analyzed to potentially trigger an alert) or a consolidated alert for each group, it is possible to reduce, or even eliminate, the aforementioned spatial, temporal, or spatiotemporal ambiguities. The false positives mentioned earlier are then effectively avoided using simple methods, notably by eliminating the need to provide multiple acoustic sensors, such as the provision of sensor / microphone arrays as described in the article by Chakraborty et al., in each private or common area of ​​the zone to be monitored.

[0025] Optionally, an acoustic monitoring system according to the invention may further include: - said at least one acoustic sensor placed in the area to be monitored; and / or - an interface for the automatic transmission of each alert from a consolidated sound extract or of each consolidated alert to a pre-registered terminal of a user able to intervene.

[0026] Optionally, the deduplication means are also designed to: - to group several sound extracts or several alerts, originating from several acoustic sensors respectively located in several distinct places within the area to be monitored, into a single group of sound extracts or alerts, when said sound extracts or said alerts also extend within the same first predetermined time window, for example a first time window of 10 to 20 seconds; and - determine the consolidated sound extract or consolidated alert by selecting, in each group formed, the sound extract or alert presenting a maximum sound level, classification confidence rate or signal-to-noise ratio.

[0027] Optionally also, the deduplication means are designed to, when several sound extracts or several alerts from the same and single acoustic sensor among said at least one acoustic sensor are grouped into the same group constituted in a second predetermined time window, for example a second time window of at least one hour, exclude from the determination of the consolidated sound extract or the consolidated alert of this group any sound extract or any alert other than the one emitted first from said same and single acoustic sensor.

[0028] Optionally, said at least one sound similarity criterion applied by the deduplication means includes at least one of the elements of the set consisting of: - a spectral proximity between sound extracts in terms of spectral envelopes; - a spectral proximity between sound extracts in terms of characteristic frequencies; - a spectral proximity between sound extracts in terms of frequency amplitudes; - a temporal proximity between sound extracts in terms of temporal envelopes; - a temporal proximity between sound extracts in terms of temporal amplitudes; and - temporal proximity between sound extracts in terms of durations.

[0029] Optionally, an acoustic monitoring system according to the invention may also include additional means for suppressing at least one alert from a consolidated sound extract or at least one consolidated alert, this suppression resulting from the application of predetermined rules and / or machine learning.

[0030] Optionally, the deduplication means are also more specifically designed to: - to group at least some of the received sound extracts according to at least one criterion of sound similarity, so as to constitute a number of group(s) of sound extracts less than the number of sound extracts received; - determine a consolidated audio extract for each group formed; and - provide each consolidated audio extract to the analysis tools rather than each audio clip received.

[0031] Optionally, the acoustic monitoring system is also designed to provide all received sound extracts to the analysis means, and the deduplication means are more specifically designed to: - to group at least part of the alerts generated by the analysis means according to said at least one sound similarity criterion, this sound similarity criterion relating to the analyzed sound extracts generating the alerts, so as to constitute a reduced number of alert group(s); - determine a consolidated alert for each group formed.

[0032] Optionally, the acoustic monitoring system is also designed for N acoustic sensors arranged in the area to be monitored and includes means for storing in memory an acoustic proximity matrix AP, of which each matrix coefficient APm>n, where l <m<N et l<n<N, comporte au moins une valeur de probabilité qu’au moins un son prédéterminé émis à proximité du capteur acoustique Cm et principalement capté par celui-ci soit également capté par le capteur acoustique Cn, les moyens de déduplication étant conçus pour exploiter cette matrice de proximité acoustique AP pour écarter de calculs de similarité sonore certains couples de capteurs acoustiques jugés trop éloignés l’un de l’autre au vu du coefficient matriciel correspondant.

[0033] An acoustic monitoring method is also proposed, comprising the following steps: - reception of a plurality of sound extracts provided by at least one acoustic sensor placed in an area to be monitored; - generation, by means of analyzing sound extracts from an acoustic monitoring system, of alerts resulting from at least a part of the sound extracts they receive; - grouping, by means of deduplication of the acoustic monitoring system, of at least a part of the sound extracts received or the alerts generated by the means of analysis according to at least one criterion of sound similarity, so as to constitute a reduced number of group(s) of sound extracts or alerts; - determination, by means of deduplication, of a consolidated sound extract or a consolidated alert for each group constituted.

[0034] Also proposed is a computer program downloadable from a communication network and / or recorded on a computer-readable medium and / or executable by a processor, comprising instructions for the execution of the steps of an acoustic monitoring process according to the invention, when said program is executed by at least one computer of an acoustic monitoring system according to the invention.

[0035] The invention will be better understood with the aid of the following description, given solely by way of example and made with reference to the attached drawings in which: - the [Fig.1] already described, illustrates a first scenario generating redundant alerts; - the [Fig.2] already described, illustrates a second scenario generating redundant alerts; - the [Fig.3] already described, illustrates a third scenario generating redundant alerts; - [Fig.4] schematically represents the general structure of an installation with an acoustic monitoring system according to a first embodiment of the invention; - [Fig.5] schematically represents the general structure of an installation with an acoustic monitoring system according to a second embodiment of the invention; - Figure 6 illustrates the successive stages of an acoustic monitoring process implemented by the system in Figure 4; and - [Fig.7] illustrates the successive steps of an acoustic monitoring process implemented by the system of [Fig.5].

[0036] The installation schematically represented in [Fig. 4] illustrates in a simplified way the area to be monitored Z and the plurality of communicating objects with acoustic sensors, generically denoted Cl, ..., CN, that it comprises and which are distributed in the common and private spaces. In accordance with the general principles of the present invention, a single acoustic sensor, for example a single microphone, can be provided in each common or private space of the area to be monitored Z. Thus, a single communicating object can advantageously be provided per common or private space. However, alternatively, the general principles of the present invention do not preclude the arrangement of several acoustic sensors and communicating objects in the same common or private space of the area to be monitored Z, particularly when it is large or consists of several rooms. Each acoustic sensor, such as sensor 20, of each communicating object Cl, ...CN is a carrier signal provider. of acoustic information. This signal is optimal in terms of content, since it is directly provided by the acoustic sensor without any post-processing. Its quality therefore depends only on the acoustic sensor itself. It is provided at each instant t, according to a predetermined optimal temporal sampling or continuously if possible. Each local computer, such as computer 22, of each communicating object Cl, ..., CN is advantageously designed to consume little energy, particularly using IoT technology. It therefore has limited computing power and memory compared to that of the remote monitoring server 18.Apart from this, it can have a completely conventional architecture with a processing unit (e.g., a processor) associated with at least one memory (e.g., RAM or other) for storing data files and computer programs whose instructions are intended to be executed by the processing unit. Each local computer of each communicating object Cl, ..., CN thus functionally includes several computer programs or several functions of the same computer program to process the signal provided by the corresponding acoustic sensor and to provide, as output from this processing, successive, time-limited sound extracts likely to contain relevant information to be analyzed by the monitoring server 18.It should be noted that a compromise must be found regarding the processing to be performed by each local computer, between the most efficient processing possible, allowing for the lightening of the transmission of sound extracts but requiring a minimum of local computing power, and the simplest possible processing, reducing the necessary local computing power but placing greater demands on the transmission capacity of network 26. Furthermore, it is entirely possible to consider a feedback loop interaction between each communicating object Cl, ..., CN and the remote monitoring server 18 to configure this local processing according to the server's analysis needs, which may evolve over time and depending on the situation.

[0037] The installation of [Fig.4] further illustrates in more detail the remote monitoring server 18, capable of interacting with the communicating objects Cl, ..., CN via the aforementioned network 26. It includes for this purpose an interface 40 for receiving a plurality of sound extracts Si provided by the acoustic sensors of the communicating objects Cl, ..., CN after processing by the local computers.

[0038] It further comprises a remote computer 42, advantageously equipped with significantly greater computing power and memory than each local computer. Apart from this, the remote computer 42 can have a completely conventional architecture with a processing unit 44 (for example, a processor) associated with at least one memory 46 (for example, RAM or other memory). It can, for example, be implemented in a computing device such as a conventional computer comprising a processor associated with one or more memories for storing data files and computer programs whose instructions are intended to be executed by the processing unit 44. As illustrated in [Fig. 4], the remote computer 42 thus functionally comprises six computer programs 48, 50, 52, 54, 56, 58, or six functions of the same computer program distributed across several functional software modules. It should be noted that the computer programs 48, 50, 52, 54, 56, 58 are presented as distinct, but this distinction is purely functional. They could just as easily be grouped in any possible combination into one or more software programs. Their functions could also be at least partially microprogrammed or micro-wired into dedicated integrated circuits.Alternatively, the computer system implementing the remote computer 42 could be replaced by an electronic device composed solely of digital circuits (without a computer program) to perform the same functions. Also alternatively, at least some of the aforementioned computer programs could be remote and accessible to the remote computer 42 via the Internet. Generally speaking, even though all the aforementioned software and memory components are presented as being integrated into the same remote computer 42, they could just as easily be dispersed across separate hardware components, even located far from each other, but interconnected via a network (data transmission bus, local area network, wide area network, Internet, etc.).

[0039] Computer programs 48 and 50 are included in a first functional software module 60 that provides means for analyzing sound clips to generate alerts. This first functional analysis module 60 is, for example, an artificial intelligence software module with at least one machine learning engine that receives the sound clips as input and provides output recognition by classification of these clips, with at least one class being provided for generating alerts. In this respect, it could be designed following the teachings of the aforementioned documents EP 3 762 942 B1 and US 2022 / 0108704 AL II. It could also integrate a rule-based expert system and an inference engine. In terms of artificial intelligence, it could also combine the advantages of at least one machine learning engine and an expert system.

[0040] Optionally but cleverly, and by way of non-limiting example, the first functional analysis module 60 of [Fig.4] actually comprises: - the computer program 48 comprising instructions which, when executed by the processing unit 44, implement at least one first supervised machine learning engine for sound detection, this first engine being capable of detecting one or more classes of sounds, for example classes of sounds such as "call for help", "scream", " "Moaning," "falling," "vomiting," etc., from the provided audio excerpts; and - the computer program 50 comprising instructions which, when executed by the processing unit 44, implement at least one second supervised machine learning engine for state detection, this second engine being capable of detecting one or more classes of states from the class or classes of sounds detected by the first engine, each class of state being representative of an environmental state (i.e. a place or any person) in the area to be monitored Z, at least one of these classes of states being representative of an alert.

[0041] This two-level arrangement of supervised machine learning engines is not the subject of the present invention, although it is advantageously combinable with its general principles, so it will not be described in further detail.

[0042] In accordance with the general principles of the present invention, the computer programs 52, 54 and 56 are included in a second functional software module 62 forming deduplication means designed to: - to group at least some of the sound extracts received or alerts generated by the first functional analysis module 60 according to at least one sound similarity criterion, so as to constitute a reduced number of group(s) of sound extracts or alerts; and - determine a consolidated audio extract or a consolidated alert for each group formed.

[0043] More specifically, in the first embodiment illustrated by [Fig.4], the first functional analysis module 60 is programmed to analyze all sound extracts Si received by the interface 40 and generate, where appropriate, alerts Aj, the number of alerts generated being less than or equal to the number of sound extracts.

[0044] These alerts Aj, potentially redundant in view of the possible scenarios in Figures 1 to 3, are provided to the second functional software deduplication module 62.

[0045] The computer program 52 of the second deduplication software functional module 62 includes instructions which, when executed by the processing unit 44, perform an accumulation, in at least one buffer memory 64, of the alerts Aj received in at least one predetermined time window: - for example, within a predetermined initial time window of 10 to 20 seconds when it is desired to reduce or eliminate spatial redundancies, this short initial time window allows for consideration of network latency variations 26 as well as variations in temporal labeling (i.e., start, duration and / or end) of sound extracts from the same sound 14 but perceived by different acoustic sensors respectively arranged in several distinct locations of the area to be monitored Z: cf. scenario of [Fig.1]; - for example in the same second predetermined time window of at least one hour when it is desired to reduce or eliminate temporal redundancies, this second longer time window allowing to take into account possible repetitions in time of the same sound captured by the same and unique acoustic sensor: cf. scenario of [Fig.2]; - for example in several distinct time windows, in particular the first and second time windows mentioned above, when it is desired to reduce or eliminate spatial, temporal and spatiotemporal redundancies: see scenario in [Fig.3].

[0046] The computer program 54 of the second deduplication software functional module 62 includes instructions which, when executed by the processing unit 44, group at least a portion of the alerts Aj accumulated by the execution of the computer program 52 according to at least one sound similarity criterion, so as to constitute a reduced number of alert group(s). This sound similarity criterion relates to the analyzed sound extracts that respectively generated the accumulated alerts. It may, for example, take the form of at least one of the following criteria: - a spectral proximity between sound extracts in terms of spectral envelopes; - a spectral proximity between sound extracts in terms of characteristic frequencies; - a spectral proximity between sound extracts in terms of frequency amplitudes; - a temporal proximity between sound extracts in terms of temporal envelopes; - a temporal proximity between sound extracts in terms of temporal amplitudes; and - a temporal proximity between sound extracts in terms of durations.

[0047] It is applied, for example, using a predetermined similarity threshold value beyond which compared alerts are considered to originate from the same sound occurrence. The similarity calculation itself can be performed using known methods such as machine learning model methods like nearest neighbors, advanced feature extraction techniques, the cosine similarity method, cross-correlation calculation, adapted filtering, etc.

[0048] Optionally, prior training can be conducted using predetermined sounds described as opportunistic, for example, high amplitudes or spectral contents assumed to be easily propagated, to estimate a priori the acoustic proximity of the respective acoustic sensors of the communicating objects Cl, ..., CN. The result of this training can be stored in memory 46 in the form of a square acoustic proximity matrix denoted AP, where each coefficient APm>n, where l <m<N et l<n<N, comporte au moins une valeur de probabilité qu’au moins un son opportuniste émis à proximité du capteur acoustique de l’objet communicant Cm et principalement capté par celui-ci soit également capté par le capteur acoustique de l’objet communicant Cn.The use of this AP matrix by the computer program 54 allows certain pairs of communicating objects deemed too far apart from each other, based on their corresponding matrix coefficient, to be excluded from similarity calculations. As a result, alerts from these pairs are not grouped together, even for similar and temporally close corresponding acoustic extracts. This simplifies the similarity calculations. Furthermore, it helps avoid false negatives when two distinct but acoustically similar sound events occur simultaneously: for example, two bottles falling and breaking at the same time in two private spaces acoustically distant from the monitored area Z.

[0049] Finally, the computer program 56 of the second deduplication software functional module 62 includes instructions which, when executed by the processing unit 44, determine a consolidated alert for each group formed by executing the computer program 54. This determination is made, for example, by selecting one of the alerts from each group, in particular the one considered the most relevant in its group because it originates from the sound extract exhibiting a maximum sound level (i.e., maximum volume), or a maximum classification rate, or a maximum signal-to-noise ratio. This selection is particularly relevant for eliminating spatial ambiguities based on the first time window defined previously.In addition, or independently, this determination can also be made, for example, by excluding from each group any alert other than the first one emitted from a single acoustic sensor. This additional selection aid is particularly relevant for eliminating temporal ambiguities based on the second time window defined previously, notably for imposing a short-duration mute (i.e., that of the second time window) on the relevant acoustic sensor after receiving a first alert. It is also possible to make the duration of this second time window variable depending on the type of alert. several similar types of alerts (for example, a "call for help" alert and a "shout" alert) dependent on the same observation time window.

[0050] The execution of the computer program 56 results in the provision, by the second functional software deduplication module 62, of consolidated CAk alerts which are a priori less numerous than the Aj alerts which it has received, and above all much less redundant or even not at all.

[0051] Of course, if only one alert accumulates in buffer 64, the grouping and selection performed by the execution of computer programs 54, 56 are reduced to their simplest expression, i.e., the rendering of this alert as is, as a consolidated alert. Similarly, if several alerts accumulate in buffer 64 but none is deemed similar to another, the grouping and selection performed by the execution of computer programs 54, 56 render these alerts as is, as consolidated alerts.

[0052] Optionally, but advantageously, it may be desirable to extend the aforementioned mute function both temporally and functionally. Thus, the computer program 58 includes instructions which, when executed by the processing unit 44, are capable of suppressing at least one consolidated alert by applying predetermined rules and / or machine learning.

[0053] By way of non-limiting example, it may be desirable to disregard certain consolidated alerts in certain very specific situations which can be defined using "if" type rules <conditions>"Then rule 1> otherwise rule 2>." In practical terms, consolidated alerts repeated over several days, or even over one or more hours, can be muted for several subsequent days, or several subsequent hours—that is, without triggering an intervention—when a resident in the monitored area Z reaches a certain number of alerts of a specific type within a predefined period. Variable threshold values ​​can be applied depending on the type of alert. For example: "If there are more than 7 'shout' type alerts per day from the same person over 4 consecutive days, then any 'shout' type alert from that person is muted from the 5th day onward for a predetermined period."

[0054] More generally, a rule-based model can be used to represent recurring situations and choose an alert that will best represent a given event during which several alerts will have been generated.

[0055] For example, starting from a normal situation, without any alerts, a generation of successive alerts can increase a measure of the severity of a situation, which is then used to decide on its handling, in accordance with: - an initial escalation phase during which sound excerpts are aggregated to determine whether an ongoing event warrants being transformed into an alert: this phase is particularly useful for aggregating weak signals which, taken individually, might have been ignored; it also allows for the processing of alerts that generate false positives (i.e., should be ignored) but whose close repetition is a sign of an ongoing event that deserves some attention (for example, respiratory distress or the sound of a machine); alerts can be grouped by their type (i.e., crying, moaning, or vocal calls), thus contributing together to the same severity measurement; - reaching a sufficient severity measurement value, then implying the designation of one of the alerts received as representing the event and its transmission, or the muting of the entire sequence of alerts in the opposite case; - a muting for an incompressible period after transmission of the alert designated as representative, this muting depending on the types of subsequent alerts, those of related type having the muting applied but not the others (for example the detection of a fall to be transmitted after several successive cries of which only the first is transmitted in the form of an alert); - taking into account the convergence over time of alerts and their types to decide whether they are part of the same event or not, when the latter is prolonged.

[0056] By way of further, but not limited to, the muting can be adaptive and scalable by using a machine learning engine, so as to automatically adjust to the habits of each resident in the monitored area Z. In other words, if a sound is recognized as habitual for a resident, for example, a frequent cough or a repetitive noise, the machine learning engine identifies it, using a learning model trained with labeled data, as non-critical and suppresses the consolidated alert resulting from this ritual sound. The system continuously learns and adapts to habitual sounds, thus minimizing irrelevant consolidated alerts over time. This type of muting is particularly well-suited to longer time scales, such as the weekly or monthly scale, which a rule-based system can hardly handle.

[0057] After execution of the computer program 58, the remote computer 42 provides consolidated and confirmed CAk' alerts which can then be transmitted to an automatic transmission interface 66 of the remote monitoring server 18, for sending each of these consolidated and confirmed CAk' alerts to a terminal pre-registered for a user able to intervene, for example a member of the intervention team 28, and for triggering an alarm (audible, visual, by message, ...) or an intervention.

[0058] It should be noted that the aforementioned processing, performed by the analysis functional module 60, the deduplication functional module 62, and the muting computer program 58, can be carried out in parallel using a distributed approach capable of reducing computational complexity. Specifically, each alert is recorded in a database of recent alerts and compared with other alerts from the same time window and the same area to be monitored. This allows for deduplication without centralizing the processing, thus improving the system's scalability.

[0059] It should also be noted that the consolidated and confirmed CAk' alerts can be transmitted, to the automatic transmission interface 66 and then to the terminals of the members of the intervention team 28, with contextual information relating in particular to other possible alerts from their respective groups. For example, in the scenario of [Fig. 1], only alert A1 can be transmitted to a member of the intervention team 28 after consolidation and confirmation, with contextual information on alerts A2 and A3 ("call for help also heard in the private area 30 and in the corridor 36"). In the scenario of [Fig. 2], only alert A1 can be transmitted to a member of the intervention team 28 after consolidation and confirmation, with contextual information on alerts A2 and A3 ("succession of repeated blows to the ground at times T2 and T3 and heard in the same private area 12"). In the scenario of [Fig.3], only alert A1 can be transmitted to a member of the intervention team 28 after consolidation and confirmation, with contextual information on alerts A2 to A9 (“disturbing cry repeated at times T2 and T3, heard in the same private area 12 but also in private area 30 and in corridor 36”). .

[0060] It should also be noted that, alternatively, all alerts Aj could be transmitted to the automatic transmission interface 66 and then to the terminals of the members of the intervention team 28, but indicating for those that are not consolidated and confirmed that a corresponding consolidated and confirmed alert is or has also been transmitted to a member of the intervention team 28. Thus, each member of the intervention team notified of such an unconsolidated and unconfirmed alert is informed of the redundancy of the alerts even if no particular action is required on their part. For example, in the scenarios of Figures 1 and 2, the three alerts A1, A2, A3 can be transmitted to three different members of the intervention team 28, only alert A1 requiring action by the person who receives it, the other two being transmitted only for redundancy information. In the scenario of [Fig.3], the nine alerts A1 to A9 can be transmitted to nine members. different from intervention team 28, only alert Al requires action by the person who receives it, the other eight are transmitted only for redundancy information.

[0061] According to a second embodiment illustrated by [Fig.5], the remote monitoring server 18 is configured in a slightly different way, so that the chain of processing carried out by executing computer programs 52, 54 and 56 is applied before that of the processing carried out by executing computer programs 48 and 50.

[0062] It follows that, according to this second embodiment, all the sound extracts Si received by the interface 40 are processed by the second software deduplication functional module 62. Thus, by executing the computer program 52, they are accumulated in the buffer 64 in one or more of the predetermined time windows mentioned above. Then, by executing the computer programs 54 and 56, consolidated sound extracts CSj are provided by the second software deduplication functional module 62, which are a priori fewer in number than the sound extracts Si it received, and above all, much less, or even no longer, redundant.

[0063] Of course, if only one sound clip accumulates in buffer 64, the grouping and selection performed by executing computer programs 54, 56 are reduced to their simplest expression, i.e., the reproduction of this sound clip as is, as a consolidated sound clip. Similarly, if several sound clips accumulate in buffer 64 but none is deemed similar to another, the grouping and selection performed by executing computer programs 54, 56 reproduce these sound clips as is, as consolidated sound clips.

[0064] The consolidated sound extracts CSj are then processed by the first functional analysis module 60 to generate, where appropriate, consolidated alerts CAk, the number of consolidated alerts generated being less than or equal to the number of consolidated sound extracts.

[0065] As in the first embodiment of [Fig. 4], optionally but advantageously, the computer program 58 can be executed to suppress at least one consolidated alert by muting. After execution of the computer program 58, the remote computer 42 provides the consolidated and confirmed CAk' alerts, which can then be transmitted to the automatic transmission interface 66 of the remote monitoring server 18.

[0066] It should be noted that the second embodiment makes it possible to advantageously lighten the processing carried out by the first functional analysis module 60 by potentially reducing significantly the number of sound extracts provided as input.

[0067] It should also be noted that, alternatively, all consolidated CAk alerts could be transmitted to the automatic transmission interface 66 and then to the terminals of the members of the intervention team 28, but indicating for those which are not confirmed that a corresponding consolidated and confirmed alert is or has also been transmitted to a member of the intervention team 28.

[0068] The operation of the remote monitoring server 18 of [Fig.4] will now be detailed with reference to [Fig.6].

[0069] During a first reception step 100 which is executed at each instant t, i.e. continuously, the interface 40 of the remote monitoring server 18 receives all the sound extracts Si provided by the acoustic sensors of the communicating objects Cl, ..., CN after processing by the local computers.

[0070] During an analysis step 102 executed at each instant and in parallel with step 100, the first functional analysis module 60 processes the sound extracts Si to generate the alerts Aj, by executing the computer programs 48 and 50.

[0071] During a subsequent deduplication step 104, the second functional deduplication module 62 processes the alerts Aj to generate the consolidated alerts CAk, by successive execution of the computer programs 52, 54 and 56. This deduplication 104 includes grouping at least a part of the alerts Aj generated in step 102 according to at least one sound similarity criterion among those mentioned previously, so as to constitute a reduced number of alert group(s), and then determining a consolidated alert for each group constituted.

[0072] During an optional muting step 106, the consolidated CAk alerts are processed by running computer program 58 for the provision of consolidated and confirmed CAk alerts.

[0073] Finally, during a transmission step 108, the consolidated and confirmed CAk' alerts are transmitted to the automatic transmission interface 66 of the remote monitoring server 18, for the sending of each of these consolidated and confirmed CAk' alerts to a pre-registered terminal of a user authorized to intervene, for example, a member of the intervention team 28, and for the triggering of an alarm (audible, visual, by message, etc.) or an intervention. As previously stated, alternatively, all alerts Aj can be sent, those that are not consolidated and confirmed being flagged as corresponding to a notified consolidated and confirmed alert and perhaps even already being handled by another team member.

[0074] The slightly different operation of the remote monitoring server 18 of [Fig.5] will now be detailed with reference to [Fig.7].

[0075] During a first reception step 200 which is executed at each instant t, i.e. continuously, the interface 40 of the remote monitoring server 18 receives all the sound extracts Si provided by the acoustic sensors of the communicating objects Cl, ..., CN after processing by the local computers.

[0076] During a deduplication step 202 executed at each instant and in parallel with step 200, the second functional deduplication module 62 processes the sound extracts Si to generate the consolidated sound extracts CSj, by successive execution of the computer programs 52, 54 and 56. This deduplication 202 includes the grouping of at least a part of the sound extracts Si according to at least one sound similarity criterion among those mentioned above, so as to constitute a reduced number of group(s) of sound extracts, then the determination of a consolidated sound extract for each group constituted.

[0077] During a subsequent analysis step 204, the first functional analysis module 60 processes the consolidated sound extracts CSj to generate the consolidated alerts CAk, by executing computer programs 48 and 50.

[0078] During an optional muting step 206, the consolidated CAk alerts are processed by running computer program 58 for the provision of consolidated and confirmed CAk alerts.

[0079] Finally, during a transmission step 208, the consolidated and confirmed CAk' alerts are transmitted to the automatic transmission interface 66 of the remote monitoring server 18, for the sending of each of these consolidated and confirmed CAk' alerts to a pre-registered terminal of a user authorized to intervene, for example, a member of the intervention team 28, and for the triggering of an alarm (audible, visual, by message, etc.) or an intervention. As previously stated, alternatively, all consolidated CAk alerts can be sent, with those that are not confirmed being flagged as corresponding to a notified consolidated and confirmed alert, and perhaps even already being handled by another team member.

[0080] It is clear that an acoustic monitoring system such as the one described above makes it possible to effectively avoid false positive type alerts, without requiring overly complex processing means.

[0081] It should also be noted that the invention is not limited to the embodiments described above. Indeed, it will be apparent to those skilled in the art that various modifications can be made to the embodiments described above, in light of the information just disclosed. In the detailed presentation of the invention given above, the terms used should not be interpreted as limiting the invention to the embodiments set forth in this description, but should be interpreted to include all equivalents. whose prediction is within the reach of the expert by applying his general knowledge to the implementation of the teaching that has just been disclosed to him.< / conditions>

Claims

Demands

1. Acoustic monitoring system comprising: - an interface (40) for receiving a plurality of sound extracts (Si) provided by at least one acoustic sensor (20; Cl, ​​..CN) disposed in a monitored area (Z); - means (60) for analyzing sound extracts (Si; CSj) for generating alerts (Aj; CAk); characterized in that it further comprises deduplication means (62) designed to: - group at least a portion of the received sound extracts (Si) or alerts (Aj) generated by the analysis means (60) according to at least one sound similarity criterion, so as to constitute a reduced number of group(s) of sound extracts or alerts; - determine a consolidated sound extract (CSj) or a consolidated alert (CAk) for each group constituted.

2. Acoustic monitoring system according to claim 1, further comprising: - said at least one acoustic sensor (20; Cl, ​​..., CN) disposed in the area to be monitored (Z); and / or - an interface (66) for automatic transmission of each alert from a consolidated sound extract (CSj) or of each consolidated alert (CAk) to a pre-registered terminal of a user able to intervene.

3. An acoustic monitoring system according to claim 1 or 2, wherein the deduplication means (62) are designed to: - group several sound extracts (Si) or several alerts (Aj), from several acoustic sensors (Cl, CN) respectively located in several distinct places (12, 32, 36) of the area to be monitored (Z), into the same group of sound extracts or alerts, when said sound extracts or said alerts extend in addition within the same first predetermined time window, for example a first time window of 10 to 20 seconds; and - determine the consolidated sound extract (CSj) or the consolidated alert (CAk) by selecting, in each group constituted, the sound extract or the alert presenting a maximum sound level, classification confidence rate or signal-to-noise ratio.

4. Acoustic monitoring system according to any one of claims 1 to 3, wherein the deduplication means (62) are designed to, when several sound extracts (Si) or several alerts (Aj) from the same and single acoustic sensor among said at least one acoustic sensor (20; Cl, ​​..., CN) are grouped into the same group constituted in a second predetermined time window, for example a second time window of at least one hour, exclude from the determination of the consolidated sound extract (CSj) or the consolidated alert (CAk) of this group any sound extract or any alert other than the one emitted first from said same and single acoustic sensor.

5. An acoustic monitoring system according to any one of claims 1 to 4, wherein said at least one sound similarity criterion applied by the deduplication means (62) comprises at least one of the elements of the assembly consisting of: - a spectral proximity between sound extracts in terms of spectral envelopes; - a spectral proximity between sound extracts in terms of characteristic frequencies; - a spectral proximity between sound extracts in terms of frequency amplitudes; - a temporal proximity between sound extracts in terms of temporal envelopes; - a temporal proximity between sound extracts in terms of temporal amplitudes; and - a temporal proximity between sound extracts in terms of durations.

6. Acoustic monitoring system according to any one of claims 1 to 5, comprising additional means (58) for suppressing at least one alert from a consolidated sound extract (CSj) or at least one consolidated alert (CAk), such suppression resulting from the application of predetermined rules and / or machine learning.

7. An acoustic monitoring system according to any one of claims 1 to 6, wherein the deduplication means (62) are more specifically designed to: - group at least some of the received sound extracts (Si) according to said at least one sound similarity criterion, so as to constitute a number of sound extract group(s) less than the number of received sound extracts (Si); - determine a consolidated sound extract (Csj) for each group constituted; and - provide each consolidated sound extract (CSj) to the analysis means (60) rather than each received sound extract (Si).

8. An acoustic monitoring system according to any one of claims 1 to 6, designed to provide all received sound extracts (Si) to the analysis means (60) and wherein the deduplication means (62) are more specifically designed to: - group at least part of the alerts (Aj) generated by the analysis means (60) according to said at least one sound similarity criterion, this sound similarity criterion relating to the analyzed sound extracts generating the alerts (Aj), so as to constitute a reduced number of alert group(s); - determine a consolidated alert (CAk) for each group constituted.

9. An acoustic monitoring system according to any one of claims 1 to 8, for N acoustic sensors (20; Cl, ​​..., CN) arranged in the area to be monitored (Z), comprising means for storing in memory (46) an acoustic proximity matrix AP, of which each matrix coefficient APm>n, where l <m<N et l<n<N, comporte au moins une valeur de probabilité qu’au moins un son prédéterminé émis à proximité du capteur acoustique Cm et principalement capté par celui-ci soit également capté par le capteur acoustique Cn, les moyens de déduplication (62) étant conçus pour exploiter cette matrice de proximité acoustique AP pour écarter de calculs de similarité sonore certains couples de capteurs acoustiques jugés trop éloignés l’un de l’autre au vu du coefficient matriciel correspondant.

10. A method for acoustic monitoring, comprising the following steps: - reception (100; 200) of a plurality of sound extracts (Si) provided by at least one acoustic sensor (20; Cl, ​​..., CN) placed in a monitored area (Z); - generation (102; 204), by means (60) of analysis of sound extracts from an acoustic monitoring system, of alerts (Aj; CAk) resulting from at least a part of sound extracts (Si; CSj) which they receive; characterized in that it also includes the following steps: - grouping (104; 202), by means of deduplication (62) of the acoustic monitoring system, of at least a part of the sound extracts (Si) received or of the alerts (Aj) generated by the analysis means (60) according to at least one criterion of sound similarity, so as to constitute a reduced number of group(s) of sound extracts or alerts; - determination (104; 202), by means of deduplication, of a consolidated sound extract (CSj) or a consolidated alert (CAk) for each group constituted.

11. Computer program (48, 50, 52, 54, 56) downloadable from a communication network and / or stored on a computer-readable medium (46) and / or executable by a processor, characterized in that it includes instructions for carrying out the steps of an acoustic monitoring method according to claim 10, when said program is executed by at least one computer (44) of an acoustic monitoring system according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • System and method for generating diagnostic health information using deep learning and sound understanding

    EP3762942A1

  • System and method for generating diagnostic health information using deep learning and sound understanding

    EP3762942B1

  • Real-time detection and alert of mental and physical abuse and maltreatment in the caregiving environment through audio and the environment parameters

    US20220108704A1