Systems and methods for contextual scene analysis
The AIS integrates acoustic sensors with human-like auditory processing models to enhance VASs' situational awareness, addressing the limitations of line-of-sight sensors by detecting obscured hazards and improving safety and efficiency.
Patent Information
- Application Number
- PCT/EP2025/058750
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-01
- Filing Date
- 2025-03-31
- Publication Date
- 2025-10-09
AI Technical Summary
Existing vehicular automation systems (VASs) rely heavily on line-of-sight sensors, which are limited by blind spots and obscured hazards, leading to safety risks and suboptimal situational awareness, while acoustic sensors are underutilized due to lack of reliability and precision in processing acoustic signals for safety-critical applications.
An Acoustic Intelligence System (AIS) that integrates acoustic sensors with machine learning algorithms, utilizing a Human Auditory Perceptual Model (HAPM) and Vector Space Model to process acoustic data, incorporating contextual metadata to enhance situational awareness by mimicking human auditory perception, thereby improving the detection of hazards beyond the line of sight.
The AIS enhances the safety and accuracy of VASs by effectively processing acoustic data to detect obscured hazards, reducing reliance on line-of-sight sensors and improving situational awareness, while being energy-efficient and cost-effective.
Smart Images

Figure EP2025058750_09102025_PF_FP_ABST
Abstract
Description
[0001] Systems and Methods for Contextual Scene Analysis
[0002] Field:
[0003] The present application is directed towards providing systems and methods for generating Acoustic Intelligence which is suitable for use with Vehicular Automation Systems (VASs). As such, the systems and methods discussed in this disclosure are preferably for use with a vehicle to control, or to assist in the control, of the vehicle. Such systems are referred to in this disclosure as VASs.
[0004] In the present disclosure, an Acoustic Intelligence System (AIS) is a system configured to receive and process acoustic information and provide an output to a VAS which is suitable for use by the VAS for the control of a vehicle.
[0005] Background:
[0006] Vehicular automation has emerged as an area of great interest. It is clear that the use, deployment, and integration of systems supporting vehicular automation will continue to increase. Indeed, such systems are becoming increasingly important and there is a continuing need to increase their accuracy and speed to improve their safety.
[0007] At least one factor in a VAS that safely reduces the need for human intervention lies in the accuracy of the sensors employed. In this disclosure, ‘sensors' refers to a sensor, sensor system, or sensor suite used to provide sensory data to a VAS. These sensors serve as crucial "perception" technologies, enabling a VAS to perceive surroundings and navigate their environment with minimal or no human oversight and intervention.
[0008] At present VASs predominantly rely on line-of-sight sensors which include visual, optical, and electromagnetic sensor suites, such as cameras, LiDAR, or radar. However, line-of sight sensors are restricted to detecting line-of-sight data. As a result, such sensors can only detect hazards in their direct visual path. Thus, the data gathered by line-of-sight sensors is inherently restricted by, and vulnerable to, blind spots. In particular, a line-of-sight sensor cannot obtain data from around a corner or behind an obstacle. Consequently, a VAS should not solely use the data provided by such sensor suites as it is insufficient to safely complete accurate navigation tasks. This is especially the case in environments shared with humans and other moving artificial systems. Thus, this over-dependence on line-of-sight sensors and data has resulted in VASs that operate without critical data. This over-dependence can have serious consequences for manufacturers, developers, and end-users. In particular, there is an ongoing risk with VASs now in use that hazards obscured by obstacles, blind spots, or adverse weather or lighting conditions will remain undetected. Such sensory limitations significantly hinder VASs reaching the level of situational awareness and perception achieved by humans. As used in this disclosure, the term ‘situational awareness’ refers to the ability of a VAS to accurately determine the presence of hazards, and the risk of collision with such hazards in the vicinity of the vehicle controlled by the VAS.
[0009] In particular, VASs at present rely on unimodal sensors (i.e. line-of-sight) whereas humans assess their surroundings based on their natural multimodal sensing capacity (i.e. line-of-sight combined with non-line-of-sight). Thus, the present disclosure is directed towards improving VASs through improving multimodal sensor integration with VASs.
[0010] Acoustic sensors can be coupled to VASs to improve their situational awareness. Through the detection and analysis of acoustic data, the detection of events or hazards beyond the line of sight can be improved.
[0011] Similar to other sensor systems, the output from the acoustic sensor system may be coupled to one or more machine learning algorithms to analyse the sensor data and extract additional situational data from the sensor data. The one or more machine learning algorithms are preferably configured to run on one or more edge-based GPUs (Graphics Processing Units) or TPUs (Tensor Processing Units). As used in this disclosure the term ‘edge-based’ means that a component is located in or on the vehicle and not located remote from the vehicle.
[0012] Additionally, and as will be apparent to those skilled in the art, various processor hardware components, such as e.g. Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), etc. may also be used. The manufacturing costs can significantly vary based on the exact selection of hardware components and the way they are integrated. However, the expenses related to powering acoustic sensor systems have been found to be consistent — and substantially lower in comparison to line-of-sight sensor systems.
[0013] In addition to manufacturing costs of producing a sensor system, there are further costs associated with running a sensor system. For example, the more data produced by a sensor system, the more memory is required for data management, and the more processing is required to distill the data acquired into a usable output, which in turn increases the power consumption of the sensor system. A further issue with line-of sight sensors is that they generate a large amount of data. As the cost of manufacturing line-of-sight sensor systems such as cameras, radars, and even LiDAR decreases, the significance of the costs of processing, storing, and transmitting the extensive data they generate, especially in the case of LiDAR, become more pronounced. In ideal conditions, the costs associated with data management may be warranted. However, VASs typically operate in less than optimal environments. As a result, the effectiveness of visual sensor systems — and consequently, their value — significantly declines. For example, a camera's range of visibility can easily drop from 50 meters on a motorway to just 5 meters on a rural road, but without a corresponding reduction in the volume of data produced.
[0014] Acoustic sensor systems advantageously have lower initial hardware costs than their line-of- sight counterparts. Furthermore, acoustic sensor systems inherently have lower energy demands than line-of-sight sensors such as electromagnetic sensors. Thus, by integrating energy-and data- efficient acoustic sensing, industries involved in machine perception and vehicle automation systems can achieve substantial benefits. Furthermore, the use of acoustic sensor systems increases the safety of VASs.
[0015] Although acoustic sensing systems offer these notable benefits, their adoption remains limited. At least one reason for this is the lack of reliability and precision of the existing systems and methods configured to analyse acoustic signals for VASs. This is particularly a problem for safety- critical applications.
[0016] At present, there are a primarily two related, and sometimes interconnected, techniques used to extract information from audio signals using machine learning:
[0017] 1. Acoustic Event Detection (AED): AED is a technique for identifying the presence or absence of specific entities through their associated sound signatures. For example, in the field of bioacoustics, detecting the distinct call of a cuckoo may signal the presence of this bird species within a given locale. The analysis can be further refined to categorise multiple entities simultaneously. This may be done by employing a predefined set of labels to distinguish between detected sounds (e.g., identifying the sounds of a cuckoo, crow, and sparrow, while noting the absence of swallows, owls, and pigeons).
[0018] 2. Acoustic Scene Classification (ASC): ACS is a technique for categorising an entire acoustic environment based on the collective sound signatures present. For example, a soundscape comprising an assortment of sounds from birds, children playing, the hum of a lawnmower, and distant traffic noise may be categorised by a machine learning tool as originating from an urban park setting.
[0019] Typically, acoustic sensing (AED and ASC) is conducted asynchronously. This means that a pre-stored audio recording is processed after it has been recorded. Occasionally, these processes are executed in real-time. However, such real-time applications typically do not fall within safety- critical scenarios. At present acoustic sensing is applied to scenarios where the functionality required is relatively basic. For example, focusing on straightforward and very discernible tasks like detecting the sound of breaking glass in an otherwise quiet room, or to detecting specific anomalies such as in acoustic-emission testing, where the sensor's sole purpose is to monitor a particular sound characteristic indicative of equipment wear.
[0020] At present an acoustic signal is typically processed by a machine-learning system as follows:
[0021] 1. Data Collection: The audio data is obtained using an audio sensor such as a microphone or one or more acoustic MEMs.
[0022] 2. Pre-processing: The collected audio data is pre-processed to improve the quality and usability of the data for detection and classification. For example, the pre-processing applied to the collected audio data may include one or more of noise reduction, normalisation, or filtering to remove background noise.
[0023] 3. Feature Extraction: In this step, relevant features are extracted from the pre- processed audio signals (the terms ‘feature’, ‘audio feature’ and ‘acoustic feature’ have been used interchangeably in the present disclosure). Relevant features may include one or more of timedomain features (e.g., zero-crossing rate, energy), frequency-domain features (e.g., spectral centroid, spectral flux), or others like Mel-frequency cepstral coefficients (MFCCs).
[0024] 4. Feature Selection: In this step, one or more features extracted are selected. The selected features are used to classify events, so preferably the process used to select an extracted feature is configured to select the most useful features for classifying events. This step can involve techniques like principal component analysis (PCA) to reduce the dimensionality of the feature set while retaining the most critical information for classification.
[0025] 5. Event Detection and / or Classification: In this step, a trained machine-learning model is used to detect signal traits in a feature that match labels mapped in the model’s training phase. Any suitable known machine-learning model known in the art may be used. For example, models derived from neural networks (especially convolutional neural networks CNNs), support vector machines (SVM), decision trees, and k-nearest neighbours (KNN) are typically used.
[0026] 6. Continuous Learning: Preferably, the machine learning system is updated with new data gained from its performance in real-world conditions. This can involve retraining models with new audio samples, such that the machine learning system is adapted to changing environments or that the machine learning system is adapted to recognise new categories of acoustic events.
[0027] As set out above, the systems which at present are used in acoustic sensing primarily focus on signal processing based on feature extraction. This results in a "signal-centric," feature analysis which maps individual features to individual categories which results in sub-optimal accuracy.
[0028] Object:
[0029] As such, there is a need for improved sensor systems that can significantly help to increase the safety of both users of autonomous systems and passersby. In addition, there is a need to address the limitations of existing sensing solutions, particularly in overcoming the constraints of line-of-sight sensors, in order to mitigate risks associated with blind spots and obscured hazards.
[0030] Summary
[0031] The present disclosure is directed to a method of audio analysis suitable for providing data to a vehicular automation system, VAS, the method comprising: obtaining a context of a VAS; obtaining one or more audio data streams from one or more audio sensors; extracting an audio feature from the one or more audio data streams; processing the audio feature with a neural network and a vector space model to determine a processing score of the audio feature, wherein: the processing of the audio feature is based on the processing score; the vector space model comprises a Human Audio Perception Model, HAPM; the HAPM is configured to mimic how an audio signal is perceived and processed by a human; and the method comprises using the HAPM to adjust the processing score of the feature based on the context of the VAS.
[0032] Preferably, the method further comprises obtaining line-of-sight or contextual data from the VAS and using the HAPM to adjust the processing score of the feature based on the data. The contextual data includes information such as weather conditions, VAS location, or the like.
[0033] Preferably, the HAPM comprises an Acoustic Signal Physics Model, ASPM and the processing score comprises a relevance score, the relevance score indicative of whether a feature is a foreground or background feature; and the method comprises using the ASPM to adjust the relevance score of a feature based on a determination of how the feature propagated from an audio source to the one or audio sensors based and the context of the VAS.
[0034] Preferably, the HAPM comprises an Ecological Auditory Dimension Model, EADM and the processing score comprises a relevance score, the relevance score indicative of whether a features is a foreground or background feature; the EADM is configured to mimic how a human perceives a sound; and the method comprises using the EADM to adjust the relevance score of a feature based on the context of the VAS.
[0035] Preferably, the HAPM comprises an Auditory Relevance Attention Model, ARAM and the processing score comprises a prioritisation score, the prioritisation score indicative of how urgently a feature needs to be processed; the ARAM is configured to mimic human attention-based tuning; and the method comprises using the ARAM to adjust the prioritisation score of a feature based on the context of the VAS.
[0036] Preferably, the HAPM comprises a Schema Auditory Experiential Model, SAEM and the processing score comprises a trend score, the trend score (e.g. a probability score) indicative of how unexpected a feature is based on previously processed features; the SAEM is configured to mimic how a human determines the context of a feature based on their experience; and the method comprises using the SAEM to adjust the trend score of a feature based on the context of the VAS.
[0037] Preferably, the method further comprises selecting a contextual template, wherein selecting the contextual template is based on the context of the VAS and the contextual template includes an expected soundscape for the context; comparing the feature to the features in the expected soundscape; and increasing the trend score of the feature if the feature does not match a feature in the expected soundscape.
[0038] Preferably, the method further comprises adjusting the contextual template based on the feature. Preferably, the method further comprises generating an alert if the processing score or the trend score of the feature is above a predetermined threshold.
[0039] Preferably, the method further comprises generating the vector space model. More preferably, generating the vector space model comprises obtaining alignment metadata for data entry to correlate data entries either by time or location.
[0040] Preferably, the alignment metadata comprises a timestamp for a data entry, or a geolocation for a data entry. Preferably, using the HARM to adjust the relevance, prioritisation, or trend score of a feature based on the context of the VAS comprising using the HARM to provide one or more cross- modal attention weights to a vector, wherein the cross-modal weighting mimics the cross-modal attention of a human.
[0041] Preferably, the neural network comprises a Bayesian Neural Network. More preferably, the HARM is processed using the Bayesian Neural Network, and the remainder of the vector space model is processed using a conventional neural network.
[0042] The present disclosure is also directed to a data processing system comprising means configured to perform any of the method steps set out above.
[0043] The present disclosure is also directed to a computer program comprising instructions which, when the program is executed by a data processing system, cause the data processing system to perform any of the method steps set out above.
[0044] The present disclosure is also directed towards a computer readable storage medium, or a computer readable data carrier, comprising instructions which, when the program is executed by a data processing system, cause the data processing system to perform any of the method steps set out above.
[0045] Brief Description:
[0046] The invention will be more clearly understood from the following description of an embodiment thereof, given by way of example only, with reference to the accompanying drawings, in which:-
[0047] Figure 1 and figure 2 are flow charts showing the operation of a method according to the present disclosure;
[0048] Figure 3 shows a system overview of a system in accordance with the present disclosure;
[0049] Figure 4 shows an implementation of a traditional neural network; and
[0050] Figure 5 shows an implementation of a hybrid neural network.
[0051] Detailed Description:
[0052] This present disclosure is directed towards systems and methods for an Acoustic Intelligence System (AIS). The AIS is designed to maintain contextual nuances in the relationships between acoustic features, which are lost in prior art solutions. As a result, metadata from the surrounding environment is maintained, which in turn increases the accuracy of an AIS of the present disclosure over existing solutions.
[0053] Thus, the use of an AIS results in improved acoustic data processing and improved contextual information. In particular, contextual metadata is used to determine how acoustic sensor information is processed. This provides a more relevant interpretation of environmental sounds for a VAS.
[0054] A Vector Space Model (VSM) 201 is preferably used to manage data. The method preferably also uses a Human Auditory Perceptual Model (HARM) 101 to process data.
[0055] In contrast to prior art systems, a system according to the present disclosure 400 is configured to use contextual metadata 102 to determine a contextual setting. Contextual metadata 102 is data that provides information from the incoming data 103 received from a sensor suite of a VAS 300 and other relevant non-VAS sources 102. In particular, the contextual metadata 102 provides information about the context of a VAS. As used in the disclosure, the term context refers to the background information and relevant details that surround and describe a dataset. For example, contextual metadata 102 for a VAS could include data indicative of: weather conditions; VAS geolocation; traffic or traffic trends; environmental topologies; road surface conditions; user behaviour-models; VAS states; pre-established operational definitions as conferred by operational safety standards; etc. Contextual metadata 102 can be used in determining the relevance of sensor readings.
[0056] As shown in figure 1, in a first step incoming data 103 is captured. In conventional systems line-of-sight data 104 is typically captured using, LiDAR, radar, cameras, optical sensors, infrared sensors, etc. In this disclosure, audio data 105 is also captured 106 using an audio sensor (for example, using one or more microphones and / or acoustic MEMS). The audio data is processed as described below in more detail and stored in HDF 150.
[0057] Contextual metadata 102 is also included in the data. In the context of the present application, metadata is data that provides information about other data, but not the content of the data itself. In particular, metadata refers to situational metadata which is data that describes the conditions under which the real-time data is acquired. For example, the contextual metadata 102 could include data indicative of the time of day, geo-location, weather conditions, road conditions, light levels, velocity, etc. Contextual metadata 102 can be extracted or inferred from data collected by existing line-of-site sensors, local weather stations, mapped topology, and predefined VAS operational definitions presently used in the art.
[0058] The AIS stores non-acoustic contextual data (which may be data, metadata or both) describing the operational environment and conditions (i.e. the context) under which the AIS operates. Preferably, both static and dynamic contextual data 102 is stored. Preferably storage is in the same HDF 150 as that used for the audio data 103. This data is used to improve the processing of acoustic data analysed, increasing the accuracy of the output of the AIS provided to a VAS.
[0059] In this disclosure, static contextual data is data or metadata about a specific environment that remains unchanged over time or that undergoes infrequent changes at extended intervals. As such, static data may include data relating to an environment’s predefined Operational Design Domain (ODD), infrastructure topology, network services, broader sensor ecosystem, etc. Thus, static contextual data is contextual data that remains unchanged for more than a predetermined period (e.g. 12 hours or more). Examples of static contextual data include data about road works, data about 5G small cells coming online, ODD regulation changes (such as changes to legal regulations), etc. In this disclosure, dynamic contextual data is data or meta data about a given environment that changes rapidly or regularly over time. Thus, dynamic contextual data is contextual data that remains unchanged for less than a predetermined period. For example, dynamic data may include data about vehicle telemetry, weather conditions, time-of-day, the behaviour of other agents / entities in the environment, etc. As such dynamic contextual data is used to monitor changes such as another agent / entity making a sudden swerve (fractions of a second), fog slowly descending on a junction (minutes), or daylight slowly shifting from dusk to night (hours). Indeed, dynamic contextual data may change in under a second and in a timeframe in the region of nanoseconds.
[0060] For example, if an autonomous vehicle is navigating a rural road, the static data that might be stored is a description file of the rural road infrastructure. The static data may be stored in text format, which is preferably Terse Resource Description Framework Triple Language format (Turtle) such as, for example, ASAM OpenXOntology. An example syntax depicting the static rural road infrastructure is as follows:
[0061] <Roads>
[0062] <RuralRoad>
[0063] <Attributes>
[0064] <roadType>Enum ( e . g . , single_lane , asphalt , unsealed) < / roadType> <condition>Enum (e.g. , wet, dry, icy )< / condi tion>
[0065] <width>Numeric (meters ) < / width>
[0066] < / Attributes>
[0067] <Relationships>
[0068] <connectsTo>Enum (RuralRoad. Describes intersections or connections)
[0069] < / connectsTo>
[0070] <parallelTo Optional="true">Enum (for roads running parallel within a certain distance) < / parallelTo>
[0071] < / Relationships>
[0072] <Elements>
[0073] <LaneMarking>
[0074] <type>Enum (e.g. , none, dashed, solid) < / type>
[0075] <color>Enum (e.g. , white, yellow) < / color>
[0076] < / LaneMarking>
[0077] <RoadsideFeature>
[0078] <type>Enum (e.g. , ditch, fence, vegetation, utilityPole) < / type>
[0079] <distanceFromRoad>Numeric (meters) < / distanceFromRoad>
[0080] < / Roads i de Feature>
[0081] < raf f icControl>
[0082] <type>Enum (e.g. , stopsign, yieldsign, wildlif eCrossing) < / type>
[0083] <location>Descriptive location or GPS coordinates< / location>
[0084] < / Traf f icControl>
[0085] < / Elements>
[0086] <Intersections>
[0087] <UncontrolledIntersection>
[0088] <Attributes>
[0089] <visibility>Enum (e.g. , obstructed, unobstructed) < / visibility>
[0090] < / Attributes>
[0091] <Relationships>
[0092] <crosses>RuralRoad< / crosses> <!-- roads that intersect ->
[0093] < / Relationships>
[0094] < / UncontrolledIntersection>
[0095] < / Intersections>
[0096] <SurroundingEnvironment>
[0097] <Natural Features >
[0098] <type>Enum (e . g . , forest, river, hill, agriculturalField) < / type>
[0099] <proximityToRoad>Numeric (meters ) < / proximityToRoad> < / NaturalFeatures>
[0100] <Wildlif eActivity>
[0101] <type>Enum ( e . g . , deer , livestock) < / type>
[0102] <activityLevel>Enum ( e . g . , high, medium, low) < / activityLevel>
[0103] < / Wildlif eActivity>
[0104] < / SurroundingEnvironment>
[0105] < / RuralRoad>
[0106] < / Roads>
[0107] As will be apparent to those skilled in the art, the above example may be expanded or modified to include additional details relevant to a specific rural road scenario, such as bridges, specific types of roadside foliage, or other unique features. The text format can be extensively customised to accurately model the complexities of real-world structures.
[0108] Preferably, the descriptor information is translated and stored for use as in a hierarchical data format. A Hierarchical Data Format (HDF) is a set of file formats (such as e.g. HDF4 or HDF5) designed to store and organize large amounts of data. HDF files are particularly suitable for storing data models that can represent very complex data objects and a wide variety of metadata.
[0109] By way of example, the descriptor information may be translated and stored using a HDF framework is as follows (where *|* signifies a level of nesting for a given entry):
[0110] / RuralRoadlnf restructure (Root Group)
[0111] I - Attributes :
[0112] As those skilled in the art will appreciate, additional groups, data, and metadata can be added following the same structure as that set out above. The additional groups, data, and metadata can be used to describe other elements like bridges, tunnels, signage, and more detailed environmental data. HDF can be implemented using a variety of methods. For example, using the h5py package, which is a Pythonic interface to the HDF5 binary data format: import h5py import numpy as np from datetime import datetime
[0113] # Define custom data types for GPS coordinates and timestamps def gps_dtype ( ) : return np.dtype( [
[0114] ( 'latitude' , np.float64) , ( 'longitude' , np.float64) , ] ) def timestamp_dtype ( ) : return h5py. string_dtype (encoding= ' ascii ' , length=20) # ISO 8601 format
[0115] # Initialise data types gps_dt = gps_dtype ( ) ts_dt = timestamp_dtype ( )
[0116] # Example of a Roads group with a specific road subgroup roads_group = rri . create_group ( ' Roads ' ) road_001 = roads_group . create_group ( ' RuralRoad_001 ' ) road_001. attrs [ ' roadType ' ] = 'asphalt' road_001. attrs [ ' condition' ] = 'wet' road_001. attrs [ ' width ' ] = 5.5 road_001. attrs [ ' gps_coordinates ' ] = np . array ( ( 51.7075 , -8.5306) , dtype=gps_dt) road_001. attrs [ ' creation_timestamp ' ] = current_iso_timestamp ( )
[0117] # LaneMarkings dataset lane_markings_type = np.dtype( [
[0118] ( ' type ' , ' S10 ' ) ,
[0119] ( ' color ' , ' S10 ' ) ,
[0120] ( 'timestamp' , ts_dt) ,
[0121] ( ' gps_latitude ' , np.float64) ,
[0122] ( ' gps_longitude ' , np.float64) , ] ) lane_markings_data = np. array ( [
[0123] (b'solid' , b'white' , current_iso_timestamp ( ) , 51.7075, -8.5306) , (b'dashed' , b'yellow' , current_iso_timestamp ( ) , 51.7075, -8.5306) , ] , dtype=lane_markings_type) road_001. create_dataset ( ' LaneMar kings ' , data=lane_markings_data)
[0124] # Trafficcontrols dataset traf f ic_controls_type = np.dtype ( [
[0125] ( ' type ' , ' S20 ' ) ,
[0126] ( 'location' , ts_dt) ,
[0127] ( ' gps_latitude ' , np.float64) ,
[0128] ( ' gps_longitude ' , np.float64) ,
[0129] ( 'timestamp' , ts_dt) , ] ) traf f ic_controls_data = np. array ( [
[0130] (b ' stopSign ' , current_iso_timestamp ( ) , 51.7075, -8.5306, current_iso_timestamp ( ) ) , ] , dtype=traffic_controls_type) road_001. create_dataset ( ' rafficcontrols ' , data=traf f ic_controls_data)
[0131] # RoadsideFeatures dataset roadside_f eatures_type = np.dtype( [
[0132] ( ' type ' , ' S20 ' ) ,
[0133] ( ' distanceFromRoad ' , np . float 64 ) ,
[0134] ( ' gps_latitude ' , np.float64) ,
[0135] ( ' gps_longitude ' , np.float64) ,
[0136] ( 'timestamp' , ts_dt) , ] ) roadside_f eatures_data = np.array( [
[0137] (b ' f armGate ' , 2.0, 51.7075, -8.5306, current_iso_timestamp ( ) ) , ] , dtype=roadside_features_type) road_001. create_dataset ( ' RoadsideFeatures ' , data=roadside_f eatures_data)
[0138] # Intersections group and datasets intersections = rri . create_group ( ' Intersections ' ) ui_001 = intersections . create_group ( ' Uncontrolledlntersection_001 ' ) ui_001. attrs [ ' visibility ' ] = 'obstructed' ui_001. attrs [ ' gps_coordinates ' ] = np . array ( ( 51.7076, -8.5307) , dtype=gps_dt) ui_001. attrs [ ' creation_timestamp ' ] = current_iso_timestamp ( )
[0139] # SurroundingEnvironment group se = rri . create_group ( ' SurroundingEnvironment ' ) wildlif e_activity_type = np.dtype( [
[0140] ( ' type ' , ' S20 ' ) ,
[0141] ( ' activityLevel ' , ' S10 ' ) ,
[0142] ( 'timestamp' , ts_dt) ,
[0143] ( ' gps_latitude ' , np.float64) ,
[0144] ( ' gps_longitude ' , np.float64) ,
[0145] ] ) wildlif e_activity_data = np.array( [
[0146] (b'fox' , b'high' , current_iso_timestamp ( ) , 51.7077, -8.5308) ,
[0147] (b ' livestock ' , b 'medium' , current_iso_timestamp ( ) , 51.7078, -8.5309) ,
[0148] ] , dtype=wildlif e_activity_type) se . create_dataset ( 'Wildlif eActivity ' , data=wildlif e_activity_data)
[0149] # Note that this code is only a foundational example, and not all road elements
[0150] # have been written out here.
[0151] By way of a further example, a HDF 150 entry for Dynamic Contextual Data may include weather data. As such, the HDF 150 could be implemented as follows (using the scenario of an autonomous vehicle navigating a road for ease of interpretation):
[0152] / WeatherConditions (Root Group)
[0153] I - Attributes : I I I |- Collection method: "Method used for collecting precipitation data
[0154] I I I |- Collection method: "Method used for measuring temperature"
[0155] I I I |- Pressure: "Atmospheric pressure data over time"
[0156] I I I - Attributes :
[0157] I I I |- Collection method: "Method used for measuring atmospheric pressure"
[0158] I I I |- Data resolution: "Time resolution of observation"
[0159] I I I |- Observation method: "Method used for determining cloud cover"
[0160] I I I | - Data resolution : "Time resolution of data collection"
[0161] I I I | - Collection method : "Method used for measuring visibility"
[0162] Similarly, the HDF 150 may be implemented using h5py.
[0163] After capture 106, an analogue audio signal is converted into a digital signal 107. The analogue to digital conversion 107 can be performed using standard sampling methods or hardware such as an analogue to digital converter (ADC). Preferably, the digital signal can then be passed through an initial processing step 108 to improve the quality of the signal. For example, noise reduction, normalization, or both can be performed on the digital signal. The initial processing step
[0164] 108 may be performed using known digital signal processing (DSP) systems and methods.
[0165] Preferably, audio features in the digital audio signal are extracted 109. Audio features are quantifiable aspects of an audio signal at a given point in time. These can be time-domain or frequency domain aspects. For example, an audio feature can be one or more of: Zero Cross Rate, Energy, Entropy of Energy, Spectral Centroid, Spectral Spread, Spectral Entropy, Spectral Flux, Spectral Roll off, Mel-frequency cepstral coefficients (MFCC), Chroma Vector, or Chroma Deviation. Feature extraction can be performed using any suitable systems or methods known in the art.
[0166] Once features have been extracted 109 from the digital audio signal, they are preferably also stored for use in a hierarchical data format 150. As such the data representing the audio features
[0167] 109 needs to be converted into HDF. This can be done using known Python libraries, audio-specific libraries (such as librosa), and libraries such as h5py. For example, if MFCC was used for feature extraction, the following HDF data may be extracted:
[0168] I | -- / Analysi sinformation ( Dataset or Attributes )
[0169] In addition, the data further includes Human Audio Perception Models (HAPMs) 101. According to the present disclosure, a HAPM is a model that represents how an audio signal is perceived and processed by a human. In particular, a HAPM is used to obtain weightings for a vector space model (VSM). These weightings define data interactions - both synergistic and antagonistic. Preferably, the HAPMs comprises four sub-models: an Acoustic Signal Physics Model (ASPM) 110, an Ecological Auditory Dimension Model (EADM) 111, an Auditory Relevance Attention Model (ARAM) 112, and a Schema Auditory Experiential Model (SAEM) 113.
[0170] Data about HAPMs 101 is also stored in HDF 150. For example, data about HAPMs 101 may be stored in the following HDF Groups:
[0171] 1. Model Description Group: The model description group stores explanatory data about a sub-model, including its theoretical background, parameters descriptions, and versions.
[0172] 2. Parameters Group: The parameters group is used to store parameters. These parameters are based on an analysis of how a human perceives sound and provide data which defines how the human auditory system processes a sound. For example, the parameters may include human auditory thresholds, reaction times, sensitivity indices, or other model-specific parameters. These parameters are related to the Results Group and Analysis Group and the parameters are used to adjust the weights in a sub-model.
[0173] 3. Stimuli Group: The Stimuli Group stores data about inputs to a sub-model that may prompt a change in the behaviour of the sub-model or generate an output from the sub-model. The Stimuli group can contain sub-groups. For example, the Stimuli Group can comprise one or more of an Audio Data Group, a Static Context Data Group, or a Dynamic Contextual Data group.
[0174] The Audio Data Group stores data about acoustic information. This may include one or more of raw audio data, pre-processed audio data, or data about extracted features. The Static Contextual Data (SCD) Group stores static contextual data (for example, data about road infrastructure features). The Dynamic Contextual Data (DCD) Group stores dynamic contextual data (for example data about weather features).
[0175] 4. Results Group: The results group is used to store data about simulations or experimental runs of a sub-model. This data can be used during training. This data can include phases, including perceptual responses, cognitive assessments, error rates, or any other metrics of interest. The results group is preferably organised into subgroups based on different experimental conditions or simulation scenarios.
[0176] 5. Analysis Group: The analysis group is used to store data obtained from sub-model training. For example, this data may include processed data output by a sub-model, statistical data, or visualisation data. The statistical data may include data summarised statistics, data indicative of comparison across conditions, and data about inferential statistical tests.
[0177] An example structure of the HAPM is as follows:
[0178] The ASPM 110 is a set of rules describing the physical interactions between an acoustic wave and the context of a VAS (such as e.g. its environment), thereby providing a model of how a sound wave will propagate from the source of the sound to an audio sensor. As these interactions significantly affect the characteristics of an acoustic signal reaching an audio sensor, the ASPM 110 is useful for determining the significance of a received sound feature. As such the output of data processed using the ASPM 110 will be based on correlations between the elements represented by non-acoustic data and acoustic data, due to the use of the non-acoustic data in determining context. The below table presents some examples of the types of interactions and their effects on the acoustic signal:
[0179] Preferably, the ASPM 110 is also stored in HDF 150. For example, an ASPM 110 may be stored as follows:
[0180] I | - Attributes :
[0181] I I I - Types of reflecting surfaces, measurement or simulation parameters
[0182] I I I - Temperature, pressure, humidity gradients causing refraction
[0183] I I I - "Datasets for diffracted sound levels, obstacle dimensions, wavelength."
[0184] I I I - Types of scattering particles or surfaces, freguency dependence
[0185] I I I - Moving source or observer details, direction of movement
[0186] I I I - Relative humidity profiles, freguency range affected
[0187] I I I - Atmospheric pressure conditions, altitudinal variations As a result, data processed using the ASPM 110 produces an output which accounts for the influence of physical context on the audio features. For example, occlusions may dampen signal amplitude, reducing the loudness of sounds as they reach the sensor, which has subsequent consequences for the other acoustic dimensions. Similarly, weather conditions - such as humidity, temperature, and wind - can affect the timbral qualities of sound, thereby altering its features. The ASPM 110 thus integrates these environmental and contextual factors from the VSM into the output provided by the AIS to the VSM.
[0188] The EADM 111 is a set of rules describing how a human perceives sound in a given context, thereby providing a model of the biases in human auditory perception. It is important to note that these biases arose in human hearing through evolution to help us survive through helping us make sense of complex acoustic information (such as e.g. the cocktail problem). The below table presents examples of some of the traits observed in ecological psychoacoustic studies.
[0189] Preferably, the EADM 111 is also stored in HDF 150. For example, an EADM 111 may be stored in
[0190] HDF 150 as set out below:
[0191] I I I - Relevant contexts: "Situations where temporal resolution is critical"
[0192] I l l i - Experiment details: "Description of experiments on pitch perception" I l l i - Conditions tested : "Different sound levels and their effects on pitch constancy"
[0193] As such, audio data processed using the EADM 111 by the AIS results in an output from the AIS that better mimics human-like auditory processing. Thus, the AIS facilitates identifying sounds which are significant for operational safety and efficiency. By better emphasising potentially important sounds over others in a crowded or complex auditory environment, the AIS will more effectively detect potential hazards.
[0194] The ARAM 112 is a set of rules describing how a human focuses on acoustic information as a function of how relevant the acoustic information is in a given context, thereby providing a model of attention-based tuning in human auditory perception. Thus, the ARAM 112 is based on the understanding that not all sounds are treated equally in auditory perception. Instead, processing is dynamically allocated to sound sources based on their perceived relevance to a current context.
[0195] The below table presents some of these traits observed in auditory attention research.
[0196] Preferably, the ARAM 112 is also stored in HDF 150. For example, an ARAM 112 may be stored in HDF 150 as follows:
[0197] I |- Description: "Overview of the auditory attention model and dataset scope"
[0198] I | - Attributes :
[0199] I I I - Parameters: "Defining the tasks, contexts, and conditions under which selective attention is measured"
[0200] I I I - Temporal structures: "Information on the temporal structure of stimuli"
[0201] I I I - Task performance metrics: "Metrics used to assess performance"
[0202] I I I - Analysi s detail s : "Detailed descriptions of analyses and key findings"
[0203] As such, the ARAM 112 in use performs a filtering process on the captured acoustic features. This filtering process is influenced by a combination of factors, including for example the properties of an acoustic feature (e.g„ volume, pitch, novelty), contextual expectations, or the current state of the VAS. As a result, a VAS using an ARAM 112 is configured to prioritise only the most relevant acoustic signals within the scene as identified by the EADM 111. This capability improves the situational awareness of a VAS equipped with ARAM 112. Further, the use of an ARAM 112 improves responsiveness and safety of a VAS.
[0204] In human hearing, auditory attention to a sound in a context is influenced by how a human perceives the sound in the context. Thus, there is a relationship between the output from the EADM 111 and the ARAM 112.
[0205] For example, processing a set of audio features with the EADM 111 will result in the audio features with increasing frequency or amplitude being emphasised over those with a decreasing frequency or amplitude. According to the ARAM 112, audio features with an increasing frequency or amplitude are sounds which are typically more important to process. In particular, sounds with an increasing frequency or amplitude often signal urgency or alarm (e.g., a car approaching). Thus, processing a set of audio features with the ARAM 112 will result in processing capacity being dynamically allocated to audio features with increasing frequency or amplitude. As such, the ARAM 112 will mimic an adaptive trait of humans that enhances the ability to respond to potential threats or important events in a given context, and preferably, the EADM 111 will assist in this process.
[0206] As a further example, if the AIS periodically receives a set of audio features at predetermined time intervals, processing these sets of audio features with the EADM 111 will result in related time-separated features being integrated - i.e. features relating to a siren will be grouped together. If these features are then processed by using the ARAM 112, the ARAM 112 will filter out background noise or intermittent distractions, focusing instead on coherent auditory stream. As such, by integrating features the EADM 111 improves the function of the ARAM 112 by assisting in the identification of sounds that are not background noise.
[0207] As a further example, a set of features processed with the EADM 111 will emphasise the amplitude of certain features based on their context. As a result, if the resultant features are to be processed by the ARAM 112, the emphasised features are more likely to be processed. For example, a relatively quiet sound might stand out and be processed in a silent context, whereas the sound might be ignored in a noisy context. Contextual effects help prioritise which sounds are relevant based on the current auditory context.
[0208] 4. EADM: Masking Patterns
[0209] Masking can significantly impact what sounds are able to be attended to. Sounds that are masked by others, particularly in noisy environments, may not be perceived at all, directing attention only to sounds that are clear and unobscured. Understanding masking patterns helps in designing technologies that improve auditory attention in challenging listening situations.
[0210] As noted above, processing audio with the EADM 111 will result in related audio features being grouped together. This in turn will result in auditory stream segregation, where audio features related to one audio source will be separated from audio features related to another source. This facilitates the ARAM 112 in filtering the audio features based on relevance. In particular, by processing the resultant features provided by the EADM 111, the ARAM 112 is better able to process the audio features from an audio stream considered relevant while ignoring features from an audio stream which is considered not relevant.
[0211] If a plurality of audio sensors is used to obtain the audio data, the processing of the extracted audio features from these audio sensors with the EADM 111 will result in the provision of directional data. When processing the resultant features, the ARAM 112 can use the directional data to better process audio features originating from an area of interest and avoid the processing of sounds originating from areas which are not relevant. As a result, the ARAM 112 can mimic a human’s directional hearing, which allows individuals to localise and focus on sound sources accurately and to focus attention spatially, to help better navigate environments and for spatially oriented tasks.
[0212] Further, by processing features using the EADM 111, the certain features will be enhanced by the precedence effect, which will further assist the ARAM 112 to process spatially relevant audio features.
[0213] The SAEM 113 are a set of rules which describe how to process sensory stimuli and contexts through the use of contextual templates. As such the SAEM 113 models how humans process and understand sensory stimuli and contexts through the use of mental templates or schemas. These stored contextual templates are preferably pre-recorded contexts which are stored in memory. A contextual template provides an expected soundscape for a given context.
[0214] In use, the SAEM 113 is configured to match audio features to a contextual template. As such, the SAEM 113 mimics how a human’s brain, when encountering sounds, will attempt to match incoming auditory stimuli with existing schemas to interpret and understand the auditory stimuli more accurately and more efficiently.
[0215] As a result, processing audio features with the SAEM 113 provides a contextual template which most closely matches a context in which audio features were previously recorded. The below table describes some features of the SAEM 113:
[0216] Preferably, the SAEM 113 can be stored in HDF 150. For example, the SAEM 113 may be stored in HDF 150 as follows:
[0217] I I I - Schema model access rates: "Influence on access rates to schema models" As noted above, contextual metadata 102 (such as time of day, location, weather data, VAS status, etc.) feeds into an AIS 400 in accordance with the present disclosure. This integration allows the system to determine a probability of an expected acoustic feature in a given context even when not yet present. For instance, if the SAEM 113 has selected a contextual template indicative of an urban context, the sound of traffic, distant sirens, and human babble on sidewalks would be expected. As a result, the AIS 400 will not be impacted by the sound of a distant siren or of humans. In contrast, if the SAEM 113 has selected a contextual template indicative of a motorway context, the relevance of a human generated sound (such as shouting) would be considered far more relevant and would impact on AIS 400 operation. Further, in a contextual template indicative of an urban context, if the sound of breaking glass is suddenly detected, the AIS 400 is notified of this anomaly, whereupon it is categorised as an unexpected (i.e. unusual or low probability) acoustic object and subsequently providing an output that preferably forces VAS action. Through this anomaly categorisation process, the SAEM 113 mimics a human’s continual experiential learning process. As such, an unexpected (i.e. a low probability) audio feature in a given context is logged so that the AIS’s 113 contextual templates adapt and evolve.
[0218] By incorporating the use of a SAEM 113, the AIS 400 better utilises non-acoustic context and system history to predict and identify the relevance of audio features more accurately. Such an approach significantly improves the situational awareness and decision-making utility of the data provided from an AIS 400 to VASs, making them more adaptable and efficient in dynamically changing environments over time.
[0219] Preferably, the processing of audio features using the SAEM 113 affects the processing of audio features using the ARAM 112. In particular, audio features are filtered based on contextual templates. This advantageously mimics human audio processing which is not solely dictated by audio sensory input received but is shaped by expectations, experiences, and predictions.
[0220] For example, the contextual templates of the SAEM 113 contain accumulated system history and (as such) predictive data about the audio features which are to be expected in a given context. This may include the type of audio sources, and their relevance to the context in question. This predictive data can be used as part of the processing of audio features using the ARAM 112 to filter incoming audio features that deviate from these predictions for processing, thus facilitating anomaly detection and increasing efficiency. Similarly, predictive data can be used as part of the processing of audio features using the ARAM 112 to filter an incoming audio feature such that the feature is ignored as background noise typical of a particular context if the audio feature fits well into a detailed contextual template associated with that context and the audio feature is not associated with a hazard or dangerous event.
[0221] Processing audio features with SAEM 113 contextual templates provides top-down processing which facilitates for rapid changes in operating contexts based on changing VAS data provided to an AIS 400 - such VAS data may be indicative of a user’s operation goals or environmental cues. For instance, by providing VAS data as input to the AIS 400, the VAS data may be processed in conjunction with incoming audio features by the SAEM 113 to change from one contextual template to another. This would result in a change in how the ARAM filters audio features, as the filtering is based on the contextual template. As such, the significance of an audio feature shifts if that audio feature now signifies a potential hazard in a newly selected contextual template. For example, moving from an urban area with enclosed recreational areas into a residential area with no such facilities, the sound of a bouncing ball will increase in significance given its greater relevance in a context associated with children playing in close proximity to a road without the safety of a fenced off court.
[0222] Thus, by processing audio features using an EADM 111, ARAM 112, and SAEM 113 modules, the processing of audio features can mimic the non-linear complexities of human perception.
[0223] As a result, an AIS 400 according to the present disclosure will focus audio processing on audio features that are novel or contextually relevant. This top-down processing reduces the processing load, allowing an AIS to provide audio navigational data for complex auditory environments more quickly and efficiently.
[0224] This further facilitates in comparing the audio features expected in a given context with those received. In turn, this facilitates the detection of errors or novelties that might indicate important environmental changes (which might require a change of context) or new information (which might require the update of a contextual template).
[0225] Further, by using the SAEM 113 to process audio features, a contextual template may indicate specific optimal responses. As such, the AIS 400 will quickly identify audio features associated with danger or other significant states. This includes the speed at which a VAS equipped with an AIS 400 of the present disclosure can respond to hazards like a car crash in the path of a vehicle controlled by a VAS. Preferably, the data, metadata, rules, and HARM parameters for the AIS 400 are vectorised and stored in a vector space model 201 for use with a neural network. Any suitable known means for vectorisation may be used. Prior to vectorisation, data files may be split into smaller pieces called chunks by the chunking algorithm. Any suitable known chunking algorithm may be used. These chunks can then be embedded into the AIS 400 using a predetermined embedding model. Preferably, semantic relationships are retained. For example, given that the primary data types for AIS 400 are audio (acoustic data) and text (scene / event descriptors), embedding models such as OpenL3 for audio and BERT for text data may be used. The use of OpenL3 for audio and BERT is preferable, as they can be configured to retain semantic relationships.
[0226] Preferably, the AIS 400 preprocess data prior to vectorisation to include metadata useful for improving the results obtained from the vectorised data. Preferably, this metadata includes alignment metadata which can be used to correlate data during vectorisation. The alignment metadata could be used to correlate data either by time or location. For example, the AIS may timestamp or geolocate data (or both) from SCD, DCD, and audio features. Timestamping data ensures the efficient incorporation and maintenance of temporal relationships in the data being processed by the AIS. Geolocating data ensures the efficient incorporation and maintenance of spatial relationships in the data being processed by the AIS 400. Thus, the AIS 400 may provide a unified data reference for machine learning analysis. For example, preprocessing and vectorisation may include:
[0227] As noted above, the AIS 400 vectorises HARM 101 parameter data. The HARM 101 parameter data is also pre-processed prior to vectorisation. Vectorising may involve several steps, for example:
[0228] Preferably, audio data, contextual data 102 (such as SCD, DCD, etc.), and HARM data are vectorised. As this data is cross-modal (i.e. it includes data from more than one type or ‘mode’ of data sensor (e.g. audio data, optical, line-of-sight, etc.), the vector database used to store the vectorised data preferably comprises some means of performing data fusion. Data fusion is the process of integrating multiple data sources to produce more consistent, accurate, and useful information than that provided by any individual data source.
[0229] In order to perform data fusion, differences between the significance of different modes of data in different contexts should be accounted for. In particular, a vectorised data entry is stored in transformer layers that comprise one or more cross-modal attention weights. This cross-model weighting is indicative of the importance of data from one mode of sensor over another. As such, the cross-modal weighting mimics the cross-modal attention of a human, which refers to the distribution of attention to different senses (where attention is the cognitive process of selectively emphasizing and ignoring sensory stimuli).
[0230] Cross-modal weighting models exist for text and image vector fusion, such as ViLBERT, as well as text and video vector fusion, such as Audio-Visual Scene-Aware Dialog (AVSD). The AIS 400 of the present disclosure offers a means for acoustic and scene-descriptor data fusion in the vector space 201 by utilising the HARM 101.
[0231] The machine learning (ML) may be performed using either Classical ML 400 (which uses one or more Neural Networks) or Hybrid ML 500 (which uses a combination of one or more Neural Networks 530 and one or more Bayesian Neural Networks 550).
[0232] Both Classical ML 400 and Hybrid ML 500 use standard vector distance calculations to determine initial relationships between data, such as Euclidean distance, dot product, and cosine distance. However, Hybrid ML is configured to use Bayesian Neural Network (BNN) 50 uncertainty with the HARM 101.
[0233] Classical ML 400 may use any suitable known neural network (NN), such as Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short Term Memory (LSTM), etc. Hybrid ML 500 combines a conventional Neural Network (NN) 530 for Acoustic Data, SCD, and DCD analytics, while using a Bayesian Neural Network (BNN) 550 for HARM 101 modelling. This is advantageous as one or more BNNs 550 are particularly suitable for handling uncertainty and complexity, especially in human perception modelling, and conventional NNs 530 can be used for efficiently processing and classifying large volumes of sensory and contextual data.
[0234] A method of Hybrid ML 500 in accordance with the present disclosure may include the following steps:
[0235] 1. Pre-processing and Vectorisation a. Temporal and Spatial Referencing: Incoming data, whether acoustic, SCD, or DCD, is tagged with temporal or spatial references (and preferably both) using timestamps, GNSS, or both. b. Vectorisation: Next, the data is transformed into a numerical format. This can be done using any suitable process, such as e.g. one-hot encoding for categorical data and normalisation for continuous data. Embedding models for text, such as BERT, and for acoustic data, such as OpenL3, may also be used. Vectorised data is then stored in a suitable structured format (e.g., matrices or tensors) compatible with ML models. c. Vector Fusion: fusion of the vector entries is implemented, using a cross-modal weightings derived from the HARM 101. d. Distance Measures: Initial vector distances are calculated based on the vector entries. e. Contextual Parameter Reduction: An initial contextual filtering of the incoming data is performed by adjusting trained model weights. The adjustment of trained model weights is based on the temporal and spatial embeddings of incoming data.
[0236] 2. Initial Neural Network Processing a. Conventional Neural Network Processing: The vectorised data is processed using conventional NNs. This processing is configured to process the acoustic data with reference to the SCD and DCD to extract features and patterns from the raw input.
[0237] 3. Bayesian Neural Network (BNN) Processing a. ASPM and SAEM processing: The output from the conventional NNs is provided to a first pretrained BNN layer 210. The first pretrained BNN layer 210 is trained such that it is tuned to the ASPM 110 and SAEM 113. The use of the ASPM and SAEM results in the first pretrained BNN layer generating Multi-Domain correlations 215, providing a context to the acoustic signals that is based on physical and previously recorded correlations. Preferably, contextual time-stamped and geolocated information (e.g. SCD and DCD) are used to adjust the output from the first pretrained BNN layer. b. EADM and SAEM processing 220: The output 215 from the first pretrained BNN layer 210 is provided to a second pretrained BNN layer 220. The second pretrained BNN 220 is trained such that it is tuned to the EADM and SAEM. The use of the EADM and SAEM results in the BNN producing Non-Linear Ecological Acoustic Correlations 225. This step mimics a human’s ecological dimension modelling, thereby increasing context. c. ARAM and SAEM Prioritisation 230: The output 225 from the second pretrained BNN layer 220 is provided to a third and a fourth pretrained BNN layer. The third pretrained BNN layer is trained such that it is tuned to the ARAM. The fourth pretrained BNN layer is trained such that it is tuned to the SAEM. As a result, the output from the third pretrained BNN layer is sorted into a Hyperfocus Foreground Stack for critical data and a Peripheral Background Stack 235 for secondary non-critical data, mimicking how humans prioritise auditory streams via human auditory attention. d. Contextual and Anomaly Analysis 240: Using the fourth pretrained BNN layer 240, the AIS is configured to determine if an audio feature is present in a contextual template representing the ‘expected soundscape’ (accounting for a current operational / environmental context), or if the audio feature is an anomaly. This is used to improve the accuracy in determining the relevance or urgency of the audio feature.
[0238] 7. VAS Alerts 250: a. If an audio feature is present in a contextual template representing the ‘expected soundscape’, it is provided to the AIS 400 as input data for normal processing. b. If an audio feature is an anomaly, it triggers an immediate alert 255 to the AIS 400. The AIS 400 is configured to perform additional processing (such as e.g. classifications, regression, and predictive analysis) to determine a response. Preferably, this is performed at the edge of the network, enabling real-time action.
[0239] 8. Acoustic Scene Descriptor Generation 260: a. As noted above, if an anomaly is identified, this can be used to refine or generate a contextual template. In particular, the AIS generates acoustic context descriptors associated with the anomaly. The acoustic context descriptor is sent to one or more remote (i.e. cloud-based) servers. The one or more servers can be processed via any suitable asynchronous, offline, and / or distributed means. Thus, an anomaly can be used to further increase the accuracy and predictive capabilities of the AIS over time.
[0240] As such the Hybrid ML model is both Trained and Adapted. Preferably, the Hybrid ML is continuously trained using feedback from real-world operations and simulations. This training involves adjusting the weights and biases within both the NNs and BNNs, refining the models based on new data and insights. Preferably, through the AIS 400 iteratively generating and improving contextual templates, the Hybrid ML is adapted to changing environments and evolving data patterns. This ensures the AIS remains effective and accurate in its auditory perception modelling and decision-making processes over time.
Claims
Claims1. A method of audio analysis suitable for providing data to a vehicular automation system, VAS, the method comprising: obtaining a context of a VAS; obtaining one or more audio data streams from one or more audio sensors; extracting an audio feature from the one or more audio data streams; processing the audio feature with a neural network and a vector space model to determine a processing score of the audio feature, and providing an output to the VAS based on the processing score, wherein: the vector space model comprises a Human Audio Perception Model, HARM; the HARM is configured to mimic how an audio signal is perceived and processed by a human; and the method comprises using the HARM to adjust the processing score of the feature based on the context of the VAS.
2. The method of claim 1 comprising obtaining line-of-sight data or contextual data from the VAS and using the HARM to adjust the processing score of the feature based on the data.
3. The method of claim 1 or 2, wherein: the HARM comprises an Acoustic Signal Physics Model, ASPM and the processing score comprises a relevance score, the relevance score indicative of whether a feature is a foreground or background feature; and the method comprises using the ASPM to adjust the relevance score of a feature based on a determination of how the feature propagated from an audio source to the one or audio sensors based and the context of the VAS.
4. The method of any preceding claim, wherein: the HARM comprises an Ecological Auditory Dimension Model, EADM and the processing score comprises a relevance score, the relevance score indicative of whether a feature is a foreground or background feature; the EADM is configured to mimic how a human perceives a sound; andthe method comprises using the EADM to adjust the relevance score of a feature based on the context of the VAS.
5. The method of any preceding claim, wherein: the HARM comprises an Auditory Relevance Attention Model, ARAM and the processing score comprises a prioritisation score, the prioritisation score indicative of how urgently a feature needs to be processed; the ARAM is configured to mimic human attention-based tuning; and the method comprises using the ARAM to adjust the prioritisation score of a feature based on the context of the VAS.
6. The method of any preceding claim, wherein: the HARM comprises a Schema Auditory Experiential Model, SAEM and the processing score comprises a trend score, the trend score indicative of the probability of a feature is based on previously processed features; the SAEM is configured to mimic how a human determines the relevance of a feature based on their experience; and the method comprises using the SAEM to adjust the trend score of a feature based on the context of the VAS.
7. The method of claim 6 comprising: selecting a contextual template, wherein selecting the contextual template is based on the context of the VAS and the contextual template includes an expected soundscape for the context; comparing the feature to the features in the expected soundscape; and increasing the trend score of the feature is the feature does not match a feature in the expected soundscape.
8. The method of claim 7, comprising adjusting the contextual template based on the feature.
9. The method of any preceding claim, comprising generating an alert if the processing score or the trend score of the feature is above a predetermined threshold.
10. The method of any preceding claim, comprising generating the vector space model.
11. The method of claim 10, wherein generating the vector space model comprises obtaining alignment metadata for data entry to correlate data entries either by time or location.
12. The method of claim 11, wherein the alignment metadata comprises a timestamp for a data entry, or a geolocation for a data entry.
13. The method of any one of claims 10 to 12, wherein using the HARM to adjust the relevance of a feature based on the context of the VAS comprising using the HARM to provide one or more cross-modal attention weights to a vector, wherein the cross-modal weighting mimics the cross- modal attention of a human.
14. The method of any preceding claim, wherein the neural network comprises a Bayesian Neural Network.
15. The method of claim 14, wherein the HARM is processed using the Bayesian Neural Network, and the remainder of the vector space model is processed using a conventional neural network.
16. A data processing system comprising means configured to perform the steps of any preceding claim.
17. A computer program comprising instructions which, when the program is executed by a data processing system, cause the data processing system to perform the steps of any one of claims 1 to 15.
18. A computer readable storage medium, or a computer readable data carrier, comprising instructions which, when the program is executed by a data processing system, cause the data processing system to perform the steps of any one of claims 1 to 15.
Citation Information
Patent Citations
Emergency response vehicle detection for autonomous driving applications
US11816987B2