Voice recognition method and system based on neural network

By constructing a neural network-based sound recognition method, the time-frequency feature group of real-time audio streams is obtained, the sound network architecture is constructed, the acoustic feature clusters are identified, the adaptive mapping network is generated, the sound source propagation path is analyzed, the frequency domain coupling coefficient is calculated, the signal-to-noise ratio index is extracted, and the robust recognition level is determined, which solves the problem of insufficient recognition accuracy in complex environments of traditional methods, and achieves higher recognition stability and accuracy.

CN120452436AActive Publication Date: 2025-08-08XIAN FULIYE MICROELECTRONICS CO LTD

Patent Information

Application Number
CN202510911087.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-08
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Traditional sound recognition methods lack recognition accuracy in complex environments, especially in noise interference or mixed scenes of multiple sound sources.

Method used

The sound recognition method based on neural network is adopted, by obtaining the time-frequency feature groups of real-time audio streams, building a sound network architecture, identifying acoustic feature clusters, generating an adaptive mapping network, analyzing the sound source propagation path, calculating the frequency domain coupling coefficient, extracting the signal-to-noise ratio index, determining the robust recognition level, reconstructing the sound source acquisition dimension, and formulating an identification plan.

Benefits of technology

It improves the sound recognition accuracy in complex environments and improves the recognition stability and accuracy in noise and multi-sound source scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452436A_ABST
    Figure CN120452436A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio processing, and discloses a voice recognition method and system based on a neural network, and the method comprises the steps: firstly obtaining a real-time audio stream of a target scene, extracting a time-frequency feature group, analyzing a pulse coding sequence, and constructing a voice network architecture; identifying an acoustic feature cluster by using the acoustic feature cluster, matching a preset framework hierarchical topological relation, and optimizing node weight to generate an adaptive mapping network; based on this, analyzing a sound source propagation path, identifying a multipath effect factor, calculating a frequency domain coupling coefficient, and performing hierarchical fusion to obtain a mixed feature tensor; querying a tensor adaptive response trajectory, extracting a phase distortion feature and a signal-to-noise ratio index, and calculating a signal-to-noise attenuation entropy; and finally, determining a robust recognition level, reconstructing a sound source acquisition dimension and formulating a sound recognition scheme. According to the invention, the voice recognition precision in a complex environment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sound recognition method and system based on a neural network, belonging to the technical field of audio processing. Background Art

[0002] Sound recognition refers to the technology of analyzing and processing sound signals through technical means to identify the content, category or source of the sound. It is widely used in speech recognition (such as speech-to-text), environmental sound classification (such as gunshot and alarm detection), speaker recognition (such as voiceprint authentication) and music information retrieval (such as song recognition) and other fields.

[0003] Currently, traditional sound recognition methods mainly rely on feature extraction algorithms such as Mel-Frequency Cepstral Coefficients (MFCCs), combined with Hidden Markov Models (HMMs) or Gaussian Mixture Models (GMMs) for classification. However, these methods have limited ability to extract sound features in complex environments, especially in noise interference or multi-source mixed scenarios, where recognition accuracy decreases significantly. Therefore, a sound recognition method based on neural networks is needed to improve sound recognition accuracy in complex environments. Summary of the Invention

[0004] The present invention provides a sound recognition method and system based on a neural network, the main purpose of which is to improve the accuracy of sound recognition in complex environments.

[0005] To achieve the above object, the present invention provides a sound recognition method based on a neural network, comprising: Acquire a real-time audio stream in a target sound source scene, extract a time-frequency feature group from the real-time audio stream, parse a pulse code sequence corresponding to the time-frequency feature group, and construct a sound network architecture corresponding to the real-time audio stream based on the pulse code sequence; Identifying acoustic feature clusters in the real-time audio stream using the sound network architecture, matching hierarchical topological relationships in a preset neural computing framework based on the acoustic feature clusters, optimizing relationship node weights corresponding to the hierarchical topological relationships, and generating an adaptive mapping network corresponding to the relationship node weights; Based on the adaptive mapping network, analyzing the sound source propagation path corresponding to the real-time audio stream, identifying the multipath effect factors in the sound source propagation path, calculating the frequency domain coupling coefficients between the multipath effect factors, and performing hierarchical fusion on the frequency domain coupling coefficients to obtain a hybrid feature tensor; querying an adaptive response trajectory of the mixed feature tensor in the target sound source scene, extracting a phase distortion feature and a signal-to-noise ratio index from the adaptive response trajectory, and calculating a signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index; Based on the signal-to-noise attenuation entropy, the robust recognition level of the real-time audio stream is determined; according to the robust recognition level, the sound source acquisition dimension corresponding to the real-time audio stream is reconstructed; and based on the sound source acquisition dimension, a sound recognition scheme corresponding to the target sound source scene is formulated.

[0006] Optionally, constructing a sound network architecture corresponding to the real-time audio stream according to the pulse code sequence includes: Analyzing frequency domain eigenvectors in the pulse code sequence; Extracting an audio frame set corresponding to the real-time audio stream; Performing time-frequency alignment on the frequency domain feature vector and the audio frame set to generate an audio alignment feature map; Identifying audio node relationships in the audio alignment feature graph; According to the audio node relationship, a sound network architecture corresponding to the real-time audio stream is constructed.

[0007] Optionally, generating the adaptive mapping network corresponding to the relationship node weights includes: Parsing the topological association code in the relationship node weights; Based on the topology association code, a preset weight mapping rule library is matched to obtain an initial mapping path set; Dynamically normalizing the initial mapping path set to obtain a normalized path sequence; removing conflicting sequence segments from the normalized path sequence to obtain optimized mapping sequence segments; Based on the optimized mapping sequence segments, an adaptive mapping network corresponding to the relationship node weights is generated.

[0008] Optionally, analyzing the sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network includes: Performing time framing on the real-time audio stream to obtain audio frame groups; Generating a time-frequency feature matrix corresponding to the audio frame group; Extracting the sound source propagation vector from the time-frequency feature matrix; Inputting the sound source propagation vector into a path analysis module in the adaptive mapping network to output a path feature sequence; Analyzing the sound source azimuth and propagation delay in the path characteristic sequence; Based on the sound source azimuth and the propagation delay, a sound source propagation path corresponding to the real-time audio stream is constructed.

[0009] Optionally, calculating the frequency domain coupling coefficient between the multipath effect factors includes: Calculate the frequency domain coupling coefficient between the multipath effect factors.

[0010] Optionally, performing layered fusion on the frequency domain coupling coefficients to obtain a hybrid feature tensor includes: Querying a frequency domain feature subset corresponding to the frequency domain coupling coefficient; Extracting local feature vectors of each frequency band of the frequency domain feature subset; Analyzing the distribution characteristics of the local feature vector in the multidimensional space; Analyzing the dimensional distribution rules corresponding to the distribution characteristics; Based on the dimensional distribution rule, the frequency domain coupling coefficients are hierarchically fused to obtain a hybrid feature tensor.

[0011] Optionally, querying the adaptive response trajectory of the mixed feature tensor in the target sound source scene includes: Analyzing the tensor distribution interval corresponding to the mixed feature tensor; Matching dynamic voiceprint segments in the target sound source scene according to the time-frequency domain distribution interval, and extracting sparse feature clusters in the voiceprint segments; Mapping the sparse feature cluster to the sound field propagation path corresponding to the target sound source scene; generating an energy attenuation gradient map corresponding to the sound field propagation path; The adaptive response trajectory in the energy decay gradient map is queried.

[0012] Optionally, the calculating the signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index includes: Calculate the signal-to-noise attenuation entropy corresponding to the real-time audio stream.

[0013] Optionally, determining the robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy includes: Analyzing the noise interference characteristics corresponding to the signal-to-noise attenuation entropy; generating frequency domain segmented data corresponding to the real-time audio stream based on the noise interference characteristics; Extracting valid segmented audio corresponding to the frequency domain segmented data; Determining, based on the valid segmented audio, a hierarchical identifier corresponding to the real-time audio stream; Based on the classification identifier, a robust recognition level corresponding to the real-time audio stream is determined.

[0014] In order to solve the above problems, the present invention further provides a sound recognition system based on a neural network, the system comprising: An architecture construction module is used to obtain a real-time audio stream in a target sound source scene, extract a time-frequency feature group from the real-time audio stream, parse a pulse code sequence corresponding to the time-frequency feature group, and construct a sound network architecture corresponding to the real-time audio stream based on the pulse code sequence; A network generation module is configured to identify acoustic feature clusters in the real-time audio stream using the sound network architecture, match hierarchical topological relationships in a preset neural computing framework based on the acoustic feature clusters, optimize relationship node weights corresponding to the hierarchical topological relationships, and generate an adaptive mapping network corresponding to the relationship node weights; a hierarchical fusion module, configured to analyze, based on the adaptive mapping network, a sound source propagation path corresponding to the real-time audio stream, identify multipath effect factors in the sound source propagation path, calculate frequency domain coupling coefficients between the multipath effect factors, and perform hierarchical fusion on the frequency domain coupling coefficients to obtain a hybrid feature tensor; an attenuation entropy calculation module, configured to query an adaptive response trajectory of the mixed feature tensor in the target sound source scene, extract a phase distortion feature and a signal-to-noise ratio index from the adaptive response trajectory, and calculate a signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index; A solution formulation module is used to determine the robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, reconstruct the sound source acquisition dimension corresponding to the real-time audio stream according to the robust recognition level, and formulate a sound recognition solution corresponding to the target sound source scene based on the sound source acquisition dimension.

[0015] Compared with the problems described in the background technology, the present invention obtains the real-time audio stream in the target sound source scene and extracts the time-frequency feature group in the real-time audio stream, which can jointly analyze the audio signal from the time domain and frequency domain dimensions, accurately capture the frequency components and time series change rules of the sound, and effectively improve the accuracy of sound recognition in complex environments. The present invention uses the sound network architecture to identify the acoustic feature clusters in the real-time audio stream, and can automatically extract feature sets with semantic associations (such as phonemes in speech, frequency combinations in environmental sounds) from complex audio signals through the hierarchical nonlinear transformation of the neural network, effectively improving the representation ability of sound features and the generalization performance of the recognition system. Furthermore, the present invention analyzes the sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network, and can use the network to dynamically associate features. Modeling capabilities can accurately analyze multipath effects such as reflection and diffraction of sound in space (such as indoor reverberation paths or outdoor obstacle reflection paths), thereby improving the accuracy of sound source positioning and identification in complex environments. Furthermore, the present invention can capture the model's dynamic processing process of complex acoustic features such as multipath effects in the scene in real time by querying the adaptive response trajectory of the mixed feature tensor in the target sound source scene, which can improve its stability and accuracy in tasks such as target sound source identification and positioning in specific scenarios. Finally, the present invention determines the robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, which can intuitively reflect the degree of audio interference and signal quality with quantitative indicators, provide a clear reference for the recognition system, and improve the system's effective recognition capability of audio in complex environments, ensuring the stability and accuracy of recognition. Therefore, the neural network-based sound recognition method and system provided in the embodiment of the present invention can improve the accuracy of sound recognition in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A flowchart of a neural network-based voice recognition method according to an embodiment of the present invention is provided; Figure 2 A schematic diagram of a sound network architecture in a sound recognition method based on a neural network provided by one embodiment of the present invention; Figure 3 A schematic diagram of modules for implementing a neural network-based voice recognition system according to an embodiment of the present invention.

[0017] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0018] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0019] The embodiments of the present application provide a neural network-based voice recognition method. The execution subject of the neural network-based voice recognition method includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiments of the present application. In other words, the neural network-based voice recognition method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. Example 1:

[0020] Reference Figure 1 FIG. 1 is a flow chart of a neural network-based voice recognition method according to an embodiment of the present invention. In this embodiment, the neural network-based voice recognition method includes: S1. Obtain a real-time audio stream in a target sound source scene, extract a time-frequency feature group from the real-time audio stream, parse a pulse code sequence corresponding to the time-frequency feature group, and construct a sound network architecture corresponding to the real-time audio stream based on the pulse code sequence.

[0021] By acquiring the real-time audio stream in the target sound source scene and extracting the time-frequency feature group in the real-time audio stream, the present invention can jointly analyze the audio signal from the time domain and frequency domain dimensions, accurately capture the frequency components and temporal variation patterns of the sound, and effectively improve the accuracy of sound recognition in complex environments.

[0022] Among them, the target sound source scene refers to a specific environment or scene where sound recognition is required, including the target sound source and the surrounding acoustic environment (such as noise, reverberation, multi-source interference, etc.). For example, when monitoring abnormal sounds in a shopping mall (such as glass breaking) in an intelligent security scenario, the spatial structure of the shopping mall, the noise of the crowd, etc. constitute the target sound source scene; when detecting machine tool equipment failures in an industrial scenario, the mechanical vibration noise and environmental background sound in the workshop also fall into this category; the real-time audio stream refers to a continuous sound signal sequence obtained in real time by an acquisition device such as a microphone, which is transmitted and processed in a digital form and has the characteristics of strong timeliness and dynamic changes. For example, the user voice command stream received in real time by a voice interaction device (such as a smart speaker) or the environmental sound stream continuously collected by a security monitoring system are both real-time audio streams and need to be analyzed in time to achieve a rapid response; the time-frequency feature group refers to a feature set extracted after time-frequency analysis (such as short-time Fourier transform) of the real-time audio stream. The fusion of time domain (time series) and frequency domain (frequency component) information can characterize the dynamic spectral characteristics of sound. For example, the time-frequency feature group of a speech signal can reflect the frequency change trajectory of a phoneme (such as the movement of the formant of a vowel), and the time-frequency feature group of an ambient sound can reflect the frequency distribution and duration characteristics of noise (such as the high-frequency pulse and time domain duration of a car horn). Optionally, the acquisition of a real-time audio stream in the target sound source scene can be achieved through an audio acquisition device and a stream processing framework, such as: collecting the original audio signal through a microphone array, combining it with the GStreamer tool for real-time streaming and packaging, and ultimately obtaining a standardized real-time audio stream; the extraction of the time-frequency feature group from the real-time audio stream can be achieved through a digital signal processing algorithm and a feature extraction library, such as: using a short-time Fourier transform or a Mel-frequency cepstral coefficient algorithm, with the help of the Librosa library for spectrum analysis and feature calculation, and ultimately obtaining a time-frequency feature group containing dimensions such as energy and frequency domain distribution.

[0023] Furthermore, by parsing the pulse code sequence corresponding to the time-frequency feature group, the present invention can effectively compress audio data and retain key time-frequency information, thereby reducing storage and transmission costs. At the same time, the discrete characteristics of the pulse code sequence facilitate subsequent pattern recognition and anomaly detection, thereby improving the efficiency of acoustic event analysis.

[0024] Among them, the pulse code sequence refers to the conversion of the time-frequency feature group into a discrete pulse signal combination, and binary coding is usually used to represent the energy intensity of different frequency bands. For example, if a time-frequency feature group contains three frequency bands: low frequency (0-500Hz), medium frequency (500-2kHz), and high frequency (above 2kHz), its pulse code can be expressed as [1,0,1], which means that the low-frequency and high-frequency energy are significant at the current moment, and the medium frequency is weak. Optionally, the analysis of the pulse code sequence corresponding to the time-frequency feature group can be achieved through signal coding algorithms and feature quantization tools, such as: using the linear predictive coding (LPC) algorithm, combined with Python's scipy.signal library to perform segmented quantization and binary conversion of time-frequency features, and finally obtaining a pulse code sequence that characterizes the energy distribution of different frequency bands.

[0025] Furthermore, the present invention constructs a sound network architecture corresponding to the real-time audio stream based on the pulse coding sequence, which can simulate the signal transmission mechanism of biological neural networks and convert audio features into discrete codes of neural-like pulses, so that the network architecture can dynamically adapt to the acoustic feature distribution of different scenarios and enhance the network's hierarchical representation capabilities for complex sound signals (such as multi-source aliasing and noise interference).

[0026] Among them, the sound network architecture refers to a neural network structure built based on the relationship between audio nodes, which simulates the hierarchical association of sound features. It usually includes an input layer, a hidden layer (such as a convolutional layer, a recurrent layer) and an output layer. The nodes correspond to feature processing units, and the connecting edges correspond to feature association weights. For example, the sound network architecture used for environmental sound classification may extract local frequency patterns in the audio alignment feature map through the convolutional layer, and then integrate cross-time features through the fully connected layer, and finally output classification results such as gunshots and alarm sounds.

[0027] As an embodiment of the present invention, constructing the sound network architecture corresponding to the real-time audio stream based on the pulse code sequence includes: analyzing the frequency domain feature vectors in the pulse code sequence; extracting the audio frame set corresponding to the real-time audio stream; performing time-frequency alignment on the frequency domain feature vectors and the audio frame set to generate an audio alignment feature graph; identifying the audio node relationships in the audio alignment feature graph; and constructing the sound network architecture corresponding to the real-time audio stream based on the audio node relationships.

[0028] Among them, the frequency domain feature vector refers to a one-dimensional numerical vector obtained after performing frequency domain analysis on the pulse code sequence, which characterizes the energy distribution or frequency characteristics of the sound signal in different frequency components. For example, the frequency domain feature vector of bird singing can highlight the characteristics of energy concentration in the high-frequency band (such as the peak of 2-8kHz), while the frequency domain feature vector of the car engine sound may show a strong energy distribution in the medium and low frequency bands (such as 500Hz-2kHz), which is used to distinguish the frequency characteristics of different sound sources; the audio frame set refers to a continuous short-time segment sequence divided into a real-time audio stream according to a fixed time window (such as every 10ms), each segment is called a frame, which contains the sound signal details within the time period. For example, a 1-second speech audio can be divided into 100 audio frames, each frame corresponds to a 10ms speech segment (such as the pronunciation of the word "ni" in "hello" can be split into multiple continuous audio frames, reflecting the transition from initial consonant to final vowel). transition process); the audio alignment feature map refers to a two-dimensional image generated by associating the frequency domain feature vector with the audio frame set through the time-frequency alignment operation. The horizontal axis is time (audio frame number), the vertical axis is frequency, and the pixel value represents the feature intensity of the corresponding time-frequency point. For example, the audio alignment feature map of speech can present the resonance peak trajectory that changes with time (such as the low-frequency resonance peak position of the vowel / a / ), which is similar to a spectrogram and intuitively displays the time-frequency dynamic characteristics of the sound; the audio node relationship refers to the association relationship between different time-frequency nodes (pixel points) in the audio alignment feature map, such as the energy transfer of adjacent frequency components, the correlation of features in different time periods, etc. For example, in piano playing audio, there is a strong correlation between the same-frequency nodes of adjacent audio frames (the energy continuation of sustained notes), while the audio node relationship of drum beats can be expressed as a sudden association of short-term high-frequency nodes (the transient characteristics of the impulsive sound source).

[0029] Furthermore, the analysis of the frequency domain feature vectors in the pulse code sequence can be achieved through a spectrum analysis algorithm, such as: using discrete Fourier transform combined with the FFT function of the NumPy library to calculate the energy distribution of each frequency band, and finally obtaining a frequency domain feature vector containing amplitude and phase information; the extraction of the audio frame set corresponding to the real-time audio stream can be achieved through a sliding window segmentation technology, such as: using the frame function of the Librosa tool to perform frame processing with a 20ms window length and a 10ms step size, and finally obtaining a time-continuous audio frame set; the time-frequency alignment of the frequency domain feature vector and the audio frame set can be achieved through a dynamic time warping algorithm, For example, the optimal alignment path is calculated in Python's dtaidistance library based on the DTW algorithm, and finally a time-frequency synchronized audio alignment feature map is obtained; the identification of the audio node relationship in the audio alignment feature map can be achieved through a graph neural network model, such as using a GAT network combined with a PyTorch framework to learn the attention weights between nodes, and finally obtaining the audio node relationship that describes the correlation of acoustic events; the construction of the sound network architecture corresponding to the real-time audio stream can be achieved through a topology generation tool, such as using the NetworkX library to construct a weighted directed graph according to the node relationship, and finally obtaining a sound network architecture that reflects the sound source interaction logic.

[0030] Specifically, to further understand the construction process of the sound network architecture in this application, you can refer to Figure 2 The image shown in is a schematic diagram of the sound network architecture related to the present invention. It should be noted that, in the present invention, Figure 2 The architectural diagram presented in the figure is related to the sound network architecture construction technology described above. The above article elaborates on the complete process from obtaining the real-time audio stream in the target sound source scene, extracting the time-frequency feature group, parsing the pulse code sequence, to building the sound network architecture. The architectural diagram in the figure shows the process of generating a short-time spectrum from the scene sound through STFT (short-time Fourier transform), then generating the Mel energy spectrum through the Mel filter and windowing the fragments, using the CNN model for two-stage training, extracting features from the CNN middle layer for random forest training, and finally realizing the audio data stream recognition. This diagram is a visual presentation of the specific implementation steps and model application of the above sound network architecture construction technology. It is a concrete expression of the above theories and methods at the actual operation process and model architecture level, and is not limited to showing all the details and relationship analysis of the sound network architecture in different actual application scenarios.

[0031] In detail, the neural network model corresponding to the sound network architecture can be divided into three hierarchical structures: the response layer, the integrated operation layer, and the representation layer. Among them, the receivable sound audio range is controlled within 20Hz~20000Hz, with a change of 50Hz as an interval (except for the first interval, which is 20Hz~50Hz). A fixed number of neurons are allocated to each interval and activated in sequence.

[0032] The specific division of the response layer and the corresponding number of neurons are as follows:

[0033] The sound spectrum can be divided into 8 large intervals, each of which contains several sub-intervals, with a total of 7,480,000 neurons.

[0034] S2. Utilize the sound network architecture to identify acoustic feature clusters in the real-time audio stream, match the hierarchical topological relationships in a preset neural computing framework based on the acoustic feature clusters, optimize the relationship node weights corresponding to the hierarchical topological relationships, and generate an adaptive mapping network corresponding to the relationship node weights.

[0035] The present invention utilizes the sound network architecture to identify acoustic feature clusters in the real-time audio stream, and can automatically extract feature sets with semantic associations (such as phonemes in speech and frequency combinations in ambient sound) from complex audio signals through hierarchical nonlinear transformations of neural networks, effectively improving the representation ability of sound features and the generalization performance of the recognition system.

[0036] Among them, the acoustic feature cluster refers to a set of acoustic features with similar attributes or semantic associations extracted by the sound network architecture, which can characterize the essential attributes or category characteristics of the sound. For example, in speech recognition, the acoustic feature cluster of "vowel / a / " may include a combination of low-frequency resonance peaks (F1 is about 700Hz), medium-frequency energy distribution (such as F2 is about 1200Hz) and time domain duration. In environmental sound classification, the feature cluster of "car horn" may include high-frequency pulses (2-4kHz), short-time burst energy and specific frequency modulation patterns (such as frequency jump trajectories), which are key feature combinations for distinguishing different sound sources. Optionally, the identification of acoustic feature clusters in the real-time audio stream can be achieved by clustering analysis methods, such as: using the K-means algorithm in combination with the scikit-learn library to perform unsupervised clustering of MFCC features, and finally obtaining feature clusters characterizing different acoustic events.

[0037] Furthermore, the present invention matches the hierarchical topological relationship in the preset neural computing framework based on the acoustic feature cluster, and optimizes the relationship node weights corresponding to the hierarchical topological relationship, which can make the network structure deeply consistent with the hierarchical correlation of sound features (such as capturing basic frequency patterns at the bottom layer and integrating semantic features at the high layer), thereby improving the accuracy of feature mapping and the adaptation efficiency of the recognition model to complex scenarios.

[0038] Among them, the preset neural computing framework refers to a pre-designed neural network structure template, which defines the type of network layer, connection method and data flow direction, and is used to support the hierarchical calculation of sound features. For example, a "convolutional layer-pooling layer-recurrent layer-fully connected layer" framework can be preset, in which the convolutional layer extracts local time-frequency features and the recurrent layer models temporal dependencies, which is suitable for speech recognition tasks; or a "self-attention layer-Transformer encoder" framework can be preset to process the global feature association of long-term audio sequences; the hierarchical topological relationship refers to the hierarchical connection relationship between each network layer and node in the preset neural computing framework, reflecting the abstraction process of features from low to high levels. For example, in the CNN architecture, the bottom convolutional layer nodes capture the basic frequency units of audio (such as single-frequency sine waves), and the middle-layer nodes combine the low-level feature shapes. The network is connected to form a complex pattern (such as a formant structure), and high-level nodes integrate cross-layer features to generate semantic-level representations (such as the categories of "human voice" and "instrumental sound"). This hierarchical connection constitutes a topological structure from low dimension to high dimension. The relationship node weight refers to the strength parameter of the connection between nodes in the neural network, which determines the transmission efficiency of the feature signal and is optimized through training to adapt to the characteristics of the input data. For example, in a speech recognition network, if a convolutional layer node is sensitive to the "voiced onset point" feature, its connection weight with the next layer node will be enhanced during training, so that this feature is prioritized in the transmission; conversely, the connection weight of irrelevant noise features will be suppressed, thereby improving the representation ability of useful signals. Optionally, the matching of the hierarchical topological relationship in the preset neural computing framework can be achieved by a graph matching algorithm, such as: using a graph isomorphism network combined with the PyTorch Geometric library to align the hierarchical structure, and finally obtaining a hierarchical topological relationship containing a connection relationship; the optimization of the relationship node weights corresponding to the hierarchical topological relationship can be achieved by a backpropagation algorithm, such as: using the Adam optimizer in the TensorFlow framework to iteratively update the node parameters to ultimately obtain the optimized relationship node weight distribution.

[0039] Furthermore, the present invention can dynamically construct a feature mapping path based on the optimized node connection strength by generating an adaptive mapping network corresponding to the relationship node weights, so that the network can flexibly adjust the signal transmission method according to the real-time audio characteristics (such as strengthening the propagation link of the key acoustic feature cluster), thereby improving the feature decoupling capability of complex scenarios such as multi-source mixing and noise interference.

[0040] Among them, the adaptive mapping network refers to a neural network structure dynamically constructed based on the optimized mapping sequence segments, and its connection paths and weights can be adjusted in real time according to the input audio characteristics. For example, in a noisy environment, the network automatically enhances the weight of the "denoising path" (such as convolutional layer + denoising autoencoder) and suppresses the "interference path"; in a multi-sound source scenario, it activates the "sound source separation path" (such as recurrent layer + attention mechanism) to achieve adaptive feature processing for different scenarios and improve recognition robustness.

[0041] As an embodiment of the present invention, the generation of an adaptive mapping network corresponding to the relationship node weight includes: parsing the topological association code in the relationship node weight; based on the topological association code, matching a preset weight mapping rule library to obtain an initial mapping path set; dynamically normalizing the initial mapping path set to obtain a normalized path sequence; removing conflicting sequence segments in the normalized path sequence to obtain an optimized mapping sequence segment; and generating an adaptive mapping network corresponding to the relationship node weight based on the optimized mapping sequence segment.

[0042] Among them, the topological association coding refers to the digital encoding of the inter-layer connection relationship implied in the relationship node weight, which represents the dependency strength and transmission direction between nodes. For example, in a recurrent neural network (RNN), the topological association coding can record the weight values of the node at the current moment and the node at the historical moment, and express "strong connection" (such as weight > 0.8 marked as 1) or "weak connection" (such as weight < 0.2 marked as 0) in binary or vector form, which intuitively reflects the association pattern of time series features; the preset weight mapping rule base refers to a pre-defined set of rules for guiding the association between relationship node weights and feature mapping paths, including network layer connection strategies, weight allocation logic and path generation constraints. For example, the rule base may include "when a node weight > 0.7, preferentially connect to the attention mechanism layer", "the low-frequency feature path should contain at least 2 layers of convolution", and other rules; in speech emotion recognition, it is also possible to set a dedicated rule that "the emotion feature path must pass through the bidirectional LSTM layer" to ensure that the generated mapping path meets the specific task requirements and provide an explainable decision basis for adaptive network construction; the initial mapping path set refers to a preliminary feature mapping path set generated by matching the preset rule library based on the topological association coding, which contains multiple possible signal transmission paths. For example, in the image sound recognition scenario, the preset rule library may define rules such as "high-frequency features preferentially pass through convolution layer A→pooling layer B", "low-frequency features preferentially pass through convolution layer C→attention layer D", and the initial mapping path set contains all path combinations that meet the rules (such as A→B→fully connected layer, C→D→fully connected layer, etc.); the normalized path sequence refers to the ordered sequence obtained after weight normalization of each path in the initial mapping path set, so that the feature transmission strength of different paths is comparable. For example, if the sum of the node weights of two paths is 3.5 and 2.1 respectively, the normalized weights are The weight ratio is 5:3, which ensures that the network will not cause feature imbalance due to differences in the original weight scale when processing multiple paths in parallel, thereby improving the synergy between paths; the optimized mapping sequence segment refers to an efficient mapping sequence formed by removing conflicting or redundant paths from the normalized path sequence. For example, if two paths simultaneously transmit the same frequency feature but in opposite directions (such as one path enhancing high frequency and the other suppressing high frequency), they are determined to be conflicting segments and one of them is deleted; if multiple paths have duplicate functions (such as processing low-frequency background noise), the representative path is retained and the redundancy is removed, thereby reducing computational redundancy and improving feature mapping efficiency.

[0043] Furthermore, the parsing of the topological association codes in the relationship node weights can be achieved through a graph neural network embedding method, such as: using the GraphSAGE algorithm in combination with the DGL framework to extract the potential association features between nodes, and ultimately obtaining a topological association code that characterizes the connection strength; the matching of the preset weight mapping rule library can be achieved through a graph pattern matching technology, such as: using the Cypher query language of the Neo4j graph database to perform rule matching, and ultimately obtaining an initial mapping path set that meets the constraint conditions; the dynamic normalization of the initial mapping path set can be achieved through an adaptive normalization algorithm, such as: implementing dynamic range adjustment in PyTorch based on the Min-Max Scaling method, and ultimately obtaining a standardized normalized path sequence; the removal of conflicting sequence segments in the normalized path sequence can be achieved through a constraint satisfaction algorithm, such as: applying a CSP solver in combination with the OR-Tools library to detect and eliminate contradictory path segments, and ultimately obtaining a conflict-free optimized mapping sequence segment; the generation of an adaptive mapping network corresponding to the relationship node weights can be achieved through a self-organizing mapping method, such as: using a Kohonen network in combination with TensorFlow to achieve dynamic topology adjustment, and ultimately obtaining a mapping network structure with adaptive capabilities.

[0044] Specifically, if the volume of the sound received in the adaptive mapping network is too high, the activity of the neurons in the adaptive mapping network needs to be suppressed to adapt to the high-sound environment. The neurons that adapt to the high-sound environment are marked as alert source neurons. The total number of alert source neurons is set to 10,000. These 10,000 neurons will release the activity of the second-layer neurons when excited, affecting the influence threshold of the neurons. The specific parameters are set as follows: High volume decibels / dB Number of activated alert neurons The threshold coefficient / threshold of the second layer neurons [80~90) 2000 80% [90~100) 5000 90% ≥100 10000 95% When the adaptive mapping network encounters excessively loud sounds, suppressing neuronal activity is a key response strategy. In addition to the established alert source neurons, this also involves regulating the refractory period of the neurons. All neurons are set to a uniform 6-second refractory period; after continuous firing, they must wait 6 seconds before responding again. This prevents excessive fatigue or disruption of neurons due to continuous high-frequency stimulation, ensuring that the entire network can continue to operate smoothly even in high-pitched environments.

[0045] S3. Based on the adaptive mapping network, analyze the sound source propagation path corresponding to the real-time audio stream, identify the multipath effect factors in the sound source propagation path, calculate the frequency domain coupling coefficients between the multipath effect factors, and perform hierarchical fusion on the frequency domain coupling coefficients to obtain a mixed feature tensor.

[0046] Based on the adaptive mapping network, the present invention analyzes the sound source propagation path corresponding to the real-time audio stream. With the help of the network's dynamic modeling ability for feature association, it can accurately analyze the multipath effects of sound in space, such as reflection and diffraction (such as indoor reverberation paths or outdoor obstacle reflection paths), thereby improving the accuracy of sound source positioning and identification in complex environments.

[0047] The sound source propagation path refers to the complete set of transmission paths from the sound source to the acquisition device, including a direct sound path and several reflection and diffraction paths. Each path corresponds to a specific azimuth, propagation delay, and energy attenuation characteristics. For example, in an outdoor scene, the sound of a car horn may reach the microphone directly through the ground (path 1) and at the same time be reflected by the building in front (path 2), forming a propagation model containing two paths, which is used to analyze the impact of multipath interference on the recognition results.

[0048] As an embodiment of the present invention, the analysis of the sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network includes: time-framing the real-time audio stream to obtain audio frame groups; generating a time-frequency feature matrix corresponding to the audio frame groups; extracting the sound source propagation vector in the time-frequency feature matrix; inputting the sound source propagation vector into a path analysis module in the adaptive mapping network to output a path feature sequence; analyzing the sound source azimuth and propagation delay in the path feature sequence; and constructing the sound source propagation path corresponding to the real-time audio stream based on the sound source azimuth and the propagation delay.

[0049] Among them, the audio frame group refers to a set of continuous short-time segments into which the real-time audio stream is divided according to a fixed time interval (such as 10-30ms), and each segment corresponds to a frame of audio signal, which is convenient for fine-grained analysis of the temporal dynamic characteristics of the sound. For example, a 2-second environmental audio can be divided into 200 audio frames (at intervals of 10ms), and each frame corresponds to a fragment of short-time sound events such as vehicle horns and pedestrian footsteps; the time-frequency feature matrix refers to a two-dimensional matrix generated after performing time-frequency transformation (such as STFT) on the audio frame group, with the horizontal axis being time (number of frames) and the vertical axis being frequency. The matrix element value represents the signal energy or amplitude at the corresponding time-frequency point, similar to the spectrogram structure. For example, The time-frequency feature matrix of sound can show the changing trajectory of vowel resonance peaks over time, while the matrix of factory noise can present a continuous distribution of medium and low-frequency energy; the sound source propagation vector refers to a one-dimensional vector extracted from the time-frequency feature matrix, which characterizes the physical characteristics of sound propagation, including parameters such as frequency attenuation coefficient and phase change rate. For example, in indoor scenes, the sound source propagation vector can reflect the energy attenuation of sound waves after reflection from the wall (such as high-frequency components attenuate faster) and phase delay (related to the reflection distance), which is used to distinguish between direct sound and reflected sound characteristics; the path analysis module refers to a subnetwork structure in the adaptive mapping network specifically used to analyze the sound source propagation characteristics, which usually includes a convolutional layer, a fully connected layer or a graph neural network. The convolution layer can extract features and model the path of the sound source propagation vector. For example, the module can capture the local propagation mode (such as burst-like reflected sound energy) in the time-frequency feature matrix through the convolution layer, and then output the path-related parameter prediction through the fully connected layer; the path feature sequence refers to the time series feature sequence output by the path analysis module after processing the sound source propagation vector, which contains the propagation path attributes corresponding to each audio frame (such as the number of multipaths and the proportion of path energy). For example, in a multi-sound source mixed scenario, the path feature sequence can mark frame by frame whether the current frame signal mainly comes from the direct sound path (accounting for 70%) or the reflected path (accounting for 30%), and record the frequency component differences of each path; the sound source azimuth is It refers to the horizontal angle of the sound source relative to the microphone array or acquisition device (usually 0° is directly in front and clockwise is positive), which is used to locate the spatial position of the sound source. For example, in the far-field speech recognition of smart speakers, the azimuth of the sound source can indicate the direction in which the user is speaking (such as 45° to the left), assisting beamforming technology to enhance the target signal; the propagation delay refers to the time difference for sound to reach the acquisition device from the sound source via different paths, which is mainly caused by the difference in propagation distance (such as the difference in distance between direct sound and reflected sound). For example, in an indoor environment, the reflected sound path is 5 meters longer than the direct sound path, and the propagation delay is about 15ms (the speed of sound is about 340m / s). This parameter can be used for multipath effect modeling and sound source localization algorithm.

[0050] Furthermore, the time framing of the real-time audio stream can be achieved through a sliding window segmentation technology, such as: using Python's Librosa library to call the frame function with a 25ms window length and a 10ms step size to perform framing processing, and finally obtaining a time-continuous audio frame group; the generation of the time-frequency feature matrix corresponding to the audio frame group can be achieved through a short-time Fourier transform, such as: using the stft function of the scipy.signal library to calculate the spectral energy distribution of each frame signal, and finally obtaining a feature matrix containing time-frequency information; the extraction of the sound source propagation vector in the time-frequency feature matrix can be achieved through a wave direction estimation algorithm, such as: calculating the phase difference information of each frequency point through NumPy based on the MUSIC algorithm, and finally obtaining the propagation vector representing the direction of the sound source. ; The path analysis module that inputs the sound source propagation vector into the adaptive mapping network can be implemented through tensor operations, such as: using the Dense layer of TensorFlow to perform feature space transformation, and finally obtaining a path feature sequence with a time series relationship; the analysis of the sound source azimuth and propagation delay in the path feature sequence can be implemented through a geometric acoustic model, such as: applying the GCC-PHAT algorithm combined with the microphone array geometric parameters to calculate the delay difference, and finally obtaining accurate sound source azimuth and propagation delay parameters; the construction of the sound source propagation path corresponding to the real-time audio stream can be implemented by a ray tracing method, such as: using the PyRoomAcoustics library to simulate the reflection path of the sound wave in the environment, and finally obtaining a complete sound source propagation path including direct sound and reflected sound.

[0051] By identifying the multipath effect factors in the sound source propagation path, the present invention can accurately capture the multipath interference characteristics (such as time delay, attenuation, and phase difference) caused by reflection and diffraction of sound in complex environments, thereby improving the sound recognition system's anti-interference ability in scenarios such as reverberation and occlusion, and reducing recognition deviations caused by multipath interference.

[0052] Among them, the multipath effect factor refers to a set of physical parameters that describe the changes in signal characteristics caused by multiple paths (direct, reflected, diffraction, etc.) during the propagation of sound, including path delay (time difference between different paths), energy attenuation coefficient (path loss), phase offset (waveform distortion) and number of paths. For example, in an indoor scene, the path of the voice signal reflected by the ceiling is 3 meters longer than the direct path, resulting in a delay of approximately 9ms, and the high-frequency energy is attenuated by 15dB. These delay and attenuation values are the multipath effect factor; in an outdoor scene, the sound of a car horn may produce a phase offset (such as a 20° waveform phase rotation) after diffraction by a wall, which also falls into the category of this factor. Optionally, the identification of the multipath effect factor in the sound source propagation path can be achieved through a multipath signal decomposition algorithm, such as: using blind source separation technology combined with the FastICA method to extract each path component from the mixed signal, and finally obtaining a multipath effect factor containing delay and attenuation characteristics.

[0053] Furthermore, the present invention can quantify the interaction intensity of signals from different propagation paths in the frequency domain (such as the degree of interference between reflected sound and direct sound at a specific frequency) by calculating the frequency domain coupling coefficient between the multipath effect factors, thereby improving the accuracy of feature analysis in a multipath environment.

[0054] The frequency domain coupling coefficient is a numerical value that measures the degree of correlation between multipath effect factors in the frequency domain. The larger the value, the stronger the interaction between multipaths in the frequency domain. For example, in an indoor acoustic environment, multiple reflection path signals interfere with each other in certain frequency bands. The frequency domain coupling coefficient can quantify the degree of this interference and assist in optimizing audio processing algorithms.

[0055] As an embodiment of the present invention, the calculating the frequency domain coupling coefficient between the multipath effect factors includes: The frequency domain coupling coefficient between the multipath effect factors is calculated using the following formula: in, represents the frequency domain coupling coefficient between the multipath effect factors, Indicates the number of sampling times corresponding to the multipath effect factor, represents the sampling index corresponding to the multipath effect factor, Indicates the The multipath effect vector of the subsample, represents the mean of the multipath effect vector of all samples, represents the smoothing coefficient, represents the highest analysis frequency, Indicates frequency The multipath effect function at Indicates the frequency change rate corresponding to the multipath effect function.

[0056] Furthermore, the multipath effect vector refers to a vector that describes the characteristics of the multipath effect at a certain sampling moment, including parameters such as time delay, attenuation, and phase. For example, for modeling outdoor car sound propagation, the multipath effect vector can represent the characteristic combination of direct sound and reflected sound at a certain moment in terms of time delay and energy attenuation; the multipath effect vector mean refers to the average characteristics of the multipath effect vector under multiple samplings, such as collecting a sound source signal indoors multiple times and calculating the multipath effect vector mean, which can reflect the average state of the multipath effect in this environment and is used to analyze common propagation characteristics; the smoothing coefficient refers to a parameter that adjusts the degree of influence of the frequency change rate on the frequency domain coupling coefficient, and its value is between 0-1. When it is close to 0, the influence of the frequency change rate is weakened; when it is close to 1, its influence is enhanced. For example, in the audio noise reduction scenario, it can be adjusted according to the surrounding The smoothing coefficient is adjusted in the environment to control the noise reduction intensity; the highest analysis frequency refers to the highest frequency value considered in the audio analysis. For example, 4kHz is often used for speech analysis, which means that only the multipath effect in the 0-4kHz frequency band is analyzed to determine the analysis frequency band range and avoid high-frequency noise interference; the multipath effect function refers to a function that uses frequency as the independent variable to describe the change of multipath effect characteristics with frequency. For example, when studying indoor reverberation, the multipath effect function can show the change of reflected sound energy, phase, etc. with frequency at different frequencies; the frequency change rate refers to the speed of change of the multipath effect function with frequency. For example, in the high frequency band, the phase of sound multipath propagation changes quickly and the frequency change rate is large; the low frequency band is relatively flat and the frequency change rate is small, reflecting the dynamic characteristics of the multipath effect in the frequency domain.

[0057] The present invention obtains a hybrid feature tensor by layering and fusing the frequency domain coupling coefficients, which can integrate the frequency domain correlation information of multipath effects at different levels, enrich the feature dimensions (such as from low-level features such as path delay and attenuation to high-level features of the overall interference pattern), and improve the accuracy of tasks such as sound recognition and positioning.

[0058] Among them, the hybrid feature tensor refers to the multidimensional data structure obtained by layering and fusing the frequency domain coupling coefficients according to the dimensional distribution rules. For example, in a three-dimensional tensor, different dimensions represent different frequency bands, multipath types, coupling strengths and other information, which can comprehensively integrate the frequency domain characteristics of the multipath effect and be used as the input of the deep learning model.

[0059] As an embodiment of the present invention, the frequency domain coupling coefficients are hierarchically fused to obtain a mixed feature tensor, including: querying the frequency domain feature subset corresponding to the frequency domain coupling coefficients; extracting the local feature vectors of each frequency band of the frequency domain feature subset; analyzing the distribution characteristics of the local feature vectors in multidimensional space; parsing the dimensional distribution rules corresponding to the distribution characteristics; and based on the dimensional distribution rules, hierarchically fusing the frequency domain coupling coefficients to obtain a mixed feature tensor.

[0060] Among them, the frequency domain feature subset refers to a feature set of a specific frequency band range divided from the entire frequency domain feature. For example, in audio processing, the voice band (300-3400Hz) can be used as a frequency domain feature subset to focus on analyzing the frequency domain coupling of speech-related multipath effect factors, which is different from the characteristics of the high-frequency noise band; the local feature vector refers to a vector extracted for each frequency band in the frequency domain feature subset that can reflect the characteristics of the frequency band. For example, in a narrowband frequency domain, a vector containing the energy, phase and other information of the multipath effect factor of the frequency band is extracted, just like the vector extracted in the 2000-2100Hz frequency band, which describes the multipath characteristics of sound propagation in this frequency band; the multidimensional space refers to an abstract space composed of multiple feature dimensions, which is used to characterize the local feature vector, for example, the multipath effect factor Features such as delay, attenuation, and phase are used as dimensions to construct a space, and each local eigenvector corresponds to a point in the space, just like a two-dimensional plane can represent the two dimensions of delay and attenuation, and can intuitively show the position of the vector; the distribution characteristics refer to the distribution laws and characteristics of local eigenvectors in multidimensional space. For example, in multidimensional space, some local eigenvectors may be concentrated in specific areas, reflecting the concentration trend of multipath effects under certain feature combinations, such as the local eigenvectors of indoor reflected sound are concentrated in areas with small delay and moderate attenuation; the dimensional distribution rules refer to the relationships and distribution criteria between dimensions summarized based on the distribution characteristics of local eigenvectors. For example, it is found that there is a negative correlation between the delay dimension and the attenuation dimension, and they are concentrated in a specific range. This is used to formulate rules to guide subsequent feature fusion, such as stipulating that the attenuation corresponding to a small delay cannot be too large.

[0061] Furthermore, the query of the frequency domain feature subset corresponding to the frequency domain coupling coefficient can be achieved through a frequency band segmentation algorithm, such as: using wavelet packet decomposition combined with the PyWavelets library to divide the signal into different frequency bands, and finally obtaining a frequency domain feature subset containing energy of a specific frequency band; the extraction of the local feature vector of each frequency band of the frequency domain feature subset can be achieved through time-frequency analysis technology, such as: using short-time Fourier transform through the scipy.signal library to calculate the spectrum centroid and bandwidth of each frequency band, and finally obtaining a feature vector that characterizes the local characteristics; the analysis of the distribution characteristics of the local feature vector in multidimensional space can be achieved through a manifold learning method, such as: applying t -The SNE algorithm uses the scikit-learn library to visualize and reduce the dimensionality of high-dimensional features, and ultimately obtains distribution characteristics that reflect the data aggregation characteristics; the analysis of the dimensional distribution rules corresponding to the distribution characteristics can be achieved through cluster analysis methods, such as: using a Gaussian mixture model to fit the feature distribution probability density through the sklearn.mixture library, and ultimately obtaining dimensional distribution rules that describe the association rules of each dimension; the hierarchical fusion of the frequency domain coupling coefficients can be achieved through tensor fusion technology, such as: using the TensorLy library based on Tucker decomposition to achieve multi-level feature interaction, and ultimately obtaining a mixed feature tensor that retains the features of each dimension.

[0062] S4. Query the adaptive response trajectory of the mixed feature tensor in the target sound source scene, extract the phase distortion feature and the signal-to-noise ratio index in the adaptive response trajectory, and calculate the signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index.

[0063] By querying the adaptive response trajectory of the mixed feature tensor in the target sound source scene, the present invention can capture the model's dynamic processing process of complex acoustic features such as the multipath effect in the scene in real time, thereby improving its stability and accuracy in tasks such as target sound source identification and positioning in specific scenarios.

[0064] Among them, the adaptive response trajectory refers to the trajectory in the energy attenuation gradient diagram that reflects the model's adaptive adjustment process to changes in sound propagation in the sound field, for example, the model's response path to changes in energy attenuation caused by sound source movement and environmental changes at different times.

[0065] As an embodiment of the present invention, the querying of the adaptive response trajectory of the mixed feature tensor in the target sound source scene includes: parsing the tensor distribution interval corresponding to the mixed feature tensor; matching the dynamic voiceprint segments in the target sound source scene according to the time-frequency domain distribution interval, and extracting sparse feature clusters in the voiceprint segments; mapping the sparse feature clusters to the sound field propagation path corresponding to the target sound source scene; generating an energy attenuation gradient map corresponding to the sound field propagation path; and querying the adaptive response trajectory in the energy attenuation gradient map.

[0066] Among them, the tensor distribution interval refers to the range set of values of the hybrid feature tensor in each dimension. For example, in the three-dimensional hybrid feature tensor, the corresponding frequency band, multipath type, and coupling strength dimensions respectively constitute the tensor distribution interval. For example, the frequency band dimension is 20-20000Hz, and the multipath type dimension includes direct sound, single reflection sound, etc., which define the range of the tensor feature; the dynamic voiceprint segment refers to the voiceprint information segment that changes with time in the target sound source scene. For example, when someone speaks, the voiceprint data sequence generated by different sentences and intonation changes contains the speaker's unique timbre, frequency changes and other characteristics, which are the effective information part with dynamic changes in voiceprint recognition. The sparse feature cluster refers to a set of features extracted from dynamic voiceprint fragments, with sparse feature distribution but representativeness. For example, in a voiceprint fragment, features such as the energy value of certain key frequency points and phase mutations at specific moments are relatively small in number but can reflect the uniqueness of the voiceprint, similar to the sparse corner features in an image. The sound field propagation path refers to the route that sound propagates in the target sound source scene space. For example, in a conference room, the path that the speaker's voice reflects from walls, ceilings, etc. to the audience position includes direct paths and multiple reflection paths. These paths determine the propagation delay, attenuation and other characteristics of the sound. The energy attenuation gradient map refers to a graph based on the sound field propagation path that depicts the change of sound energy as the propagation path attenuates. For example, in an indoor scene, the attenuation degree of sound energy on different paths is marked with color depth or numerical value, intuitively showing the decreasing trend of energy in each propagation direction from the sound source, facilitating the analysis of propagation characteristics.

[0067] Furthermore, the analysis of the tensor distribution interval corresponding to the mixed feature tensor can be achieved through a tensor decomposition method, such as: using CP decomposition combined with the TensorLy library to extract the distribution range of each modal feature, and finally obtaining a set of intervals describing the distribution of tensor data; the matching of dynamic voiceprint segments in the target sound source scene can be achieved through a dynamic time warping algorithm, such as: using Python's dtw-package to calculate the similarity path between the voiceprint template and the real-time signal, and finally obtaining the best matching voiceprint segment; the extraction of sparse feature clusters in the voiceprint segment can be achieved through sparse coding technology, such as: applying the KSVD algorithm to learn the sparse representation of voiceprint features through the scikit-learn library, and finally obtaining a discriminative sparse feature cluster. ; Mapping the sparse feature cluster to the sound field propagation path corresponding to the target sound source scene can be achieved through acoustic ray tracing, such as: using the PyRoomAcoustics library to simulate the propagation trajectory of the feature cluster in the environment, and finally obtaining the sound field propagation path including the reflection path; generating the energy attenuation gradient map corresponding to the sound field propagation path can be achieved through an acoustic energy attenuation model, such as: calculating the energy attenuation rate of each point on the path through SciPy based on the sound wave propagation equation, and finally obtaining a visual attenuation gradient map; querying the adaptive response trajectory in the energy attenuation gradient map can be achieved through a path optimization algorithm, such as: using the Dijkstra algorithm to find the optimal energy transfer path in the gradient map, and finally obtaining an adaptive response trajectory that conforms to the acoustic characteristics.

[0068] By extracting the phase distortion characteristics and signal-to-noise ratio index from the adaptive response trajectory, the present invention can accurately quantify the degree of waveform distortion (such as phase rotation and waveform broadening) and signal purity caused by multipath interference during sound propagation, effectively improving the intelligibility of sound signals and the robustness of the recognition system in complex environments.

[0069] Among them, the phase distortion feature refers to the feature of the phase deviation of the sound signal from the ideal state due to factors such as multipath effect and medium inhomogeneity during the propagation process, which manifests as nonlinear distortion of the phase spectrum, phase mutation or phase difference. For example, the phase difference between the reflected sound and the direct sound in the room may cause the composite signal to have phase cancellation at a specific frequency point (such as energy attenuation when the phase difference is 180°). This phase anomaly can be quantified by the phase distortion feature and used to evaluate the degree of damage to the signal waveform caused by multipath interference; the signal-to-noise ratio index refers to the ratio of the target sound signal energy to the background noise energy, which is used to measure the purity of the audio signal. For example, in the in-vehicle voice interaction scenario, if Engine noise causes the signal-to-noise ratio index of the microphone-collected signal to drop from 20dB to 10dB, indicating a significant increase in noise energy, which may lead to an increase in the speech recognition error rate. By monitoring this index in real time, the noise reduction algorithm parameters can be dynamically adjusted to optimize signal quality. Optionally, the extraction of phase distortion features and signal-to-noise ratio index in the adaptive response trajectory can be achieved through signal analysis technology, such as: using Hilbert transform in combination with Python's scipy.signal library to calculate the instantaneous phase offset, and using the power spectral density estimation method to obtain the signal-to-noise energy ratio, ultimately obtaining the phase distortion features that characterize the signal distortion and the accurate signal-to-noise ratio index.

[0070] Furthermore, the present invention calculates the signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion characteristics and the signal-to-noise ratio index. The uncertainty of the signal under noise interference can be quantified through the information entropy theory, thereby dynamically optimizing the signal processing strategy in a complex acoustic environment, improving the separability of the target sound and the robustness of the recognition system.

[0071] Among them, the signal-to-noise attenuation entropy refers to a quantitative indicator that comprehensively measures the degree to which the signal in the real-time audio stream is affected by noise interference and phase distortion. By integrating phase distortion characteristics and signal-to-noise ratio index related parameters, the concept of information entropy is used to reflect the uncertainty and quality loss of audio signals in complex environments. The larger the value, the more serious the signal interference and the more obvious the quality degradation.

[0072] As an embodiment of the present invention, calculating the signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index includes: The signal-to-noise attenuation entropy corresponding to the real-time audio stream is calculated using the following formula: in, represents the signal-to-noise attenuation entropy corresponding to the real-time audio stream, Indicates the total number of dimensions of the distortion dimension corresponding to the phase distortion feature, Indicates the dimension index corresponding to the distortion dimension, Indicates the The distortion quantization value corresponding to the distortion dimension is Indicates the The distortion weight corresponding to the distortion dimension, Indicates the total number of dimensions of the index dimension corresponding to the signal-to-noise ratio index, Indicates the total number of dimensions corresponding to the index dimension, Indicates the The signal-to-noise ratio quantization value corresponding to the exponential dimension, Indicates the The signal-to-noise ratio weight corresponding to the exponential dimension is Indicates the length of time selected when analyzing real-time audio streams. represents the attenuation influence function at time t.

[0073] Specifically, the distortion dimension refers to a dimension used to describe different aspects or attributes of phase distortion characteristics. For example, it may include dimensions such as the degree of nonlinear phase offset, the frequency range of phase mutations, and the amplitude of phase fluctuations. Phase distortion is characterized from multiple perspectives to more comprehensively analyze signal distortion. The distortion quantization value refers to a numerical value that quantifies the degree of phase distortion under a specific distortion dimension. It converts the actual phase distortion into a specific numerical value to facilitate calculation and analysis in the above formula. The distortion weight refers to a parameter that reflects the relative importance of different distortion dimensions in calculating the signal-to-noise attenuation entropy. Different distortion dimensions have different degrees of impact on audio signal quality. By assigning corresponding weights, the formula can more reasonably comprehensively consider the role of each dimension. The exponential dimension refers to a dimension that describes different attributes or components of the signal-to-noise ratio index. For example, it may include dimensions such as the signal-to-noise ratio in different frequency bands, the change in the signal-to-noise ratio at different times in the time domain, and the impact of noise type on the signal-to-noise ratio. The signal-to-noise ratio quantization value refers to a numerical value that quantifies the signal-to-noise ratio under a specific exponential dimension. The SNR value is converted into a specific numerical value to accurately measure the degree to which the signal is affected by noise in that dimension. The SNR weight is a weighting coefficient assigned to each exponential dimension based on the importance of the different exponential dimensions on the audio signal quality. If, in a specific environment, the time domain SNR fluctuation has a more critical impact on speech recognition accuracy, then the SNR weight corresponding to the time domain will be set higher to highlight its role in the calculation of the SNR. The time length is a period of time selected for calculating the SNR when analyzing a real-time audio stream. It determines the time range of the calculation. For example, if 10 seconds of audio stream data is selected for analysis, these 10 seconds are the time length. Different time length settings may affect the final calculated SNR result. The attenuation influence function is a function of time t that is used to describe the attenuation effect on the audio signal at different times. For example, in an echo-filled indoor environment, the sound signal will produce energy attenuation over time due to factors such as reflection. A(t) can be used to establish a mathematical model based on the reflection path, attenuation coefficient, etc. to quantitatively describe the degree of signal attenuation at each moment.

[0074] S5. Based on the signal-to-noise attenuation entropy, determine the robust recognition level of the real-time audio stream, reconstruct the sound source acquisition dimension corresponding to the real-time audio stream according to the robust recognition level, and formulate a sound recognition scheme corresponding to the target sound source scene based on the sound source acquisition dimension.

[0075] The present invention determines the robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, can intuitively reflect the degree of audio interference and signal quality with quantitative indicators, provide a clear reference for the recognition system, improve the system's effective audio recognition capability in complex environments, and ensure the stability and accuracy of recognition.

[0076] Among them, the robust recognition level refers to the classification level of the degree to which real-time audio streams can be accurately recognized in a noisy environment. For example, it is divided into four levels: Level 1 (Excellent): extremely low signal-to-noise attenuation entropy (e.g., <0.3 bits), effective segmentation ratio >80%, and recognition accuracy >95%; Level 2 (Good): medium entropy value (0.3-0.6 bits), effective segmentation ratio 60%-80%, and recognition accuracy ratio 80%-95%; Level 3 (Fair): high entropy value (0.6-0.9 bits), effective segmentation ratio 40%-60%, and enhanced algorithm-assisted recognition is required; Level 4 (Poor): extremely high entropy value (>0.9 bits), effective segmentation ratio <40%, recognition accuracy ratio <50%, and manual intervention is required. This level is used to dynamically adjust the recognition strategy (e.g., starting multi-microphone array beamforming at Level 4) to improve system adaptability.

[0077] As an embodiment of the present invention, determining the robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy includes: parsing the noise interference characteristics corresponding to the signal-to-noise attenuation entropy; generating frequency domain segmented data corresponding to the real-time audio stream based on the noise interference characteristics; extracting valid segmented audio corresponding to the frequency domain segmented data; determining the hierarchical identifier corresponding to the real-time audio stream based on the valid segmented audio; and determining the robust recognition level corresponding to the real-time audio stream based on the hierarchical identifier.

[0078] Among them, the noise interference feature refers to a set of characteristic parameters that are analyzed from the signal-to-noise attenuation entropy and reflect the impact of noise on the audio signal. For example, through the frequency distribution characteristics of the signal-to-noise attenuation entropy, parameters such as the center frequency of the noise (such as the interference peak at 2kHz), bandwidth (the frequency range covered by the noise), and energy distribution pattern (such as Gaussian noise or impulse noise) can be extracted; in the time domain, information such as the duration and burst frequency of the noise can be obtained; the frequency domain segmented data refers to a data set formed by dividing the spectrum of the real-time audio stream into multiple frequency bands according to specific rules. For example, based on the frequency distribution of the noise interference feature, the audio spectrum is divided into a low-noise frequency band (such as 0-500Hz, noise energy <10dB), a medium-noise frequency band (500-2kHz, noise energy 10-20dB) and a high-noise frequency band (>2kHz, noise energy>20dB); the Valid segmented audio refers to audio segments selected from frequency domain segmented data whose signal-to-noise ratio meets recognition requirements. For example, in a speech recognition scenario, if the signal-to-noise ratio value of a frequency band is higher than a preset threshold (e.g., 15 dB) and the phase distortion is lower than a critical value (e.g., phase deviation <30°), the frequency band is determined to be a valid segment. Valid segmented audio retains key features of the original signal (e.g., the formant structure of speech) and can be used for subsequent grading, eliminating interference from invalid frequency bands dominated by noise. The grading identifier is a numerical or symbolic identifier calculated based on the characteristics of the valid segmented audio and used to characterize the robustness of the audio. For example, by integrating parameters such as the proportion of valid segments (e.g., the valid frequency band accounts for 70% of the total frequency band), the average signal-to-noise ratio (e.g., 20 dB), and phase stability (e.g., phase variance <0.5), a multidimensional vector is generated as a grading identifier (e.g., [0.7, 20, 0.5]). This identifier is associated with the robust recognition level through a preset mapping rule (e.g., a decision tree classifier) to achieve automated grading.

[0079] Furthermore, the analysis of the noise interference characteristics corresponding to the signal-to-noise attenuation entropy can be achieved through an entropy decomposition algorithm, such as: using a spectral entropy analysis method in combination with Python's librosa library to calculate the noise energy ratio of each frequency band, and ultimately obtaining a noise interference characteristic that characterizes the interference intensity; the generation of frequency domain segmented data corresponding to the real-time audio stream can be achieved through time-frequency segmentation technology, such as: using short-time Fourier transform through the scipy.signal library to divide the spectrum with a fixed frame length, and ultimately obtaining time-frequency aligned frequency domain segmented data; the extraction of valid segmented audio corresponding to the frequency domain segmented data can be achieved through energy threshold screening, such as: based on the dynamic threshold method, using NumPy to eliminate low-energy noise segments, and ultimately obtaining pure valid segmented audio; the determination of the hierarchical identifier corresponding to the real-time audio stream can be achieved through a clustering analysis method, such as: using the K-means algorithm with the help of the scikit-learn library to automatically classify according to the signal-to-noise ratio characteristics, and ultimately obtaining a digital hierarchical identifier; the determination of the robust recognition level corresponding to the real-time audio stream can be achieved through a decision tree model, such as: using the XGBoost classifier to comprehensively evaluate the signal-to-noise ratio and distortion indicators, and ultimately obtaining a quantified robust recognition level.

[0080] Based on the robust recognition level, the present invention reconstructs the sound source acquisition dimensions corresponding to the real-time audio stream, and can optimize the acquisition configuration in a targeted manner (such as increasing the number of microphones or adjusting the array layout at a low level) according to the audio interference status and the difficulty of recognition. This can effectively enhance the sound source signal strength, suppress noise interference, improve the audio acquisition quality, and provide higher-quality data for subsequent recognition, analysis, and other processing.

[0081] The sound source acquisition dimension refers to a dimensional system composed of various factors involved in collecting sound source signals, covering aspects such as space, time, and frequency. For example, in the spatial dimension, the layout (linear, circular, etc.) and spacing of the microphone array determine the ability to perceive the direction of the sound source; the temporal dimension includes the sampling frequency and sampling duration, which affects the temporal resolution of the signal; and the frequency dimension involves the frequency range of the acquisition. For example, voice acquisition often focuses on the 300-3400Hz frequency band. Optionally, reconstructing the sound source acquisition dimension corresponding to the real-time audio stream can be achieved through adaptive filtering methods, such as using the RLS algorithm through the scipy.signal library to dynamically adjust the array weight coefficients, ultimately obtaining an optimized sound source acquisition dimension that is resistant to interference.

[0082] Furthermore, based on the sound source acquisition dimensions, the present invention formulates a sound recognition scheme corresponding to the target sound source scene, which can optimize the acquisition parameters (such as microphone spacing and sampling frequency band) according to the scene characteristics (such as strong indoor reverberation and complex outdoor noise), thereby improving signal quality from the source, effectively reducing the impact of multipath interference and noise, and improving the accuracy of tasks such as sound source localization and speech recognition.

[0083] Among them, the sound recognition scheme refers to a full-process adaptive recognition strategy formed by analyzing parameters such as the signal-to-noise attenuation entropy and phase distortion characteristics of real-time audio, and dynamically adjusting the sound source acquisition method (such as microphone array layout, sampling frequency band), model inference strategy (such as noise reduction algorithm, feature extraction network) and environmental adaptation parameters (such as multipath effect compensation coefficient). For example, in a smart car scenario, if the system detects that the signal-to-noise attenuation entropy increases due to the mixture of highway wind noise and engine noise, it will automatically switch to a directional pickup microphone to focus on the human voice frequency band, enable a noise suppression model based on a generative adversarial network (GAN), and optimize the frequency domain feature extraction layer of a convolutional neural network (CNN) to achieve high-precision recognition of voice commands. Optionally, the sound recognition scheme corresponding to the target sound source scenario can be achieved through transfer learning technology, such as using a pre-trained VGGish model through the Keras library for domain adaptive fine-tuning, ultimately obtaining an optimized sound recognition scheme adapted to the specific scenario.

[0084] Compared with the problems described in the background technology, the present invention obtains the real-time audio stream in the target sound source scene and extracts the time-frequency feature group in the real-time audio stream, which can jointly analyze the audio signal from the time domain and frequency domain dimensions, accurately capture the frequency components and time series change rules of the sound, and effectively improve the accuracy of sound recognition in complex environments. The present invention uses the sound network architecture to identify the acoustic feature clusters in the real-time audio stream, and can automatically extract feature sets with semantic associations (such as phonemes in speech, frequency combinations in environmental sounds) from complex audio signals through the hierarchical nonlinear transformation of the neural network, effectively improving the representation ability of sound features and the generalization performance of the recognition system. Furthermore, the present invention analyzes the sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network, and can use the network to dynamically associate features. Modeling capabilities can accurately analyze multipath effects such as reflection and diffraction of sound in space (such as indoor reverberation paths or outdoor obstacle reflection paths), thereby improving the accuracy of sound source positioning and identification in complex environments. Furthermore, the present invention can capture the model's dynamic processing process of complex acoustic features such as multipath effects in the scene in real time by querying the adaptive response trajectory of the mixed feature tensor in the target sound source scene, which can improve its stability and accuracy in tasks such as target sound source identification and positioning in specific scenarios. Finally, the present invention determines the robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, which can intuitively reflect the degree of audio interference and signal quality with quantitative indicators, provide a clear reference for the recognition system, and improve the system's effective recognition capability of audio in complex environments, ensuring the stability and accuracy of recognition. Therefore, the neural network-based sound recognition method and system provided in the embodiment of the present invention can improve the accuracy of sound recognition in complex environments.

[0085] Example 2: like Figure 3 FIG. 1 is a functional module diagram of a sound recognition system based on a neural network according to the present invention.

[0086] The neural network-based voice recognition system 200 described in the present invention can be installed in an electronic device. Depending on the functionality implemented, the neural network-based voice recognition system can include an architecture building module 201, a network generation module 202, a layered fusion module 203, an attenuated entropy calculation module 204, and a solution formulation module 205. The modules described in the present invention, also known as units, refer to a series of computer program segments that can be executed by an electronic device processor and perform a fixed function, and are stored in the electronic device's memory.

[0087] In the embodiment of the present invention, the functions of each module / unit are as follows: The architecture construction module 201 is used to obtain a real-time audio stream in a target sound source scene, extract a time-frequency feature group from the real-time audio stream, parse a pulse code sequence corresponding to the time-frequency feature group, and construct a sound network architecture corresponding to the real-time audio stream based on the pulse code sequence; The network generation module 202 is configured to identify acoustic feature clusters in the real-time audio stream using the sound network architecture, match hierarchical topological relationships in a preset neural computing framework based on the acoustic feature clusters, optimize the relationship node weights corresponding to the hierarchical topological relationships, and generate an adaptive mapping network corresponding to the relationship node weights; The hierarchical fusion module 203 is configured to analyze the sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network, identify multipath effect factors in the sound source propagation path, calculate frequency domain coupling coefficients between the multipath effect factors, and perform hierarchical fusion on the frequency domain coupling coefficients to obtain a hybrid feature tensor; The attenuation entropy calculation module 204 is used to query the adaptive response trajectory of the mixed feature tensor in the target sound source scene, extract the phase distortion feature and the signal-to-noise ratio index in the adaptive response trajectory, and calculate the signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index; The solution formulation module 205 is used to determine the robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, reconstruct the sound source acquisition dimension corresponding to the real-time audio stream according to the robust recognition level, and formulate a sound recognition solution corresponding to the target sound source scene based on the sound source acquisition dimension.

[0088] In detail, the modules in the neural network-based voice recognition system 200 according to the embodiment of the present invention are used in the same manner as above. Figure 1The technical means are similar to the sound recognition method based on neural network described in , and can produce the same technical effects, so I will not go into details here.

[0089] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A sound recognition method based on a neural network, characterized in that: The method comprises: Acquire a real-time audio stream in a target sound source scene, extract a time-frequency feature group from the real-time audio stream, parse a pulse code sequence corresponding to the time-frequency feature group, and construct a sound network architecture corresponding to the real-time audio stream based on the pulse code sequence; Identifying acoustic feature clusters in the real-time audio stream using the sound network architecture, matching hierarchical topological relationships in a preset neural computing framework based on the acoustic feature clusters, optimizing relationship node weights corresponding to the hierarchical topological relationships, and generating an adaptive mapping network corresponding to the relationship node weights; Based on the adaptive mapping network, analyzing the sound source propagation path corresponding to the real-time audio stream, identifying the multipath effect factors in the sound source propagation path, calculating the frequency domain coupling coefficients between the multipath effect factors, and performing hierarchical fusion on the frequency domain coupling coefficients to obtain a hybrid feature tensor; querying an adaptive response trajectory of the mixed feature tensor in the target sound source scene, extracting a phase distortion feature and a signal-to-noise ratio index from the adaptive response trajectory, and calculating a signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index; Based on the signal-to-noise attenuation entropy, the robust recognition level of the real-time audio stream is determined; according to the robust recognition level, the sound source acquisition dimension corresponding to the real-time audio stream is reconstructed; and based on the sound source acquisition dimension, a sound recognition scheme corresponding to the target sound source scene is formulated.

2. A sound recognition method based on a neural network as claimed in claim 1, characterized in that: The step of constructing a sound network architecture corresponding to the real-time audio stream according to the pulse code sequence includes: Analyzing frequency domain eigenvectors in the pulse code sequence; Extracting an audio frame set corresponding to the real-time audio stream; Performing time-frequency alignment on the frequency domain feature vector and the audio frame set to generate an audio alignment feature map; Identifying audio node relationships in the audio alignment feature graph; According to the audio node relationship, a sound network architecture corresponding to the real-time audio stream is constructed.

3. A sound recognition method based on a neural network as claimed in claim 1, characterized in that: Generating the adaptive mapping network corresponding to the relationship node weights includes: Parsing the topological association code in the relationship node weights; Based on the topology association code, a preset weight mapping rule library is matched to obtain an initial mapping path set; Dynamically normalizing the initial mapping path set to obtain a normalized path sequence; removing conflicting sequence segments from the normalized path sequence to obtain optimized mapping sequence segments; Based on the optimized mapping sequence segments, an adaptive mapping network corresponding to the relationship node weights is generated.

4. A method for sound recognition based on a neural network as claimed in claim 1, characterized in that: The analyzing the sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network includes: Performing time framing on the real-time audio stream to obtain audio frame groups; Generating a time-frequency feature matrix corresponding to the audio frame group; Extracting the sound source propagation vector from the time-frequency feature matrix; Inputting the sound source propagation vector into a path analysis module in the adaptive mapping network to output a path feature sequence; Analyzing the sound source azimuth and propagation delay in the path characteristic sequence; Based on the sound source azimuth and the propagation delay, a sound source propagation path corresponding to the real-time audio stream is constructed.

5. A sound recognition method based on a neural network as claimed in claim 1, characterized in that: The calculating the frequency domain coupling coefficient between the multipath effect factors includes: Calculate the frequency domain coupling coefficient between the multipath effect factors.

6. A method for sound recognition based on a neural network as claimed in claim 1, characterized in that: The layered fusion of the frequency domain coupling coefficients to obtain a hybrid feature tensor includes: Querying a frequency domain feature subset corresponding to the frequency domain coupling coefficient; Extracting local feature vectors of each frequency band of the frequency domain feature subset; Analyzing the distribution characteristics of the local feature vector in the multidimensional space; Analyzing the dimensional distribution rules corresponding to the distribution characteristics; Based on the dimensional distribution rule, the frequency domain coupling coefficients are hierarchically fused to obtain a hybrid feature tensor.

7. A sound recognition method based on a neural network as claimed in claim 1, characterized in that: The querying of the adaptive response trajectory of the mixed feature tensor in the target sound source scene includes: Analyzing the tensor distribution interval corresponding to the mixed feature tensor; Matching dynamic voiceprint segments in the target sound source scene according to the time-frequency domain distribution interval, and extracting sparse feature clusters in the voiceprint segments; Mapping the sparse feature cluster to the sound field propagation path corresponding to the target sound source scene; generating an energy attenuation gradient map corresponding to the sound field propagation path; The adaptive response trajectory in the energy decay gradient map is queried.

8. A method for sound recognition based on a neural network as claimed in claim 1, characterized in that: The calculating, based on the phase distortion feature and the signal-to-noise ratio index, the signal-to-noise attenuation entropy corresponding to the real-time audio stream includes: Calculate the signal-to-noise attenuation entropy corresponding to the real-time audio stream.

9. A method for sound recognition based on a neural network as claimed in claim 1, characterized in that: The determining, based on the signal-to-noise attenuation entropy, a robust recognition level of the real-time audio stream includes: Analyzing the noise interference characteristics corresponding to the signal-to-noise attenuation entropy; generating frequency domain segmented data corresponding to the real-time audio stream based on the noise interference characteristics; Extracting valid segmented audio corresponding to the frequency domain segmented data; Determining, based on the valid segmented audio, a hierarchical identifier corresponding to the real-time audio stream; Based on the classification identifier, a robust recognition level corresponding to the real-time audio stream is determined.

10. A sound recognition system based on a neural network, characterized in that: The system comprises: An architecture construction module is used to obtain a real-time audio stream in a target sound source scene, extract a time-frequency feature group from the real-time audio stream, parse a pulse code sequence corresponding to the time-frequency feature group, and construct a sound network architecture corresponding to the real-time audio stream based on the pulse code sequence; A network generation module is configured to identify acoustic feature clusters in the real-time audio stream using the sound network architecture, match hierarchical topological relationships in a preset neural computing framework based on the acoustic feature clusters, optimize relationship node weights corresponding to the hierarchical topological relationships, and generate an adaptive mapping network corresponding to the relationship node weights; a hierarchical fusion module, configured to analyze, based on the adaptive mapping network, a sound source propagation path corresponding to the real-time audio stream, identify multipath effect factors in the sound source propagation path, calculate frequency domain coupling coefficients between the multipath effect factors, and perform hierarchical fusion on the frequency domain coupling coefficients to obtain a hybrid feature tensor; an attenuation entropy calculation module, configured to query an adaptive response trajectory of the mixed feature tensor in the target sound source scene, extract a phase distortion feature and a signal-to-noise ratio index from the adaptive response trajectory, and calculate a signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index; A solution formulation module is used to determine the robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, reconstruct the sound source acquisition dimension corresponding to the real-time audio stream according to the robust recognition level, and formulate a sound recognition solution corresponding to the target sound source scene based on the sound source acquisition dimension.

Citation Information

Patent Citations

  • Method and device for constructing voice identification model capable of automatically searching parameters

    CN115083421A

  • Streaming speech recognition method, device and equipment based on adaptive AI large model

    CN119252234A

  • Hybrid speech processing method, electronic equipment and computer readable medium

    CN120236599A

  • Self-development type voice language pattern recognition system, and method and program for structuring self-organizing neural network structure used for same system

    JP2006171714A

  • Method for training a speech recognition model and method for speech recognition

    US20230031733A1

Cited By

  • Directional pickup method and device, computer equipment and medium

    CN121768413A