A neural network-based sound recognition method and system

By using a neural network-based approach, the time-frequency feature set of real-time audio streams is obtained, a sound network architecture is constructed, the sound source propagation path is analyzed, and the signal-to-noise attenuation entropy is calculated. This solves the problem of insufficient accuracy of traditional sound recognition in complex environments and achieves higher recognition accuracy and stability.

CN120452436BActive Publication Date: 2026-02-06XIAN FULIYE MICROELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510911087.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2026-02-06
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Traditional sound recognition methods lack accuracy in complex environments, especially in noisy or multi-source mixed scenarios where accuracy drops significantly.

Method used

A neural network-based sound recognition method is adopted. By acquiring the time-frequency feature groups of real-time audio streams, a sound network architecture is constructed. An adaptive mapping network is used to analyze the sound source propagation path, calculate the signal-to-noise attenuation entropy, determine the robust recognition level, reconstruct the sound source acquisition dimensions, and formulate a recognition scheme.

Benefits of technology

It improves the accuracy of sound recognition in complex environments, enhances the stability and accuracy of target sound source identification and localization in specific scenarios, and strengthens the system's effective recognition capability in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452436B_ABST
    Figure CN120452436B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of audio processing, and discloses a sound recognition method and system based on a neural network, which comprises the following steps: firstly, obtaining a target scene real-time audio stream, extracting a time-frequency feature group, analyzing a pulse code sequence, and constructing a sound network architecture; then, recognizing acoustic feature clusters by using the sound network architecture, matching a preset framework hierarchical topological relationship, optimizing node weights to generate an adaptive mapping network; based on the adaptive mapping network, analyzing a sound source propagation path, recognizing a multipath effect factor, calculating a frequency domain coupling coefficient, and hierarchically fusing to obtain a hybrid feature tensor; querying a tensor adaptive response track, extracting a phase distortion feature and a signal-to-noise ratio index, and calculating a signal-to-noise attenuation entropy; finally, judging a robust recognition level according to the signal-to-noise attenuation entropy, reconstructing a sound source collection dimension, and formulating a sound recognition scheme. The application can improve the sound recognition precision in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a sound recognition method and system based on a neural network and belongs to the technical field of audio processing. BACKGROUND

[0002] Sound recognition refers to a technology of analyzing and processing sound signals by technical means to recognize the content, category or source of the sound, which is widely applied to fields such as speech recognition (such as speech-to-text), environmental sound classification (such as gunshot and siren detection), speaker recognition (such as voiceprint authentication) and music information retrieval (such as song recognition).

[0003] At present, traditional sound recognition methods mainly rely on feature extraction algorithms such as mel-frequency cepstral coefficients (MFCC) and combine hidden Markov models (HMM) or Gaussian mixture models (GMM) for classification. However, these methods have limited sound feature extraction capability in complex environments, especially in noisy interference or multi-source mixed scenes, and the recognition accuracy significantly decreases. Therefore, a sound recognition method based on a neural network is needed to improve the sound recognition accuracy in complex environments. SUMMARY

[0004] The application provides a sound recognition method and system based on a neural network, which aims to improve the sound recognition accuracy in complex environments.

[0005] To achieve the above-mentioned purpose, the application provides a sound recognition method based on a neural network, which comprises:

[0006] Obtaining a real-time audio stream in a target sound source scene, extracting a time-frequency feature group in the real-time audio stream, analyzing a pulse code sequence corresponding to the time-frequency feature group, constructing a sound network architecture corresponding to the real-time audio stream according to the pulse code sequence;

[0007] Identifying acoustic feature clusters in the real-time audio stream by using the sound network architecture, matching a hierarchical topological relationship in a preset neural computing framework according to the acoustic feature clusters, optimizing relationship node weights corresponding to the hierarchical topological relationship, and generating an adaptive mapping network corresponding to the relationship node weights;

[0008] Analyzing a sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network, identifying multi-path effect factors in the sound source propagation path, calculating frequency domain coupling coefficients between the multi-path effect factors, performing hierarchical fusion on the frequency domain coupling coefficients, and obtaining a mixed feature tensor;

[0009] querying an adaptive response track of the mixed feature tensor in the target sound source scene, extracting a phase distortion feature and a signal-to-noise ratio index in the adaptive response track, calculating a signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index;

[0010] determining a robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, reconstructing a sound source collection dimension corresponding to the real-time audio stream according to the robust recognition level, and formulating a sound recognition scheme corresponding to the target sound source scene based on the sound source collection dimension.

[0011] Optionally, the constructing the sound network architecture corresponding to the real-time audio stream according to the pulse coding sequence comprises:

[0012] analyzing a frequency domain feature vector in the pulse coding sequence;

[0013] extracting an audio frame set corresponding to the real-time audio stream;

[0014] performing time-frequency alignment on the frequency domain feature vector and the audio frame set to generate an audio alignment feature map;

[0015] recognizing an audio node relationship in the audio alignment feature map;

[0016] constructing the sound network architecture corresponding to the real-time audio stream according to the audio node relationship.

[0017] Optionally, the generating the adaptive mapping network corresponding to the relationship node weight comprises:

[0018] parsing a topological association code in the relationship node weight;

[0019] matching a preset weight mapping rule library based on the topological association code to obtain an initial mapping path set;

[0020] dynamically normalizing the initial mapping path set to obtain a normalized path sequence;

[0021] removing a conflict segment in the normalized path sequence to obtain an optimized mapping sequence segment;

[0022] generating the adaptive mapping network corresponding to the relationship node weight based on the optimized mapping sequence segment.

[0023] Optionally, the analyzing the sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network comprises:

[0024] time-framing the real-time audio stream to obtain an audio framing group;

[0025] generate a time-frequency feature matrix corresponding to the audio frame group;

[0026] extract a sound source propagation vector in the time-frequency feature matrix;

[0027] input the sound source propagation vector into a path analysis module in the adaptive mapping network to output a path feature sequence;

[0028] analyze a sound source azimuth angle and a propagation delay in the path feature sequence;

[0029] based on the sound source azimuth angle and the propagation delay, construct a sound source propagation path corresponding to the real-time audio stream.

[0030] Optionally, the calculation of the frequency domain coupling coefficients between the multipath effect factors comprises:

[0031] calculating the frequency domain coupling coefficients between the multipath effect factors.

[0032] Optionally, the hierarchical fusion of the frequency domain coupling coefficients to obtain a hybrid feature tensor comprises:

[0033] query a frequency domain feature subset corresponding to the frequency domain coupling coefficients;

[0034] extract a local feature vector of each frequency band of the frequency domain feature subset;

[0035] analyze the distribution characteristics of the local feature vector in a multi-dimensional space;

[0036] analyze the dimension distribution rule corresponding to the distribution characteristics;

[0037] based on the dimension distribution rule, hierarchically fuse the frequency domain coupling coefficients to obtain a hybrid feature tensor.

[0038] Optionally, the query of the adaptive response trajectory of the hybrid feature tensor in the target sound source scene comprises:

[0039] analyze the tensor distribution interval corresponding to the hybrid feature tensor;

[0040] according to the time-frequency domain distribution interval, match a dynamic voiceprint segment in the target sound source scene, and extract a sparse feature cluster in the voiceprint segment;

[0041] map the sparse feature cluster to a sound field propagation path corresponding to the target sound source scene;

[0042] generate an energy attenuation gradient graph corresponding to the sound field propagation path;

[0043] query an adaptive response trajectory in the energy attenuation gradient graph.

[0044] Optionally, the calculating the signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index comprises:

[0045] The signal-to-noise attenuation entropy corresponding to the real-time audio stream is calculated.

[0046] Optionally, the determining the robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy comprises:

[0047] The noise interference feature corresponding to the signal-to-noise attenuation entropy is analyzed;

[0048] The frequency domain segmentation data corresponding to the real-time audio stream is generated based on the noise interference feature;

[0049] The effective segmented audio corresponding to the frequency domain segmentation data is extracted;

[0050] The hierarchical identifier corresponding to the real-time audio stream is determined based on the effective segmented audio;

[0051] The robust recognition level corresponding to the real-time audio stream is determined based on the hierarchical identifier.

[0052] To solve the above problems, the application further provides a sound recognition system based on a neural network, which comprises:

[0053] An architecture construction module is configured to acquire a real-time audio stream in a target sound source scene, extract a time-frequency feature group in the real-time audio stream, analyze a pulse code sequence corresponding to the time-frequency feature group, and construct a sound network architecture corresponding to the real-time audio stream according to the pulse code sequence;

[0054] A network generation module is configured to identify an acoustic feature cluster in the real-time audio stream by using the sound network architecture, match a hierarchical topological relationship in a preset neural calculation framework according to the acoustic feature cluster, optimize a relationship node weight corresponding to the hierarchical topological relationship, and generate an adaptive mapping network corresponding to the relationship node weight;

[0055] A hierarchical fusion module is configured to analyze a sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network, identify a multipath effect factor in the sound source propagation path, calculate a frequency domain coupling coefficient between the multipath effect factors, perform hierarchical fusion on the frequency domain coupling coefficient, and obtain a hybrid feature tensor;

[0056] An attenuation entropy calculation module is configured to query an adaptive response trajectory of the hybrid feature tensor in the target sound source scene, extract a phase distortion feature and a signal-to-noise ratio index in the adaptive response trajectory, and calculate a signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index;

[0057] The scheme formulation module is configured to determine a robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, reconstruct a sound source collection dimension corresponding to the real-time audio stream according to the robust recognition level, and formulate a sound recognition scheme corresponding to the target sound source scene based on the sound source collection dimension.

[0058] Compared with the problems described in the background art, the present application can jointly analyze audio signals from the time domain and the frequency domain by obtaining a real-time audio stream in a target sound source scene and extracting a time-frequency feature group in the real-time audio stream, accurately capture the frequency component and timing change rule of sound, and effectively improve the accuracy of sound recognition in a complex environment. By using the sound network architecture to recognize the acoustic feature cluster in the real-time audio stream, the present application can automatically extract a feature set with semantic association (such as phonemes in speech and frequency combinations in environmental sound) from complex audio signals through hierarchical nonlinear transformation of the neural network, effectively improve the representation ability of sound features and the generalization performance of the recognition system. Further, based on the adaptive mapping network, the present application analyzes the sound source propagation path corresponding to the real-time audio stream, accurately analyzes the multipath effect (such as indoor reverberation path or outdoor obstacle reflection path) of sound in space by means of the dynamic modeling capability of the network on feature association, thereby improving the accuracy of sound source positioning and recognition in a complex environment. Further, by querying the adaptive response trajectory of the mixed feature tensor in the target sound source scene, the present application can capture the dynamic processing process of the model for complex acoustic features such as multipath effect in the scene in real time, improve the stability and accuracy of the model for target sound source recognition, positioning, and other tasks in a specific scene. Finally, based on the signal-to-noise attenuation entropy, the present application determines the robust recognition level of the real-time audio stream, which can intuitively reflect the degree of audio interference and signal quality with quantitative indicators, provide clear reference for the recognition system, and improve the effective recognition ability of the system for audio in a complex environment, thereby ensuring the stability and accuracy of recognition. Therefore, the neural network-based sound recognition method and system provided by the embodiments of the present application can improve the sound recognition accuracy in a complex environment. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 A flowchart of a neural network-based sound recognition method according to an embodiment of the present application is shown in FIG. 1.

[0060] Figure 2 A sound network architecture diagram in a neural network-based sound recognition method according to an embodiment of the present application is shown in FIG. 2.

[0061] Figure 3 A module diagram of a neural network-based sound recognition system according to an embodiment of the present application is shown in FIG. 3.

[0062] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0063] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0064] This application provides a neural network-based voice recognition method. The executing entity of this neural network-based voice recognition method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the neural network-based voice recognition method can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. Example 1:

[0065] Reference Figure 1 The diagram shown is a flowchart illustrating a neural network-based sound recognition method according to an embodiment of the present invention. In this embodiment, the neural network-based sound recognition method includes:

[0066] S1. Obtain the real-time audio stream in the target sound source scene, extract the time-frequency feature group in the real-time audio stream, parse the pulse coding sequence corresponding to the time-frequency feature group, and construct the sound network architecture corresponding to the real-time audio stream based on the pulse coding sequence.

[0067] This invention acquires real-time audio streams from a target sound source scene and extracts time-frequency feature groups from the real-time audio streams. This allows for joint analysis of audio signals in both the time and frequency domains, accurately capturing the frequency components and temporal variation patterns of sound, and effectively improving the accuracy of sound recognition in complex environments.

[0068] The target sound source scene refers to a specific environment or scene that needs to be identified, including target sound sources and surrounding acoustic environment (such as noise, reverberation, multi-source interference, etc.), for example, in the intelligent security scene, monitoring the abnormal sound (such as glass breaking sound) in the shopping mall, the space structure of the shopping mall, the crowd noise, etc. form a target sound source scene; in the industrial scene, detecting machine tool equipment failure, mechanical vibration noise in the workshop, environmental background sound, etc. also belong to this category; the real-time audio stream refers to a continuous sound signal sequence obtained by a microphone or other acquisition device in real time, transmitted and processed in digital form, with strong timeliness and dynamic change characteristics, for example, the user voice instruction stream received by the voice interaction device (such as smart speaker) in real time, or the environmental sound stream continuously acquired by the security monitoring system, which is a real-time audio stream and needs to be analyzed in time to realize fast response; the time-frequency feature group refers to the feature set extracted after time-frequency analysis (such as short-time Fourier transform) of the real-time audio stream, which integrates time domain (time sequence) and frequency domain (frequency component) information, and can represent the dynamic spectral characteristics of sound, for example, the time-frequency feature group of the voice signal can reflect the frequency change trajectory of the phoneme (such as the formant movement of the vowel), and the time-frequency feature group of the environmental sound can reflect the noise frequency distribution and time length characteristics (such as the high-frequency pulse and time domain duration of the car horn), optionally, the real-time audio stream in the target sound source scene can be obtained through an audio acquisition device and a stream processing framework, such as: acquiring the original audio signal through a microphone array, combining GStreamer tools for real-time streaming and packaging, and finally obtaining a standardized real-time audio stream; the time-frequency feature group in the real-time audio stream can be obtained through a digital signal processing algorithm and a feature extraction library, such as: using short-time Fourier transform or mel frequency cepstral coefficient algorithm, with the help of Librosa library for spectrum analysis and feature calculation, and finally obtaining a time-frequency feature group containing energy, frequency domain distribution and other dimensions.

[0069] Further, by analyzing the pulse code sequence corresponding to the time-frequency feature group, the audio data can be effectively compressed and the key time-frequency information can be retained, the storage and transmission costs can be reduced, and at the same time, the discretization characteristics of the pulse code sequence facilitate subsequent pattern recognition and anomaly detection, and the efficiency of acoustic event analysis is improved.

[0070] The pulse coding sequence refers to converting a time-frequency feature group into a discrete pulse signal combination, and a binary code is usually used to represent the energy intensity of different frequency bands. For example, if a time-frequency feature group contains low frequency (0-500Hz), medium frequency (500-2kHz), and high frequency (above 2kHz), the pulse coding can be represented as [1, 0, 1], which represents that the low frequency and high frequency energy is significant at the current time, and the medium frequency is weak. Optionally, the pulse coding sequence corresponding to the time-frequency feature group can be obtained by using a signal coding algorithm and a feature quantization tool, such as using a linear predictive coding (LPC) algorithm, combining the scipy.signal library of Python to segment and quantize the time-frequency features and convert them into binary codes, and finally obtaining the pulse coding sequence representing the energy distribution of different frequency bands.

[0071] Further, the present application constructs a sound network architecture corresponding to the real-time audio stream according to the pulse coding sequence, which can simulate the biological neural network signal transmission mechanism, convert the audio features into discrete coding of neural-like pulses, and enable the network architecture to dynamically adapt to the acoustic feature distribution of different scenes, thereby enhancing the hierarchical representation capability of the network for complex sound signals (such as multi-source aliasing and noise interference).

[0072] The sound network architecture refers to a neural network structure constructed based on the relationship between audio nodes, which simulates the hierarchical association of sound features, and usually includes an input layer, a hidden layer (such as a convolutional layer and a recurrent layer), and an output layer. The nodes correspond to feature processing units, and the connection edges correspond to feature association weights. For example, the sound network architecture for environmental sound classification can extract local frequency patterns in the audio alignment feature map through a convolutional layer, integrate cross-time period features through a fully connected layer, and finally output classification results such as gunshots and alarm sounds.

[0073] As an embodiment of the present application, the construction of the sound network architecture corresponding to the real-time audio stream according to the pulse coding sequence includes: analyzing the frequency domain feature vector in the pulse coding sequence; extracting an audio frame set corresponding to the real-time audio stream; performing time-frequency alignment on the frequency domain feature vector and the audio frame set to generate an audio alignment feature map; identifying the audio node relationship in the audio alignment feature map; and constructing the sound network architecture corresponding to the real-time audio stream according to the audio node relationship.

[0074] The frequency domain feature vector refers to a one-dimensional numerical vector obtained by frequency domain analysis of the pulse coding sequence, representing the energy distribution or frequency characteristics of the sound signal at different frequency components. For example, the frequency domain feature vector of bird chirping sound can highlight the characteristics of high frequency energy concentration (such as peak value at 2-8 kHz), while the frequency domain feature vector of car engine sound may present strong energy distribution at low and medium frequency bands (such as 500 Hz-2 kHz), which is used to distinguish the frequency characteristics of different sound sources. The audio frame set refers to a sequence of continuous short-time segments obtained by dividing the real-time audio stream into fixed time windows (such as every 10 ms). Each segment is called a frame and contains the details of the sound signal in that time period. For example, a 1-second voice audio can be divided into 100 audio frames, each corresponding to a 10-ms voice segment (such as the pronunciation of the word "you" in "hello" can be split into multiple consecutive audio frames, reflecting the transition from initial to final). The audio alignment feature map refers to a two-dimensional image generated by associating the frequency domain feature vector with the audio frame set through time-frequency alignment operation. The horizontal axis is time (audio frame number), the vertical axis is frequency, and the pixel value represents the feature intensity of the corresponding time-frequency point. For example, the audio alignment feature map of speech can present the formant trajectory changing with time (such as the low-frequency formant position of vowel / a / ), similar to spectrogram, which directly shows the time-frequency dynamic characteristics of sound. The audio node relationship refers to the correlation between different time-frequency nodes (pixels) in the audio alignment feature map, such as energy transfer between adjacent frequency components, correlation between features at different time periods, etc. For example, in the audio of piano playing, there is strong correlation between the same frequency nodes of adjacent audio frames (energy continuation of sustained notes), while the audio node relationship of drum sound can be represented as a burst-like association of short-time high-frequency nodes (transient characteristics of impulsive sound sources).

[0075] Further, the analysis of the frequency domain feature vector in the pulse coding sequence can be implemented by a spectrum analysis algorithm, such as: using the FFT function of the NumPy library to calculate the energy distribution of each frequency band by discrete Fourier transform, and finally obtaining the frequency domain feature vector containing amplitude and phase information; the extraction of the audio frame set corresponding to the real-time audio stream can be implemented by a sliding window segmentation technology, such as: using the frame function of the Librosa tool to perform frame processing with a 20ms window length and a 10ms step, and finally obtaining a time-continuous audio frame set; the time-frequency alignment of the frequency domain feature vector and the audio frame set can be implemented by a dynamic time warping algorithm, such as: calculating the optimal alignment path based on the DTW algorithm in the dtaidistance library of Python, and finally obtaining an audio alignment feature map synchronized in time and frequency; the identification of the audio node relationship in the audio alignment feature map can be implemented by a graph neural network model, such as: using the GAT network combined with the PyTorch framework to learn the attention weight between nodes, and finally obtaining the audio node relationship describing the correlation of acoustic events; the construction of the sound network architecture corresponding to the real-time audio stream can be implemented by a topology generation tool, such as: using the NetworkX library to construct a weighted directed graph according to the node relationship, and finally obtaining a sound network architecture reflecting the logic of sound source interaction.

[0076] Specifically, to further intuitively understand the construction process of the sound network architecture in the present application, reference can be made to the image shown in Figure 2 , which is a schematic diagram of the sound network architecture related to the present application. It should be noted that in the present application, Figure 2 , the architecture schematic diagram presented therein is associated with the sound network architecture construction technology described above. The above detailed the complete process from obtaining the real-time audio stream in the target sound source scene, extracting time-frequency feature groups, analyzing pulse coding sequences, to constructing a sound network architecture. The architecture schematic diagram in the figure shows the process of generating a short-time spectrum from scene sound through STFT (Short-Time Fourier Transform), generating a Mel energy spectrum through a Mel filter and window shifting fragments, using a CNN model for two-stage training, extracting features from the intermediate layer of the CNN for random forest training, and finally realizing audio data stream recognition. The schematic diagram is a visual representation of the sound network architecture construction technology described above in terms of specific implementation steps and model use. It is a concrete expression of the theory and method described above in terms of actual operation process and model architecture. It is not limited to showing all the details and relationship analysis of the sound network architecture in different application scenarios.

[0077] In detail, the neural network model corresponding to the sound network architecture can be divided into a response layer, an integration operation layer, and a representation layer, which are three hierarchical structures, wherein the sound audio range that can be received is controlled in 20Hz-20000Hz, with a change of 50Hz as an interval (except the first interval, which is 20Hz-50Hz), a fixed number of neurons are allocated to each interval, and the neurons are activated in sequence.

[0078] The specific division of the response layer and the corresponding number of neurons are as follows:

[0079]

[0080] The spectrum of the sound can be divided into 8 large intervals, each large interval contains a plurality of sub-intervals, and a total of 7480000 neurons are included.

[0081] S2, identifying an acoustic feature cluster in the real-time audio stream by using the sound network architecture, matching a hierarchical topological relationship in a preset neural computing framework according to the acoustic feature cluster, optimizing a relationship node weight corresponding to the hierarchical topological relationship, and generating an adaptive mapping network corresponding to the relationship node weight.

[0082] By using the sound network architecture to identify the acoustic feature cluster in the real-time audio stream, the present application can automatically extract a feature set with semantic association (such as phonemes in speech and frequency combinations in environmental sound) from complex audio signals through hierarchical nonlinear transformation of the neural network, effectively improving the representation ability of the sound feature and the generalization performance of the recognition system.

[0083] The acoustic feature cluster refers to a set of acoustic features with similar attributes or semantic associations extracted by the sound network architecture, which can represent the essential attributes or category features of the sound. For example, in speech recognition, the acoustic feature cluster of the vowel “ / a / ” may include the combination of low-frequency formant (F1 is about 700Hz), medium-frequency energy distribution (such as F2 is about 1200Hz), and time-domain duration; in environmental sound classification, the feature cluster of “car siren” can include high-frequency pulse (2-4kHz), short-time burst energy, and specific frequency modulation mode (such as frequency jump trajectory), which are key feature combinations for distinguishing different sound sources. Optionally, the identification of the acoustic feature cluster in the real-time audio stream can be realized by clustering analysis methods, such as using the K-means algorithm combined with the scikit-learn library for unsupervised clustering of MFCC features, and finally obtaining feature clusters representing different acoustic events.

[0084] Further, the application matches the hierarchical topological relationship in the preset neural computing framework according to the acoustic feature cluster, and optimizes the relationship node weight corresponding to the hierarchical topological relationship, so that the network structure and the hierarchical correlation of the sound features are deeply matched (such as capturing the basic frequency mode at the bottom layer and integrating the semantic features at the high layer), thereby improving the accuracy of feature mapping and the adaptation efficiency of the recognition model to complex scenes.

[0085] The preset neural computing framework refers to a pre-designed neural network structure template, which defines the type, connection mode and data flow direction of the network layer, and is used to support hierarchical calculation of sound features. For example, a framework of "convolution layer-pooling layer-recurrent layer-full connection layer" can be preset, wherein the convolution layer extracts local time-frequency features, and the recurrent layer models the time sequence dependence, which is suitable for speech recognition tasks; or a "self-attention layer-Transformer encoder" framework is preset for processing global feature correlation of long-time audio sequences. The hierarchical topological relationship refers to the hierarchical connection relationship between the network layers and nodes in the preset neural computing framework, which embodies the abstraction process of features from low layer to high layer. For example, in the CNN architecture, the bottom layer convolution layer node captures the basic frequency unit of the audio (such as a single frequency sine wave), the middle layer node combines the low layer features to form a composite mode (such as a formant structure), and the high layer node integrates the cross-layer features to generate semantic level representation (such as "human voice" and "instrument sound" categories), which constitutes a topological structure from low dimension to high dimension. The relationship node weight refers to the strength parameter of the connection between nodes in the neural network, which determines the transmission efficiency of the feature signal. It is optimized through training to adapt to the characteristics of the input data. For example, in a speech recognition network, if a convolution layer node is sensitive to "voiced onset point" features, the connection weight between the node and the next layer node will be enhanced during training, so that the features are preferentially processed during transmission. Conversely, the connection weight of irrelevant noise features is suppressed, thereby improving the representation ability of useful signals. Optionally, the matching of the hierarchical topological relationship in the preset neural computing framework can be realized by a graph matching algorithm, such as using a graph isomorphism network combined with a PyTorch Geometric library to align the hierarchical structure, and finally obtaining a hierarchical topological relationship containing connection relationships. The optimization of the relationship node weight corresponding to the hierarchical topological relationship can be realized by a back propagation algorithm, such as using an Adam optimizer to iteratively update the node parameters in a TensorFlow framework, and finally obtaining the optimal relationship node weight distribution.

[0086] Further, the application generates an adaptive mapping network corresponding to the relationship node weight, dynamically constructs a feature mapping path based on the optimized node connection strength, so that the network can flexibly adjust the signal transmission mode according to the real-time audio characteristics (such as strengthening the propagation link of key acoustic feature clusters), and improves the feature decoupling ability to complex scenes such as multi-source mixing and noise interference.

[0087] The adaptive mapping network refers to a neural network structure dynamically constructed based on an optimized mapping sequence segment, a connection path and a weight of which can be adjusted in real time according to input audio characteristics, for example, in a noisy environment, the network automatically enhances the weight of a "noise removal path" (such as a convolution layer + a noise removal autoencoder) and suppresses an "interference path"; in a multi-sound source scene, a "sound source separation path" (such as a recurrent layer + an attention mechanism) is activated to realize adaptive feature processing on different scenes and improve recognition robustness.

[0088] As an embodiment of the present application, the adaptive mapping network corresponding to the relationship node weight is generated by analyzing the topological association code in the relationship node weight, matching a preset weight mapping rule library based on the topological association code to obtain an initial mapping path set, dynamically normalizing the initial mapping path set to obtain a normalized path sequence, removing conflict sequence segments in the normalized path sequence to obtain an optimized mapping sequence segment, and generating the adaptive mapping network corresponding to the relationship node weight based on the optimized mapping sequence segment.

[0089] The topological correlation coding refers to the digital coding of the connection relationship between network layers implied in the relationship node weight, representing the dependence strength and transmission direction between nodes. For example, in a recurrent neural network (RNN), the topological correlation coding can record the weight values of the current time node and the historical time node in binary or vector form to represent "strong connection" (e.g., weight > 0.8 marked as 1) or "weak connection" (e.g., weight < 0.2 marked as 0), directly reflecting the association mode of the time sequence characteristics. The pre-defined weight mapping rule library refers to a set of rules that guide the association between relationship node weight and feature mapping path, including network layer connection strategy, weight distribution logic and path generation constraint conditions. For example, the rule library can include rules such as "when the weight of a certain node is > 0.7, preferentially connect to the attention mechanism layer", "low-frequency feature path should contain at least 2 convolution layers", etc. In speech emotion recognition, exclusive rules such as "emotion feature path needs to pass through bidirectional LSTM layer" can also be set to ensure that the generated mapping path meets the specific task requirements and provides an interpretable decision basis for adaptive network construction. The initial mapping path set refers to the preliminary feature mapping path set generated by matching the topological correlation coding with the pre-defined rule library, containing multiple possible signal transmission paths. For example, in the image sound recognition scene, the pre-defined rule library can define rules such as "high-frequency features prefer to pass through convolution layer A → pooling layer B" and "low-frequency features prefer to pass through convolution layer C → attention layer D". The initial mapping path set contains all path combinations that meet the rules (e.g., A → B → fully connected layer, C → D → fully connected layer, etc.). The normalized path sequence refers to the ordered sequence obtained by performing weight normalization on each path in the initial mapping path set, making the feature transmission strength of different paths comparable. For example, if the total weight of two paths is 3.5 and 2.1 respectively, the normalized weight ratio is 5:3, ensuring that the network does not cause feature imbalance when processing multiple paths in parallel due to the difference in original weight scale, and improving the collaboration between paths. The optimized mapping sequence segment refers to the efficient mapping sequence formed by removing conflicting or redundant paths from the normalized path sequence. For example, if two paths transmit the same frequency feature but in opposite directions (e.g., one path enhances high frequency and the other path suppresses high frequency), it is determined as a conflict sequence segment and one of them is deleted. If multiple paths have redundant functions (e.g., all processing low-frequency background noise), the representative path is retained and the redundancy is removed, thereby reducing calculation redundancy and improving feature mapping efficiency.

[0090] Further, the parsing of the topological correlation coding in the relationship node weight can be realized by a graph neural network embedding method, such as: using the GraphSAGE algorithm combined with the DGL framework to extract the potential correlation characteristics between nodes, and finally obtaining the topological correlation coding representing the connection strength; the matching of the preset weight mapping rule library can be realized by a graph pattern matching technology, such as: using the Cypher query language of the Neo4j graph database for rule matching, and finally obtaining the initial mapping path set meeting the constraint conditions; the dynamic normalization of the initial mapping path set can be realized by an adaptive normalization algorithm, such as: based on the Min-Max Scaling method to realize dynamic range adjustment in PyTorch, and finally obtaining the normalized normalized path sequence; the removal of the conflict sequence in the normalized path sequence can be realized by a constraint satisfaction algorithm, such as: applying a CSP solver combined with an OR-Tools library to detect and eliminate contradictory path segments, and finally obtaining the conflict-free optimized mapping sequence segment; the generation of the adaptive mapping network corresponding to the relationship node weight can be realized by a self-organizing mapping method, such as: using the Kohonen network combined with TensorFlow to realize dynamic topology adjustment, and finally obtaining the mapping network structure with adaptive ability.

[0091] Specifically, if the sound volume received in the adaptive mapping network is too high, the activity of neurons in the adaptive mapping network needs to be suppressed to adapt to the high-pitched environment. The neurons adapted to the high-pitched environment are marked as alert source neurons, and the total number of alert source neurons is set to 10000. These 10000 neurons will release an influence on the activity of the second layer neurons when excited, and the influence threshold of the neurons is as follows:

[0092] High volume decibel number / decibel Number of alerted neurons / neuron Second layer neuron threshold coefficient / threshold [80~90) 2000 80% [90~100) 5000 90% ≥100 10000 95%

[0093] When the adaptive mapping network encounters a sound with excessively high volume, suppressing the activity of neurons is a key coping strategy. In addition to the established alert source neuron setting, the regulation of the neuron refractory period is also involved. All neurons are uniformly set with a 6-second refractory period, that is, after continuous discharge, a 6-second interval is required before responding to external stimuli again. This is to prevent neurons from being excessively fatigued or disordered due to continuous high-frequency stimulation, and to ensure that the entire network can still operate in an orderly manner in a high-pitched environment.

[0094] S3, based on the adaptive mapping network, analyzing the sound source propagation path corresponding to the real-time audio stream, identifying the multipath effect factor in the sound source propagation path, calculating the frequency domain coupling coefficient between the multipath effect factors, and performing hierarchical fusion on the frequency domain coupling coefficient to obtain a mixed feature tensor.

[0095] Based on the adaptive mapping network, the sound source propagation path corresponding to the real-time audio stream is analyzed, and with the dynamic modeling capability of the network for feature association, the multipath effect (such as indoor reverberation path or outdoor obstacle reflection path) of sound in space can be accurately analyzed, so that the accuracy of sound source positioning and identification in complex environment is improved.

[0096] The sound source propagation path refers to a complete transmission path set of sound from a sound source to a collection device, including a direct sound path and a plurality of reflection and diffraction paths, each path corresponding to a specific azimuth angle, propagation delay and energy attenuation characteristic, for example, in an outdoor scene, a car horn sound may be directly transmitted to the microphone through the ground (path 1), and at the same time, it is reflected by the front building to arrive (path 2), forming a propagation model containing two paths, which is used to analyze the influence of multipath interference on the identification result.

[0097] As an embodiment of the present application, based on the adaptive mapping network, the sound source propagation path corresponding to the real-time audio stream is analyzed, including: time framing the real-time audio stream to obtain an audio framing group; generating a time-frequency feature matrix corresponding to the audio framing group; extracting a sound source propagation vector in the time-frequency feature matrix; inputting the sound source propagation vector into a path analysis module in the adaptive mapping network to output a path feature sequence; analyzing the sound source azimuth and propagation delay in the path feature sequence; based on the sound source azimuth and the propagation delay, the sound source propagation path corresponding to the real-time audio stream is constructed.

[0098] The audio frame group refers to a set of continuous short-time segments obtained by dividing a real-time audio stream into fixed time intervals (such as 10-30 ms), each segment corresponding to a frame of audio signal, facilitating fine analysis of the timing dynamic characteristics of sound. For example, a 2-second environmental audio can be divided into 200 audio frames (with an interval of 10 ms), each frame corresponding to a segment of a short-time sound event such as vehicle horn or pedestrian footsteps. The time-frequency feature matrix is a two-dimensional matrix generated after time-frequency transformation (such as STFT) of the audio frame group, with the horizontal axis representing time (frame number) and the vertical axis representing frequency. The matrix element value represents the signal energy or amplitude of the corresponding time-frequency point, similar to the structure of a spectrogram. For example, the time-frequency feature matrix of speech can show the trajectory of vowel formant changes over time, while the matrix of factory noise can present a continuous distribution of medium and low frequency energy. The sound source propagation vector is a one-dimensional vector extracted from the time-frequency feature matrix, representing the physical characteristics of sound propagation, including frequency attenuation coefficient, phase change rate, etc. For example, in an indoor scene, the sound source propagation vector can reflect the energy attenuation (such as faster attenuation of high frequency components) and phase delay (related to reflection distance) of sound waves after reflection by walls, which can be used to distinguish between direct sound and reflected sound characteristics. The path analysis module is a sub-network structure in the adaptive mapping network specifically used to analyze the propagation characteristics of sound sources, usually including convolution layers, fully connected layers or graph neural network layers, which can extract features and model paths for sound source propagation vectors. For example, this module can capture local propagation patterns (such as burst reflection sound energy) in the time-frequency feature matrix through convolution layers, and then output parameter predictions related to the path through fully connected layers. The path feature sequence is a time sequence of features output by the path analysis module after processing the sound source propagation vector, including the propagation path attributes (such as the number of multipath and path energy proportion) corresponding to each audio frame. For example, in a multi-source mixed scene, the path feature sequence can mark whether the current frame signal mainly comes from the direct sound path (70% proportion) or the reflection path (30% proportion) frame by frame, and record the frequency component differences of each path. The sound source azimuth angle is the horizontal direction angle of the sound source relative to the microphone array or the collection device (usually with 0° as the front direction and clockwise as positive), which is used to locate the spatial position of the sound source. For example, in the far-field speech recognition of a smart speaker, the sound source azimuth angle can indicate the direction of the user's speech (such as 45° to the left), which can assist beamforming technology to enhance the target signal. The propagation delay is the time difference of sound propagation from the sound source to the collection device through different paths, mainly caused by the difference in propagation distance (such as the difference in distance between direct sound and reflected sound). For example, in an indoor environment, the reflection path is 5 meters longer than the direct sound path, resulting in a propagation delay of about 15 ms (sound speed is about 340 m / s). This parameter can be used for multipath effect modeling and sound source positioning algorithms.

[0099] Further, the time framing of the real-time audio stream can be achieved by a sliding window segmentation technique, such as using the Librosa library of Python to call the frame function for framing processing with a window length of 25 ms and a step length of 10 ms, and finally obtaining a time-continuous audio framing group; the generation of the time-frequency feature matrix corresponding to the audio framing group can be achieved by short-time Fourier transform, such as using the stft function of the scipy.signal library to calculate the spectral energy distribution of each frame of signal, and finally obtaining a feature matrix containing time-frequency information; the extraction of the sound source propagation vector in the time-frequency feature matrix can be achieved by a direction of arrival estimation algorithm, such as calculating the phase difference information of each frequency point based on the MUSIC algorithm through NumPy, and finally obtaining a propagation vector representing the direction of the sound source; the input of the sound source propagation vector into the path analysis module in the adaptive mapping network can be achieved by tensor operation, such as using the Dense layer of TensorFlow for feature space transformation, and finally obtaining a path feature sequence with time sequence relationship; the analysis of the sound source azimuth and propagation delay in the path feature sequence can be achieved by a geometric acoustics model, such as applying the GCC-PHAT algorithm combined with the geometric parameters of the microphone array to calculate the time delay difference, and finally obtaining accurate sound source azimuth and propagation delay parameters; the construction of the sound source propagation path corresponding to the real-time audio stream can be achieved by a ray tracing method, such as using the PyRoomAcoustics library to simulate the reflection path of sound waves in the environment, and finally obtaining a complete sound source propagation path containing direct sound and reflected sound.

[0100] By identifying the multipath effect factor in the sound source propagation path, the application can accurately capture the multi-path interference characteristics (such as time delay, attenuation, phase difference) of sound caused by reflection and diffraction in a complex environment, thereby improving the anti-interference ability of the sound recognition system in reverberation, shielding and other scenes, and reducing the recognition deviation caused by multipath interference.

[0101] The multipath effect factor refers to a set of physical parameters describing the changes in signal characteristics caused by multiple paths (direct, reflection, diffraction, etc.) during sound propagation. These parameters include path delay (time difference between different paths), energy attenuation coefficient (path loss), phase shift (waveform distortion), and the number of paths. For example, in an indoor scene, the path of a speech signal reflected from the ceiling is 3 meters longer than the direct path, resulting in a delay of about 9ms and a 15dB high-frequency energy attenuation. These delay and attenuation values ​​are the multipath effect factor. In an outdoor scene, the sound of a car horn may experience a phase shift (such as a 20° waveform phase rotation) after diffraction through a wall, which also falls under the category of this factor. Optionally, the identification of the multipath effect factor in the sound source propagation path can be achieved through a multipath signal decomposition algorithm, such as using blind source separation technology combined with the FastICA method to extract each path component from the mixed signal, ultimately obtaining the multipath effect factor containing delay and attenuation characteristics.

[0102] Furthermore, by calculating the frequency domain coupling coefficient between the multipath effect factors, this invention can quantify the interaction strength of signals from different propagation paths in the frequency domain (such as the degree of interference between reflected sound and direct sound at a specific frequency), thereby improving the feature resolution accuracy in multipath environments.

[0103] The frequency domain coupling coefficient refers to a value that measures the degree of correlation between multipath effect factors in the frequency domain. The larger the value, the stronger the interaction between multipaths in the frequency domain. For example, in an indoor acoustic environment, multiple reflected path signals interfere with each other in certain frequency bands. The frequency domain coupling coefficient can quantify the degree of this interference and help optimize audio processing algorithms.

[0104] As an embodiment of the present invention, the calculation of the frequency domain coupling coefficient between the multipath effect factors includes:

[0105] The frequency domain coupling coefficients between the multipath effect factors are calculated using the following formula: in, This represents the frequency domain coupling coefficient between the multipath effect factors. This indicates the number of samplings corresponding to the multipath effect factor. This represents the sampling index corresponding to the multipath effect factor. Indicates the first Multipath effect vector of subsampled This represents the mean of the multipath effect vector across all samples. Represents the smoothing coefficient. Indicates the highest analysis frequency. Represents frequency The multipath effect function at the location, This represents the rate of change of frequency corresponding to the multipath effect function.

[0106] Further, the multipath effect vector refers to a vector describing the characteristics of the multipath effect at a certain sampling time, containing parameters such as time delay, attenuation, phase, etc. For example, for outdoor car sound propagation modeling, the multipath effect vector can represent the characteristic combination of direct sound and reflected sound in time delay and energy attenuation at a certain time; the multipath effect vector mean refers to the average characteristics of the multipath effect vector under multiple samplings, such as collecting the signal of a certain sound source in a room multiple times, calculating the multipath effect vector mean, which can reflect the average state of the multipath effect in this environment, and is used for analyzing common propagation characteristics; the smoothing coefficient refers to a parameter adjusting the influence degree of the frequency change rate on the frequency domain coupling coefficient, which is between 0 and 1, close to 0, and weakens the influence of the frequency change rate; close to 1, and enhances the influence, for example, in the audio noise reduction scene, the smoothing coefficient can be adjusted according to the environment to control the noise reduction strength; the highest analysis frequency refers to the highest frequency value considered in audio analysis, for example, speech analysis is usually taken as 4kHz, which means that only the multipath effect in the frequency band of 0-4kHz is analyzed, which determines the analysis frequency band range and avoids high-frequency noise interference; the multipath effect function refers to a function taking frequency as the independent variable, describing the characteristics of the multipath effect changing with frequency, for example, in the study of indoor reverberation, the multipath effect function can show the changes of reflected sound energy, phase, etc. with frequency at different frequencies; the frequency change rate refers to the speed of the multipath effect function changing with frequency, for example, in the high frequency band, the phase change of sound multipath propagation is fast, and the frequency change rate is large; in the low frequency band, it is relatively flat, and the frequency change rate is small, reflecting the dynamic characteristics of the multipath effect in the frequency domain.

[0107] The application can integrate the frequency domain correlation information of multipath effects at different levels by layering and fusing the frequency domain coupling coefficients to obtain a mixed feature tensor, enrich the feature dimension (such as from the bottom layer features such as path time delay and attenuation to the high layer features such as the overall interference mode), and improve the accuracy of sound recognition, positioning and other tasks.

[0108] The mixed feature tensor refers to a multi-dimensional data structure obtained by layering and fusing the frequency domain coupling coefficients according to the dimension distribution rule, for example, in a three-dimensional tensor, different dimensions represent different frequency bands, multipath types and coupling strengths, etc. The frequency domain characteristics of the multipath effect can be comprehensively integrated, and the input of the deep learning model can be used.

[0109] As an embodiment of the application, the layering and fusion of the frequency domain coupling coefficients to obtain a mixed feature tensor comprises: querying the frequency domain feature subset corresponding to the frequency domain coupling coefficient; extracting the local feature vector of each frequency band of the frequency domain feature subset; analyzing the distribution characteristics of the local feature vector in the multi-dimensional space; analyzing the dimension distribution rule corresponding to the distribution characteristics; based on the dimension distribution rule, layering and fusing the frequency domain coupling coefficients to obtain a mixed feature tensor.

[0110] The frequency domain feature subset refers to a feature set of a specific frequency band range divided from the entire frequency domain feature. For example, in audio processing, a speech frequency band (300-3400 Hz) can be used as a frequency domain feature subset for focusing on analyzing the frequency domain coupling of the speech-related multipath effect factor, which is different from the features of a high-frequency noise frequency band. The local feature vector refers to a vector extracted for each frequency band in the frequency domain feature subset, which can reflect the characteristics of the frequency band. For example, in a certain narrow frequency domain, a vector containing the energy, phase, and other information of the multipath effect factor of the frequency band is extracted, such as a vector extracted in a 2000-2100 Hz frequency band, which describes the multipath characteristics of sound propagation in this frequency band. The multidimensional space refers to an abstract space composed of multiple feature dimensions, which is used to represent the local feature vector. For example, a space is constructed with the time delay, attenuation, phase, and other characteristics of the multipath effect factor as dimensions. Each local feature vector corresponds to a point in the space. For example, a two-dimensional plane can represent the time delay and attenuation dimensions, which can intuitively display the vector position. The distribution characteristics refer to the distribution rules and characteristics of the local feature vector in the multidimensional space. For example, in the multidimensional space, some local feature vectors may be concentrated in a specific area, reflecting the concentration trend of the multipath effect under certain characteristic combinations. For example, the local feature vectors of indoor reflected sound are concentrated in an area with small time delay and moderate attenuation. The dimension distribution rule refers to the relationship and distribution criteria between dimensions summarized according to the distribution characteristics of the local feature vector. For example, it is found that there is a negative correlation between the time delay dimension and the attenuation dimension, and they are concentrated in a specific range. Based on this, a rule is formulated to guide the subsequent feature fusion, such as the rule that small time delay corresponds to not too large attenuation.

[0111] Further, the query of the frequency domain feature subset corresponding to the frequency domain coupling coefficient can be realized by a frequency band segmentation algorithm, such as: using wavelet packet decomposition combined with the PyWavelets library to divide the signal into different frequency bands, and finally obtaining a frequency domain feature subset containing the energy of a specific frequency band; the extraction of the local feature vector of each frequency band of the frequency domain feature subset can be realized by time-frequency analysis technology, such as: using short-time Fourier transform to calculate the spectral centroid and bandwidth of each frequency band through the scipy.signal library, and finally obtaining a feature vector representing local characteristics; the analysis of the distribution characteristics of the local feature vector in a multi-dimensional space can be realized by manifold learning method, such as: applying t-SNE algorithm to perform high-dimensional feature visualization dimension reduction with the help of scikit-learn library, and finally obtaining distribution characteristics reflecting data aggregation characteristics; the analysis of the dimension distribution rule corresponding to the distribution characteristics can be realized by clustering analysis method, such as: using Gaussian mixture model to fit the feature distribution probability density through the sklearn.mixture library, and finally obtaining a dimension distribution rule describing the association rule of each dimension; the hierarchical fusion of the frequency domain coupling coefficient can be realized by tensor fusion technology, such as: using TensorLy library based on Tucker decomposition to realize multi-level feature interaction, and finally obtaining a mixed feature tensor that preserves the characteristics of each dimension.

[0112] S4, query the adaptive response trajectory of the mixed feature tensor in the target sound source scene, extract the phase distortion feature and signal-to-noise ratio index in the adaptive response trajectory, and calculate the signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index.

[0113] By querying the adaptive response trajectory of the mixed feature tensor in the target sound source scene, the present application can capture the dynamic processing process of the model for complex acoustic characteristics such as multipath effect in the scene in real time, and can improve the stability and accuracy of the model in specific scenes for tasks such as target sound source identification and positioning.

[0114] Among them, the adaptive response trajectory refers to a trajectory in the energy attenuation gradient graph that reflects the adaptive adjustment process of the model for sound propagation changes in the sound field, for example, the response path of the model for energy attenuation changes caused by sound source movement and environmental changes at different times.

[0115] As an embodiment of the present application, the query of the adaptive response trajectory of the mixed feature tensor in the target sound source scene comprises: analyzing the tensor distribution interval corresponding to the mixed feature tensor; matching a dynamic voiceprint segment in the target sound source scene according to the time-frequency domain distribution interval, and extracting a sparse feature cluster in the voiceprint segment; mapping the sparse feature cluster to a sound field propagation path corresponding to the target sound source scene; generating an energy attenuation gradient graph corresponding to the sound field propagation path; and querying the adaptive response trajectory in the energy attenuation gradient graph.

[0116] The tensor distribution interval refers to a range set of values of the mixed feature tensor in each dimension, for example, in a three-dimensional mixed feature tensor, corresponding to the frequency band, multipath type, and coupling strength dimensions, respectively, and the value range constitutes the tensor distribution interval, such as the frequency band dimension being 20-20000 Hz, the multipath type dimension having direct sound and first-order reflected sound, etc., defining the range of the tensor feature; the dynamic voiceprint segment refers to a voiceprint information segment that changes over time in the target sound source scene. For example, voiceprint data sequences generated by different sentences and intonation changes when a person is speaking, which contain the unique timbre and frequency change characteristics of the speaker, are effective information parts that change dynamically in voiceprint recognition; the sparse feature cluster refers to a feature set that is sparse in feature distribution but representative, extracted from the dynamic voiceprint segment, for example, in a voiceprint segment, the energy value of certain key frequency points, phase mutation at a specific time, etc., are relatively small in number but can reflect the uniqueness of the voiceprint, similar to the sparse corner point features in an image; the sound field propagation path refers to the route of sound propagation in the target sound source scene space, for example, in a conference room, the path of the speaker's voice reflected to the listener's position through the wall, ceiling, etc., including the direct path and the multiple reflection path, which determines the propagation delay, attenuation, etc. of the sound; the energy attenuation gradient graph refers to a graph that depicts the change of sound energy with the attenuation of the propagation path based on the sound field propagation path, for example, in an indoor scene, the attenuation degree of sound energy on different paths is marked by color depth or numerical value, which intuitively shows the decreasing trend of energy in each propagation direction from the sound source, facilitating the analysis of the propagation characteristics.

[0117] Further, the parsing of the tensor distribution interval corresponding to the mixed feature tensor can be realized by a tensor decomposition method, such as: using CP decomposition combined with TensorLy library to extract the distribution range of each modal feature, and finally obtaining the interval set describing the distribution of the tensor data; the matching of the dynamic voiceprint segment in the target sound source scene can be realized by a dynamic time warping algorithm, such as: using the dtw-package of Python to calculate the similarity path of the voiceprint template and the real-time signal, and finally obtaining the best matching voiceprint segment; the extraction of the sparse feature cluster in the voiceprint segment can be realized by a sparse coding technology, such as: applying the KSVD algorithm to learn the sparse representation of the voiceprint feature through the scikit-learn library, and finally obtaining the discriminative sparse feature cluster; the mapping of the sparse feature cluster to the sound field propagation path corresponding to the target sound source scene can be realized by acoustic ray tracing, such as: using the PyRoomAcoustics library to simulate the propagation trajectory of the feature cluster in the environment, and finally obtaining the sound field propagation path containing the reflection path; the generation of the energy attenuation gradient graph corresponding to the sound field propagation path can be realized by an acoustic energy attenuation model, such as: calculating the energy attenuation rate of each point on the path based on the acoustic wave propagation equation through SciPy, and finally obtaining the visualized attenuation gradient graph; the query of the adaptive response trajectory in the energy attenuation gradient graph can be realized by a path optimization algorithm, such as: using Dijkstra algorithm to find the optimal energy transmission path in the gradient graph, and finally obtaining the adaptive response trajectory conforming to the acoustic characteristics.

[0118] By extracting the phase distortion features and signal-to-noise ratio indexes in the adaptive response trajectory, the present application can accurately quantify the waveform distortion degree (such as phase rotation, waveform broadening) and signal purity caused by multipath interference in the sound propagation process, effectively improving the intelligibility of sound signals and the robustness of recognition systems in complex environments.

[0119] The phase distortion feature refers to a characteristic that a sound signal deviates from an ideal state due to factors such as multipath effect and medium non-uniformity during propagation, and is manifested as nonlinear distortion of a phase spectrum, phase mutation or abnormal phase difference, for example, the phase difference between indoor reflected sound and direct sound may cause phase cancellation of a synthesized signal at a specific frequency point (such as energy attenuation when the phase difference is 180°), and this phase anomaly can be quantified by a phase distortion feature, which is used to evaluate the damage degree of multipath interference on a signal waveform; the signal-to-noise ratio index refers to a ratio index of target sound signal energy and background noise energy, which is used to measure the purity of an audio signal, for example, in a vehicle voice interaction scene, if the engine noise reduces the signal-to-noise ratio index of the microphone collected signal from 20 dB to 10 dB, it indicates that the noise energy has increased significantly, which may cause the voice recognition error rate to rise; by monitoring the index in real time, the signal quality can be optimized by dynamically adjusting the noise reduction algorithm parameters, and optionally, the extraction of the phase distortion feature and the signal-to-noise ratio index in the adaptive response trajectory can be realized by signal analysis technology, such as: using Hilbert transform combined with Python's scipy.signal library to calculate the instantaneous phase shift, and using the power spectral density estimation method to obtain the signal-to-noise energy ratio, and finally obtaining the phase distortion feature representing signal distortion and the accurate signal-to-noise ratio index.

[0120] Further, based on the phase distortion feature and the signal-to-noise ratio index, the signal-to-noise attenuation entropy corresponding to the real-time audio stream is calculated, which can quantify the uncertainty of the signal under noise interference by information entropy theory, so as to dynamically optimize the signal processing strategy in a complex acoustic environment and improve the separability of the target sound and the robustness of the recognition system.

[0121] The signal-to-noise attenuation entropy refers to a quantitative index that comprehensively measures the influence degree of noise interference and phase distortion on the signal in the real-time audio stream, which integrates the phase distortion feature and the signal-to-noise ratio index related parameters to reflect the uncertainty and quality loss of the audio signal in a complex environment by the concept of information entropy. The greater the value, the more serious the signal interference and the more obvious the quality decline.

[0122] As an embodiment of the present application, the calculation of the signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index comprises:

[0123] The signal-to-noise attenuation entropy corresponding to the real-time audio stream is calculated by the following formula: wherein, represents the signal-to-noise attenuation entropy corresponding to the real-time audio stream, represents the total number of dimensions of the distortion dimension corresponding to the phase distortion feature, represents the dimension index corresponding to the distortion dimension, represents the first a distortion quantization value corresponding to the distortion dimension, a distortion weight corresponding to the distortion dimension, a distortion weight corresponding to the distortion dimension, a dimension total number representing the index dimension, a dimension total number representing the index dimension, a distortion weight corresponding to the distortion dimension, a distortion weight corresponding to the distortion dimension, a distortion weight corresponding to the distortion dimension, a distortion weight corresponding to the distortion dimension, a time length selected when analyzing the real-time audio stream, a decay influence function at time t.

[0124] In detail, the distortion dimension refers to a dimension for describing different aspects or attributes of the phase distortion characteristics, for example, can include the degree of non-linear phase shift, the frequency range of phase mutation, the amplitude of phase fluctuation, etc. The distortion dimension describes the phase distortion from multiple angles to analyze the distortion of the signal more comprehensively. The distortion quantization value refers to a numerical value for quantifying the degree of phase distortion in a specific distortion dimension. The distortion quantization value converts the actual situation of the phase distortion into a specific numerical value, which is convenient for calculation and analysis in the above formula. The distortion weight refers to a parameter reflecting the relative importance of different distortion dimensions in calculating the signal-to-noise attenuation entropy. Different distortion dimensions have different effects on the quality of the audio signal. By assigning appropriate weights, the formula can more reasonably consider the role of each dimension. The index dimension refers to a dimension for describing different attributes or components of the signal-to-noise ratio, for example, can include the signal-to-noise ratio of different frequency bands, the signal-to-noise ratio change at different times in the time domain, the influence of noise type on the signal-to-noise ratio, etc. The index dimension analyzes the signal-to-noise ratio from multiple perspectives. The signal-to-noise ratio quantization value refers to a numerical value for quantifying the signal-to-noise ratio in a specific index dimension. The signal-to-noise ratio quantization value converts the actual signal-to-noise ratio into a specific numerical value to accurately measure the degree of noise interference of the signal in that dimension. The signal-to-noise ratio weight refers to a weighting coefficient assigned to each index dimension according to the importance of the index dimension in affecting the quality of the audio signal. If the time-domain signal-to-noise ratio fluctuation has a more critical effect on the accuracy of speech recognition in a specific environment, the signal-to-noise ratio weight corresponding to the time domain will be set higher to highlight its role in the calculation of the signal-to-noise attenuation entropy. The time length refers to a time length selected for calculating the signal-to-noise attenuation entropy when analyzing real-time audio streams. The time length determines the time range of the calculation, for example, selecting 10 seconds of audio stream data for analysis. The 10 seconds is the time length. Different time length settings may affect the final signal-to-noise attenuation entropy result. The attenuation influence function refers to a function of time t, which is used to describe the attenuation effect of the audio signal at different times. For example, in an indoor environment with echo, the sound signal will attenuate over time due to factors such as reflection. A(t) can be established based on the reflection path, attenuation coefficient, etc. to quantitatively describe the attenuation degree of the signal at each time.

[0125] S5, based on the signal-to-noise attenuation entropy, determining the robust recognition level of the real-time audio stream, reconstructing the sound source acquisition dimension corresponding to the real-time audio stream according to the robust recognition level, and formulating a sound recognition scheme corresponding to the target sound source scene based on the sound source acquisition dimension.

[0126] Based on the signal-to-noise attenuation entropy, the application determines the robust recognition level of the real-time audio stream, which can intuitively reflect the degree of audio interference and signal quality with quantitative indicators, providing clear reference for the recognition system, and improving the effective recognition ability of the system in complex environments, ensuring the stability and accuracy of the recognition.

[0127] The robust recognition level refers to a classification level of the degree of accurate recognition of the real-time audio stream in a noise environment, for example, four levels: Level 1 (excellent): the signal-to-noise attenuation entropy is extremely low (such as <0.3 bits), the effective segment proportion is >80%, and the recognition accuracy is >95%; Level 2 (good): the entropy value is medium (0.3-0.6 bits), the effective segment proportion is 60%-80%, and the recognition accuracy is 80%-95%; Level 3 (medium): the entropy value is relatively high (0.6-0.9 bits), the effective segment proportion is 40%-60%, and the algorithm needs to be enhanced to assist in recognition; and Level 4 (poor): the entropy value is extremely high (>0.9 bits), the effective segment proportion is <40%, and the recognition accuracy is <50%, and manual intervention is needed, which is used for dynamically adjusting the recognition strategy (such as starting a multi-microphone array beamforming when Level 4) to improve the adaptability of the system.

[0128] As an embodiment of the present application, the robust recognition level of the real-time audio stream is determined based on the signal-to-noise attenuation entropy, which includes: analyzing the noise interference features corresponding to the signal-to-noise attenuation entropy; generating the frequency domain segment data corresponding to the real-time audio stream based on the noise interference features; extracting the effective segment audio corresponding to the frequency domain segment data; determining the hierarchical identifier corresponding to the real-time audio stream based on the effective segment audio; and determining the robust recognition level corresponding to the real-time audio stream based on the hierarchical identifier.

[0129] The noise interference feature refers to a set of characteristic parameters reflecting the influence of noise on the audio signal, which is parsed from the signal-to-noise attenuation entropy. For example, through the frequency distribution characteristics of the signal-to-noise attenuation entropy, the center frequency of the noise (such as the interference peak at 2 kHz), the bandwidth (the frequency range covered by the noise), the energy distribution mode (such as Gaussian noise or impulse noise), and other parameters can be extracted. In the time domain, the duration of the noise and the burst frequency can be obtained. The frequency domain segmented data refers to a set of data formed after the frequency spectrum of the real-time audio stream is divided into multiple frequency bands according to a specific rule. For example, based on the frequency distribution of the noise interference feature, the audio frequency spectrum is divided into a low-noise frequency band (such as 0-500Hz, noise energy <10dB), a medium-noise frequency band (500-2kHz, noise energy 10-20dB), and a high-noise frequency band (>2kHz, noise energy >20dB). The effective segmented audio refers to an audio segment that meets the identification requirements of the signal-to-noise ratio, which is selected from the frequency domain segmented data. For example, in the speech recognition scenario, if the signal-to-noise ratio of a certain frequency band is higher than a preset threshold (such as 15dB) and the phase distortion is lower than a critical value (such as a phase deviation <30°), it is determined that the frequency band is an effective segment. The effective segmented audio retains the key features of the original signal (such as the formant structure of speech), which can be used for subsequent hierarchical determination and exclusion of invalid frequency band interference dominated by noise. The hierarchical identifier refers to a digital or symbolic identifier calculated according to the characteristics of the effective segmented audio, which is used to represent the robustness of the audio. For example, by integrating the proportion of effective segments (such as 70% of the total frequency band), the average signal-to-noise ratio (such as 20dB), and the phase stability (such as a phase variance <0.5), a multi-dimensional vector is generated as a hierarchical identifier (such as [0.7, 20, 0.5]). The identifier is associated with the robust recognition level through a preset mapping rule (such as a decision tree classifier), realizing automatic classification.

[0130] Furthermore, the analysis of the noise interference features corresponding to the signal-to-noise attenuation entropy can be achieved through entropy decomposition algorithms, such as using spectral entropy analysis combined with Python's librosa library to calculate the noise energy ratio of each frequency band, ultimately obtaining noise interference features characterizing the interference intensity; the generation of frequency domain segmented data corresponding to the real-time audio stream can be achieved through time-frequency segmentation techniques, such as using short-time Fourier transform with the scipy.signal library to divide the spectrum with a fixed frame length, ultimately obtaining time-frequency aligned frequency domain segmented data; the extraction of effective segmented audio corresponding to the frequency domain segmented data can be achieved through energy threshold filtering, such as using a dynamic threshold method with NumPy to remove low-energy noise segments, ultimately obtaining clean and effective segmented audio; the determination of the hierarchical identifier corresponding to the real-time audio stream can be achieved through clustering analysis methods, such as using the K-means algorithm with the scikit-learn library to automatically classify according to the signal-to-noise ratio feature, ultimately obtaining a digitized hierarchical identifier; the determination of the robust recognition level corresponding to the real-time audio stream can be achieved through a decision tree model, such as using an XGBoost classifier to comprehensively evaluate the signal-to-noise ratio and distortion index, ultimately obtaining a quantified robust recognition level.

[0131] Based on the robust recognition level, this invention reconstructs the sound source acquisition dimension corresponding to the real-time audio stream. Depending on the audio interference and the difficulty of recognition, the acquisition configuration can be optimized (e.g., increasing the number of microphones or adjusting the array layout at lower levels). This can effectively enhance the sound source signal strength, suppress noise interference, improve audio acquisition quality, and provide better data for subsequent recognition, analysis, and other processing.

[0132] The sound source acquisition dimension refers to the dimensional system composed of various factors involved in acquiring sound source signals, covering aspects such as space, time, and frequency. For example, in the spatial dimension, the layout (linear, circular, etc.) and spacing of the microphone array determine the ability to perceive the location of the sound source; the temporal dimension includes sampling frequency and sampling duration, affecting the temporal resolution of the signal; the frequency dimension involves the frequency range of acquisition, such as the 300-3400Hz frequency band often focused on in speech acquisition. Optionally, the sound source acquisition dimension corresponding to the reconstructed real-time audio stream can be achieved through adaptive filtering methods, such as using the RLS algorithm to dynamically adjust the array weight coefficients through the scipy.signal library, ultimately obtaining an interference-resistant optimized sound source acquisition dimension.

[0133] Furthermore, based on the aforementioned sound source acquisition dimension, this invention formulates a sound recognition scheme corresponding to the target sound source scene. It can optimize acquisition parameters (such as microphone spacing and sampling frequency band) for scene characteristics (such as strong indoor reverberation and complex outdoor noise), improve signal quality from the source, effectively reduce multipath interference and noise impact, and improve the accuracy of tasks such as sound source localization and speech recognition.

[0134] wherein the sound recognition scheme refers to a full-process adaptive recognition strategy formed by dynamically adjusting sound source acquisition modes (such as microphone array layout, sampling frequency band), model inference strategies (such as noise reduction algorithm, feature extraction network) and environment adaptation parameters (such as multipath effect compensation coefficient) by analyzing parameters such as signal-to-noise attenuation entropy and phase distortion characteristics of real-time audio. For example, in the intelligent vehicle scene, if the system detects that the signal-to-noise attenuation entropy is increased due to the mixing of highway wind noise and engine noise, it will automatically switch to a directional pickup microphone to focus on the human voice frequency band, enable a noise suppression model based on a generative adversarial network (GAN), and optimize the frequency domain feature extraction layer of a convolutional neural network (CNN), thereby achieving high-precision recognition of voice commands. Optionally, the sound recognition scheme corresponding to the target sound source scene can be implemented through a transfer learning technique, such as using a pre-trained VGGish model to perform domain adaptation fine-tuning through the Keras library, and finally obtaining an optimized sound recognition scheme adapted to a specific scene.

[0135] Compared with the problems described in the background art, the present application can accurately capture the frequency components and timing variation rules of sound by obtaining real-time audio streams in the target sound source scene and extracting time-frequency feature groups in the real-time audio streams, thereby effectively improving the accuracy of sound recognition in complex environments. By using the sound network architecture to recognize acoustic feature clusters in the real-time audio stream, the present application can automatically extract a feature set with semantic association (such as phonemes in speech and frequency combinations in environmental sound) from complex audio signals through hierarchical nonlinear transformation of the neural network, thereby effectively improving the representation ability of sound features and the generalization performance of the recognition system. Furthermore, based on the adaptive mapping network, the present application can analyze the sound source propagation path corresponding to the real-time audio stream, and by virtue of the dynamic modeling capability of the network with respect to feature association, accurately analyze the multipath effects (such as indoor reverberation paths or outdoor obstacle reflection paths) of sound in space, thereby improving the accuracy of sound source positioning and recognition in complex environments. Furthermore, by querying the adaptive response trajectory of the mixed feature tensor in the target sound source scene, the present application can capture the dynamic processing process of the model with respect to complex acoustic features such as multipath effects in the scene in real time, thereby improving the stability and accuracy of the model in recognizing and positioning target sound sources in specific scenes. Finally, based on the signal-to-noise attenuation entropy, the present application can determine the robust recognition level of the real-time audio stream, which can directly reflect the degree of audio interference and signal quality with quantitative indicators, thereby providing a clear reference for the recognition system and improving the effective recognition ability of the system in complex environments, thereby ensuring the stability and accuracy of recognition. Therefore, the sound recognition method and system based on neural networks provided by the embodiments of the present application can improve the accuracy of sound recognition in complex environments.

[0136] Embodiment 2:

[0137] As Figure 3 shown in FIG. 1, it is a function module diagram of a sound recognition system based on a neural network according to the present application.

[0138] The sound recognition system based on a neural network 200 according to the present application can be installed in an electronic device. According to the implemented functions, the sound recognition system based on a neural network can include an architecture construction module 201, a network generation module 202, a hierarchical fusion module 203, an attenuation entropy calculation module 204, and a scheme formulation module 205. The modules according to the present application can also be referred to as units, which refer to a series of computer program segments that can be executed by an electronic device processor and can complete a fixed function, which are stored in the memory of the electronic device.

[0139] In the embodiments of the present application, the functions of each module / unit are as follows:

[0140] The architecture construction module 201 is configured to obtain a real-time audio stream in a target sound source scene, extract a time-frequency feature group in the real-time audio stream, analyze a pulse code sequence corresponding to the time-frequency feature group, and construct a sound network architecture corresponding to the real-time audio stream according to the pulse code sequence.

[0141] The network generation module 202 is configured to identify an acoustic feature cluster in the real-time audio stream by using the sound network architecture, match a hierarchical topological relationship in a preset neural computing framework according to the acoustic feature cluster, optimize a relationship node weight corresponding to the hierarchical topological relationship, and generate an adaptive mapping network corresponding to the relationship node weight.

[0142] The hierarchical fusion module 203 is configured to analyze a sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network, identify a multipath effect factor in the sound source propagation path, calculate a frequency domain coupling coefficient between the multipath effect factors, and perform hierarchical fusion on the frequency domain coupling coefficient to obtain a hybrid feature tensor.

[0143] The attenuation entropy calculation module 204 is configured to query an adaptive response trajectory of the hybrid feature tensor in the target sound source scene, extract a phase distortion feature and a signal-to-noise ratio index in the adaptive response trajectory, calculate a signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index.

[0144] The scheme formulation module 205 is configured to determine a robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, reconstruct a sound source acquisition dimension corresponding to the real-time audio stream according to the robust recognition level, and formulate a sound recognition scheme corresponding to the target sound source scene based on the sound source acquisition dimension.

[0145] In detail, the modules in the sound recognition system 200 in the embodiment of the present application adopt the same technical means as the sound recognition method based on neural network in the above-mentioned Figure 1 and can produce the same technical effects, which will not be described here again.

[0146] It is obvious for those skilled in the art that the present application is not limited to the details of the above-mentioned exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application but not limit the present application, and although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A neural network-based voice recognition method, characterized by, The method comprises: acquiring a real-time audio stream in a target sound source scene, extracting a time-frequency feature group in the real-time audio stream, analyzing a pulse coding sequence corresponding to the time-frequency feature group, and constructing a sound network architecture corresponding to the real-time audio stream according to the pulse coding sequence; identifying an acoustic feature cluster in the real-time audio stream by using the sound network architecture, matching a hierarchical topological relationship in a preset neural computing framework according to the acoustic feature cluster, optimizing a relationship node weight corresponding to the hierarchical topological relationship, and generating an adaptive mapping network corresponding to the relationship node weight; analyzing a sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network, identifying a multipath effect factor in the sound source propagation path, calculating a frequency domain coupling coefficient between the multipath effect factors, performing hierarchical fusion on the frequency domain coupling coefficient, and obtaining a hybrid feature tensor; querying an adaptive response trajectory of the hybrid feature tensor in the target sound source scene, extracting a phase distortion feature and a signal-to-noise ratio index in the adaptive response trajectory, calculating a signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion feature and the signal-to-noise ratio index, determining a robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, reconstructing a sound source collection dimension corresponding to the real-time audio stream according to the robust recognition level, and formulating a sound recognition scheme corresponding to the target sound source scene based on the sound source collection dimension. The method comprises:

2. The neural network-based voice recognition method of claim 1, wherein, analyzing a frequency domain feature vector in the pulse coding sequence; extracting an audio frame set corresponding to the real-time audio stream; performing time-frequency alignment on the frequency domain feature vector and the audio frame set to generate an audio alignment feature map; identifying an audio node relationship in the audio alignment feature map; constructing a sound network architecture corresponding to the real-time audio stream according to the audio node relationship. The method comprises:

3. The neural network-based voice recognition method of claim 1, wherein, analyzing a topological association code in the relationship node weight; matching a preset weight mapping rule library based on the topological association code to obtain an initial mapping path set; performing dynamic normalization on the initial mapping path set to obtain a normalized path sequence; removing a conflict sequence segment in the normalized path sequence to obtain an optimized mapping sequence segment; generating an adaptive mapping network corresponding to the relationship node weight based on the optimized mapping sequence segment. The method comprises:

4. The neural network-based voice recognition method of claim 1, wherein, performing time framing on the real-time audio stream to obtain an audio framing group; generating a time-frequency feature matrix corresponding to the audio framing group; extracting a sound source propagation vector in the time-frequency feature matrix; inputting the sound source propagation vector into a path analysis module in the adaptive mapping network to output a path feature sequence; analyzing a sound source azimuth angle and a propagation delay in the path feature sequence; constructing a sound source propagation path corresponding to the real-time audio stream based on the sound source azimuth angle and the propagation delay. ​ 5. The neural network-based voice recognition method of claim 1, wherein, The hierarchical fusion of the frequency domain coupling coefficients is performed to obtain a mixed feature tensor, including: querying a frequency domain feature subset corresponding to the frequency domain coupling coefficients; extracting local feature vectors of each frequency band of the frequency domain feature subset; analyzing distribution characteristics of the local feature vectors in a multi-dimensional space; analyzing dimension distribution rules corresponding to the distribution characteristics; based on the dimension distribution rules, the frequency domain coupling coefficients are hierarchically fused to obtain a mixed feature tensor.

6. The neural network-based voice recognition method of claim 1, wherein, The adaptive response trajectory of the mixed feature tensor in the target sound source scene is queried, including: analyzing a tensor distribution interval corresponding to the mixed feature tensor; according to the tensor distribution interval, matching a dynamic voiceprint segment in the target sound source scene, and extracting a sparse feature cluster in the voiceprint segment; mapping the sparse feature cluster to a sound field propagation path corresponding to the target sound source scene; generating an energy attenuation gradient graph corresponding to the sound field propagation path; querying the adaptive response trajectory in the energy attenuation gradient graph.

7. The neural network-based voice recognition method of claim 1, wherein, Based on the signal-to-noise attenuation entropy, the robust recognition level of the real-time audio stream is determined, including: analyzing noise interference features corresponding to the signal-to-noise attenuation entropy; based on the noise interference features, generating frequency domain segmented data corresponding to the real-time audio stream; extracting effective segmented audio corresponding to the frequency domain segmented data; based on the effective segmented audio, determining a hierarchical identifier corresponding to the real-time audio stream; based on the hierarchical identifier, determining the robust recognition level corresponding to the real-time audio stream.

8. A neural network based voice recognition system, characterized by, The system includes: an architecture construction module for obtaining a real-time audio stream in a target sound source scene, and extracting a time-frequency feature group in the real-time audio stream, analyzing a pulse coding sequence corresponding to the time-frequency feature group, and constructing a sound network architecture corresponding to the real-time audio stream according to the pulse coding sequence; a network generation module for identifying acoustic feature clusters in the real-time audio stream using the sound network architecture, matching a hierarchical topological relationship in a preset neural computing framework according to the acoustic feature clusters, optimizing relationship node weights corresponding to the hierarchical topological relationship, and generating an adaptive mapping network corresponding to the relationship node weights; a hierarchical fusion module for analyzing a sound source propagation path corresponding to the real-time audio stream based on the adaptive mapping network, identifying multipath effect factors in the sound source propagation path, calculating frequency domain coupling coefficients between the multipath effect factors, and hierarchically fusing the frequency domain coupling coefficients to obtain a mixed feature tensor; an attenuation entropy calculation module for querying an adaptive response trajectory of the mixed feature tensor in the target sound source scene, extracting phase distortion features and signal-to-noise ratio indexes in the adaptive response trajectory, and calculating a signal-to-noise attenuation entropy corresponding to the real-time audio stream based on the phase distortion features and the signal-to-noise ratio indexes; a scheme development module for determining a robust recognition level of the real-time audio stream based on the signal-to-noise attenuation entropy, reconstructing a sound source acquisition dimension corresponding to the real-time audio stream according to the robust recognition level, and developing a sound recognition scheme corresponding to the target sound source scene based on the sound source acquisition dimension.

Citation Information

Patent Citations

  • Hybrid speech processing method, electronic equipment and computer readable medium

    CN120236599A

  • Self-development type voice language pattern recognition system, and method and program for structuring self-organizing neural network structure used for same system

    JP2006171714A