A voiceprint recognition processing method for scheduling a phone call

By constructing an orthogonal projection operator for the channel subspace basis vector group and spectral entropy weighting, combined with silent anchor point dynamic residual elimination, the problem of voiceprint feature drift caused by nonlinear encoding and decoding distortion in dispatch telephones was solved, and stable identity recognition under multi-protocol channels was achieved.

CN122177124AActive Publication Date: 2026-06-09INFORMATION & COMM CO OF STATE GRID SHAANXI ELECTRIC POWER CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INFORMATION & COMM CO OF STATE GRID SHAANXI ELECTRIC POWER CO LTD
Filing Date
2026-05-11
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing technologies cannot effectively decouple nonlinear encoding and decoding distortion in dispatching telephones, resulting in a decline in the recognition performance of voiceprint features under multi-protocol channels, especially under low bit rate encoding and decoding protocols, where feature distribution drift and noise are difficult to separate.

Method used

An orthogonal projection operator based on the channel subspace basis vector group is constructed. By weighting the spectral entropy value and eliminating the dynamic residual of the silent anchor point, the voiceprint features and channel distortion are decoupled. A deep neural network model is used for feature extraction and calibration.

Benefits of technology

Maintaining the stability of voiceprint features under multi-protocol channels ensures the accuracy and robustness of identity recognition, adapts to extreme channel conditions, and prevents performance degradation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122177124A_ABST
    Figure CN122177124A_ABST
Patent Text Reader

Abstract

This invention relates to the field of speech signal processing technology and discloses a voiceprint recognition processing method for dispatching telephone calls. The method includes: acquiring digital speech signals from dispatching calls and converting them into frequency domain acoustic feature sequences; statistically aggregating and generating hybrid feature vectors using a deep neural network model; calculating the spectral entropy values ​​of each frequency band and constructing a feature weight matrix; performing weighted multiplication on the hybrid feature vectors using the weight matrix to obtain a weighted voiceprint vector; constructing an orthogonal projection operator using a preset channel subspace basis vector set; performing matrix multiplication on the weighted voiceprint vector and the orthogonal projection operator, filtering out channel projection components while using weight distribution to block the spread of high spectral entropy frequency band noise, resulting in a decoupled voiceprint vector; calculating similarity and outputting instructions. This invention solves the feature manifold shrinkage problem caused by nonlinear quantization distortion through the synergy of spectral entropy physical gating and algebraic geometric projection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a voiceprint recognition processing method for dispatching telephones, belonging to the field of speech signal processing technology. Background Technology

[0002] Current dispatch telephone networks carry high-time-efficiency command transmission tasks. Voiceprint recognition systems need to anchor speaker identity under dynamically changing channel conditions. Existing technologies typically use deep neural network architectures to extract Mel frequency cepstral coefficient acoustic features, map variable-length speech sequences into fixed-dimensional embedding vectors, and use cosine similarity to confirm identity. Existing technologies rely on large-scale datasets to train models to cover various types of environmental noise and use data augmentation techniques to suppress environmental background noise or room impulse response interference.

[0003] However, the scheduling of voice signals is constrained by the nonlinear distortion introduced by low bit-rate encoding and decoding protocols. For example, Chinese invention patent CN116935859A discloses a voiceprint recognition processing method and system, which collects voice data through wearable devices, uploads it to the cloud, extracts voiceprint features for comparison and unlocking, and analyzes environmental voice to classify abnormal sound waves. Although such solutions improve the business process of voiceprint recognition in secure interaction and environmental event monitoring, the core feature processing still follows the traditional linear extraction-comparison mode. Environmental voice processing focuses on macro-level event classification, such as identifying the type of argument or noise, without delving into the micro-level of the signal to solve the nonlinear distortion problem of the channel itself, and lacks the introduction of channel fingerprints for specific communication protocols. Mathematical decoupling mechanisms exist; however, the scheduling of communication speech signals is constrained by nonlinear distortions introduced by low bit-rate encoding and decoding protocols. In order to achieve bandwidth compression, G.729 and AMR standard speech encoders perform lossy quantization and reconstruction of speech signals, destroying the fine spectral structure and formant details, resulting in irreversible nonlinear collapse of voiceprint features in the manifold space. The encoding and decoding protocols generate quantization distortions that are coupled with the speech content, manifesting as feature distribution drift and intra-class distance divergence. This makes it difficult to separate identity features extracted based on physiological differences from quantization noise introduced by the channel. Conventional data augmentation methods cannot exhaustively enumerate the encoders and decoders and parameter combinations to generate complex distortion patterns, leading to a decrease in the recognition performance of the model in cross-channel or multi-protocol interoperability scenarios.

[0004] Therefore, how to construct a voiceprint recognition processing mechanism that decouples nonlinear encoding and decoding distortion, maintains feature stability and purity under multi-protocol channels and short voice conditions, and realizes the identity recognition of dispatch telephone speakers has become the technical problem to be solved by this invention. Summary of the Invention

[0005] To address the problems mentioned in the background art, the technical solution of the present invention is as follows: A voiceprint recognition processing method for dispatching telephones, comprising the following steps: Collect digital voice signals from dispatch calls and convert them into frequency domain acoustic feature sequences; A deep neural network model is used to statistically aggregate frequency domain acoustic feature sequences to generate a hybrid feature vector. Calculate the spectral entropy value of each frequency band in the frequency domain acoustic feature sequence. The spectral entropy value is used to characterize the randomness of the signal in the vector space of the corresponding frequency band. A characteristic weight matrix is ​​constructed based on the spectral entropy value, wherein the weight coefficients of each dimension in the characteristic weight matrix have a negative correlation with the corresponding spectral entropy value. The weighted voiceprint vector is obtained by performing a dimension-wise weighted multiplication operation on the mixed feature vector using the feature weight matrix; A preset channel subspace basis vector set is obtained. The channel subspace basis vector set is obtained by performing principal component analysis to reduce the dimensionality of the channel fingerprint vectors of multiple communication codec protocols. It is used to characterize the principal component direction of nonlinear quantization distortion in the feature space. An orthogonal projection operator is constructed using the basis vector set of the channel subspace. The orthogonal projection operator defines a linear transformation rule that maps the input vector to the complement space orthogonal to the channel subspace. The weighted acoustic vector is subjected to matrix multiplication with the orthogonal projection operator. While filtering out the projection components of the weighted acoustic vector in the channel subspace, the weight distribution determined by the spectral entropy value in the weighted acoustic vector is used to eliminate the numerical diffusion of noise components in the high spectral entropy frequency band to the low spectral entropy frequency band through matrix multiplication, thus obtaining the decoupled acoustic vector. Calculate the cosine similarity between the decoupled voiceprint vector and the pre-stored voiceprint template, and output the recognition command based on the cosine similarity.

[0006] Preferably, the method further includes a dynamic residual elimination step based on silent anchor points: using endpoint detection logic to identify non-human voice segments in digital speech signals, extracting acoustic feature sequences of non-human voice segments and encoding them using a deep neural network model to generate an environment embedding vector; calculating the projection component of the decoupled voiceprint vector in the direction of the environment embedding vector; and subtracting the projection component from the decoupled voiceprint vector to obtain the final voiceprint feature vector after environment calibration.

[0007] Preferably, the step of constructing an orthogonal projection operator using the channel subspace basis vector set specifically includes constructing the orthogonal projection operator based on the following relationship. : ,in, It is the identity matrix. This is the basis vector set of the channel subspace. Let be the transpose of the basis vector set of the channel subspace. The inverse matrix of the product of the channel subspace basis vector set and its transpose; orthogonal projection operator. Used to perform a linear transformation on the input vector to filter out the channel subspace projection components.

[0008] Preferably, the step of constructing a characteristic weight matrix based on spectral entropy values ​​includes: mapping spectral entropy values ​​to normalized weight values ​​through a nonlinear mapping function to construct a diagonal matrix corresponding to the dimensions of the mixed feature vector; the nonlinear mapping function is configured to perform the following logic: when the spectral entropy value exceeds a preset randomness threshold, the corresponding weight coefficient is set to zero to form an algebraic shielding layer for invalid high-noise frequency bands in the characteristic weight matrix.

[0009] Preferably, the steps for obtaining the channel subspace basis vector set include: collecting non-human voice silence segment data containing G.711, G.729, GSM and Opus encoding formats; extracting the long-term average acoustic features of the non-human voice silence segment data and performing principal component analysis; selecting the top few principal components whose cumulative variance contribution rate exceeds a preset contribution threshold as the channel subspace basis vector set, so as to lock the common quantization distortion direction introduced by different encoding and decoding protocols.

[0010] Preferably, in the step of statistically aggregating the frequency domain acoustic feature sequence using a deep neural network model, a geometric attention aggregation strategy based on projected residual energy is adopted: the magnitude of the orthogonal complement component of each frame feature vector in the frequency domain acoustic feature sequence on the space spanned by the basis vector group in the channel subspace is calculated, and the magnitude is used as the residual energy value of the frame feature vector; the attention weight of each frame feature vector is calculated according to the residual energy value, wherein the residual energy value is positively correlated with the attention weight; the frequency domain acoustic feature sequence is weighted and aggregated using the attention weight to generate a hybrid feature vector.

[0011] Preferably, the dynamic residual elimination step based on silent anchor points is performed after matrix multiplication, and the environment embedding vector system is obtained by encoding the background noise segment within the same call using a deep neural network model, which is used to characterize the transient environmental distortion unique to the current call session.

[0012] Preferably, the method further includes a confidence verification step based on orthogonality breaking: calculating the difference vector between the mixed feature vector and the decoupled voiceprint vector as the channel residual component; calculating the absolute value of the cosine similarity between the decoupled voiceprint vector and the channel residual component as the orthogonality breaking index; if the orthogonality breaking index exceeds a preset safety threshold, it is determined that the current call channel has undergone nonlinear manifold collapse, and the value of the cosine similarity is reduced.

[0013] Preferably, the step of calculating the spectral entropy value of each frequency band in the frequency domain acoustic feature sequence includes: normalizing the spectral energy of the frequency domain acoustic feature sequence to obtain the probability density function; calculating the spectral flatness of each frequency band using the Shannon entropy formula, and using the spectral flatness as the spectral entropy value characterizing whether the frequency band contains effective formant information or pure random quantization noise.

[0014] Preferably, the step of outputting the recognition instruction further includes: calculating the signal-to-noise ratio of the digital speech signal before calculating the cosine similarity, constructing a correction coefficient that is positively correlated with the signal-to-noise ratio; using the correction coefficient to perform a weighted summation of the decoupled voiceprint vector and the mixed feature vector to obtain the corrected voiceprint vector; and using the corrected voiceprint vector to perform a similarity comparison with the pre-stored voiceprint template.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. In voiceprint recognition for dispatch telephones, this invention constructs a basis vector set for the channel subspace based on standard communication codec protocols. An orthogonal projection operator is then constructed using this basis vector set to address the nonlinear quantization distortion problem introduced by low-bit-rate codecs in dispatch telephone networks. The distortion introduced by the communication protocol exhibits a low-dimensional subspace distribution in the feature space. The orthogonal projection operator directly acts on the original voiceprint embedding vector, filtering out the projected components of the vector in the channel subspace. This separates the identity features from the strong coupling between identity and channel. Based on the decoupling mechanism of linear algebraic geometry, the processed voiceprint feature vector reverts to an intrinsic manifold only related to the speaker's physiological structure, eliminating feature distribution drift caused by different communication protocols. The system does not require retraining an end-to-end model for each new channel environment; only a single matrix operation is needed to achieve zero-sample adaptation to multi-protocol channel distortion, ensuring intra-class compactness of identity recognition in cross-channel dispatch scenarios.

[0016] 2. This invention combines non-human voice silence segments during a call as a dynamic environment probe. Based on filtering fixed protocol distortions using a static channel basis, it constructs a dual orthogonal feature cleaning mechanism. It uses endpoint detection logic to capture the current session silence frame in real time and extract the environment embedding vector as a negative sample representing the current transient acoustic environment and terminal characteristics. It performs vector space residual elimination operation and subtracts the direction projection component of the environment embedding vector from the feature vector after processing by the static projection operator. The static basis and dynamic anchor point work together to solve the common quantization distortion caused by the communication protocol and eliminate the background noise and frequency response differences of the call site. This mechanism transforms the silence segments, which are considered invalid data in traditional technologies, into a tool to improve feature purity calibration, and achieves adaptive resistance to full-spectrum interference in complex scheduling sites, ensuring that the final output voiceprint features retain only pure identity information.

[0017] 3. Before performing projection decoupling, this invention introduces a frequency band spectral entropy weighting mechanism to solve the problem of feature decoupling failure under extreme channel conditions such as extremely low signal-to-noise ratio or severe packet loss. It calculates the spectral entropy value of each frequency band to quantify the randomness of the signal. Based on the higher feature weight of the high signal-to-noise ratio frequency band, it suppresses the high entropy noise frequency band. The pre-processing of information density perception blocks the invalid noise dimension before the feature enters the linear projection operation, preventing random noise from spreading to the effective identity feature dimension through matrix multiplication. This mechanism ensures that the subsequent orthogonal projection operation only acts on the feature component containing effective human voice information, avoiding the erroneous removal of effective information or the expansion of noise pollution due to blind projection. It ensures that the system maintains a minimum identity discrimination capability under poor communication links and prevents catastrophic degradation of recognition performance. Attached Figure Description

[0018] Figure 1 This is a flowchart of the voiceprint recognition process that integrates spectral entropy weighting and orthogonal projection decoupling according to the present invention. Figure 2 This is a schematic diagram of the nonlinear mapping and suppression mechanism of the characteristic weight matrix of the present invention; Figure 3 This is a system interaction timing diagram of the channel subspace basis vector group construction process of the present invention. Detailed Implementation

[0019] The technical solutions of the present invention will now be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0020] This invention provides a voiceprint recognition processing method for dispatching telephone calls, executed by a computing device such as a digital signal processor or cloud computing server. This method addresses the nonlinear quantization distortion problem introduced by low-bit-rate codecs in dispatching communication networks. Based on algebraic geometry principles, it decouples identity features from channel distortion by constructing a subspace projection operator orthogonal to channel characteristics; and constructs a channel subspace basis vector set. Collect non-human voice silence segment data or standard test tone data containing G.711, G.729, GSM-FR, and Opus standard communication coding formats. Extract long-term average acoustic features from the above data, namely Mel-frequency cepstral coefficients (MFCC) or Fbank features. Perform principal component analysis on the feature set and select the features whose cumulative variance contribution rate exceeds a preset threshold. The principal components serve as the basis vector set of the channel subspace. The preset threshold is set to 95% or other empirical values. For dimension The matrix, Indicates the dimension of voiceprint features. The matrix represents the channel subspace dimension and is permanently stored in the system memory. During real-time processing, digital voice signals from the scheduled call are acquired, and analog signals are converted to digital signals at a sampling rate of 8kHz or 16kHz. Frame-by-frame windowing with a frame length of 25ms and a frame shift of 10ms is performed. The time-domain signal is converted to the frequency-domain signal using a Fast Fourier Transform, and the frequency-domain acoustic feature sequence is extracted. To suppress noise propagation in low signal-to-noise ratio bands, the spectral entropy value of each frequency band in the frequency-domain acoustic feature sequence is calculated. The spectral energy of each frame is normalized to obtain the discrete probability density function. ,in For frequency band indexing, use the Shannon entropy formula. Calculate the spectral entropy value and construct a characteristic weight matrix based on the spectral entropy value. This matrix is ​​a diagonal matrix, and the diagonal elements are negatively correlated with the corresponding spectral entropy values. Use the inverse Sigmoid function to map high spectral entropy values ​​to weights close to zero and low spectral entropy values ​​to weights close to 1. Perform a dimension-wise weighted multiplication using the characteristic weight matrix and the corresponding frequency domain eigenvectors.

[0021] A deep neural network model is used to encode and statistically aggregate frequency domain acoustic feature sequences, generating a fixed-dimensional hybrid feature vector, i.e., the original speaker embedding vector. This deep neural network model adopts the ECAPA-TDNN or ResNet architecture. For short instruction scenarios, it uses a geometric attention aggregation strategy based on projected residual energy and calls a pre-set channel subspace basis vector set. Calculate the feature vector for each frame The projection modulus along the orthogonal complement direction of this subspace, i.e., the residual energy. The residual energy The attention weights are calculated based on the residual energy to represent the channel-independent amount of human voice information contained in the current frame. The calculation formula is: Using weights The feature sequences are weighted and summed to generate a mixed feature vector. Perform orthogonal decoupled projection using a pre-set set of channel subspace basis vectors. Constructing orthogonal projection operators , The mathematical expression is ,in For the identity matrix, this operator Define a linear transformation that maps the input vector to the complement space orthogonal to the channel subspace, and mixes the feature vectors. With orthogonal projection operator Perform matrix multiplication, that is thus filtering out The component parallel to the channel distortion direction is used to obtain the decoupled acoustic vector. .

[0022] Dynamic residual elimination based on silent anchors is performed to filter out transient environmental noise. Endpoint detection logic is used to identify non-human voice segments in the current speech signal, and the acoustic feature sequence of these segments is extracted and input into the same deep neural network model to generate an environment embedding vector. Calculate the decoupled acoustic vector Embedded vector in this environment The projection components in the direction are calculated using vector subtraction, and the formula is as follows: This yields the final voiceprint feature vector after environmental calibration. Perform confidence verification based on orthogonality breaking degree and calculate the mixed feature vector. With decoupled voiceprint vector The difference vector between them, i.e., the channel residual components ,calculate and The absolute value of the cosine similarity between them is used as the degree of orthogonality breaking. The calculation formula is: ,like If the value exceeds the preset safety threshold of 0.2, the feature separation is deemed abnormal and the final recognition score is reduced or an untrusted alarm is triggered. Finally, the cosine similarity between the processed voiceprint vector and the pre-stored voiceprint template is calculated, and a scheduling recognition instruction is output based on the similarity.

[0023] Example 1: In a cross-standard emergency command and dispatch network scenario, the command center dispatch console communicates with frontline police terminals via an IP network. The voice signal needs to be compressed and transmitted through multiple low-bit-rate codecs. The received voice signal suffers from nonlinear quantization distortion, resulting in a high coupling between voiceprint features and channel features. The system offline constructs a channel subspace basis vector group. By collecting non-human voice silence segment data containing G.711, G.729, GSM-FR, and Opus encoding formats, MFCC features were extracted, and principal component analysis was performed. The top features with a cumulative variance contribution rate exceeding 95% were selected. Each principal component constitutes When the command center receives a dispatch call, the system collects digital voice signals in real time and converts them into frequency domain acoustic feature sequences. In response to strong background noise at the scene, the system calculates the spectral entropy value of each frequency band, constructs a feature weight matrix using the inverse Sigmoid function, and performs a dimension-wise weighted multiplication on the frequency domain feature vector.

[0024] The system uses a deep neural network model to extract the original voiceprint embedding vector. To address the short duration of scheduling instructions, a geometric attention aggregation strategy based on projected residual energy is adopted. The system utilizes... Calculate the feature vector for each frame Residual energy in the orthogonal complement direction of the channel subspace And calculate attention weights accordingly. Aggregation generated The system mainly utilizes effective frames with large residual energy. Constructing orthogonal projection operators And perform matrix multiplication. This operation filters out The system projects the feature vectors of components parallel to the channel subspace onto an identity manifold orthogonal to the channel distortion. It then performs dynamic residual elimination based on silent anchors, utilizes endpoint detection to identify non-human voice segments, and extracts their features to generate an environment embedding vector. and from Subtract the projection component in that direction; finally, the system calculates the orthogonality breaking degree. ,like If the similarity is less than the preset threshold of 0.2, the cosine similarity between the processed voiceprint vector and the pre-stored voiceprint template is calculated, and a recognition instruction is output.

[0025] Example 2: This example provides a performance verification test report for the above-mentioned voiceprint recognition processing method. The test platform is built on a computing server equipped with an NVIDIA A100 graphics processing unit and an Intel Xeon Scalable processor. The dataset used in the test is based on the VoxCeleb1 test set, which contains speech samples from 1251 speakers. To simulate a real dispatch communication environment, the test uses the ITU-TG.191 software toolkit to compress and decompress the original 16kHz sampling rate speech signal through G.711 (64kbps), G.729 (8kbps) and AMR-NB (12.2kbps) standard codecs to introduce nonlinear quantization distortion. Using the NOISEX-92 noise library, siren sounds and radio background white noise with signal-to-noise ratios (SNR) of 10dB, 5dB and 0dB are superimposed on the processed speech to construct a test sample set with gradient difficulty.

[0026] The core parameters for the experiment are set as follows: channel subspace dimension The setting of the value directly affects the channel feature filtering rate and the identity information retention rate. The experiment determined the value through principal component analysis. The value is defined as the number of principal components corresponding to a cumulative variance contribution rate of 95%. The value, specifically ranging from 12 to 18, represents the orthogonality breaking degree. The safety threshold was set at 0.2, based on the statistical distribution characteristics of a large number of normally decoupled samples. The experimental design included the following three sample groups for multi-dimensional comparative verification: Control group 1 (baseline group): only the ECAPA-TDNN model was used to extract voiceprint features, without performing projection decoupling or spectral entropy weighting; Control group 2 (partially missing group): the orthogonal projection operator was applied after the ECAPA-TDNN model. The matrix multiplication operation is performed, but the step of weighting the characteristic weight matrix based on spectral entropy is missing. The experimental group of this invention executes a complete processing flow including spectral entropy weighting, orthogonal projection decoupling, and dynamic residual elimination. The processed voiceprint vectors of each group are compared with the distortion-free registration template using cosine similarity, and the equal error rate (EER) is calculated, while the orthogonality breaking degree is monitored. Key intermediate data such as channel residual energy magnitude.

[0027] Table 1: Performance Comparison Data under Key Operating Conditions

[0028] Referring to Table 1, under low bit rate (G.729 / AMR-NB) and strong noise (0dB / 5dB) conditions, the EER of control group 1 increased to 18.50%, indicating that direct feature extraction fails under nonlinear distortion. Control group 2, under 0dB conditions... The value reached 0.35, exceeding the safety threshold of 0.2, and the channel residual energy modulus was high, indicating that in the absence of spectral entropy weighting, noise energy diffused during the projection process, leading to the destruction of the feature space geometry. The experimental group of this invention tested this under all operating conditions. The value remained below 0.11, and the EER dropped to 6.80% under the extreme condition of 0dB.

[0029] Example 3: This example combines Figures 1 to 3 This describes a voiceprint recognition processing method for dispatching telephone calls, such as... Figure 1As shown, the overall execution flow starts with digital voice signal acquisition, converting the input raw data of the dispatch call into a frequency domain acoustic feature sequence. The flow is divided into two parallel processing branches. The first branch calculates the spectral entropy value of each frequency band to characterize the randomness of the signal and constructs a feature weight matrix accordingly. The second branch uses a deep neural network (DNN) model to perform statistical aggregation to generate a hybrid feature vector. The two branches converge at the step of dimensional weighted multiplication. The weight matrix is ​​used to block the noise diffusion of the high spectral entropy frequency band and output a weighted voiceprint vector. A pre-set channel subspace basis vector group is introduced to construct an orthogonal projection operator. Matrix multiplication is performed on the weighted voiceprint vector to filter out the channel projection component. The flow further combines the non-human voice silence anchor point extraction step to generate an environment embedding vector. A dynamic residual elimination step is performed to subtract the environment embedding projection component. Finally, confidence verification is performed based on the orthogonality breaking degree, and a recognition command containing similarity and control instructions is output according to the calculation results.

[0030] like Figure 2 As shown, this chart uses the feature dimension as the horizontal axis and the eigenvalue or weighting coefficient as the vertical axis, intuitively presenting the numerical mapping relationship between the eigenvalues ​​before weighting, the spectral entropy weighting coefficient, and the eigenvalues ​​after weighting. The spectral entropy weighting coefficient curve exhibits a specific non-linear decreasing trend with changes in the feature dimension. Multiplying this coefficient by the corresponding points on the eigenvalue curve before weighting generates the eigenvalue curve after weighting. Especially in the region with higher feature dimensions, the decay of the weighting coefficient causes the weighted eigenvalues ​​to approach zero. Figure 3 As shown, the system pre-executes the channel subspace basis vector group construction process. This timing process involves the interaction between the data acquisition system, codec group, feature extractor, principal component analysis module, and system memory. Specifically, the process begins with the data acquisition system acquiring non-human voice silent segment data. The data is processed through various encoding formats such as G.711, G.729, GSM-FR, and Opus before being transmitted to the feature extractor. Long-time average acoustic features are extracted, and Mel-frequency cepstral coefficients (MFCC) or Fbank features are calculated. The generated feature set is transmitted to the principal component analysis module for dimensionality reduction operations. The feature value sequence is calculated, and principal components with a cumulative variance contribution rate exceeding 95% are selected. Finally, the generated channel subspace basis vector group is stored in the system memory for use in the real-time processing stage.

[0031] Example 4: This example illustrates the system initialization and adaptive calibration procedure for the core algebraic and statistical parameters in the above-described voiceprint recognition processing method. This procedure is executed during the initial deployment of the system or when there are significant changes in the channel environment. It transforms the setting of key parameters from empirical values ​​to optimal solutions based on the statistical characteristics of the current acoustic environment, and performs channel subspace dimension calibration. The system adaptively defines and collects silence or test tone data containing all typical codec protocols (G.711, G.729, GSM-FR, Opus) in the target network, extracts the long-time average acoustic feature covariance matrix, and calculates the eigenvalue sequence of this matrix. ,in Given the total feature dimension, and considering that channel distortion typically manifests as high-energy principal components while random noise exhibits a gentle tail, the system calculates the logarithmic slope rate of change of the eigenvalue sequence and searches for the eigenvalue index. To maximize the curvature of the eigenvalue distribution curve at that point, i.e., to satisfy the second-order difference maxima condition, the system sets the index of this inflection point as the channel subspace dimension. In the calibration of the police network, the eigenvalues ​​show a steep decrease in the first 14 components, entering a flat zone. Based on this, the system... The value is locked at 14 to avoid relying on a fixed variance contribution rate. Environmental adaptation is performed using the spectral entropy weighting function. Background noise in different scheduling environments has different spectral entropy distribution characteristics. The system collects pure background noise segments and standard human voice segments from the current environment, calculates frame-level spectral entropy values ​​separately, constructs a bimodal histogram of the spectral entropy distribution, and uses the maximum inter-class variance method to calculate the optimal spectral entropy segmentation threshold for distinguishing noise patterns from speech patterns. Based on this threshold, the system defines the central parameter of the Sigmoid weight mapping function. And set the slope parameter , making in ( Within the range of the standard deviation of the noise spectral entropy, the weighting coefficients rapidly roll off from 0.9 to 0.1, performing soft shielding for specific noise frequency bands in the current environment, and constructing a characteristic weighting matrix based on the spectral entropy value.

[0032] To distinguish between high-frequency background noise and voiceless consonants, a joint energy and entropy value discrimination gating logic is executed for any frequency band. Calculate the spectral entropy value and normalized logarithmic energy value The nonlinear mapping function logic is set as follows: when the frequency band spectral entropy value Greater than the preset randomness threshold And the normalized logarithmic energy value Below the preset energy limit When the frequency band is determined to be a pure noise channel, the corresponding weighting coefficient is set to zero; if Greater than but Higher than The frequency band is determined to contain high-frequency resonant peaks of fricative or plosive sounds. The original weights are maintained, or a linear attenuation coefficient of 0.6 to 0.8 is applied. and The joint decision parameter is used to define the boundary between the channel noise floor and the speech energy-entropy plane distribution. The engineering calibration of the joint decision parameter adopts an offline calculation procedure based on statistical distribution, collecting no less than 500 hours of data without human voice background noise in a specific scheduling scenario, calculating the full-band spectral entropy distribution histogram, and selecting the value of the 95th percentile of the histogram as the randomness threshold. To define the lower limit of noise entropy, speech samples containing standard Mandarin voiceless consonant pronunciations were collected, and the minimum energy value in the high-frequency band from 2kHz to 4kHz was extracted and set as the energy lower limit. In actual deployments of DSP digital signal processors, the logic is implemented using a lookup table method, with a pre-built... and To index the two-dimensional weight mapping table, the quantized weight values ​​are directly read, avoiding complex floating-point conditional judgment operations in real-time call streams and meeting the microsecond-level processing latency requirements of scheduling communication. To address the nonlinear distortion differences introduced by the different encoding and decoding protocols of G.729 and AMR, a protocol-specific adaptive compensation strategy is adopted. After identifying the current call signaling encoding type, pre-stored protocol compensation factors are automatically invoked. The calculated spectral entropy value is corrected using the following formula: ,in The average quantization distortion rate of the current encoding protocol relative to the G.711 baseline protocol is directly calculated from the decrease in signal-to-noise ratio of the standard test audio. For the high compression rate G.729 protocol, Taking positive values ​​and appropriately amplifying the spectral entropy value enhances sensitivity to quantization noise; low compression rate protocol. Taking zeros maps quantization noise from different protocols to the same baseline entropy plane, so that subsequent weight matrix construction does not require separate parameter training for each protocol.

[0033] Execute orthogonality breaking threshold For safety boundary determination, the system constructs a calibration dataset containing matched and unmatched sample pairs, and the system scans the entire range. Values, and calculate in different The effectiveness of feature decoupling under the threshold is monitored by the system when... When a certain value is exceeded, does the cosine similarity of the decoupled voiceprint vectors of the matched sample pairs show a statistical decrease? The system plots a correlation curve between orthogonality breaking and recognition accuracy, and marks the point at which the accuracy begins to show an inflection point decline. The value is set as the system's safety circuit breaker threshold. In actual deployment, this dynamic calibration value is typically between 0.18 and 0.25. For the selection of silent anchor points, the system sets a minimum stability constraint: the environment embedding vector is only triggered when the length of continuously detected non-human voice frames exceeds 200ms and the cosine similarity variance of the feature vectors is less than 0.05. Update.

[0034] Example 5: This example describes the standardized offline calibration and field calibration procedures for the core parameters and models of the aforementioned voiceprint recognition processing method. These procedures address the deterministic source of initial model parameters and ensure a consistent performance baseline across different deployment environments. The offline calibration and data filling procedures for the channel feature base library are executed. A controlled communication protocol simulation environment is built, and pure channel test signals under four standard coding formats—G.711, G.729, GSM-FR, and Opus—are generated using a programmable attenuator and a protocol analyzer. For each protocol, at least 100 hours of silent segment data are collected and divided into training and validation sets according to time sequence. For the training set data, Mel-frequency cepstral coefficients with a frame length of 25ms are extracted to construct a feature covariance matrix. Eigenvalue decomposition is performed on this matrix to extract the feature value sequence. Using the scree plot inflection point determination method, the second difference of the eigenvalue sequence is calculated. Select The index at which the maximum value is reached serves as the channel subspace dimension under this protocol. The validation set data is used to validate the selected... The validity of the value is confirmed by calculating the reconstruction error to verify whether the coverage of channel features reaches more than 95%. This calibration process generates... Values ​​and their corresponding basis vectors It is embedded in the system's read-only memory as the factory default configuration.

[0035] Before deploying the system on-site, a pre-deployment spectral entropy weight calibration procedure was executed. After the system connected to the specific scheduling network on-site, an adaptive learning mode for environmental noise was initiated. The system continuously collected pure background noise segments from the on-site environment for a total duration of 300 seconds. The collected noise data was segmented into frames, and the spectral entropy value of each frame was calculated to construct a histogram of the on-site noise spectral entropy distribution. The optimal segmentation threshold for this histogram was calculated using the maximum inter-class variance method. This threshold characterizes the statistical boundary between noise and potential speech signals at the information entropy level in the current environment. The system maps the central parameter of the Sigmoid weighting function. Set as And based on the standard deviation of the noise spectral entropy Set slope parameters This ensures that the roll-off rate of the weighting function in the transition zone matches the fluctuation range of the environmental noise. After the parameter update is completed, the system automatically saves the configuration and exits the adaptive mode to enter the normal working state.

[0036] Example 6: This example provides a standardized pre-deployment calibration and model building procedure for a specific scheduling scenario. This procedure addresses the performance fluctuations of models caused by differences in acoustic environments under different application scenarios. Through a standardized parameter calibration process, it enables rapid adaptation to new environments. Before the system connects to a new scheduling network environment, initial calibration of environmental noise characteristics is performed. During non-call periods, the system automatically starts the environmental noise acquisition mode, continuously recording background noise data for no less than 600 seconds. The acquired noise data is then subjected to spectral analysis using Fast Fourier Transform to calculate the noise power spectral density across the entire frequency band. Based on the distribution characteristics of the noise power spectral density, the system automatically adjusts the coefficients of the pre-emphasis filter. For environments with strong high-frequency noise, the system automatically reduces the high-frequency boost coefficient to avoid signal-to-noise ratio degradation. For environments with strong low-frequency humming, the system enables a high-pass filter with a cutoff frequency set to 1.2 times the main energy frequency of the noise. This process ensures that the preprocessing parameters of the input signal in the frequency domain match the current environmental noise characteristics.

[0037] Dynamic calibration of the endpoint detection threshold is performed. The system plays a set of standard test speech signals with different signal-to-noise ratios (SNRs), ranging from 0dB to 20dB. The system monitors the output status of the endpoint detection module in real time and records the detection errors at the start and end points of the speech. Using a binary search strategy, the energy threshold and zero-crossing rate threshold of the endpoint detection module are adjusted until the average error of speech endpoint detection is less than 50ms under all test SNRs, ensuring that the system can accurately capture speech segments under different noise levels. Adaptive optimization of channel equalization parameters is performed. The system sends a set of test audio signals containing a full-band sweep signal. After transmission through the scheduling network, the test signal is acquired at the receiving end. The system calculates the spectral difference between the received signal and the transmitted signal to obtain the channel frequency response curve. Based on the frequency response curve, the coefficients of a set of inverse filters are automatically calculated to compensate for the frequency selective attenuation caused by the channel. The order of the inverse filters is adaptively determined according to the fluctuation of the frequency response curve to ensure that excessive computational delay is not introduced while compensating for channel distortion. This step provides a flatter spectral input for subsequent feature extraction through physical-level channel equalization.

[0038] The system performs hyperparameter fine-tuning and validation of the execution model. Based on the calibrated environmental parameters, it uses a pre-set small-scale adaptation dataset to perform transfer learning fine-tuning of the voiceprint recognition model, employing a relatively small learning rate during the fine-tuning process. The number of iteration rounds is set to 10 to avoid overfitting. After fine-tuning, the system automatically runs a set of standard test cases to verify the model's recognition accuracy and response time in the current environment. If all indicators meet the preset acceptance criteria, the system calibration is considered complete; otherwise, the system will prompt engineers to check the hardware connection or adjust the environment layout.

[0039] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0040] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A voiceprint recognition processing method for dispatching telephones, characterized in that, Includes the following steps: Collect digital voice signals from dispatch calls and convert them into frequency domain acoustic feature sequences; A deep neural network model is used to statistically aggregate frequency domain acoustic feature sequences to generate a hybrid feature vector. Calculate the spectral entropy value of each frequency band in the frequency domain acoustic feature sequence. The spectral entropy value is used to characterize the randomness of the signal in the vector space of the corresponding frequency band. A characteristic weight matrix is ​​constructed based on the spectral entropy value, wherein the weight coefficients of each dimension in the characteristic weight matrix have a negative correlation with the corresponding spectral entropy value. The weighted voiceprint vector is obtained by performing a dimension-wise weighted multiplication operation on the mixed feature vector using the feature weight matrix; A preset channel subspace basis vector set is obtained. The channel subspace basis vector set is obtained by performing principal component analysis to reduce the dimensionality of the channel fingerprint vectors of multiple communication codec protocols. It is used to characterize the principal component direction of nonlinear quantization distortion in the feature space. An orthogonal projection operator is constructed using the basis vector set of the channel subspace. The orthogonal projection operator defines a linear transformation rule that maps the input vector to the complement space orthogonal to the channel subspace. The weighted acoustic vector is subjected to matrix multiplication with the orthogonal projection operator. While filtering out the projection components of the weighted acoustic vector in the channel subspace, the weight distribution determined by the spectral entropy value in the weighted acoustic vector is used to eliminate the numerical diffusion of noise components in the high spectral entropy frequency band to the low spectral entropy frequency band through matrix multiplication, thus obtaining the decoupled acoustic vector. Calculate the cosine similarity between the decoupled voiceprint vector and the pre-stored voiceprint template, and output the recognition command based on the cosine similarity.

2. The voiceprint recognition processing method for dispatching telephones according to claim 1, characterized in that, The method also includes a dynamic residual elimination step based on silent anchors: using endpoint detection logic to identify non-human voice segments in digital speech signals, extracting acoustic feature sequences of non-human voice segments and encoding them using a deep neural network model to generate an environment embedding vector; calculating the projection component of the decoupled voiceprint vector in the direction of the environment embedding vector; and subtracting the projection component from the decoupled voiceprint vector to obtain the final voiceprint feature vector after environment calibration.

3. The voiceprint recognition processing method for dispatching telephones according to claim 1, characterized in that, The steps for constructing an orthogonal projection operator using the channel subspace basis vector set specifically include constructing the orthogonal projection operator based on the following relationship. : ,in, It is the identity matrix. This is the basis vector set of the channel subspace. Let be the transpose of the basis vector set of the channel subspace. The inverse matrix of the product of the channel subspace basis vector set and its transpose; orthogonal projection operator. Used to perform a linear transformation on the input vector to filter out the channel subspace projection components.

4. The voiceprint recognition processing method for dispatching telephones according to claim 1, characterized in that, The steps for constructing a characteristic weight matrix based on spectral entropy values ​​include: mapping spectral entropy values ​​to normalized weight values ​​through a nonlinear mapping function to construct a diagonal matrix corresponding to the dimensions of the mixed feature vector; the nonlinear mapping function is configured to perform the following logic: when the spectral entropy value exceeds a preset randomness threshold, the corresponding weight coefficient is set to zero to form an algebraic shielding layer for invalid high-noise frequency bands in the characteristic weight matrix.

5. The voiceprint recognition processing method for dispatching telephones according to claim 1, characterized in that, The steps for obtaining the channel subspace basis vector set include: collecting non-human voice silence segment data containing G.711, G.729, GSM and Opus encoding formats; extracting the long-term average acoustic features of the non-human voice silence segment data and performing principal component analysis; selecting the top few principal components whose cumulative variance contribution rate exceeds the preset contribution threshold as the channel subspace basis vector set, so as to lock the common quantization distortion direction introduced by different encoding and decoding protocols.

6. The voiceprint recognition processing method for dispatching telephones according to claim 1, characterized in that, In the step of statistically aggregating the frequency domain acoustic feature sequence using a deep neural network model, a geometric attention aggregation strategy based on projection residual energy is adopted: the magnitude of the orthogonal complement component of each frame feature vector in the frequency domain acoustic feature sequence on the space spanned by the basis vector group in the channel subspace is calculated, and the magnitude is used as the residual energy value of the feature vector of that frame; the attention weight of each frame feature vector is calculated based on the residual energy value, where the residual energy value is positively correlated with the attention weight; the frequency domain acoustic feature sequence is weighted and aggregated using the attention weight to generate a hybrid feature vector.

7. The voiceprint recognition processing method for dispatching telephones according to claim 2, characterized in that, The dynamic residual elimination step based on silent anchors is performed after matrix multiplication, and the environment embedding vector system is obtained by encoding background noise segments within the same scheduled call using a deep neural network model, which is used to characterize the transient environmental distortions unique to the current call session.

8. The voiceprint recognition processing method for dispatching telephones according to claim 1, characterized in that, The method also includes a confidence verification step based on orthogonality breaking: calculating the difference vector between the mixed feature vector and the decoupled acoustic vector as the channel residual component; The absolute value of the cosine similarity between the decoupled voiceprint vector and the channel residual components is calculated as an orthogonality breaking index. If the orthogonality breaking index exceeds the preset safety threshold, it is determined that the current call channel has experienced nonlinear manifold collapse, and the value of the cosine similarity is reduced.

9. A voiceprint recognition processing method for dispatching telephones according to claim 1, characterized in that, The steps for calculating the spectral entropy value of each frequency band in the frequency domain acoustic feature sequence include: normalizing the spectral energy of the frequency domain acoustic feature sequence to obtain the probability density function; calculating the spectral flatness of each frequency band using the Shannon entropy formula, and using the spectral flatness as the spectral entropy value characterizing whether the frequency band contains effective formant information or pure random quantization noise.

10. A voiceprint recognition processing method for dispatching telephones according to claim 1, characterized in that, The steps for outputting recognition instructions also include: calculating the signal-to-noise ratio (SNR) of the digital speech signal before calculating the cosine similarity, constructing a correction coefficient that is positively correlated with the SNR; using the correction coefficient to perform a weighted summation of the decoupled voiceprint vector and the mixed feature vector to obtain the corrected voiceprint vector; and comparing the similarity between the corrected voiceprint vector and the pre-stored voiceprint template.

Citation Information

Patent Citations

  • Voiceprint recognition processing method and system

    CN116935859A

  • Voiceprint authentication system and method for rapid channel compensation

    CN102129859A

  • Method for constructing source tracing model of original speaker of forged voice

    CN120260579A

  • Voice processing method and device based on voiceprint feature screening, equipment and medium

    CN120526776A

  • Voiceprint recognition method and device, electronic equipment and storage medium

    CN121999786A