Optical fiber distributed voiceprint recognition method based on multi-scale decomposition and hybrid recombination

The voiceprint recognition method based on multi-scale decomposition and hybrid recombination solves the problem of insufficient recognition performance in complex scenarios in existing technologies. It realizes multi-timescale modeling and feature extraction of voiceprint signals, thereby improving the accuracy and stability of recognition.

CN121938375APending Publication Date: 2026-04-28NORTHEASTERN UNIV AT QINHUANGDAO +2
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTHEASTERN UNIV AT QINHUANGDAO
Filing Date
2026-03-18
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing voiceprint recognition methods lack multi-timescale modeling capabilities and refined temporal dimension feature selection mechanisms in complex scenarios, leading to a decline in recognition performance.

Method used

A multi-scale decomposition and hybrid recombination method is adopted. The voiceprint signal is converted into a basic voiceprint feature representation through a multi-scale decomposition network, multi-scale long-term trend features and short-term detail information are extracted, and adaptive weighted hybrid recombination is performed to form hybrid recombination features.

Benefits of technology

It improves the accuracy and stability of voiceprint recognition, fully characterizes identity features in complex scenarios, suppresses the influence of noise, and enhances the accuracy and robustness of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121938375A_ABST
    Figure CN121938375A_ABST
Patent Text Reader

Abstract

The invention discloses an optical fiber distributed voiceprint recognition method based on multi-scale decomposition and hybrid recombination, and belongs to the technical field of distributed optical fiber sensing and artificial intelligence, and the method comprises the following steps: constructing a DAS system, and employing the DAS system to demodulate and restore the phase and intensity signal of backward Rayleigh scattering light to obtain a voiceprint signal; converting the voiceprint signal into basic voiceprint feature representation by using a multi-scale decomposition and hybrid recombination network; performing multi-scale decomposition on the basic voiceprint feature representation, and extracting multi-scale long-term trend features and short-term detail information; performing feature extraction and fusion on the multi-scale short-term detail information to obtain multi-scale short-term detail features; carrying out adaptive weighting on the multi-scale short-term detail features, and carrying out hybrid recombination on the multi-scale short-term detail features and the long-term trend features to obtain hybrid recombination features; and outputting an identification result through a plurality of layers of linear mapping and a Softmax function on the hybrid recombined features. By adopting the method, the accuracy and the stability of voiceprint recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical fields of distributed optical fiber sensing and artificial intelligence, and in particular to an optical fiber distributed voiceprint recognition method based on multi-scale decomposition and hybrid recombination. Background Technology

[0002] Voiceprint recognition technology, as a biometric authentication and information identification method, has been widely used in many key fields such as security monitoring, financial payment, smart access control, and judicial evidence collection due to its inherent advantages such as non-contact data collection, strong anti-counterfeiting capabilities, and ease of operation. However, with the continuous improvement of security requirements and the widespread adoption of IoT technology, traditional voiceprint recognition systems are gradually revealing significant shortcomings in terms of adaptability to complex scenarios, recognition accuracy, and coverage. Therefore, there is an urgent need for innovative integration of new sensing and recognition technologies to overcome these bottlenecks.

[0003] Traditional electroacoustic sensors suffer from drawbacks when acquiring sound signals, including susceptibility to electromagnetic interference, limited detection range, and difficulty operating normally in harsh environments such as high temperature and high humidity. Fiber optic distributed acoustic sensing (DAS) systems, using optical fiber as both the sensing medium and signal transmission carrier, utilize the characteristic that the phase of backscattered Rayleigh light in the fiber is proportional to the vibration of external sound waves. By detecting phase changes in the Rayleigh scattered light, they sense the sound wave vibration information along the fiber. DAS systems offer advantages such as long sensing distance, strong anti-interference capability, immunity to electromagnetic interference, long lifespan, and long-term stable operation in various harsh environments. Currently, DAS systems are widely used in fields such as oil and gas pipeline leak monitoring, natural disaster monitoring, and structural health monitoring. Extending them to the field of voiceprint recognition has significant practical implications and application value.

[0004] However, applying DAS systems to voiceprint recognition faces unique technical challenges: the voiceprint-related signals acquired by DAS are essentially vibration phase and light intensity variation data along the optical fiber, characterized by high signal dimensionality, complex time-frequency characteristics, and a large amount of environmental noise (such as wind vibration and mechanical vibration) unrelated to the target voiceprint. Traditional voiceprint feature extraction methods struggle to effectively extract key identity information. Existing methods fall into two categories: one based on traditional signal processing features and statistical modeling, such as extracting acoustic features like short-time energy, spectral features, or Mel-frequency cepstral coefficients, combined with Gaussian mixture models or support vector machines for recognition. This type of method relies on manually designed features and is typically analyzed within a fixed time window, making it difficult to characterize the changes in voiceprint signals across different time scales. Recognition performance tends to degrade when DAS signals are noisy or highly non-stationary. Another type of method is based on end-to-end voiceprint recognition models using deep learning, such as convolutional neural networks or recurrent neural networks, to directly model the voiceprint signal. Although these methods have improved the automatic feature learning ability to some extent, most of them adopt a single scale or fixed receptive field structure, focusing on local or global feature modeling. They are difficult to simultaneously take into account the long-term stable identity features and short-term changing detailed features in the voiceprint signal, resulting in insufficient utilization of identity discrimination information.

[0005] In summary, existing voiceprint recognition methods generally lack the ability to model multiple time scales for the characteristics of DAS voiceprint signals and the mechanism for selecting refined time-dimensional features, making it difficult to achieve stable and accurate voiceprint recognition in complex application scenarios. Summary of the Invention

[0006] The purpose of this invention is to provide a fiber-optic distributed voiceprint recognition method based on multi-scale decomposition and hybrid recombination, which overcomes the shortcomings of existing voiceprint recognition technologies in terms of adaptability to complex scenarios, DAS signal feature mining capabilities, and recognition performance, thereby improving the accuracy and stability of voiceprint recognition.

[0007] To achieve the above objectives, this invention provides a fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction, comprising the following steps: S1. Construct a DAS system and use the DAS system to demodulate and restore the phase and intensity signals of the backscattered Rayleigh light carrying external sound signal information to obtain the voiceprint signal. S2. Use a multi-scale decomposition and hybrid reconstruction network to convert the voiceprint signal into a basic voiceprint feature representation; S3. Perform multi-scale decomposition on the basic voiceprint feature representation to extract multi-scale long-term trend features and multi-scale short-term detail information; S4. Multi-scale feature extraction and fusion are performed on multi-scale short-term detail information to obtain multi-scale short-term detail features; S5. Adaptively weight the multi-scale short-term detail features and mix and recombine them with long-term trend features to obtain mixed recombination features; S6. The hybrid recombination features are processed through several layers of linear mapping and the Softmax function to output the recognition results corresponding to different people.

[0008] Preferably, the DAS system includes an ultra-narrow linewidth fiber laser source, a pulsed light modulator, an erbium-doped fiber amplifier, a fiber coupler, a fiber circulator, and a single-mode fiber connected in sequence. The pulsed light modulator is connected to a signal driver, and the data acquisition and demodulation unit is connected to both the fiber coupler and the fiber circulator.

[0009] Preferably, the DAS system further includes a high-order random fiber laser amplification unit, which includes a ytterbium-doped random fiber laser, a wavelength division multiplexer (WDM) I, a wavelength division multiplexer II, and a point feedback unit. WDM II and WDM I are connected sequentially between the fiber circulator and the single-mode fiber. WDM II is connected to the point feedback unit, and WDM I is connected to the ytterbium-doped random fiber laser.

[0010] Preferably, in step S2, the voiceprint signal is initially mapped using a one-dimensional convolutional layer, and then converted into a basic voiceprint feature representation suitable for deep feature analysis.

[0011] Preferably, step S3 specifically includes the following steps: S31. Through several average pooling operations at different scales, the voiceprint features are decomposed in the time dimension to extract multi-scale long-term trend features, resulting in a multi-scale long-term trend feature set: ; in, The number of channels representing voiceprint characteristics. Represents the length of voiceprint features in the time dimension; S32. By performing difference or residual calculations between long-term trend features at different scales and the original voiceprint features, the corresponding multi-scale short-term detail information is extracted, resulting in a multi-scale short-term detail information set: .

[0012] Preferably, in step S31, one way to set the pooling kernel size set is {25,19,13,7}.

[0013] Preferably, step S4 specifically includes the following steps: S41. To extract features from multi-scale short-term detail information, the network is configured with one-dimensional convolutional layers of different receptive field sizes to obtain multi-scale speaker texture details at different temporal resolutions: ; in, This represents the kernel size of a one-dimensional convolutional layer; S42. Concatenate the detailed features after convolution at each scale along the channel dimension to form multi-scale short-term detailed features: .

[0014] Preferably, in step S41, the kernel size satisfies One way to set the kernel size is as follows: .

[0015] Preferably, step S5 specifically includes the following steps: S51. An adaptive weighting mechanism is introduced into the spliced ​​multi-scale short-term detail feature representation in the time dimension to generate time-weighted features. The generated time-weighted features are multiplied element-wise with the concatenated multi-scale short-term detail features, and then summed along the channel dimension to obtain the weighted detail features: ; in, , Represents a scale index. Represents the channel index. An index representing the time position; S52. Add the weighted detail features to the long-term trend features at the corresponding scale element by element to obtain the mixed recombination features: .

[0016] Therefore, the fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction described above has the following advantages: (1) In this invention, by decomposing the voiceprint signal into long-term trend information and short-term detail information of multiple scales, and performing targeted modeling of features at different time scales, the stable features and personalized detail features of the person's identity can be characterized simultaneously and fully, avoiding the problem of insufficient discrimination information caused by relying only on single-scale features in traditional methods, and improving the discriminativeness and robustness of voiceprint features.

[0017] (2) In this invention, by adaptively weighting short-term detail information of multiple scales in the time dimension and mixing and recombining it with long-term trend information of the corresponding scale, the voiceprint segments that contribute significantly to the identification of persons can be automatically enhanced, and the influence of noise and redundant information can be suppressed, thereby effectively improving the accuracy and stability of voiceprint recognition in complex DAS application scenarios.

[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid recombination of the present invention. Figure 2 This is a network framework diagram based on multi-scale decomposition and hybrid recombination according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the DAS system framework based on sixth-order random fiber laser amplification according to an embodiment of the present invention; Figure 4 This is a simulation diagram of the power distribution of the sensing signal based on sixth-order random fiber laser amplification in an embodiment of the present invention; Figure 5 This is a sub-category confusion matrix diagram of the voiceprint recognition results in an embodiment of the present invention; Figure 6 This is a ROC curve of the voiceprint recognition results in an embodiment of the present invention. Reference numerals in the attached figures: 1. Ultra-narrow linewidth fiber laser source; 2. Pulse light modulator; 3. Signal driver; 4. Erbium-doped fiber amplifier; 5. Fiber coupler; 6. Fiber circulator; 7. Single-mode fiber; 8. Data acquisition and demodulation unit; 9. Ytterbium-doped random fiber laser; 10. Wavelength division multiplexer one; 11. Wavelength division multiplexer two; 12. Point feedback unit. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Specific model specifications need to be selected and determined according to the actual specifications of the device, etc. The specific selection calculation method adopts existing technology in the art, and therefore will not be described in detail.

[0021] Example like Figures 1-2 As shown, this invention provides a fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction, comprising the following steps: S1. Construct a DAS system and use the DAS system to demodulate and restore the phase and intensity signals of the backscattered Rayleigh light carrying external sound signal information to obtain the voiceprint signal. S2. Using a multi-scale decomposition and hybrid recombination network, a one-dimensional convolutional layer is used to perform preliminary feature mapping on the voiceprint signal, converting it into a basic voiceprint feature representation suitable for deep feature analysis. S3. Perform multi-scale decomposition on the basic voiceprint feature representation to extract multi-scale long-term trend features and multi-scale short-term detail information; S31. By performing average pooling operations at four different scales, the voiceprint features are decomposed in the time dimension to extract multi-scale long-term trend features, resulting in a multi-scale long-term trend feature set: ; in, The number of channels representing voiceprint characteristics. The length of the voiceprint features in the time dimension is represented by the pooling kernel size set to {25,19,13,7} in this embodiment; the long-term trend features are gradually enhanced, reflecting the overall change features of the voiceprint over a longer time scale. S32. By performing difference or residual calculations between long-term trend features at different scales and the original voiceprint features, the corresponding multi-scale short-term detail information is extracted, resulting in a multi-scale short-term detail information set: ; This process yields features that progressively refine local details, with the information becoming increasingly detailed. Through this decomposition, a structured representation of the voiceprint signal at different time scales is achieved, enabling explicit modeling of long-term identity characteristics and short-term personality details.

[0022] S4. Multi-scale feature extraction and fusion are performed on multi-scale short-term detail information to obtain multi-scale short-term detail features; S41. To extract multi-scale short-term detail information, the network uses four one-dimensional convolutional layers with different receptive field sizes to obtain multi-scale speaker texture details at different temporal resolutions: ; in, This represents the kernel size of a one-dimensional convolutional layer, and the kernel size satisfies... In this embodiment, the kernel size is set to ; S42. Concatenate the detailed features after convolution at each scale along the channel dimension to form multi-scale short-term detailed features: .

[0023] S5. Adaptively weight the multi-scale short-term detail features and mix and recombine them with long-term trend features to obtain mixed recombination features; S51. An adaptive weighting mechanism is introduced into the spliced ​​multi-scale short-term detail feature representation in the time dimension, and time-weighted features are generated through the Sigmoid function. The generated time-weighted features are multiplied element-wise with the concatenated multi-scale short-term detail features, and then summed along the channel dimension to obtain the weighted detail features: ; in, , Represents a scale index. Represents the channel index. An index representing the time position; S52. Add the weighted detail features to the long-term trend features at the corresponding scale element by element to obtain the mixed recombination features: .

[0024] S6. The hybrid recombination features are processed through two layers of linear mapping and the Softmax function to output the recognition results corresponding to different people.

[0025] like Figure 3 As shown, the DAS system includes an ultra-narrow linewidth fiber laser source 1, a pulsed light modulator 2, an erbium-doped fiber amplifier 4, a fiber coupler 5, a fiber circulator 6, and a single-mode fiber 7 connected in sequence. The pulsed light modulator 2 is connected to a signal driver 3, and the data acquisition and demodulation unit 8 is connected to both the fiber coupler 5 and the fiber circulator 6. The ultra-narrow linewidth fiber laser source 1 generates a narrow linewidth coherent optical signal, which is modulated into a pulsed signal by the pulsed light modulator 2 driven by the signal driver 3, and then amplified by the erbium-doped fiber amplifier 4. The amplified pulsed signal light is split by the fiber coupler 5; one path is used as a probe light and injected into the single-mode fiber 7 via the fiber circulator 6, while the other path is used as a local oscillator signal input to the data acquisition and demodulation unit 8. The probe light input to the single-mode fiber 7 generates a backscattered Rayleigh signal during its propagation along the fiber. Since the phase of the backscattered Rayleigh signal is proportional to the magnitude of the external sound wave vibration signal, the external sound wave signal can be quantitatively demodulated by demodulating the phase information of the Rayleigh scattered light. By further utilizing the return time of pulsed light, accurate positioning of acoustic vibration signals can be achieved, thereby realizing distributed acoustic sensing.

[0026] To extend the sensing distance of the DAS system, a sixth-order random fiber laser amplification technique can be used to provide gain for the signal light. The DAS system based on the sixth-order random fiber laser amplification also includes a high-order random fiber laser amplification unit. The high-order random fiber laser amplification unit includes a ytterbium-doped random fiber laser 9, a wavelength division multiplexer (WDM) 10, a wavelength division multiplexer (WDM) 2 11, and a point feedback unit 12. The WDM 2 11 and the WDM 10 are connected sequentially between the fiber circulator 6 and the single-mode fiber 7. The WDM 2 11 is connected to the point feedback unit 12, and the WDM 10 is connected to the ytterbium-doped random fiber laser 9.

[0027] A ytterbium-doped random fiber laser 9 serves as the pump source, injected into a single-mode fiber 7 through the transmission port of wavelength division multiplexer 10. Under the influence of the ytterbium-doped random fiber laser 9, the single-mode fiber 7 provides the stimulated Raman scattering gain and distributed Rayleigh scattering feedback required for laser lasing. Wavelength division multiplexer 11 separates the 1550nm band probe signal light from the first to fifth order cascaded random fiber lasers. The transmission port of wavelength division multiplexer 11 is connected to a point feedback unit 12, which provides broadband feedback for the 1.1-1.4μm band cascaded random fiber lasers. The 1.45μm band fifth-order random fiber laser generated in the single-mode fiber 7 can be used as the direct pump for the 1550nm band probe signal light, resulting in distributed amplification. The power distribution of the signal light after amplification by the sixth-order random fiber laser is shown below. Figure 4 As shown, the detection optical signal is effectively amplified, and the power fluctuation is only 8.5dB within a 100km fiber optic range, thus extending the sensing distance of the DAS system.

[0028] In this embodiment, the DAS system was used to collect voice signals from three subjects. 826, 803, and 808 samples were collected from each subject, respectively, for a total of 2437 samples. The training, validation, and test sets were then divided into training, validation, and test sets at a ratio of 70.0%, 15.0%, and 15.0%, respectively, containing 1705, 366, and 366 samples. The voice signals collected by the DAS system were then used for training, employing a batch iteration method with a batch size of 32, for a total of 50 iterations. The learning rate was set to 0.0005, and the Adam optimizer was used. The network was considered converged when the accuracy did not improve within the set number of training iterations or within three iterations, and the trained network parameters were saved at this point.

[0029] Table 1. Quantitative comparison of the identification method of this invention with other identification methods.

[0030] Table 2. Quantitative comparison of the identification method of this invention with other identification methods on the human subclass.

[0031] The quantitative comparison results of the identification method of this invention and other identification methods are shown in Table 1. The quantitative comparison results of the identification method of this invention and other identification methods on three person subclasses are shown in Table 2. Accuracy, AUC, Precision, Recall, F1-Score, and Specificity were all compared, demonstrating that the identification method of this invention performs better in all indicators. The confusion matrices of the subclasses of the method of this invention, the ordinary CNN method, and the LSTM method are shown in Table 2. Figure 5As shown, the ROC curve of the identification results of this invention is as follows. Figure 6 As shown, all of them demonstrate a good balance between different categories.

[0032] Therefore, this invention adopts a fiber-optic distributed voiceprint recognition method based on multi-scale decomposition and hybrid recombination to overcome the shortcomings of existing voiceprint recognition technologies in terms of adaptability to complex scenarios, DAS signal feature mining capabilities, and recognition performance, thereby improving the accuracy and stability of voiceprint recognition.

[0033] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction, characterized in that: Includes the following steps: S1. Construct a DAS system and use the DAS system to demodulate and restore the phase and intensity signals of the backscattered Rayleigh light carrying external sound signal information to obtain the voiceprint signal. S2. Use a multi-scale decomposition and hybrid reconstruction network to convert the voiceprint signal into a basic voiceprint feature representation; S3. Perform multi-scale decomposition on the basic voiceprint feature representation to extract multi-scale long-term trend features and multi-scale short-term detail information; S4. Multi-scale feature extraction and fusion are performed on multi-scale short-term detail information to obtain multi-scale short-term detail features; S5. Adaptively weight the multi-scale short-term detail features and mix and recombine them with long-term trend features to obtain mixed recombination features; S6. The hybrid recombination features are processed through several layers of linear mapping and the Softmax function to output the recognition results corresponding to different people.

2. The fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction according to claim 1, characterized in that: The DAS system includes an ultra-narrow linewidth fiber laser source, a pulsed light modulator, an erbium-doped fiber amplifier, a fiber coupler, a fiber circulator, and a single-mode fiber, connected in sequence. The pulsed light modulator is connected to a signal driver, and the data acquisition and demodulation unit is connected to both the fiber coupler and the fiber circulator.

3. The fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction according to claim 2, characterized in that: The DAS system also includes a high-order random fiber laser amplification unit, which includes a ytterbium-doped random fiber laser, a wavelength division multiplexer (WDM) I, a wavelength division multiplexer II, and a point feedback unit. WDM II and WDM I are connected sequentially between the fiber circulator and the single-mode fiber. WDM II is connected to the point feedback unit, and WDM I is connected to the ytterbium-doped random fiber laser.

4. The fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction according to claim 1, characterized in that: In step S2, the voiceprint signal is initially mapped using a one-dimensional convolutional layer, converting it into a basic voiceprint feature representation suitable for deep feature analysis.

5. The fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction according to claim 4, characterized in that: Step S3 specifically includes the following steps: S31. Through several average pooling operations at different scales, the voiceprint features are decomposed in the time dimension to extract multi-scale long-term trend features, resulting in a multi-scale long-term trend feature set: ; in, The number of channels representing voiceprint characteristics. Represents the length of voiceprint features in the time dimension; S32. By performing difference or residual calculations between long-term trend features at different scales and the original voiceprint features, the corresponding multi-scale short-term detail information is extracted, resulting in a multi-scale short-term detail information set: 。 6. The fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction according to claim 5, characterized in that: In step S31, one way to set the pooling kernel size set is {25,19,13,7}.

7. The fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction according to claim 6, characterized in that: Step S4 specifically includes the following steps: S41. To extract features from multi-scale short-term detail information, the network is configured with one-dimensional convolutional layers of different receptive field sizes to obtain multi-scale speaker texture details at different temporal resolutions: ; in, This represents the kernel size of a one-dimensional convolutional layer; S42. Concatenate the detailed features after convolution at each scale along the channel dimension to form multi-scale short-term detailed features: ; in, This represents the multi-scale short-term detail features after splicing.

8. The fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction according to claim 7, characterized in that: In step S41, the kernel size satisfies One way to set the kernel size is as follows: .

9. The fiber-optic distributed acoustic signature recognition method based on multi-scale decomposition and hybrid reconstruction according to claim 8, characterized in that: Step S5 specifically includes the following steps: S51. An adaptive weighting mechanism is introduced into the stitched multi-scale short-term detail feature representation in the time dimension to generate time-weighted features. The generated time-weighted features are multiplied element-wise with the concatenated multi-scale short-term detail features, and then summed along the channel dimension to obtain the weighted detail features: ; in, , Represents a scale index. Represents the channel index. An index representing the time position; S52. Add the weighted detail features to the long-term trend features at the corresponding scale element by element to obtain the mixed recombination features: 。

Citation Information

Patent Citations

  • Optical fiber distributed sound wave sensing signal high-precision classification and identification method based on model fusion

    CN112985574A

  • Voice extraction method and system based on hourglass structure and self-attention mechanism

    CN116665655A

  • Speech emotion recognition method and system based on multi-scale adaptive feature fusion

    CN120356487A

  • Transformer state identification method and device based on multi-scale time-frequency characteristics

    CN121565199A

  • Classification-Based Frame Loss Concealment for Audio Signals

    US20080033718A1