Quantum-derived newton-raphson optimal fractional order spectrogram generation method and system

By using a quantum-derived Newton-Raphson optimal fractional-order spectrogram generation method, the problems of insufficient time-frequency resolution and hyperparameter rigidity in traditional audio signal processing are solved. The generated spectrogram has concentrated energy and clear features, which significantly improves speech recognition and classification performance.

CN122290569APending Publication Date: 2026-06-26FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUZHOU UNIV
Filing Date
2026-02-06
Publication Date
2026-06-26

Smart Images

  • Figure CN122290569A_ABST
    Figure CN122290569A_ABST
Patent Text Reader

Abstract

This invention provides a quantum-derived Newton-Raphson optimal fractional-order spectrogram generation method, comprising the following steps: Step 1: Acquire the original audio signal and construct a fractional-order spectrogram based on fractional Fourier transform (FRFT); Step 2: Perform nonlinear scaling compression on the fractional-order spectrogram using a Mel filter bank to generate a fractional-order Mel spectrogram; Step 3: Construct an adaptive optimization framework with information entropy minimization as the objective function to measure the information fidelity between the spectrogram and the original signal; Step 4: Use the quantum-derived Newton-Raphson optimization algorithm (QNRBO) to globally optimize the fractional-order order, frame length, and frame shift hyperparameters to generate the optimal fractional-order spectrogram; Step 5: Input the optimal fractional-order spectrogram into a downstream speech recognition model. This technical solution aims to systematically solve core problems such as insufficient traditional time-frequency representation capabilities, rigid hyperparameter configuration, limited optimization algorithm performance, and feature-task disconnect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing technology, and in particular to a method and system for generating quantum-derived Newton-Raphson optimal fractional-order spectrograms. Background Technology

[0002] With the rapid development of modern information technology, audio signal processing technology is increasingly widely used in fields such as speech recognition, sentiment analysis, and music information retrieval. Traditional audio signal processing methods, such as Short-Time Fourier Transform (STFT) and Wavelet Transform, while providing a certain time-frequency resolution, often fall short when dealing with audio signals with complex frequency modulation characteristics or nonlinear phase structures. Especially in speech recognition, accurately capturing subtle changes in the sound signal is crucial for improving classification accuracy. Therefore, finding a more effective method for representing audio features has become a research hotspot. With the rapid development of deep learning technology, models such as convolutional neural networks and visual Transformers in the image domain have achieved revolutionary breakthroughs in tasks such as image classification, object detection, and semantic segmentation. Inspired by this, researchers have begun to explore applying deep learning methods to audio signal processing. However, audio signals are essentially one-dimensional time-series data and cannot be directly input into two-dimensional deep neural networks. Therefore, converting audio signals into time-spectrum graphs has become a key bridge. Spectrograms map one-dimensional audio signals into two-dimensional time-frequency energy distribution maps through short-time Fourier transform, enabling them to be effectively processed by visual deep learning models and widely used in tasks such as speech recognition and sentiment analysis.

[0003] However, traditional spectrograms, based on fixed window functions and linear time-frequency decomposition, are constrained by the Heisenberg uncertainty principle, resulting in inherent contradictions in their time-frequency resolution: poor time resolution in the high-frequency range and poor frequency resolution in the low-frequency range, making it difficult to simultaneously and accurately capture transient impulses and steady-state harmonics. Furthermore, for audio signals with complex frequency modulation or non-stationary characteristics, the spectrograms generated by STFT often exhibit energy diffusion, ambiguity, or cross-term interference, leading to fragmented feature representation and affecting subsequent recognition performance. Although the Mel filter bank can simulate human auditory perception to some extent and improve speech recognition, its fixed nonlinear frequency scale still cannot adaptively match the dynamic changes of the signal. Against this backdrop, the Fractional Fourier Transform (FRFT) can effectively characterize the energy distribution of a signal in the time-frequency domain and has gradually become a new approach to solving these problems. FRFT extends the concept of the traditional Fourier transform to rotations of arbitrary angles, allowing signals to be analyzed in a wider time-frequency space, making it particularly suitable for processing audio signals with complex frequency characteristics. Despite its theoretically superior performance, FRFT still faces challenges in practical applications: how to select appropriate fractional-order parameters to adapt to different types of audio signals has become a key issue limiting its widespread use. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a quantum-derived Newton-Raphson optimal fractional-order spectrogram generation method and system, aiming to systematically solve the core problems such as insufficient traditional time-frequency representation capabilities, rigid hyperparameter configuration, limited performance of optimization algorithms, and disconnect between features and tasks.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a quantum-derived Newton-Raphson optimal fractional-order spectrogram generation method, comprising the following steps:

[0006] Step 1: Acquire the original audio signal and construct a fractional Fourier transform (FRFT) spectrum, where the fractional order 'a' is an adjustable parameter used to rotate the signal by any angle in the time-frequency plane to optimize the signal energy distribution.

[0007] Step 2: The fractional-order spectrogram is nonlinearly scaled using a Mel filter bank to generate a fractional-order Mel spectrogram, in order to integrate the characteristics of human auditory perception.

[0008] Step 3: Using the minimization of information entropy as the objective function, construct an adaptive optimization framework to measure the information fidelity between the spectrogram and the original signal;

[0009] Step 4: Use the quantum-derived Newton-Raphson optimization algorithm QNRBO to globally optimize the fractional order order, frame length, and frame shift hyperparameter to generate the optimal fractional order spectrogram;

[0010] Step 5: Input the optimal fractional-order spectrogram into the downstream speech recognition model to achieve speech emotion classification, environmental sound recognition, or pronunciation quality assessment tasks.

[0011] In a preferred embodiment, in step 1, the integral formula for the fractional Fourier transform is defined as shown in equation (1):

[0012]

[0013] in, It is a fractional Fourier transform operator; It is the kernel function of the fractional Fourier transform, which depends on the order p, and its expression is shown in equation (2):

[0014]

[0015] in, , The rotation angle of the time-frequency plane determines the degree of transformation; kernel function. It has orthogonality, different The kernel functions of the values ​​are orthogonal to each other, ensuring that different orders do not interfere with each other, and the original signal can be accurately recovered through inverse transform; FRFT covers all intermediate states between the pure time domain and the pure frequency domain, and can be adjusted... The value optimizes the signal energy distribution in a wider time-frequency space.

[0016] In a preferred embodiment, the audio signal is converted into a spectrogram using short-time Fourier transform, as shown in equation (3), where x(t) is the input audio signal, w(t-τ) is the window function, and X(τ,ω) is the time-frequency representation of the signal in time τ and frequency ω.

[0017]

[0018] First, the audio signal is divided into multiple overlapping segments using a window function. A Fourier transform is performed on the signal segment within each window to obtain the frequency domain representation of the signal segment. The power spectrum S(τ,ω) of each signal segment is then calculated, as shown in equation (4).

[0019]

[0020] The power spectra of all time segments are arranged in chronological order. The horizontal axis represents time and the vertical axis represents frequency. The color intensity represents the energy magnitude at the time-frequency point, and finally, a time-frequency energy distribution spectrogram is obtained.

[0021] In a preferred embodiment, step 4, the quantum-derived Newton-Raphson optimization algorithm QNRBO, specifically includes:

[0022] 1) Initialize the population using quantum encoding and expand the search space through quantum superposition states;

[0023] 2) The quantum rotation gate dynamically adjusts the quantum state of individuals, guiding the population to evolve towards the global optimum;

[0024] 3) Quantum mutation and periodic catastrophe mechanisms;

[0025] 4) Combine Lévy flight with simulated annealing strategy to avoid premature convergence.

[0026] In a preferred embodiment, quantum encoding is used for population initialization, as shown in Equation (5). Based on the upper and lower bounds of each dimension and the required precision, the number of binary encoding bits for each dimension is calculated, and the initial population is quantum encoded based on the number of binary encoding bits.

[0027] (5)

[0028] in, It is the relative phase of the signal. It is the rotation angle update rate of the quantum rotating gate. It is a diversity maintenance factor. It is the quantum rotation angle. The unit modulus complex number in the complex plane is used to construct the transformation kernel. Based on the number of quantum bits and the population size, the entire quantum population is represented as a matrix, and the elements of the matrix are generated using equation (6).

[0029] (6)

[0030] When a quantum individual is measured, each quantum bit will collapse into a classical bit 0 / 1 with the probability shown in Equation (7), and the final binary population matrix is ​​shown in Equation (8).

[0031] (7)

[0032] (8)

[0033] in, Let represent the j-th quantum rotation angle of the i-th element. Decode the binary population into actual variable values ​​according to equation (9), and map them to the domain [LB, UB] of the actual problem. Let the d-th dimension be allocated... If the binary code is 1 bit, then:

[0034] (9)

[0035] in, It is the decimal integer corresponding to the binary string of the i-th individual in the d-th dimension; These are the corresponding actual variable values. It is the lower bound of the domain of dimension d. It is the upper bound of the d-th dimension domain. The number of binary code bits is allocated in the d-th dimension.

[0036] In a preferred embodiment, the rotating door update strategy combines two key factors: the difference between the current individual and the globally optimal individual, which guides the individual toward the optimal direction; and the random difference between the current individual and other individuals, which introduces diversity to avoid getting trapped in local optima; assuming the update formula for the i-th individual at the j-th qubit angle is as follows:

[0037] (10)

[0038] in, It is the perspective of the currently globally optimal individual in the j-th dimension. It represents the angle of a randomly selected individual in the j-th dimension. WEP(t) is the weight exponent parameter, which changes dynamically with the number of iterations t; TDR(t) is the rate of decrease, which also changes with the number of iterations. It is the angle change, used to drive the individual to evolve towards the optimal direction; WEP and TDR are defined as shown in equation (11):

[0039] (11)

[0040] In the early stages of iteration, a larger WEP value drives individuals to move quickly toward the optimal direction, while in the later stages of iteration, a smaller TDR value enhances the algorithm's local search capability.

[0041] In a preferred embodiment, a quantum mutation operation is introduced to randomly perturb the individual optimization, triggered with a fixed probability, simulating the gene mutation behavior in biological evolution: each quantum bit of the individual has a certain probability of being randomly adjusted; the mutation amplitude changes dynamically with the iteration process, with stronger mutations in the early stage and weaker mutations in the later stage; according to the above three rules, the quantum mutation operation is as shown in equation (12):

[0042] (12)

[0043] in, , , The maximum variation range is represented by DF, which is the variation control factor that determines whether the variation is biased towards directional or random variation.

[0044] A quantum catastrophe mechanism is introduced: some individuals are randomly replaced; the best individuals are retained and the rest are reinitialized; the mutation intensity is increased; when the algorithm fails to find a better solution for several consecutive generations, a catastrophe operation is triggered to perform a large-scale reset of part or all of the population, simulating a "catastrophic event" in nature, thereby reactivating population diversity.

[0045] In a preferred embodiment, the simulated annealing mechanism is used: at high temperatures, particle motion is intense, and the system exhibits strong randomness. As the temperature decreases, the system gradually tends towards a stable state, and the probability of accepting a poor solution also decreases, as shown in equation (13):

[0046] (13)

[0047] Where T represents the current temperature, This represents the change in fitness, where T0 is the initial temperature and β is the cooling coefficient ranging from [0,1].

[0048] The Lévy flight mechanism is introduced by introducing a random step size that follows a Lévy distribution, as shown in equation (14):

[0049] (14)

[0050] Where u and v are random numbers following a standard normal distribution; γ is the Lévy exponent; and s is the generation step size, which is updated according to the step size as shown in equation (15):

[0051] (15)

[0052] in, This is the step size scaling factor, which is dynamically adjusted with each iteration.

[0053] In a preferred embodiment, the QNRBO algorithm is used to optimize the hyperparameters of the fractional-order spectrogram generation, and information entropy is selected as the objective function. Information entropy is usually used to measure the uncertainty or complexity of a signal in its transform domain. For the discrete signal x[n], the squared amplitude obtained after transformation is regarded as a probability distribution p. i Where i represents different components of the transform domain; the information entropy H is defined as shown in equation (16):

[0054] (16)

[0055] in, X[i] represents the proportion of the energy of the i-th transform coefficient to the total energy, and X[i] represents the i-th coefficient of the signal x[n] in the transform domain.

[0056] This invention also provides a quantum-derived Newton-Raphson optimal fractional-order spectrogram generation system, comprising a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executed by the processor; characterized in that, when the system is running, the processor and the memory communicate via the bus, and the machine-readable instructions executed by the processor are as described in the quantum-derived Newton-Raphson optimal fractional-order spectrogram generation method.

[0057] Compared with existing technologies, this invention has the following advantages: This paper proposes an optimal fractional-order spectrogram generation method based on a quantum-derived Newton-Raphson algorithm to improve the performance of speech recognition and classification tasks. Addressing the problems of high computational complexity, lack of diversity maintenance, and susceptibility to local optima in the Newton-Raphson algorithm, a quantum-derived Newton-Raphson optimization algorithm (QNRBO) is proposed. To address the difficulty of balancing time-frequency concentration and cross-term suppression in traditional spectrograms, and the limitation of providing only a fixed time-frequency resolution, a fractional-order spectrogram is generated through fractional-order Fourier transform, and then scaled using a Mel filter to obtain a fractional-order Mel spectrogram. Using the minimization of information entropy as the objective function, the QNRBO algorithm is used to adaptively optimize hyperparameters such as order, frame length, and frame shift to obtain the optimal fractional-order spectrogram. Experimental results show that the QNRBO algorithm has stronger global optimization capabilities than existing optimization algorithms and higher stability in solving high-dimensional complex problems; the generated optimal fractional-order spectrogram can effectively focus signal energy and enhance feature separability. Attached Figure Description

[0058] Figure 1 This is a schematic diagram of a fractional-order spectrogram generation method according to a preferred embodiment of the present invention;

[0059] Figure 2 This is a schematic diagram of the optimal fractional-order spectrogram generation process according to a preferred embodiment of the present invention;

[0060] Figure 3 The diagram shows the iteration curves of different optimization algorithms in a preferred embodiment of the present invention. (a) is the F1 iteration curve, (b) is the F2 iteration curve, (c) is the F3 iteration curve, (d) is the F4 iteration curve, (e) is the F5 iteration curve, (f) is the F6 iteration curve, (g) is the F7 iteration curve, (h) is the F8 iteration curve, (i) is the F9 iteration curve, (j) is the F10 iteration curve, (k) is the F11 iteration curve, and (l) is the F12 iteration curve.

[0061] Figure 4Box plots of different optimization algorithms in preferred embodiments of the present invention are shown, wherein (a) is an F1 box plot, (b) is an F2 box plot, (c) is an F3 box plot, (d) is an F4 box plot, (e) is an F5 box plot, (f) is an F6 box plot, (g) is an F7 box plot, (h) is an F8 box plot, (i) is an F9 box plot, (j) is an F10 box plot, (k) is an F11 box plot, and (l) is an F12 box plot;

[0062] Figure 5 This is an example of the optimal spectrogram of happy emotion in a preferred embodiment of the present invention, wherein (a) is QNRBO, (b) is NRBO, (c) is QPSO, (d) is IVY, (e) is ANT, and (f) is CUCKOO;

[0063] Figure 6 This is an example of the optimal spectrogram of sadness emotion in a preferred embodiment of the present invention, wherein (a) is QNRBO, (b) is NRBO, (c) is QPSO, (d) is IVY, (e) is ANT, and (f) is CUCKOO;

[0064] Figure 7 The above are the recognition confusion matrices of three datasets in a preferred embodiment of the present invention, wherein (a) is the RAVDESS dataset, (b) is the UrbanSound8K dataset, and (c) is the singing vowel pronunciation dataset. Detailed Implementation

[0065] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0066] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0067] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0068] A quantum-derived Newton-Raphson optimal fractional-order spectrogram generation method, referenced Figure 1-7To address the numerous bottlenecks in existing speech signal processing technologies for high-precision speech recognition and evaluation tasks, this invention proposes a speech enhancement method based on quantum-derived optimal fractional-order Mel spectrograms. This method aims to systematically solve core problems such as insufficient traditional time-frequency representation capabilities, rigid hyperparameter configuration, limited performance of optimization algorithms, and the disconnect between features and tasks.

[0069] First, to overcome the limitations of the traditional Short Time Fourier Transform (STFT) due to the Heisenberg uncertainty principle and its difficulty in achieving both time and frequency resolution, this invention introduces the Fractional Fourier Transform (FRFT). FRFT maps the signal to an arbitrary rotational domain between the time and frequency domains through an adjustable fractional order α, thus allowing for more flexible focusing of speech energy with complex frequency modulation characteristics. Based on this, nonlinear scaling compression is performed using a Mel filter bank to generate a fractional Mel spectrogram. This retains the superior representation ability of FRFT for non-stationary signals while incorporating the characteristics of human auditory perception, significantly improving the physical meaning and discriminative power of the features.

[0070] Secondly, addressing the shortcomings of existing methods where key hyperparameters such as fractional-order parameters, frame length, and frame shift rely on manual experience and lack adaptive optimization mechanisms, this invention constructs an adaptive optimization framework with information entropy minimization as the objective function. This objective function measures the information fidelity between the generated spectrogram and the original speech signal; the lower the entropy value, the more concentrated and complete the spectrogram reflects the energy distribution of the original signal, and the clearer the feature representation. By minimizing this objective function, the system can be driven to automatically search for the optimal hyperparameter combination, achieving "data-driven" spectrogram customization.

[0071] More importantly, to efficiently solve this complex high-dimensional, non-convex, and multi-modal optimization problem, this invention designs and implements a quantum-derived Newton-Raphson optimization algorithm (QNRBO). This algorithm integrates quantum computing concepts with classical numerical optimization methods: it uses quantum encoding to initialize the population, expanding the initial search space; it utilizes an adaptive quantum rotation gate to guide individuals towards the global optimum; and it introduces quantum mutation and periodic catastrophe mechanisms, embedding Lévy flight and simulated annealing strategies to effectively maintain population diversity, avoid premature convergence, and significantly enhance the algorithm's global exploration capability and local escape robustness. Compared to traditional optimizers, QNRBO can stably converge to a high-quality solution in a shorter time, providing a feasible path for real-time or large-scale speech processing.

[0072] Finally, the optimal fractional-order Mel spectrogram generated by this invention is directly applied to downstream recognition tasks (such as multimodal evaluation of singing pronunciation quality), ensuring a high degree of alignment between feature extraction and the final target. Experimental verification shows that, compared to fixed-parameter FRFT or STFT-Mel spectrograms, the spectrogram generated by this method has more concentrated energy and clearer texture, significantly improving classification accuracy, recall, and F1 score under the same network depth architecture, fully demonstrating its effectiveness and advancement in improving the performance of speech recognition systems. In summary, this invention, through a three-pronged technical approach of "theoretical innovation—algorithmic breakthrough—task-oriented," comprehensively solves the key defects of existing speech enhancement technologies, providing solid support for high-precision intelligent speech analysis.

[0073] The detailed technical solution is as follows:

[0074] (I) Construction of fractional-order spectrograms

[0075] (1) Short-time fractional Fourier transform

[0076] The fractional Fourier transform (FRFT) can be understood as a representation of a signal in the fractional Fourier domain formed by rotating the coordinate axes counterclockwise around the origin in the time-frequency plane by any angle. It integrates information from both the time and frequency domains, making it a time-frequency analysis method. FRFT is the expansion of a signal over a set of orthogonal chirp bases, and the FRFT of a chirp signal at a certain order is also a delta function. Therefore, FRFT has excellent focusing properties for chirp signals. The integral formula for the fractional Fourier transform is defined as shown in equation (1).

[0077]

[0078] in, It is a fractional Fourier transform operator. It is a kernel function that depends on the order p, and its expression is shown in equation (2).

[0079]

[0080] in, , The rotation angle of the time-frequency plane determines the degree of transformation; kernel function. It has orthogonality, different The kernel functions of the values ​​are orthogonal to each other, ensuring that different orders do not interfere with each other, and the original signal can be accurately recovered through inverse transform. FRFT covers all intermediate states between the pure time domain and the pure frequency domain, and can be adjusted... The value can optimize the signal energy distribution in a wider time-frequency space, providing better time-frequency resolution, thereby adapting to different types of signal characteristics.

[0081] (2) Generation of fractional-order Mel spectrograms

[0082] A spectrogram is a visualization tool used to represent the energy distribution of a signal in two dimensions: time and frequency. It is particularly suitable for analyzing non-stationary signals such as audio. The audio signal is converted into a spectrogram using a short-time Fourier transform, as shown in Equation (3), where x(t) is the input audio signal, w(t-τ) is the window function, and X(τ,ω) is the time-frequency representation of the signal in time τ and frequency ω.

[0083]

[0084] First, the audio signal is divided into multiple overlapping segments using a window function. Fourier transform is performed on the signal segment within each window to obtain the frequency domain representation of the signal segment. The power spectrum S(τ,ω) of each signal segment is then calculated, as shown in Equation (4).

[0085]

[0086] The power spectra of all time segments are arranged in chronological order. The horizontal axis represents time and the vertical axis represents frequency. The color intensity indicates the energy magnitude at the time-frequency point, resulting in a time-frequency energy distribution spectrogram. The construction of the spectrogram is constrained by the Heisenberg uncertainty principle, making it difficult to achieve a balance between time-frequency concentration and cross-term suppression. Furthermore, it only provides a fixed time-frequency resolution, making it difficult to adapt to different types of signals.

[0087] This invention addresses the limitations of traditional spectrograms by incorporating fractional Fourier transform into the generation of fractional-order spectrograms. This allows for optimization of signal energy distribution across a wider time-fractional frequency space, thereby improving time-frequency resolution. The generation of fractional-order spectrograms combines the theories of fractional Fourier transform and traditional spectrograms, and uses a Mel filter to compress the scale of the fractional-order spectrogram. The specific implementation process is as follows: Figure 1 As shown.

[0088] The generation quality of fractional-order spectrograms under different tasks is significantly affected by parameters such as the fractional-order number, frame length, and frame shift. Therefore, this invention proposes a quantum-derived Newton-Raphson optimization algorithm to optimize the hyperparameters for fractional-order spectrogram generation and obtain the optimal fractional-order spectrogram.

[0089] (II) Quantum-derived optimal fractional-order acoustic spectrogram

[0090] (1) Quantum Newton-Raphson algorithm

[0091] The traditional Newton-Raphson algorithm initiates the search for the optimal solution by generating an initial random population within the boundary of candidate solutions. The random population is generated by equation (5), and each population consists of a fuzzy decision vector.

[0092]

[0093] in, Let i represent the i-th dimension of the n-th population, rand represents a random number between [0,1], and lb and ub represent the lower and upper bounds of the population range.

[0094] The algorithm uses the Newton-Raphson search strategy to update the optimal position of the population. This strategy uses the Taylor series expansion of the function to perform linear approximation near the current point and find a position closer to the true root, as shown in Equation (6).

[0095]

[0096] in, Indicates the current solution position. This indicates the updated solution position. To prevent getting trapped in local optima, the Newton-Raphson algorithm introduces a trap avoidance strategy, as shown in equation (7).

[0097]

[0098] Where θ1 is the random direction coefficient, controlling movement towards the optimal direction; θ2 is the population perturbation coefficient, introducing the influence of the population's average position; μ1 and μ2 are weighting factors, dynamically adjusting the influence of the optimal solution and the current solution; δ is the dynamic decay factor, ensuring that later perturbations weaken and tend towards convergence; x b X is the current optimal solution. n IT For the current solution, X n IT+1 An updated solution that does not introduce trap avoidance strategies.

[0099] The Newton-Raphson algorithm suffers from high computational complexity and lacks diversity maintenance, making it prone to getting trapped in local optima. Therefore, this invention combines quantum computing with optimization ideas to propose a quantum-derived Newton-Raphson optimization algorithm.

[0100] This invention deeply integrates quantum computing concepts on the basis of NRBO, specifically in the synergistic design of the three core mechanisms of quantum coding (QE), which jointly ensure the balance between population diversity and convergence efficiency.

[0101] 1) Quantum population initialization

[0102] The population is initialized using quantum encoding, as shown in Equation (8). The number of binary encoding bits for each dimension is calculated based on the upper and lower bounds of each dimension and the required precision. The initial population is then quantum encoded based on the number of binary encoding bits.

[0103]

[0104] Based on the number of qubits and the population size, the entire quantum population can be represented as a matrix, and the elements of the matrix are generated by equation (9).

[0105]

[0106] When a quantum individual is measured, each quantum bit collapses into a classical bit 0 / 1 with the probability shown in Equation (10), and the final binary population matrix is ​​shown in Equation (11).

[0107]

[0108]

[0109] The binary population is decoded into actual variable values ​​according to equation (12) and mapped to the domain [LB, UB] of the actual problem. Let the d-th dimension be allocated... If the binary code is 1 bit, then:

[0110]

[0111] in, It is the decimal integer corresponding to the binary string of the i-th individual in the d-th dimension; These are the corresponding actual variable values.

[0112] 2) Quantum Revolving Door

[0113] This invention introduces a quantum rotation gate mechanism to dynamically adjust the quantum state representation of individuals, thereby guiding the population to evolve towards a better solution. The quantum rotation gate changes the measurement probability distribution by adjusting the angle of the qubits. In optimization algorithms, this can simulate the process of "individuals moving towards the optimal direction," and the balance between exploration and exploitation is controlled by parameters. The rotation gate update strategy of this invention combines two key factors: the difference between the current individual and the globally optimal individual, causing the individual to move towards the optimal direction; and the random difference between other individuals, introducing diversity to avoid getting trapped in local optima. Assume the update formula for the angle of the j-th qubit of the i-th individual is as follows:

[0114]

[0115] in, It is the perspective of the currently globally optimal individual in the j-th dimension. It represents the angle of a randomly selected individual in the j-th dimension. WEP(t) is the weight exponent parameter, which changes dynamically with the number of iterations t; TDR(t) is the rate of decrease, which also changes with the number of iterations. It is the change in angle, used to drive the individual to evolve towards the optimal direction. WEP and TDR are defined as shown in equation (5):

[0116]

[0117] In the early stages of iteration, a larger WEP value drives individuals to move quickly toward the optimal direction. In the later stages of iteration, a smaller TDR value can enhance the algorithm's local search capability.

[0118] 3) Quantum mutation and catastrophe

[0119] To maintain population diversity and avoid premature convergence during optimization, a quantum mutation operation is introduced to randomly perturb individual optimizations. This is triggered with a fixed probability, simulating gene mutation behavior in biological evolution: each quantum bit of an individual has a certain probability of being randomly adjusted; the mutation amplitude changes dynamically with the iteration process, with stronger mutations in the early stages and weaker mutations in the later stages; the direction of mutation can be towards the optimal solution or a random perturbation. Based on the above three rules, the quantum mutation operation is shown in equation (15):

[0120]

[0121] in, , , The maximum variation range is represented by DF, which is the variation control factor that determines whether the variation is biased towards directional or random variation.

[0122] In addition, to further enhance the algorithm's ability to escape local optima, a quantum catastrophe mechanism is introduced: randomly replacing some individuals; retaining the best individuals and reinitializing the rest; and increasing the mutation intensity. When the algorithm fails to find a better solution for several consecutive generations, a catastrophe operation is triggered, performing a large-scale reset of part or all of the population, simulating a "catastrophic event" in nature, thereby reactivating population diversity.

[0123] 4) Hybrid escape strategy combining Lévy flight and simulated annealing

[0124] Although quantum mechanisms enhance exploration capabilities, algorithms can still be trapped in local minima in highly rugged, multi-peaked optimization terrains (such as information entropy surfaces). To enhance the algorithm's ability to escape local optima, simulated annealing is introduced. This mechanism allows the algorithm to accept a worse solution with a certain probability, thereby avoiding local convergence traps. The principle of this mechanism is derived from the physical annealing process: at high temperatures, particle motion is intense, and the system has strong randomness. As the temperature decreases, the system gradually tends to a stable state, and the probability of accepting a worse solution also decreases, as shown in Equation (16).

[0125]

[0126] Where T represents the current temperature, The fitness change is represented by T0, the initial temperature, and β, the cooling coefficient, which is in the range [0,1].

[0127] To further enhance global search capabilities, a Lévy flight mechanism is introduced. By introducing a random step size that follows a Lévy distribution, individuals can make non-uniform long-distance jumps in the search space, which helps to discover new regions far away from the current population. The Lévy distribution is shown in Equation (17).

[0128]

[0129] Where u and v are random numbers that follow a standard normal distribution; γ is the Lévy exponent, which is usually taken as 1.5; s is the generation step size, which is updated according to the step size as shown in equation (18).

[0130]

[0131] in, This is the step size scaling factor, which is dynamically adjusted with each iteration.

[0132] (2) Generation of quantum-derived optimal fractional-order spectrograms

[0133] To obtain the optimal fractional-order spectrogram, the QNRBO algorithm is used to optimize the hyperparameters for generating the fractional-order spectrogram. Information entropy is chosen as the objective function, which is typically used to measure the uncertainty or complexity of a signal in its transform domain. For a discrete signal x[n], the squared amplitude obtained after transformation can be regarded as a probability distribution p. i , where i represents different components of the transform domain. The information entropy H is defined as shown in equation (19):

[0134]

[0135] in, X[i] represents the proportion of the energy of the i-th transform coefficient to the total energy, where X[i] represents the i-th coefficient of signal x[n] in the transform domain. The overall optimization process is as follows: Figure 2 As shown, the specific steps are as follows:

[0136] Step 1: Acquire the raw audio signal and initialize the parameters. Calculate the number of bits for binary encoding based on the upper and lower bounds of each dimension and the required precision, and initialize the quantum population matrix to ensure that the population has sufficient diversity to cover the solution space.

[0137] Step 2: Generate fractional-order spectrograms and evaluate fitness. Compare individual fitness values ​​with the global best fitness value. If the current individual is better, update the global best position.

[0138] Step 3: Determine if the termination condition has been met. If so, proceed to Step 6; otherwise, proceed to Step 4.

[0139] Step 4: Update the population state by randomly perturbing the population using quantum rotation operations and resetting the population using quantum mutation and catastrophe.

[0140] Step 5: Introduce simulated annealing and Lévy flight mechanisms to enhance global search capabilities, then proceed to Step 2.

[0141] Step 6: Output the optimal fractional-order spectrogram.

[0142] (III) Setup of the experimental environment and construction of the dataset

[0143] The hardware configuration for all experiments in this invention was a high-performance workstation equipped with an Intel(R) Core(TM) i9-13900K processor (3.0GHz, 24 cores), 64GB DDR5 memory, and an NVIDIA GeForce RTX 4090 (24GB) graphics card. The software environment was based on Matlab 2022 and Python 3.9 programming languages, with PyTorch 1.13.1 as the core deep learning framework.

[0144] To verify the effectiveness of the method of the present invention in various speech recognition enhancement tasks, a comprehensive experimental test object was constructed, which integrates multiple representative speech datasets, covering natural dialogue, emotional speech and professional singing pronunciation scenarios.

[0145] The UrbanSound8K dataset is a widely used public benchmark dataset for environmental sound classification research. It covers 10 categories of urban environmental sounds, including air conditioners, car horns, dog barks, and drilling. It is often used to evaluate the classification performance and robustness of audio feature extraction methods, machine learning, and deep learning models in real and varied acoustic environments.

[0146] The RAVDESS dataset contains audio and video data of eight basic emotions performed by 24 professional actors, covering spoken sentences and singing clips, and is widely used in speech emotion recognition research. This experiment uses its audio portion to validate an emotion classification task.

[0147] A self-built vowel singing pronunciation dataset: This study, in collaboration with the School of Music at Fujian Normal University, constructed a professional vocal training speech dataset. The dataset consists of 2,000 video files recorded by 50 music students during classroom vocal exercises. Each audio clip is 10 seconds long and contains sustained pronunciation of a single vowel ( / a / , / e / , / i / , / o / , / u / ). Based on multiple vocal evaluation dimensions, including vocal range, timbre, volume, pitch, breath control, and resonance area, professional teachers comprehensively scored the pronunciation quality and categorized it into six classes: "Male A," "Male B," "Male C," "Female A," "Female B," and "Female C" (i.e., three levels for each gender). This dataset focuses on fine-grained feature analysis of singing speech to evaluate the model's ability to classify professional pronunciation quality.

[0148] (1) Performance test experiment of quantum-derived optimization algorithm

[0149] To systematically verify the effectiveness of each core mechanism and its incremental contribution in the proposed quantum-derived Newton-Raphson optimization algorithm (QNRBO), this invention adopts a bottom-up ablation strategy. Starting from a purely classical Newton-Raphson optimizer, it gradually introduces components such as quantum encoding (QE), quantum rotation gate (QRG), quantum mutation and catastrophe (MC), Lévy flight and simulated annealing (SA) to construct five progressive variants A0–A4.

[0150] Among them, A0 adopts traditional real number random initialization and retains only the Newton-Raphson search rule as a baseline without any quantum mechanism; A1 introduces quantum binary encoding, initializes the population through quantum superposition, and generates the initial solution through measurement collapse; A2 adds a quantum rotation gate on the basis of A1 to guide the quantum state to evolve in the optimal direction; A3 further integrates quantum mutation and periodic catastrophe operations to enhance population diversity; A4 introduces the Lévy flight mechanism and integrates simulated annealing strategy to improve global exploration capability and local escape robustness.

[0151] All variants were run independently 30 times on the CEC2022 standard test function set, as shown in Table 1. The contribution of each component to the overall algorithm performance was quantified by comparing their average fitness value, standard deviation, convergence curve, and statistical significance. The convergence curve for obtaining the minimum value is shown in Table 1. Figure 3 As shown in the figure, the experimental results are shown in Table 2.

[0152] surface 2022 CEC Standard Test Functions

[0153] Table 2. Experimental results of QNRBO algorithm ablation

[0154] Table 2 shows that the performance of all functions is significantly improved from the classical baseline A0 to A1 with quantum encoding, confirming that quantum superposition initialization effectively enhances population diversity. Further addition of the quantum rotation gate (A2) reduces the convergence value of the single-peak function F1 by 18.3%, indicating its efficient guidance of the search direction. Integrating quantum mutation and catastrophe (A3) significantly improves stability, reducing the standard deviation of F11 by 42.6%, effectively suppressing premature convergence. Finally, fusing Lévy flight and simulated annealing to form the complete QNRBO (A4) further reduces the mean of the high-dimensional combinatorial function F11 by 28.0%, demonstrating excellent global exploration capabilities. Among these, QNRBO performs slightly worse than A3 on F6, due to the mismatch between F6's narrow optimal region and Lévy's large step size, which easily leads to overshoot. However, it outperforms all other 11 functions, showing a significant overall advantage. In summary, the components are progressively integrated and complementary in function, which together give QNRBO an excellent balance between convergence speed, exploration capability and robustness, providing solid support for its effectiveness in practical tasks such as speech hyperparameter optimization.

[0155] To verify the performance of the QNRBO algorithm of this invention, its results were compared and analyzed with those of the Projection-Iterative-Methods-based Optimizer (PIMO), Fatamorgana algorithm (FATA), Quantum-behaved Particle Swarm Optimization (QPSO), Ant Colony Optimization (ANT), Ivy Algorithm (IVY), and Cuckoo Search Algorithm (CUCKOO). In this experiment, the population size was fixed at 50, and the number of iterations was 1000. To avoid random results, the same algorithm was tested 30 times, and the average of the results was taken as the final result.

[0156] The iterative curves of the compared optimization algorithms on 12 standard test functions are as follows: Figure 3 As shown in Table 3, the statistical results of the evaluation metrics for the 12 standard test functions are presented. The box plots on the 2022 standard test function set are shown below. Figure 4 As shown.

[0157] from Figure 3 It is evident that, on certain test functions, the optimal solution of the proposed QNRBO algorithm is slightly inferior to that of comparative algorithms such as IVY or NRBO. However, the comprehensive performance evaluation of optimization algorithms should consider optimality, stability, and robustness, rather than relying solely on the minimum value of a single run. As shown in Table 2, although IVY achieves a minimum value close to the theoretical lower bound on F1, its mean and standard deviation are much higher than QNRBO, indicating that its results are highly dependent on random initialization and have low reliability in practical applications. Conversely, from... Figure 4 The box plots show that QNRBO exhibits the lowest or second-lowest standard deviation across all 12 CEC2022 test functions, particularly in high-dimensional multimodal (F3), mixed (F5), and combined (F11) functions, where its mean significantly outperforms all comparable algorithms. This is attributed to the algorithm's integration of the directional search capability of quantum rotating gates, the long-distance jump mechanism of Lévy flight, and the escape strategy from local minima using simulated annealing, thus achieving a better trade-off between exploration-exploration balance and convergence stability.

[0158] Table 3 Performance indicators of different optimization algorithms

[0159] To further evaluate the computational efficiency of the proposed QNRBO algorithm, this invention recorded the average execution time (in seconds) of each ablation variant (A0–A4) and mainstream comparison algorithms on the CEC2022 test function (dimension dim=10). The results are shown in the last row of Table 2. The experimental environment was an Intel(R) Core(TM) i9-13900K processor and 64GB of memory. As can be seen, the baseline method A0 takes an average of 57.4 seconds because it needs to approximate the Hessian matrix in each iteration. After introducing quantum encoding (A1), the parallelism in the initialization phase slightly increases the overhead, but with the quantum rotation gate (A2) guiding the efficient search direction, the total number of iterations is reduced, and the time is reduced to 53.6 seconds. Although the addition of quantum mutation and catastrophe (A3) introduces additional perturbation operations, it effectively avoids redundant iterations caused by premature convergence, and the time is stabilized at 52.3 seconds. The final version A4 (which combines Lévy flight and simulated annealing, although slightly more complex in a single iteration, has significantly enhanced global exploration capabilities, accelerated convergence speed, and further reduced the average time to 51.2 seconds, which is comparable to the original NRBO) is not much different.

[0160] In comparison, lightweight algorithms such as IVY and QPSO have lower time overhead, but their optimization accuracy is significantly lower than QNRBO; while the ANT algorithm is the most time-consuming because it needs to maintain a global pheromone matrix. In summary, QNRBO achieves significant performance improvements with only a slight increase in computational cost, and is particularly suitable for offline processing scenarios such as speech emotion analysis and pronunciation quality assessment where recognition accuracy is critical.

[0161] (2) Quantum-derived optimal fractional-order spectrogram generation experiment

[0162] To verify the effectiveness of the QNRBO algorithm in optimizing fractional-order spectrograms, this invention uses the RAVDESS sentiment classification public dataset for experiments. Fractional-order spectrograms were generated for emotions such as happiness, sadness, anger, calmness, surprise, frustration, and fear, with the information entropy of the generated spectrograms used as the evaluation metric. The population size was set to 50, the maximum number of iterations was set to 1000, and different optimization algorithms were optimized 10 times, with the average value taken as the final parameter metric.

[0163] Table 4 shows the optimal parameters for the fractional-order Mel spectrograms of each algorithm. It can be seen that the QNRBO algorithm of this invention performs exceptionally well in optimizing the fractional-order Mel spectrograms of audio signals representing seven emotions. Compared to other optimization algorithms such as NRBO, BKA, FATA, QPSO, IVY, ANT, and CUCKOO, QNRBO not only achieves a lower minimum information entropy but also provides better choices in frame shift, frame length, and order parameter settings, thereby generating spectrograms with higher clarity and energy concentration. Importantly, this performance advantage does not come at the expense of computational efficiency. As shown in the last row of Table 4, the average optimization time of QNRBO is comparable to that of lightweight algorithms. This indicates that QNRBO effectively avoids high-complexity computations through a quantum derivation mechanism, achieving efficient parameter search while maintaining high accuracy, laying a solid foundation for its application in practical speech processing systems. Optimal fractional-order spectrogram generation algorithms for happiness and sadness emotions, such as... Figure 5 and Figure 6 As shown, the spectrogram optimized by QNRBO can more accurately reflect the emotional characteristics of the original audio signal and effectively reduce noise interference.

[0164] Table 4 Optimization results of spectrograms for different emotions

[0165] In summary, the fractional-order spectrogram optimization method based on the quantum-derived Newton-Raphson algorithm proposed in this invention has significant advantages, demonstrating good results in both algorithm performance and practical applications. It is particularly suitable for audio signal processing tasks that require adaptive parameter adjustment to achieve the best representation effect.

[0166] (3) Experiment on speech recognition enhancement effect

[0167] To verify the enhancement effect of the optimal fractional-order spectrogram on speech recognition in this invention, the MultimodalCNN model, which has been the most effective in speech classification and recognition in recent years, was used to perform speech emotion recognition, speaker recognition, and singing vowel pronunciation quality level evaluation on the RAVDESS dataset, UrbanSound8K dataset, and a self-constructed singing vowel pronunciation dataset, respectively. The effectiveness of the method of this invention was verified by the improvement in speech recognition accuracy. A control group was set up in this experiment, and recognition experiments were conducted using MFCC, ordinary spectrogram, single-order fractional-order spectrogram, and optimal fractional-order spectrogram as input features of the MultimodalCNN model.

[0168] This invention's optimal fractional-order spectrogram, through adaptive optimization of the fractional-order parameter α, achieves efficient focusing of signal energy in the optimal transform domain, significantly enhancing the discriminative power of key features such as emotional prosody, environmental noise patterns, and vocal quality. The confusion matrices for the three datasets are shown below. Figure 7 As shown, the method of this invention exhibits high classification accuracy and low cross-class misclassification rate on three datasets, especially showing outstanding advantages in easily confused category pairs such as "sadness and neutral" and "drill v and crusher". This fully demonstrates its robustness and generalization ability in complex and fine speech classification tasks, and is a high-precision speech feature representation framework that is superior to traditional methods.

[0169] Table 5 shows the statistical results of recognition performance for different audio features across the three datasets. Traditional spectrograms and MFCC features generally suffer from inter-class confusion when handling different speech classification tasks, especially when distinguishing between categories with similar timbres, overlapping spectra, or complex time-frequency dynamics. Although single-order fractional spectrograms improve feature representation to some extent by introducing fractional transformations, their fixed parameters limit their adaptability to different speech characteristics, resulting in no significant performance improvement. The optimal fractional-order spectrogram optimized by QNRBO in this invention significantly improves recognition accuracy in multiple audio classification tasks: In the RAVDESS emotion recognition task, its F1 score reaches 0.882, nearly 5 percentage points higher than the second-best method, effectively reducing misjudgments of easily confused emotions; In the UrbanSound8K environmental sound classification task, the F1 score is 0.772, better than traditional spectrograms and MFCC, demonstrating its strong ability to distinguish spectrally similar sounds in noisy environments; In the singing vowel pronunciation quality assessment task, the F1 score reaches 0.834, with an accuracy of 0.855, demonstrating excellent characterization of subtle acoustic features in professional speech, significantly better than unoptimized single-order fractional-order spectrograms and other traditional features, fully verifying the superiority and universality of this method in complex audio analysis.

[0170] Table 5 Experimental results for different audio features

[0171] Table5Experimentalresultsofdifferentaudiofeatures

[0172] Experimental results show that the optimal fractional-order spectrogram based on QNRBO optimization proposed in this invention achieves the best performance in three tasks: RAVDESS emotion recognition, UrbanSound8K environmental sound classification, and vocal vowel pronunciation quality assessment. This method, through adaptive optimization of the fractional-order parameter α, effectively focuses signal energy and significantly enhances feature discriminativity. It comprehensively surpasses traditional spectrograms, MFCC, and single-order fractional-order spectrograms with fixed parameters in terms of accuracy, recall, and F1 score. Its advantages are particularly prominent in distinguishing easily confused categories, fully verifying its superiority and universality in complex audio analysis.

[0173] The method of this invention outperforms traditional feature methods in multiple speech classification tasks, especially in distinguishing easily confused categories. Compared with the suboptimal single-order fractional spectrogram method, this method improves the accuracy by 2.2 percentage points in the RAVDESS emotion recognition task, 3.3 percentage points in the UrbanSound8K environmental sound classification task, and 4.7 percentage points in the singing vowel pronunciation quality assessment task, with an average classification accuracy improvement of 3.4 percentage points. Simultaneously, its F1 score is improved by 4.9, 1.5, and 2.2 percentage points on the RAVDESS, UrbanSound8K, and singing vowel datasets, respectively, compared to the suboptimal method. These significant performance improvements fully validate the superiority and universality of this method in complex audio analysis, providing a new paradigm of high precision and robustness for speech feature extraction, and demonstrating promising application prospects.

Claims

1. A method for generating quantum-derived Newton-Raphson optimal fractional-order spectrograms, characterized in that, Includes the following steps: Step 1: Acquire the original audio signal and construct a fractional Fourier transform (FRFT) spectrum, where the fractional order 'a' is an adjustable parameter used to rotate the signal by any angle in the time-frequency plane to optimize the signal energy distribution. Step 2: The fractional-order spectrogram is nonlinearly scaled using a Mel filter bank to generate a fractional-order Mel spectrogram, in order to integrate the characteristics of human auditory perception. Step 3: Using the minimization of information entropy as the objective function, construct an adaptive optimization framework to measure the information fidelity between the spectrogram and the original signal; Step 4: Use the quantum-derived Newton-Raphson optimization algorithm QNRBO to globally optimize the fractional order order, frame length, and frame shift hyperparameter to generate the optimal fractional order spectrogram; Step 5: Input the optimal fractional-order spectrogram into the downstream speech recognition model to achieve speech emotion classification, environmental sound recognition, or pronunciation quality assessment tasks.

2. The method for generating a quantum-derived Newton-Raphson optimal fractional-order spectrogram according to claim 1, characterized in that, In step 1, the integral formula for the fractional Fourier transform is defined as shown in equation (1): in, It is a fractional Fourier transform operator; It is the kernel function of the fractional Fourier transform, which depends on the order p, and its expression is shown in equation (2): in, , The rotation angle of the time-frequency plane determines the degree of transformation; kernel function. It has orthogonality, different The kernel functions of the values ​​are orthogonal to each other, ensuring that different orders do not interfere with each other, and the original signal can be accurately recovered through inverse transform; FRFT covers all intermediate states between the pure time domain and the pure frequency domain, and can be adjusted... The value optimizes the signal energy distribution in a wider time-frequency space.

3. The method for generating a quantum-derived Newton-Raphson optimal fractional-order spectrogram according to claim 1, characterized in that, The audio signal is converted into a spectrogram by using short-time Fourier transform, as shown in equation (3), where x(t) is the input audio signal, w(t-τ) is the window function, and X(τ,ω) is the time-frequency representation of the signal in time τ and frequency ω. First, the audio signal is divided into multiple overlapping segments using a window function. A Fourier transform is performed on the signal segment within each window to obtain the frequency domain representation of the signal segment. The power spectrum S(τ,ω) of each signal segment is then calculated, as shown in equation (4). The power spectra of all time segments are arranged in chronological order. The horizontal axis represents time and the vertical axis represents frequency. The color intensity represents the energy magnitude at the time-frequency point, and finally, a time-frequency energy distribution spectrogram is obtained.

4. The method for generating a quantum-derived Newton-Raphson optimal fractional-order spectrogram according to claim 1, characterized in that, In step 4, the quantum-derived Newton-Raphson optimization algorithm QNRBO specifically includes: 1) Initialize the population using quantum encoding and expand the search space through quantum superposition states; 2) The quantum rotation gate dynamically adjusts the quantum state of individuals, guiding the population to evolve towards the global optimum; 3) Quantum mutation and periodic catastrophe mechanisms; 4) Combine Lévy flight with simulated annealing strategy to avoid premature convergence.

5. The method for generating a quantum-derived Newton-Raphson optimal fractional-order spectrogram according to claim 4, characterized in that, Quantum encoding is used for population initialization, as shown in Equation (5). Based on the upper and lower bounds of each dimension and the required precision, the number of binary encoding bits for each dimension is calculated, and the initial population is quantum encoded based on the number of binary encoding bits. (5) in, It is the relative phase of the signal. It is the rotation angle update rate of the quantum rotating gate. It is a diversity maintenance factor. It is the quantum rotation angle. The unit modulus complex number on the complex plane is used to construct the transformation kernel; the entire quantum population is represented as a matrix according to the number of quantum bits and the population size, and the elements of the matrix are generated by equation (6); (6) When a quantum individual is measured, each quantum bit will collapse into a classical bit 0 / 1 with the probability shown in Equation (7), and the final binary population matrix is ​​shown in Equation (8). (7) (8) in, Represents the j-th quantum rotation angle of the i-th element; the binary population is decoded into actual variable values ​​according to equation (9), and mapped to the domain [LB, UB] of the actual problem. Let the d-th dimension be allocated... If the binary code is 1 bit, then: (9) in, It is the decimal integer corresponding to the binary string of the i-th individual in the d-th dimension; These are the corresponding actual variable values. It is the lower bound of the domain of dimension d. It is the upper bound of the d-th dimension domain. The number of binary code bits is allocated in the d-th dimension.

6. The method for generating a quantum-derived Newton-Raphson optimal fractional-order spectrogram according to claim 4, characterized in that, The rotating door update strategy combines two key factors: the difference between the current individual and the globally optimal individual, which drives the individual towards the optimal direction; and the random difference between the current individual and other individuals, which introduces diversity to avoid getting trapped in local optima. Assume the update formula for the i-th individual at the j-th qubit angle is as follows: (10) in, It is the perspective of the currently globally optimal individual in the j-th dimension. It represents the angle of a randomly selected individual in the j-th dimension. WEP(t) is the weight exponent parameter, which changes dynamically with the number of iterations t; TDR(t) is the rate of decrease, which also changes with the number of iterations. It is the angle change, used to drive the individual to evolve towards the optimal direction; WEP and TDR are defined as shown in equation (11): (11) In the early stages of iteration, a larger WEP value drives individuals to move quickly toward the optimal direction, while in the later stages of iteration, a smaller TDR value enhances the algorithm's local search capability.

7. The method for generating a quantum-derived Newton-Raphson optimal fractional-order spectrogram according to claim 4, characterized in that, Quantum mutation operation is introduced to randomly perturb individual optimization, triggered with a fixed probability, to simulate gene mutation behavior in biological evolution: each quantum bit of an individual has a certain probability of being randomly adjusted; the mutation amplitude changes dynamically with the iteration process, with stronger mutations in the early stage and weaker mutations in the later stage; according to the above three rules, the quantum mutation operation is as shown in equation (12): (12) in, , , The maximum variation range is represented by DF, which is the variation control factor that determines whether the variation is biased towards directional or random variation. A quantum catastrophe mechanism is introduced: some individuals are randomly replaced; the best individuals are retained and the rest are reinitialized; the mutation intensity is increased; when the algorithm fails to find a better solution for several consecutive generations, a catastrophe operation is triggered to perform a large-scale reset of part or all of the population, simulating a "catastrophic event" in nature, thereby reactivating population diversity.

8. The method for generating a quantum-derived Newton-Raphson optimal fractional-order spectrogram according to claim 4, characterized in that, Simulated annealing mechanism: At high temperatures, particle motion is intense, and the system exhibits strong randomness. As the temperature decreases, the system gradually tends towards a stable state, and the probability of accepting a poor solution also decreases, as shown in equation (13): (13) Where T represents the current temperature, This represents the change in fitness, where T0 is the initial temperature and β is the cooling coefficient ranging from [0,1]. The Lévy flight mechanism is introduced by introducing a random step size that follows a Lévy distribution, as shown in equation (14): (14) Where u and v are random numbers following a standard normal distribution; γ is the Lévy exponent; and s is the generation step size, which is updated according to the step size as shown in equation (15): (15) in, This is the step size scaling factor, which is dynamically adjusted with each iteration.

9. The method for generating a quantum-derived Newton-Raphson optimal fractional-order spectrogram according to claim 4, characterized in that, The QNRBO algorithm is used to optimize the hyperparameters of fractional-order spectrogram generation. Information entropy is chosen as the objective function. Information entropy is usually used to measure the uncertainty or complexity of a signal in its transform domain. For a discrete signal x[n], the squared amplitude obtained after transformation is regarded as a probability distribution p. i Where i represents different components of the transform domain; the information entropy H is defined as shown in equation (16): (16) in, X[i] represents the proportion of the energy of the i-th transform coefficient to the total energy, and X[i] represents the i-th coefficient of the signal x[n] in the transform domain.

10. A quantum-derived Newton-Raphson optimal fractional-order spectrogram generation system, comprising a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executed by the processor; characterized in that, When the system is running, the processor communicates with the memory via a bus, and the machine-readable instructions are executed by the processor as described in any one of claims 1 to 9.