NBI Doppler frequency shift spectrum fitting method and system for optimizing high-value region stability
By constructing a hybrid peak shape model and using differentiated parameter constraints, the problems of poor model fit, inaccurate background estimation, and easy oscillation in high-value regions in NBI Doppler frequency shift spectral fitting were solved, achieving higher spectral fitting accuracy and stability.
Patent Information
- Application Number
- CN202511738533.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-01-09
AI Technical Summary
Existing NBI Doppler frequency shift spectral fitting methods suffer from problems such as insufficient adaptability of peak shape models, limited background fitting accuracy, and poor fitting stability in high-value regions when dealing with complex spectra with large signal-to-noise ratio differences.
A hybrid peak model (combining Gaussian, Lorentz, and pseudo-Voigt components) was used for fitting, and a two-step background fitting method and differential parameter constraints were combined to optimize the fitting process in the high-value region through vectorized calculation.
It improves the accuracy of spectral fitting and the stability of high-value regions, reduces systematic bias and peak height calculation errors, and enhances the efficiency and reliability of spectral data processing.
Smart Images

Figure CN121301904A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of spectral data processing technology, and in particular to an NBI Doppler frequency shift spectral fitting method and system for optimizing the stability of high-value regions. Background Technology
[0002] Neutral beam injection (NBI) Doppler shift spectroscopy is a key technique in the diagnostics of magnetically confined nuclear fusion plasmas. By analyzing the spectral shifts generated after the interaction between neutral beam particles and the plasma, it can effectively retrieve key physical parameters of the plasma, such as ion temperature, rotational velocity, and density. In this technique, the accuracy and stability of spectral fitting directly determine the reliability of the final extracted physical parameters. Its core task is to accurately separate characteristic peak signals representing different energy state particle populations from raw spectral data that is affected by noise, stray light, and complex background radiation, and to precisely quantify their peak positions, peak heights, and peak shapes using mathematical models. Existing NBI Doppler shift spectral fitting methods still face a series of technical bottlenecks when dealing with complex real-world spectra with large signal-to-noise ratio differences, mainly in the following aspects:
[0003] Peak shape models lack adaptability. Peak shapes in NBI spectra are influenced by multiple physical mechanisms, including Doppler broadening, instrument functions, and plasma collision broadening. Actual peak shapes often exhibit a hybrid characteristic between Gaussian and Lorentz types. Traditional fitting methods often use a single Gaussian or Lorentz model for approximation, which is insufficient to flexibly and accurately describe this complex hybrid peak shape. This leads to systematic deviations between the fitted curve and the true signal, consequently affecting the extraction accuracy of key parameters such as peak area and peak width.
[0004] Background fitting has limited accuracy. Spectral background is typically composed of multiple factors, including bremsstrahlung, and exhibits a slowly changing trend. Existing methods usually perform polynomial fitting directly on the entire spectral region, including characteristic peaks, without effectively distinguishing between signal and pure background regions. This results in the background curve fitting process being severely interfered with by the characteristic peak signal, causing the final background curve to be "inflated" in the peak interval, failing to reflect the true background trend. This biased background estimation further propagates to the net signal after background subtraction, causing distortion in subsequent peak height calculations.
[0005] The fitting stability of high-value regions is poor. High-value peaks usually carry the core physical information of plasma, and their fitting accuracy requirements are the highest. However, existing methods generally adopt uniform and relatively loose parameter boundary constraints and equal residual weights for all peaks. This causes significant problems for high-value peaks: loose boundaries easily lead to large oscillations or even divergence in the fitting parameters of high-value peaks (especially peak height and peak position) during the optimization process, getting trapped in local optima or physically unreasonable solutions; while equal weights fail to highlight the criticality of high-value regions in the optimization objective, making it impossible to concentrate resources to ensure the accuracy of the most important signals in the fitting process. In addition, fixed regularization strategies cannot dynamically adjust the constraint strength according to the peak intensity, often resulting in insufficient constraints on high-value peaks and overfitting oscillations, or excessive constraints on low-value peaks and obscuring their weak signal details.
[0006] In summary, this application proposes an NBI Doppler frequency shift spectrum fitting method and system for optimizing the stability of high-value regions. Summary of the Invention
[0007] The purpose of this invention is to address the problems in the existing NBI Doppler frequency shift spectral fitting methods in the background art, such as insufficient adaptability of peak shape models, limited accuracy of background fitting, and poor stability of fitting in high-value regions when dealing with complex actual spectra with large signal-to-noise ratio differences. The invention proposes an NBI Doppler frequency shift spectral fitting method and system that optimizes the stability of high-value regions.
[0008] The technical solution of this invention: A method for fitting NBI Doppler frequency shift spectra to optimize stability in high-value regions, comprising the following steps:
[0009] S1. Data loading and preprocessing: Read the NBI Doppler frequency shift spectrum file, obtain the original wavelength data and original intensity data, and perform preprocessing to suppress noise and correct the baseline; identify the characteristic peaks in the preprocessed spectral data, and extract the peak height, peak position and peak width parameters of each characteristic peak;
[0010] S2. Background fitting: The final background curve is obtained by combining two polynomial fittings with pure background region screening. The pure background region is defined based on the residual of the initial background fitting.
[0011] S3. High-value region definition: Calculate the maximum value of the spectral signal after deducting the final background curve, set the high-value region threshold according to the preset ratio, and divide the characteristic peak into high-value peaks and low-value peaks.
[0012] S4. Hybrid peak modeling: Construct a hybrid peak model containing Gaussian, Lorentz, and pseudo-Voigt components, and use vectorized calculation to superimpose and fit the multi-peak signal.
[0013] S5. Optimize fitting parameters by setting differentiated parameter boundaries for high-value peaks and low-value peaks respectively; construct a fitting objective function that includes weighted mean square error and dynamic regularization term, in which higher weights are given to high-value regions and stronger regularization constraints are applied to the peak height parameter of high-value peaks.
[0014] S6. Iterative Fitting and Result Output: The objective function is iteratively optimized using an optimization algorithm to obtain the best fitting parameters, and the fitting result, including the total fitting curve, peak parameters, and high-value region markers, is output.
[0015] Optionally, in step S1, the identification of characteristic peaks is achieved by peak detection on the preprocessed spectral data, including locating the peak position and calculating the initial peak parameters by combining first derivative, second derivative analysis and intensity threshold screening.
[0016] Optionally, in step S2, background fitting specifically includes:
[0017] A polynomial was used to fit the preprocessed spectral data to the initial background to obtain an initial background curve. Pure background regions were selected based on the residuals between the initial background curve and the original intensity data. Polynomial fitting was then performed again based on the data of the pure background regions to obtain the final background curve and fitting coefficients.
[0018] The order of the polynomial fitting is selected between order 2 and 5 based on the complexity of the spectral background; the pure background region is defined by determining whether the absolute value of the residual is less than the standard deviation of the residual of the initial background curve.
[0019] Optionally, in step S4, the mixed peak shape model is a weighted sum of the peak-forming components, and the total peak shape signal of a single peak is... Represented as:
[0020] ,
[0021] in, Gaussian components, Lorentz component, It is a pseudo-Voigt component. , , These are the proportion coefficients of the corresponding components, and satisfy the following conditions: , , , ∈ [0, 1].
[0022] Optionally, in step S5, the differentiated parameter boundaries are specifically:
[0023] The peak height boundary of the high-value peak is set to 0.8 to 1.2 times the initial peak height, the peak position boundary is set to ±0.5 wavelength units of the initial peak position, and the peak width boundary is set to 0.8 to 1.2 times the initial peak width.
[0024] The peak height boundary of the low-value peak is set to 0.5 to 1.5 times the initial peak height, the peak position boundary is set to the initial peak position ± 1.0 wavelength unit, and the peak width boundary is set to 0.5 to 1.5 times the initial peak width.
[0025] Optionally, in step S5, the weighted mean square error The calculation formula is:
[0026]
[0027] in, The total number of spectral data points. The spectral signal after background removal, For the i-th wavelength point, The total fitted signal, As the weighting coefficient, in the high value region In the low value area ;
[0028] The dynamic regularization term The calculation formula is:
[0029]
[0030] in, and These are sets of high-value peaks and low-value peaks, respectively. Represents a set of high-value peaks One of the peaks in Represents the set of low-value peaks One of the peaks in and These are the peak height parameters for the high-value peak and the low-value peak, respectively. and These are the corresponding regularization coefficients, and .
[0031] Optionally, in step S6, the optimization algorithm is the L-BFGS-B algorithm, and the maximum number of iterations and convergence accuracy are set; before outputting the fitting result, the residual is smoothed to obtain the residual correction term, and the total peak signal, the residual correction term and the background curve are superimposed to generate a complete spectral fitting curve.
[0032] In a second aspect, the present invention provides an NBI Doppler frequency shift spectral fitting system for implementing the method described in the first aspect, comprising:
[0033] The data processing module is used to read NBI Doppler frequency shift spectrum files, obtain raw wavelength data and raw intensity data, and perform preprocessing to suppress noise and correct baseline; identify characteristic peaks in the preprocessed spectral data, and extract the peak height, peak position and peak width parameters of each characteristic peak;
[0034] The background fitting module uses a polynomial to perform initial background fitting on the preprocessed spectral data to obtain an initial background curve; pure background regions are selected based on the residual between the initial background curve and the original intensity data; and polynomial fitting is performed again on the data of the pure background regions to obtain the final background curve and fitting coefficients.
[0035] The high-value region definition module is used to calculate the maximum value of the spectral signal after deducting the final background curve, set the high-value region threshold according to a preset ratio, and divide the characteristic peaks into high-value peaks and low-value peaks.
[0036] The mixed peak modeling module is used to construct a mixed peak model containing Gaussian, Lorentz and pseudo-Voigt components, and uses a vectorized calculation method to superimpose and fit the multi-peak signal.
[0037] The parameter optimization module is used to set differentiated parameter boundaries for high-value peaks and low-value peaks respectively; it constructs a fitting objective function that includes weighted mean square error and dynamic regularization term, in which higher weights are given to high-value regions and stronger regularization constraints are applied to the peak height parameter of high-value peaks;
[0038] The iterative fitting module uses an optimization algorithm to iteratively optimize the objective function to obtain the best fitting parameters;
[0039] The results output module is used to output the fitting results, which include the overall fitted curve, peak shape parameters, and high-value region markers.
[0040] Compared with the prior art, this application includes at least one of the following beneficial technical effects:
[0041] By constructing a hybrid peak shape model that combines Gaussian, Lorentz, and pseudo-Voigt, the problem of poor adaptability of a single model to complex actual peak shapes is effectively overcome. This model can more accurately characterize the hybrid peaks formed by the combined action of multiple physical mechanisms in NBI spectra, reduce systematic bias at the model level, and provide a more reliable data foundation for subsequent physical parameter inversion.
[0042] A two-step background fitting strategy, consisting of initial fitting, pure background region screening, and secondary fitting, effectively eliminates the interference of characteristic peak signals on background trend calculation, thereby obtaining a fitting curve that more closely approximates the true background. This fundamentally reduces the calculation errors in peak height and peak area caused by background estimation bias, improving the accuracy of the net spectral signal.
[0043] For high-value peaks carrying key physical information, this strategy achieves focused optimization of key signals by tightening the parameter boundaries of peak height, peak position, and peak width, doubling the weight of high-value regions in the objective function, and applying stronger dynamic regularization constraints. This strategy effectively suppresses the oscillation and divergence trends of high-value peak parameters during the fitting process, ensuring the robustness and repeatability of core physical parameter extraction.
[0044] By employing a hybrid peak shape model and vectorized calculation, multi-peak superposition is achieved, avoiding inefficient point-by-point cycling. At the same time, the entire process, from data preprocessing to result output, is fully integrated and automated, reducing human intervention. While ensuring high accuracy, it significantly improves the processing efficiency of spectral data, making it suitable for large-scale plasma diagnostic applications with high real-time requirements.
[0045] In summary, this invention systematically solves the technical bottlenecks in traditional NBI Doppler frequency shift spectral fitting, such as poor model fit, inaccurate background estimation, and easy oscillation in high-value regions, by introducing a mixed peak shape model, a two-step background fitting method, and differentiated parameter constraints and weighted optimization strategies for high-value peaks. This significantly improves the overall accuracy of spectral fitting, the stability of high-value regions, and the processing efficiency. Attached Figure Description
[0046] Figure 1 This is a flowchart of the fitting process in Embodiment 1 of the present invention;
[0047] Figure 2 This is a comparison diagram of the background fitting process in Embodiment 2 of the present invention;
[0048] Figure 3 This is a comparison chart of the traditional fitting and the fitting after high-value optimization in Embodiment 3 of the present invention;
[0049] Figure 4 This is the final peak fitting result diagram of Embodiment 4 of the present invention. Detailed Implementation
[0050] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0051] Example 1
[0052] This invention proposes an NBI Doppler frequency shift spectral fitting method to optimize stability in high-value regions. Based on the raw spectral digital signal output from the hardware acquisition terminal, it sequentially completes data loading, background estimation, peak shape parameter initialization, mixed peak shape modeling, weighted residual optimization, smoothing correction, and visualization output through a multi-stage processing chain. (See attached...) Figure 1 Fitting flowchart.
[0053] I. Data Loading and Preprocessing
[0054] 1. Loading Spectral Data
[0055] The system reads the original file storing the NBI Doppler frequency shift spectral signal. This file must contain two core data sequences: a wavelength sequence and the corresponding intensity sequence. The wavelength sequence must meet the requirement of equal-interval sampling, with the sampling interval set to 0.001-0.01 nm based on the spectrometer resolution to ensure complete capture of characteristic peak details. The intensity sequence includes background signal, characteristic peak signal, and noise components, and the data format must be compatible with conventional numerical calculations. After loading, the system parses the data to obtain the original wavelength data and the original intensity data, both of which must be of the same length.
[0056] 2. Data Preprocessing Operations
[0057] Preprocessing aims to reduce noise interference and correct data anomalies, specifically including:
[0058] (1) Noise suppression
[0059] A sliding window filtering method is employed, with the window size set to 3-15 data points (an odd number) depending on the noise characteristics. For each data point, the mean or median of the intensity data within its window is calculated. This statistical value is then used to replace the intensity value of the original data point, achieving smoothing of high-frequency random noise while preserving the contour information of characteristic peaks. The window size can be determined through preliminary experiments: the stronger the noise, the larger the window can be, but excessive smoothing should be avoided to prevent peak distortion.
[0060] (2) Anomaly Correction
[0061] Outliers were identified using local standard deviation analysis. The intensity sequence was segmented, with each segment containing 20-50 consecutive data points. The mean and standard deviation of each segment were calculated, and data points deviating from the mean by more than three times the standard deviation were identified as outliers. Linear interpolation was then used to correct outliers: based on the trend of 5-10 normal data points before and after each outlier, interpolation was performed to replace the outlier, ensuring the data sequence remained continuous and smooth.
[0062] (3) Characteristic peak identification and parameter extraction
[0063] Peak location: Characteristic peak location is achieved by combining first-order and second-order derivative analysis. The first and second derivatives of the intensity sequence are calculated, and points where the first derivative changes from positive to negative and the second derivative is negative are selected as candidate peak locations. Simultaneously, an intensity threshold is introduced to filter valid peaks: the threshold is set to 10%–20% of the maximum intensity after baseline subtraction, eliminating false peaks below the threshold, and finally determining the peak coordinates of the characteristic peaks.
[0064] Peak parameter calculation: For each identified characteristic peak, three core parameters are calculated: peak height is the difference between the intensity value at the peak position and the baseline; peak width is calculated using full width at half maximum (FWHM), which is the distance between two wavelength points corresponding to half the peak height; peak shape symmetry is judged by the ratio of the left and right half peak widths.
[0065] Example 2
[0066] This embodiment, based on Embodiment 1, also provides a background fitting method. The background fitting process is described below. Figure 2 This demonstrates the process from initial fitting to selecting a pure background region to final fitting.
[0067] II. Background Fitting Implementation Process
[0068] 1. Initial Background Model Construction
[0069] The core objective of the initial background model is to quickly establish a basic framework that can reflect the overall background trend of the spectrum, avoid computational redundancy caused by model complexity, and provide a reference benchmark for subsequent screening of pure background regions.
[0070] A polynomial model is used to fit the background signal. The polynomial order is selected from 2 to 5 depending on the background complexity: 2-3 orders are chosen when the background changes gradually, and 4-5 orders are chosen when the background contains significant fluctuations. The preprocessed intensity sequence (y) is fitted with a polynomial using the least squares method to minimize the mean square error between the fitted curve and the original intensity data, thus obtaining the initial background curve. During the fitting process, it is necessary to ensure that the order of magnitude of the polynomial coefficients matches the background intensity to avoid distortion of the background curve due to excessively large coefficients.
[0071] 2. Pure background area filtering
[0072] Calculate the residual sequence between the initial background curve and the intensity sequence (y), and calculate the standard deviation (σ) of the residuals. Regions with absolute residual values less than σ are defined as pure background regions (these regions contain no characteristic peak signals, only background and noise). If the number of data points in the pure background region is less than the polynomial order + 1 (i.e., insufficient to meet the fitting degrees of freedom requirement), the screening threshold is relaxed to 1.5σ until sufficient data points are obtained. Pure background regions should be evenly distributed at both ends of the spectrum or between characteristic peaks to avoid concentration in a single interval, which could lead to fitting bias.
[0073] 3. Final Background Determined
[0074] The selected pure background region data (wavelength and corresponding intensity values) are refitted with a polynomial to obtain the optimized background curve and fitting coefficients. If, after threshold adjustment, sufficient pure background data points are still not obtained (e.g., characteristic peaks densely cover most areas of the spectrum), the coefficients of the initial background fitting are used as the final background coefficients. The final background curve must meet the following requirements: the deviation from the measured data in the pure background region is less than 2σ, and the fitted value in the characteristic peak region does not exceed 50% of the peak bottom intensity, ensuring that it does not interfere with the extraction of the peak shape signal.
[0075] Figure 2 The complete process of background fitting is presented intuitively. It can be seen that after the process of "initial modeling → region selection → final background determination", the final background fit in the pure background region is significantly improved, and peak signal interference is effectively eliminated.
[0076] Example 3
[0077] This embodiment, based on Embodiment 2, further provides a method for optimizing high-value regions, including:
[0078] III. Delineation of High-Value Regions and Classification of Peak Shapes
[0079] 1. Background Removal
[0080] The core of background subtraction is to separate the background signal from the preprocessed spectral data to obtain the net signal reflecting the characteristic peaks. Let the preprocessed intensity sequence be y(x), and the final background curve obtained through polynomial fitting be B(x). Then, the mathematical expression for the background-subtracted spectral signal S(x) is: S(x) = y(x) - B(x). Where B(x) is an nth-order polynomial function (n = 2, 3, 4, 5), in the form: .Mode , , , The background fitting coefficients are determined through quadratic fitting optimization in the pure background region to ensure that the deviation of B(x) from the measured data in the region without characteristic peaks is less than twice the residual standard deviation, in order to avoid interference from the background signal on the intensity of the characteristic peaks. The signal S(x) after removing the background contains only the characteristic peak signal and the residual signal, and its physical meaning is the net radiation intensity, which is the basic input for subsequent peak shape classification.
[0081] 2. Threshold setting and quantitative standards for peak shape classification
[0082] The core of high-value region and peak shape classification is based on the quantization of signal intensity. The specific steps are as follows:
[0083] (1) Extraction of maximum signal value
[0084] Calculate the global maximum value of the background signal S(X). ,Right now: In the formula The wavelength range of the spectrum. The peak value of the highest intensity characteristic peak in the corresponding spectrum is the benchmark for dividing the information region.
[0085] (2) Mathematical definition of high value threshold
[0086] Set the scaling factor The threshold T for the high-value region is dynamically adjusted based on the signal-to-noise ratio (SNR), taking a smaller value when the SNR is high and a larger value when the SNR is low. The physical meaning of this threshold is to define the region where the intensity reaches 20%-40% of its strongest peak as a high-value region, ensuring that strong signals carrying key physical information are accurately identified.
[0087] (3) Rules for determining peak shape
[0088] The identified feature set is m is the number of peaks, and each peak The peak intensity is The classification rule is as follows:
[0089] like ,but It was determined to be a high-value peak;
[0090] like ,but It was determined to be a low-value peak.
[0091] Meanwhile, the boundary of the high-value region in the wavelength dimension is defined as follows: for high-value peaks Its corresponding high-value wavelength range is ,in and They are respectively Both sides of the peak The corresponding wavelength point satisfies: In the formula for The peak wavelength. This interval accurately defines the wavelength range of high-value signals, providing a spatial boundary for subsequent differential fitting.
[0092] IV. Construction of Hybrid Peak Shape Model
[0093] 1. Definition of basic peak formation components
[0094] (1) Gaussian components
[0095] Applicable to symmetrical peak shapes, the expression is: In the formula, h is the peak height. Peak position, is the Gaussian half-width parameter, and x is the wavelength variable.
[0096] (2) Lorentz components
[0097] For asymmetric peaks dominated by collision broadening, the expression is: In the formula This is the Lorentz half-width parameter; the other parameters are defined the same as those for the Gaussian component.
[0098] (3) Pseudo-Voigt component
[0099] Adapted to a mixed peak shape between Gaussian and Lorentz, using an equal-weighted superposition format: .
[0100] 2. Weighted calculation of mixed peak shapes
[0101] (1) Percentage parameter constraints
[0102] Let the Gaussian proportion be The proportion of pseudo-Voigt is Lorenz's percentage was ,satisfy: , This constraint ensures that the contribution weights of each component are normalized, and its physical meaning is the relative contribution ratio of different broadening mechanisms to the peak shape.
[0103] (2) Single-peak mixing model
[0104] The total peak shape signal P(x) of a single characteristic peak is a weighted sum of the three components: The parameter set in the formula { } represents the fitting parameters to be optimized, corresponding to peak height, peak position, Gaussian half-width, Lorentz half-width, Gaussian proportion, and pseudo-Voigt proportion, respectively.
[0105] (3) Multi-peak superposition calculation
[0106] For a spectrum containing m characteristic peaks, the total peak shape signal The superposition of all single-peak signals: In the formula Given the mixed peak shape signal of the i-th characteristic peak, the multi-peak signal is quickly superimposed through vectorized matrix operations, avoiding point-by-point iterative calculation, improving computational efficiency, and ensuring numerical accuracy.
[0107] V. Fitting Parameter Optimization
[0108] 1. Parameter boundary constraints
[0109] To ensure the physical rationality of the fitted parameters and improve the stability of high-value peaks, differentiated parameter boundary constraints are set for high-value peaks and low-value peaks, as follows:
[0110] (1) High peak parameter boundary
[0111] The initial parameter for the high-value peak is the peak height. Peak Peak width Then the parameter boundary in the optimization process is defined as:
[0112] Peak height boundary: ;
[0113] Peak boundary: ;
[0114] Peak width boundary: ;
[0115] Component proportion boundary: Gaussian proportion The proportion of pseudo-Voigt And satisfy Lorenz percentage .
[0116] The aforementioned boundary tightens the fluctuation range of high-value peak parameters, ensuring that the core parameters converge stably within a physically reasonable range and avoiding deviation from the true value due to over-optimization.
[0117] (2) Low peak parameter boundary
[0118] The initial parameter for the low-value peak is the peak height. Peak Peak width Then the parameter boundary is defined as:
[0119] Peak height boundary: ;
[0120] Peak boundary: ;
[0121] Peak width boundary: ;
[0122] Component proportion boundary: Gaussian proportion The proportion of pseudo-Voigt And satisfy Lorenz percentage .
[0123] The low-value peaks adopt relatively loose boundaries to retain more optimization space to capture the detailed features of weak signals. At the same time, non-negative constraints ensure the physical meaning of the proportion parameters, and the contribution ratio of each broadening mechanism cannot be negative.
[0124] 2. Construction of the objective function
[0125] The objective function achieves a balance between fitting accuracy and parameter stability by integrating the weighted mean square error (MSE) and the dynamic regularization term.
[0126] (1) Weighted mean square error (MSE)
[0127] The spectral signal S(x) after background subtraction is output by the mixing peak shape model. Based on the residuals, regional weights are introduced to enhance the fitting accuracy in high-value regions: In the formula, N is the total number of spectral data points. For the i-th wavelength point, the weights are... T is the threshold for the high-value region. By doubling the weight of the high-value region, the optimization process focuses more on the fitting accuracy of the key signal and reduces the interference of noise on the high-value peak parameters.
[0128] (2) Dynamic regularization terms
[0129] To suppress parameter fluctuations, a regularization constraint is applied only to the peak height parameter, and a differentiation coefficient is used between high-value peaks and low-value peaks. In the formula, H represents the set of high-value peaks, and L represents the set of low-value peaks. These are the peak height parameters for the high-value peak and the low-value peak, respectively. By increasing the regularization coefficient of the high-value peak, the constraint strength on its peak height parameter is enhanced, avoiding numerical oscillations caused by overfitting.
[0130] (3) Overall objective function
[0131] The final objective function is the superposition of the weighted MSE and the dynamic regularization term, and the optimization objective is to minimize this function. By using this objective function, while ensuring the overall fitting accuracy, we can achieve stable convergence of high-peak parameters and effective preservation of low-peak details, providing a clear mathematical optimization direction for subsequent iterative optimization.
[0132] Figure 3 The results show a comparison between the optimized fitting for high values and the traditional undifferentiated peak shape fitting method. The traditional fitting method is insufficient in the high value region, while the optimized fitting method significantly improves the parameter stability of high value peaks and performs far better than the traditional method overall.
[0133] Example 4
[0134] This embodiment, based on embodiment 3, also provides a demonstration of the final fitting result, see... Figure 4 As shown.
[0135] VI. Iterative Fitting and Result Output
[0136] Iterative fitting focuses on optimizing peak shape parameters, reducing the deviation between the fitted curve and the real signal through multiple rounds of calculation, especially focusing on accuracy in high-value regions. First, peak types are distinguished by signal intensity: using 30% of the maximum signal value after background subtraction as a threshold, high-value peaks above the threshold have their parameter boundaries tightened (peak height 0.8-1.2 times the initial value, peak position ±0.5nm offset), while low-value peaks retain a loose range (peak height 0.5-1.5 times, peak position ±1.0nm), ensuring stable fitting of the core peaks.
[0137] During iteration, a mixed-peak model is used to calculate the fitted signal. Combined with smoothing residual correction, a weighted strategy is employed—the weight of high-value regions is doubled to amplify the influence of deviations in the core region. Simultaneously, a dynamic regularization term is introduced, with the regularization intensity of high-value peaks being twice that of low-value peaks, to avoid overfitting. A strict convergence threshold of 300 iterations is set, and the parameters corresponding to the minimum residual are continuously updated.
[0138] The output integrates all information, including wavelength, original signal, background, and fitting curves for each stage; it records peak parameters (peak height, position, width, and peak proportion), residual values, and high-value region markers. The results are presented intuitively through charts and graphs, and structured data is saved simultaneously, providing a complete basis for subsequent analysis.
[0139] This invention addresses the systematic bias problem of traditional single-model approaches when describing spectral peaks with coexisting complex broadening mechanisms by employing a hybrid peak shape model combining Gaussian, Lorentz, and pseudo-Voigt components. This model more accurately characterizes peak shape features formed by multiple physical processes in real plasma environments, thus reducing fitting errors at the source. Simultaneously, by implementing a two-step background fitting process of "initial fitting - pure background region screening - secondary fitting," the interference of characteristic peak signals on background trend calculations is effectively eliminated, significantly improving the accuracy of background curve estimation. This directly reduces the calculation errors of peak height and peak area caused by background shift, improving the signal-to-noise ratio of the net spectral signal after background subtraction.
[0140] It is worth noting that, to address the issue of unstable fitting in high-value regions, this invention implements a differentiated parameter constraint strategy. By setting stricter boundary conditions for the peak height, position, and width parameters of high-value peaks, and using a weighted mean square error objective function and a dynamically adjusted regularization term, the optimization focus and numerical stability control of key signal regions are strengthened. This method effectively suppresses the oscillation and divergence of high-value peak parameters during the optimization process, ensuring the repeatability and reliability of the extracted core physical parameters. Furthermore, the model calculation is implemented using vectorization, avoiding inefficient loop operations, thus improving data processing throughput while maintaining accuracy, enabling it to adapt to the batch processing and real-time analysis needs of large-scale spectral data.
[0141] The above specific embodiments are merely several optional embodiments of the present invention. Based on the technical solutions of the present invention and the relevant teachings of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.
Claims
1. A method for NBI Doppler frequency shift spectral fitting to optimize stability in high-value regions, characterized in that, Includes the following steps: S1. Data loading and preprocessing: Obtain the wavelength and intensity data of the NBI Doppler frequency shift spectrum, perform preprocessing, identify the characteristic peaks in the spectrum, and extract their peak height, peak position, and peak width parameters. S2. Background fitting: The final background curve is obtained by combining two polynomial fittings with pure background region screening. The pure background region is defined based on the residual of the initial background fitting. S3. High-value region definition: Calculate the maximum value of the spectral signal after deducting the final background curve, set the high-value region threshold according to the preset ratio, and divide the characteristic peak into high-value peaks and low-value peaks. S4. Hybrid peak modeling: Construct a hybrid peak model containing Gaussian, Lorentz, and pseudo-Voigt components, and use vectorized calculation to superimpose and fit the multi-peak signal. S5. Optimize fitting parameters by setting differentiated parameter boundaries for high and low peaks and constructing a fitting objective function that includes weighted mean square error and dynamic regularization term. S6. Iterative Fitting and Result Output: The objective function is iteratively optimized using an optimization algorithm to obtain the best fitting parameters, and the fitting result, including the total fitting curve, peak parameters, and high-value region markers, is output.
2. The NBI Doppler frequency shift spectrum fitting method for optimizing stability in high-value regions according to claim 1, characterized in that, In step S1, the characteristic peaks in the spectrum are identified by peak detection of the preprocessed spectral data, including locating the peak position and calculating the initial peak parameters by combining first derivative, second derivative analysis and intensity threshold screening.
3. The NBI Doppler frequency shift spectral fitting method for optimizing stability in high-value regions according to claim 1, characterized in that, In step S2, background fitting specifically includes: A polynomial was used to fit the preprocessed spectral data to the initial background to obtain an initial background curve. Pure background regions were selected based on the residuals between the initial background curve and the original intensity data. Polynomial fitting was then performed again based on the data of the pure background regions to obtain the final background curve and fitting coefficients. The order of the polynomial fitting is selected between order 2 and 5 based on the complexity of the spectral background; the pure background region is defined by determining whether the absolute value of the residual is less than the standard deviation of the residual of the initial background curve.
4. The NBI Doppler frequency shift spectrum fitting method for optimizing stability in high-value regions according to claim 1, characterized in that, In step S4, the mixed peak shape model is a weighted sum of the peak-forming components, and the total peak shape signal of a single peak is... Represented as: , in, Gaussian components, Lorentz component, It is a pseudo-Voigt component. , , These are the proportion coefficients of the corresponding components, and satisfy the following conditions: , , , ∈ [0, 1].
5. The NBI Doppler frequency shift spectral fitting method for optimizing stability in high-value regions according to claim 1, characterized in that, In step S5, the differentiated parameter boundaries are specifically as follows: The peak height boundary of the high-value peak is set to 0.8 to 1.2 times the initial peak height, the peak position boundary is set to ±0.5 wavelength units of the initial peak position, and the peak width boundary is set to 0.8 to 1.2 times the initial peak width. The peak height boundary of the low-value peak is set to 0.5 to 1.5 times the initial peak height, the peak position boundary is set to the initial peak position ± 1.0 wavelength unit, and the peak width boundary is set to 0.5 to 1.5 times the initial peak width.
6. The NBI Doppler frequency shift spectrum fitting method for optimizing stability in high-value regions according to claim 1, characterized in that, In step S5, the weighted mean square error The calculation formula is: , in, The total number of spectral data points. The spectral signal after background removal, For the i-th wavelength point, The total fitted signal, As the weighting coefficient, in the high value region In the low value area ; The dynamic regularization term The calculation formula is: , in, and These are sets of high-value peaks and low-value peaks, respectively. Represents a set of high-value peaks One of the peaks in Represents the set of low-value peaks One of the peaks in and These are the peak height parameters for the high-value peak and the low-value peak, respectively. and These are the corresponding regularization coefficients, and ; In this approach, higher weights are assigned to high-value regions, and stronger regularization constraints are applied to the peak height parameter of high-value peaks.
7. The NBI Doppler frequency shift spectrum fitting method for optimizing stability in high-value regions according to claim 1, characterized in that, In step S6, the optimization algorithm is the L-BFGS-B algorithm, and the maximum number of iterations and convergence accuracy are set. Before outputting the fitting result, the residual is smoothed to obtain the residual correction term. The total peak shape signal, the residual correction term and the background curve are superimposed to generate a complete spectral fitting curve.
8. An NBI Doppler frequency shift spectral fitting system for implementing the method as described in any one of claims 1-7, characterized in that, include: The data processing module is used to read NBI Doppler frequency shift spectrum files, obtain raw wavelength data and raw intensity data, and perform preprocessing to suppress noise and correct baseline. Identify the characteristic peaks in the preprocessed spectral data and extract the peak height, peak position, and peak width parameters for each characteristic peak; The background fitting module uses a polynomial to perform initial background fitting on the preprocessed spectral data to obtain an initial background curve; pure background regions are then selected based on the residual between the initial background curve and the original intensity data. Based on the data of the pure background region, a polynomial fitting is performed again to obtain the final background curve and fitting coefficients. The high-value region definition module is used to calculate the maximum value of the spectral signal after deducting the final background curve, set the high-value region threshold according to a preset ratio, and divide the characteristic peaks into high-value peaks and low-value peaks. The mixed peak modeling module is used to construct a mixed peak model containing Gaussian, Lorentz and pseudo-Voigt components, and uses a vectorized calculation method to superimpose and fit the multi-peak signal. The parameter optimization module is used to set differentiated parameter boundaries for high-value peaks and low-value peaks respectively; it constructs a fitting objective function that includes weighted mean square error and dynamic regularization term, in which higher weights are given to high-value regions and stronger regularization constraints are applied to the peak height parameter of high-value peaks; The iterative fitting module uses an optimization algorithm to iteratively optimize the objective function to obtain the best fitting parameters; The results output module is used to output the fitting results, which include the overall fitted curve, peak shape parameters, and high-value region markers.