Six-dimensional data decomposition method, terminal device, and storage medium
By combining the alternating hexagonal method and six-dimensional parallel factor analysis, the noise tolerance and speed issues in six-dimensional data decomposition are solved, enabling fast and stable six-dimensional data decomposition and analysis, and improving qualitative and quantitative capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-31
- Publication Date
- 2026-03-10
AI Technical Summary
Existing six-dimensional data decomposition algorithms are insufficient in tolerating high-noise data and in terms of running speed. They are also sensitive to too many sets of fractions and are difficult to effectively resolve real six-dimensional chemical data arrays.
The alternating hexlinear approach is adopted to quickly optimize random initial values by constructing a hexlinear component model and the alternating hexlinear algorithm (AHLD), and to decompose the six-dimensional data array by combining six-way parallel factor analysis (Six-way PARAFAC-ALS) to minimize the influence of noise.
It achieves fast and stable six-dimensional data decomposition, can tolerate high noise and too many groups, provides stable qualitative and quantitative analysis results, and has stronger anti-collinearity and high sensitivity.
Smart Images

Figure CN115795247B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of qualitative and quantitative analysis in analytical chemistry, and in particular to a six-dimensional data decomposition method, terminal equipment, and storage medium. Background Technology
[0002] With the widespread adoption of second- and higher-order analytical instruments, an increasing number of instrument response signals consist of high-dimensional chemical data composed of numerous data points. Currently commonly used second- and higher-order analytical instruments include excitation-emission matrix fluorescence (EEMF), high-performance liquid chromatography (HPLC) with diode array detectors (DADs), liquid chromatography with EEMF-fluorescence detectors (LCs), two-dimensional liquid chromatography with DDAs, and full three-dimensional gas chromatography-mass spectrometry (GC-MS). It is easy to imagine that with the increasing dimensionality of instruments and the growing complexity of research systems, more and more higher-order analytical instruments may emerge in the future. Combining higher-order analytical instruments with higher-order correction algorithms not only yields significant "second-order advantages"—the ability to successfully extract qualitative and quantitative information of target analytes even in the presence of unknown or known interferences—but also offers additional advantages, namely "higher-order advantages," such as providing richer information, overcoming matrix effects, improving the algorithm's ability to resolve challenging collinear data, and enhancing the method's sensitivity and selectivity. Therefore, researching analytical methods for high-dimensional instrument data has great potential value and can provide guidance for the future development of analytical instruments.
[0003] Due to numerous theoretical and experimental challenges, there is very little theoretical and applied research on five-dimensional and ultra-high-dimensional arrays. Existing research on five-dimensional arrays primarily involves introducing additional dimensions into second- or third-order instruments. Furthermore, current research on five-dimensional and six-dimensional arrays is limited to studying the relationship between matter and experimental conditions, focusing more on optimizing conditions rather than analyzing actual chemical arrays.
[0004] Currently, algorithms used for analyzing five-dimensional arrays include expanded partial least squares combined with residual quadlinear decomposition, alternating five-linear decomposition, alternating fitting weighted residual five-linear decomposition, and five-dimensional parallel factor analysis. These are mainly extensions of two types of algorithms. One type is based on residual multilinear algorithms, which are more suitable for analyzing nonlinear data. The other type is iterative algorithms based on the alternating least squares principle. This type of algorithm can decompose a multidimensional array into a relatively unique solution with clear chemical meaning. Currently, this type of algorithm is the most popular. Based on the characteristics and wide application of iterative algorithms, iterative algorithms can be further developed into ultra-high-dimensional correction algorithms. Existing hexalinear decomposition algorithms only use six-dimensional parallel factor analysis to optimize experimental conditions and do not analyze real chemical data arrays. Theoretically, six-dimensional parallel factor analysis has similar characteristics to three-dimensional parallel factor analysis; it can tolerate high-noise data, but its running speed is slow and it is sensitive to too many components. In reality, the data information contained in a six-dimensional array is much richer than that in a low-dimensional array, which poses new challenges to the performance of the algorithm, especially its running speed. Therefore, there is an urgent need to develop a superior algorithm that can better analyze six-dimensional data arrays. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a six-dimensional data decomposition method, terminal device and storage medium that, under the premise of tolerating high noise data, fast running speed and insensitivity to too many components, realize the analysis of real six-dimensional data and a series of six-dimensional analog arrays with different noise.
[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a six-dimensional data decomposition method, comprising the following steps:
[0007] S1. Construct a hexadecimal component model to obtain a six-dimensional matrix of size I×J×K×L×M×N. X I×J×K×L×M×N The expression for the hexagonal component model is as follows:
[0008]
[0009] Where x ijklmn For X I×J×K×L×M×N The elements in the expression represent the response intensity of the nth sample in the i-th, j-th, k-th, l-th, and m-th channels; a io ,b jo ,c ko ,d lo ,e mo and f no These represent the elements in the first through sixth contour matrices A, B, C, D, E, and F, respectively; e ijklmn It is a residual arrayE The elements in the table are: O represents the number of components in the study system; I represents the number of excitation wavelength channels; J represents the number of emission wavelength channels; K represents the number of pH levels; L represents the number of dilution levels; M represents the number of voltage levels; and N represents the total number of samples.
[0010] S2. Randomly initialize the contour matrices A, B, C, D, and E; X I×J×K×L×M×N Six different three-dimensional matrices are obtained along different dimensions in the form of incompletely expanded matrices. X I×JKLM×N , X J×KLMN×I , X K×LMNI×J , X L×MNIJ×K , X M ×NIJK×L , X N×IJKL×M ;extract X I×JKLM×N The nth front slice matrix, X J×KLMN×I The i-th front slice matrix, X K×LMNI×J The j-th front slice matrix, X L×MNIJ×K The k-th front slice matrix, X M×NIJK×L The l-th front slice matrix, X N×IJKL×M The m-th front slice matrix is obtained respectively. With three-dimensional array X I×J×K For example, X ..1 Represents a three-dimensional array X I×J×K The first front slice, X ..2 This indicates the second front slice, and so on.
[0011] S3. Based on the initialized A, B, C, D, and E, use the formula Calculate F; based on the calculated F, initialize B, C, D, and E, and use the formula... Calculate A; based on the calculated F and A, and the initialized C, D, and E, use the formula Calculate B; based on the calculated F, A, and B, and the initialized D and E, use the formula Calculate C; based on the calculated F, A, B, C, and the initialized E, use the formula... Calculate D; based on the calculated F, A, B, C, and D, use the formula... Calculate E; where diagm() represents extracting diagonal elements to form a column vector; + represents the Moore-Penrose generalized inverse of the matrix; ⊙ represents the Khatri-Rao product; f (n) a (i) b (j) c (k) d (l) e (m) Let F represent the nth row of F, the ith row of A, the jth row of B, the kth row of C, the lth row of D, and the mth row of E, respectively.
[0012] S4. Determine whether the first convergence condition is met: If true, the optimized six contour matrices are obtained; otherwise, the contour matrices obtained in this iteration are used as initial values, and the process returns to step S3; where SSR (m) Represents the residual matrix after m iterations. E The sum of squares of all elements in the set, where ε1 is the first threshold.
[0013] This invention extends the alternating trilinear method to the alternating hexagonal method, rapidly optimizes random initial values, achieves the decomposition of real six-dimensional data, can tolerate high-noise data, runs fast, and is insensitive to too many components, thus enabling qualitative and quantitative analysis through the analysis of six-dimensional data.
[0014] In this invention, to minimize the impact of noise in the six-dimensional data and further improve the tolerance to noise data, when the first convergence condition is met, the following further step is taken:
[0015] S5, will X I×J×K×L×M×N Different two-dimensional matrices X are obtained according to the form of the fully extended matrix. I×JKLMN X J×KLMNI X K×LMNIJ X L×MNIJK X M×NIJKL X N×IJKLM ;
[0016] S6. Use the optimized second to sixth contour matrices obtained in step S4 as initialization values;
[0017] S7. Substitute the initial values into the formula A = X I×JKLMN [(F⊙E⊙D⊙C⊙B) T ] + We obtain the first contour matrix A for the current optimization; we substitute the initial values of the first contour matrix A and the third to sixth contour matrices into the formula B = X. J×KLMNI [(A⊙F⊙E⊙D⊙C) T ] +We obtain the currently optimized second contour matrix B; substitute the initial values of the currently optimized first contour matrix A, the currently optimized second contour matrix B, and the fourth to sixth contour matrices into the formula C = X. K×LMNIJ [(B⊙A⊙F⊙E⊙D) T ] + We obtain the currently optimized third contour matrix C; substitute the initial values of the currently optimized first contour matrix A, the currently optimized second contour matrix B, the third contour matrix C, and the fifth to sixth contour matrices into the formula D = X. L×MNIJK [(C⊙B⊙A⊙F⊙E) T ] + This yields the currently optimized fourth contour matrix D; the initial values of the currently optimized first contour matrix A, the currently optimized second contour matrix B, the currently optimized third contour matrix C, the currently optimized fourth contour matrix D, and the sixth contour matrix are substituted into the formula E = X. M×NIJKL [(D⊙C⊙B⊙A⊙F) T ] + This yields the currently optimized fifth contour matrix E; substituting the currently optimized first contour matrix A, second contour matrix B, third contour matrix C, fourth contour matrix D, and fifth contour matrix E into the formula F = X N×IJKLM [(E⊙D⊙C⊙B⊙A) T ] + This yields the currently optimized sixth contour matrix F;
[0018] S8. Determine whether the second convergence condition is met: If true, the process ends, and the final first to sixth contour matrices are output; otherwise, the second to sixth contour matrices obtained in this iteration are used as initial values, and the process returns to step S7; SSR (n) Represents the residual matrix after the nth iteration. E The sum of squares of all elements in the equation, where ε2 is the set second threshold.
[0019] In this invention, ε1 can be set to 10. -3 ε2 can be set to 10. -6 .
[0020] As an inventive concept, the present invention also provides a terminal device, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the steps of the method described above.
[0021] As an inventive concept, the present invention also provides a computer-readable storage medium having a computer program / instructions stored thereon; when the computer program / instructions are executed by a processor, they implement the steps of the method described above.
[0022] Compared with existing technologies, the beneficial effects of this invention are as follows: This invention successfully realizes the analysis of real six-dimensional data and a series of six-dimensional simulated arrays with different noise levels. The six modes of the real data's six-dimensional array include excitation wavelength, emission wavelength, pH, dilution ratio, detection voltage, and sample. The overall performance of the method of this invention is excellent; it not only has a fast convergence speed and is insensitive to excessive components and noise, but also obtains stable and accurate qualitative and quantitative results even in the presence of known interference, unknown interference, peak overlap, and strong collinearity. Compared with three-dimensional correction methods, the method of this invention also exhibits "higher-order advantages," extracting richer information, possessing strong anti-collinearity capabilities, and higher sensitivity and selectivity. This invention not only provides data analysis tools for future high-order instruments but also provides real data support and methodological reference for theoretical research in higher-order tensor algebra. Attached Figure Description
[0023] Figure 1 This is a graphical representation of the hexalinear component model in an embodiment of the present invention.
[0024] Figure 2 The three-dimensional array is an embodiment of the present invention. X I×J×K For example, a graphical representation in the form of slices.
[0025] Figure 3 This is a flowchart of the Six-way ACM operation in an embodiment of the present invention.
[0026] Figure 4 This is an example of an embodiment of the invention showing the undisturbed normalized contour plots of the three components in a six-dimensional analog array across five different dimensions.
[0027] Figure 5 When O=5 in this embodiment of the invention, the decomposition diagrams obtained by the two algorithms from parsing the six-dimensional analog array (noise level is 1%) are shown.
[0028] Figure 6 The following is an excitation-emission fluorescence spectrum of an embodiment of the present invention, wherein: (a) thiabendazole; (b) pirimicarb; (c) predicted spiked water sample PR03.
[0029] Figure 7 This is a decomposition diagram of a six-dimensional data array with true chemical significance, obtained by Six-way ACM analysis when O=2, according to an embodiment of the present invention. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] In this document, the terms "first," "second," and other similar words are not intended to imply any order, quantity, or importance, but are merely used to distinguish different elements. The terms "one," "a," and other similar words are not intended to indicate the existence of only one of the stated things, but rather that the description pertains to only one of the two stated things, which may include one or more. The terms "comprising," "including," and other similar words are intended to indicate a logical relationship, not a spatial relationship. For example, "A includes B" means that logically B belongs to A, not that spatially B is located inside A. Furthermore, the meanings of the terms "comprising," "including," and other similar words should be considered open-ended, not closed. For example, "A includes B" means that B belongs to A, but B does not necessarily constitute all of A; A may also include other elements such as C, D, and E.
[0032] Example 1
[0033] This embodiment proposes a hexaneline decomposition method for analyzing six-dimensional chemical data, namely the Six-way ACM method. The specific implementation process includes:
[0034] S1. Propose a six-linear component model.
[0035] Specifically, based on the trilinear component model, the following hexalinear component model is proposed:
[0036]
[0037] Where a io ,b jo ,c ko ,d lo ,e mo and f no These represent elements in the six contour matrices A (I×O), B (J×O), C (K×O), D (L×O), E (M×O), and F (N×O), respectively. ijklmn It is a residual array EThe elements in (I×J×K×L×M×N) represent the number of components in the study system, including all components capable of generating a response signal, such as the analyte of interest, known interfering substances, unknown interfering substances, and noise. I, J, K, L, M, and N represent the dimensionality of different modes of the constructed array. For example, in this embodiment, I represents the number of excitation wavelength channels; J represents the number of emission wavelength channels; K represents the number of pH levels; L represents the number of dilution levels; M represents the number of voltage levels; and N represents the total number of samples.
[0038] S2. Based on an algorithm fusion strategy, a six-way ACM (Six-line decomposition method) is proposed.
[0039] Specifically, based on the advantages and disadvantages of the six-linear component model proposed in S1 and the published six-linear algorithm (Six-way Parallel Factor Analysis (S6-way PARAFAC-ALS, see Bro Rasmus. Chemometrics and Intelligent Laboratory Systems, 38(1997) 149-171), a new six-linear decomposition algorithm (six-dimensional algorithm combination method, Six-way ACM) is proposed, which mainly consists of the following two stages:
[0040] The first stage is the initialization stage. The alternating trilinear algorithm (ATLD, see Wu Hai-Long, Shibukawa Masami, Oguma Koichi, Journal of Chemometrics, 12(1998)1-26) is directly extended to the alternating hexalinear algorithm (AHLD) as the objective function to quickly optimize the random initial values. The objective function expression is as follows:
[0041]
[0042]
[0043]
[0044]
[0045]
[0046]
[0047] Where, diagm() represents extracting diagonal elements and converting them into column vectors; + represents the Moore-Penrose generalized inverse of the matrix; ⊙ represents the Khatri-Rao product; and the resulting six-dimensional matrix is of size I×J×K×L×M×N. X I×J×K×L×M×NSix different three-dimensional matrices were obtained by applying incompletely extended matrices along different dimensions. X I×JKLM×N , X J×KLMN×I , X K×LMNI×J , X L×MNIJ×K , X M×NIJK×L , X N×IJKL×M Then extract X I×JKLM×N The nth front slice matrix, X J×KLMN×I The i-th front slice matrix, X K×LMNI×J The j-th front slice matrix, X L×MNIJ×K The k-th front slice matrix, X M×NIJK×L The l-th front slice matrix, X N×IJKL×M The m-th front slice matrix is obtained respectively Here, the pre-slice matrix refers to the matrix obtained by vertically slicing a three-dimensional matrix along its third dimension; using a three-dimensional matrix... X I×J×K For example, the pre-slice matrix is the matrix obtained by vertically slicing along the K-dimensional axis, where X ..1 Represents a three-dimensional array X I×J×K The first front slice, X ..2 This represents the second pre-slice, and so on. The pre-slice form is as follows: Figure 2 As shown. By randomly initializing A, B, C, D, and E, F is solved using Equation 2. Then, A is solved using Equation 3 based on B, C, D, E, and F, and so on. B, C, D, and E can be solved using Equations 4-7 respectively. Equations 2-7 are repeated until the sum of squared residuals satisfies the following convergence criterion:
[0048]
[0049] SSR (m) Let represent the sum of squares of all elements in the residual matrix after m iterations in the first stage. When the above convergence criterion is met, the first stage stops iterating and the second stage begins.
[0050] The second stage uses Six-way PARAFAC-ALS as the final objective function to minimize the impact of noise. The initial values for this stage are derived from the initial values optimized in the previous stage.
[0051] A = XI×JKLMN [(F⊙E⊙D⊙C⊙B) T ] + (8)
[0052] B = X J×KLMNI [(A⊙F⊙E⊙D⊙C) T ] + (9)
[0053] C = X K×LMNIJ [(B⊙A⊙F⊙E⊙D) T ] + (10)
[0054] D = X L×MNIJK [(C⊙B⊙A⊙F⊙E) T ] + (11)
[0055] E=X M×NIJKL [(D⊙C⊙B⊙A⊙F) T ] + (12)
[0056] F = X N×IJKLM [(E⊙D⊙C⊙B⊙A) T ] + (13)
[0057] A six-dimensional array of size I×J×K×L×M×N X I×J×K×L×M×N Different two-dimensional matrices X are obtained according to the form of the fully extended matrix. I×JKLMN X J×KLMNI X K×LMNIJ X L×MNIJK X M×NIJKL X N×IJKLM Then, based on the initial values (B, C, D, E, and F) from the previous stage, solve for A using Equation 8. Then, solve for B using Equation 9 using A, C, D, E, and F, and so on. Solve for C, D, E, and F using Equations 10-13 respectively. Repeat Equations 8-13 until the sum of squared residuals reaches the following convergence criterion:
[0058]
[0059] SSR (n) Let represent the sum of squares of all elements in the residual matrix after n iterations in the second stage. The iteration stops when the above convergence criterion is met.
[0060] Specifically, to explore the performance of the algorithm proposed in S2, a six-dimensional simulated array of size 40×30×25×10×3×12 was simulated. The relevant MATLAB code for constructing the six-dimensional simulated array is shown in Table 1. To make the simulated data closer to the real instrument signal and to evaluate the algorithm's performance under different noise levels, different levels of homoscedastic noise were added to the six-dimensional simulated array before each algorithm run. This is just one example; however, simulations can be performed according to specific circumstances.
[0061] Specifically, in order to explore the performance of the algorithm proposed by S2, taking the determination of the bactericide thiabendazole (TBZ) in environmental water samples as an example, a six-dimensional data array with real chemical significance was constructed. First, the six dimensions of the six-dimensional data were defined as excitation wavelength, emission wavelength, pH, dilution ratio, detection voltage and number of samples.
[0062] Then, 0.2 mol L solutions with pH values of 5.5, 6.0, 6.5, and 7.0 were prepared. -1 Na₂HPO₄-NaH₂PO₄ buffer solution was prepared. Thiamethoxam stock solution and working solution, pirimicarb stock solution and working solution, and 13 samples (7 calibration samples, 3 Xiangjiang River water samples, and 3 spiked Xiangjiang River water samples) were then prepared. Excitation-emission matrix fluorescence was used in conjunction with three auxiliary methods (different pH buffers, dilution ratios, and detection voltages) to measure the above samples, constructing a true six-dimensional chemical matrix. Furthermore, interpolation (see Bahram Morteza, Bro Rasmus, Stedmon Colin, Afkhami Abbas, Journal of Chemometrics, 2006, 20(3-4):99-105) and blank subtraction (i.e., subtracting the average of three blank sample matrix data from each sample matrix) were used to eliminate the effects of Rayleigh scattering and Raman scattering in the fluorescence spectrum. Finally, the pretreated arrays are stacked along the sample dimensions to obtain a six-dimensional array with dimensions of 55×41×4×3×3×13 (excitation×emission×pH×dilution ratio×voltage×sample number). This is just one example; however, it is not limited to this and can be prepared and designed according to actual needs.
[0063] Specifically, the six-dimensional algorithm combination method proposed in S2 and the existing six-dimensional parallel factor analysis algorithm are used to analyze the six-dimensional simulated number array and the six-dimensional real chemical number array constructed in S3 and S4, and the methods are compared, thereby proving the qualitative and quantitative advantages brought by the method proposed in this embodiment through the analysis of six-dimensional data.
[0064] The real three-dimensional array was analyzed using a three-dimensional correction algorithm (alternating trilinear algorithm, ATLD), and the results were compared with those obtained by the proposed algorithm.
[0065] The following detailed embodiments illustrate a six-linear decomposition method of the present invention, namely the Six-way ACM method, which analyzes real six-dimensional chemical data.
[0066] 1. Theory
[0067] 1.1 Hexadecimal Component Model
[0068] If a single sample can yield a fifth-order tensor of size I×J×K×L×M, then stacking a series of samples along the sample dimension can form a six-dimensional data array of size I×J×K×L×M×N. X h .
[0069] In this context, each element x of the six-dimensional data array ijklmn It can be represented in the following form:
[0070]
[0071] Where a io ,b jo ,c ko ,d lo ,e mo and f no These represent elements in the six potential matrices A (I×O), B (J×O), C (K×O), D (L×O), E (M×O), and F (N×O), respectively. ijklmn It is a residual array E The elements in (I×J×K×L×M×N) represent the number of components in the study system, including all components capable of generating a response signal, such as the analyte of interest, known interfering substances, unknown interfering substances, and noise. I, J, K, L, M, and N represent the dimensionality of different modes of the constructed matrix. For example, in this embodiment, I represents the number of excitation wavelength channels; J represents the number of emission wavelength channels; K represents the number of pH levels; L represents the number of dilution levels; M represents the number of voltage levels; and N represents the total number of samples. A graphical representation of the hexalinear component model is shown below. Figure 1 As shown.
[0072] 1.2 Six-dimensional parallel factor analysis
[0073] Six-way PARAFAC-ALS is an extension of 3D to 5D PARAFAC-ALS. It primarily designs the objective function based on the following fully extended matrix form and the alternating least squares principle. Therefore, the least squares solutions for matrices A, B, C, D, E, and F are as follows:
[0074] A = X I×JKLMN [(F⊙E⊙D⊙C⊙B) T ] +(8)
[0075] B = X J×KLMNI [(A⊙F⊙E⊙D⊙C) T ] + (9)
[0076] C = X K×LMNIJ [(B⊙A⊙F⊙E⊙D) T ] + (10)
[0077] D = X L×MNIJK [(C⊙B⊙A⊙F⊙E) T ] + (11)
[0078] E=X M×NIJKL [(D⊙C⊙B⊙A⊙F) T ] + (12)
[0079] F = X N×IJKLM [(E⊙D⊙C⊙B⊙A) T ] + (13)
[0080] Where + denotes the Moore-Penrose generalized inverse of the matrix. Because this algorithm follows strict least squares, it minimizes the impact of noise. (F⊙E⊙D⊙C⊙B) T (A⊙F⊙E⊙D⊙C) T ,(B⊙A⊙F⊙E⊙D) T ,(C⊙B⊙A⊙F⊙E) T (D⊙C⊙B⊙A⊙F) T and (E⊙D⊙C⊙B⊙A) T The large matrix size of the six-dimensional parallel factor consumes a significant amount of memory, resulting in slow convergence and low computational efficiency. Furthermore, similar to the three-dimensional parallel factor, the six-dimensional parallel factor is also prone to computational "swamps" and is affected by two-factor degradation.
[0081] 1.3 Six-Dimensional Algorithm Combination Method
[0082] It is mainly divided into the following two stages.
[0083] The first stage is the initialization stage. The Alternating Trilinear Decomposition (ATLD) algorithm is directly extended to the Alternating Hexagonal Decomposition (AHLD) algorithm as the objective function to quickly optimize random initial values. The objective function is based on a six-dimensional incompletely extended matrix. Based on the least squares solution principle, A, B, C, D, E, and F are updated as follows:
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090] The second stage uses Six-way PARAFAC-ALS as the final objective function to minimize the impact of noise. The objective function is shown in Equation 8-13:
[0091] After optimization, the first-stage convergence criterion was set to 10. -3 The second-stage convergence criterion is set at 10. -6 Therefore, the convergence thresholds for the first and second stages are:
[0092]
[0093]
[0094] SSR (m) and SSR (n) Let represent the sum of squares of all elements in the residual matrix after m and n iterations in the first and second stages, respectively. Furthermore, the maximum allowed number of iterations for the first stage is set to 200, and the maximum allowed number of iterations for the second stage is set to 500.
[0095] The Six-way ACM operation steps are as follows: Figure 3 As shown, the process mainly consists of two stages. The first stage uses AHLD to quickly optimize random initial values, resulting in fast convergence and less susceptibility to excessive sample sizes and 2FD. The second stage uses Six-way PARAFAC-ALS, which ensures stable convergence and excellent performance even under low signal-to-noise ratio and high collinearity conditions. Therefore, the Six-way ACM in this embodiment not only combines the advantages of the two algorithms mentioned above but also possesses greater versatility and better overall performance.
[0096] 2. Experimental Design
[0097] 2.1 Construction of a Six-Dimensional Simulation Array
[0098] A six-dimensional simulated data array of size 40×30×25×10×3×12 was constructed to explore the performance of the proposed algorithm. The simulated data array has three components. The normalized profile plots of the three components for each mode are shown below. Figure 4 As shown. From Figure 3 As can be seen, there is a certain degree of overlap between the various components under different modes, which is frequently encountered in experiments. The sixth dimension represents the concentration dimension, which records the concentration information of the three components in the 12 samples. The simulated concentration values are randomly generated between 0.1 and 1. The first 7 samples out of the 12 samples are the calibration set without unknown interference; the last 5 samples are the prediction set. Among them, component 1 (blue), component 2 (red), and component 3 (green) represent the target analyte, the corrected interference, and the uncorrected interference, respectively.
[0099] Furthermore, to make the simulated data more closely resemble real instrument signals and to evaluate the algorithm's performance under different noise levels, different levels of homoscedastic noise were added to the simulated dataset before each algorithm run. The simulation equations are as follows:
[0100]
[0101] Where n = 1, 2, ..., N, and N = 12. Here, randn(I, J, K, L, M) represents generating a five-dimensional matrix of size I × J × K × L × M, where each element follows a standard normal distribution; σ represents three different noise levels (low (0.1%), medium (1%), and high (5%)). It is a noise-free simulated five-dimensional data array; This represents the average value of the five-dimensional data array for the nth sample. The MATLAB code for the six-dimensional simulated array is shown in Table 1.
[0102] Table 1. MATLAB code for a six-dimensional simulated number array.
[0103]
[0104]
[0105] a Function y=peak(a,s,x,I); y=zeros(I,1); for i=1:I; y(i)=a*exp(-(ix).^2 / s.^2); end.
[0106] b l = 1:10.
[0107] c m = 1:3.
[0108] 2.2 Methods for obtaining true six-dimensional hexalinear chemical data
[0109] 2.2.1 Experimental Apparatus and Materials
[0110] Instruments: F-7000 fluorescence spectrophotometer; KQ-250TDB high-frequency CNC ultrasonic cleaner; TGL-16A benchtop high-speed refrigerated centrifuge. Data analysis and program execution were performed using MATLAB software.
[0111] Materials: Thiamethoxam and pirimicarb standards (≥99%) were purchased from Aladdin (Shanghai, China). Chromatographic grade methanol was purchased from Sigma-Aldrich (St. Louis, USA). Ultrapure water was provided by a Milli-Q ultrapure water system (Millipore, MA, USA). 0.2 mol / L solutions with pH values of 5.5, 6.0, 6.5, and 7.0 were prepared. -1 Na₂HPO₄-NaH₂PO₄. River water samples were collected from the Xiangjiang River (Changsha, China). Before fluorescence analysis, the environmental water samples were filtered through a 0.45 μm aqueous filter membrane. Considering that both thiabendazole and imidacloprid are insoluble in water, and that methanol solvent has different effects on different pH buffers, the stock solutions of thiabendazole and imidacloprid were obtained by diluting the stock solutions with pure methanol, and the working solutions were obtained by diluting the stock solutions with ultrapure water.
[0112] 2.2.2 Parameter Settings
[0113] Fluorescence data were collected using an F-7000 fluorometer equipped with a 150W xenon lamp. The instrument parameters were set as follows: excitation wavelength range: 200-400nm (5nm sampling interval); emission wavelength range: 270-540nm (5nm sampling interval); scan speed: 30000nm / min. -1 The excitation slit width was 5 nm, and the emission slit width was 5 nm. Three detector voltages (600 V, 620 V, and 640 V) were used. All samples were scanned using the same 1 cm quartz cuvette.
[0114] 2.2.3 Sample preparation and determination
[0115] A total of 13 samples were prepared, including 7 calibration samples (C01-C07), 3 Xiangjiang River water samples (R01-R03), and 3 spiked Xiangjiang River water samples (PR01-PR03) for subsequent analysis. The thiamethoxam content in each sample is listed in Table 2. Imidacloprid was added to samples R01-R03 and PR01-PR03 as a known interfering agent. Each sample was diluted with three different dilution ratios (1:1, 2:1, 3:1), and 1 mL of the prepared phosphate buffer solutions at different pH values (5.5, 6.0, 6.5, 7.0) was added to all samples. Finally, the volume was adjusted to 10 mL with ultrapure water. For the preparation of spiked and real water samples with different dilution ratios, in addition to adding phosphate buffer solutions at different pH values, a certain volume of the known interfering agent imipod and Xiangjiang River water also needed to be added. It is worth noting that samples with the same dilution ratio should have the same volume of imipod and Xiangjiang River water added. Samples at different dilution ratios were added to different volumes of pirimicarb and Xiangjiang River water, with the added volume ratio consistent with the dilution ratio. The above samples were measured using excitation-emission matrix fluorescence combined with three auxiliary methods (different pH buffers, dilution ratios, and detection voltages) to construct a realistic six-dimensional hexagonal linear chemical matrix.
[0116] Table 2. Concentration Designs of Thiamethoxam and Pirimicarb in the Calibration Set, Prediction Set, and Spiked Prediction Set
[0117]
[0118] 3 Results and Discussion
[0119] 3.1 Simulation Data Analysis
[0120] 3.1.1 Algorithm Parameters and Evaluation Criteria
[0121] The above six-dimensional simulation matrix was analyzed simultaneously using the Six-way ACM and Six-way PARAFAC-ALS methods, and the results were compared. To avoid randomness and improve the reliability of the results, both algorithms were run 100 times. The maximum number of iterations allowed by the Six-way ACM algorithm was 200, and the maximum number of iterations allowed by the Six-way PARAFAC-ALS algorithm was 500. The following seven parameters were calculated to evaluate the performance of these algorithms: (1) average recovery rate; (2) root mean square error of prediction (RMSEP); (3) profile consistency (Cons); (4) underfit (LOF1 and LOF2); (5) number of iterations; (6) iteration time; and (7) number of failed runs. Detailed information and calculation formulas for these parameters are shown in Table 3. The algorithm is considered successful when the results simultaneously meet the following conditions: (1) the average recovery rate of the target analyte is between 90% and 110%; (2) the consistency is between 0.9 and 1.1; (3) LOF1 and LOF2 should be less than 0.1; and (4) the number of iterations does not exceed the maximum number of iterations mentioned above. Otherwise, the algorithm fails.
[0122] Table 3 shows the performance parameters of the two algorithms for processing six-dimensional analog arrays with different noise levels when the set fraction O=3. a
[0123]
[0124] a The results in the table are the average of the correct results obtained during 100 runs of the algorithm. Only the recovery rate, RMSEP, Cons, LOF1, LOF2, number of iterations, and iteration time of analyte 1 were calculated.
[0125] b Here, P refers to the number of predicted samples, y p and Let represent the true concentration and relative predicted value of the P-th sample in the prediction set, respectively.
[0126] c in and These are the analytical contours for the first five dimensions, while a, b, c, d, and e are the actual contours for the first five dimensions.
[0127] d For LOF1, and x ijklmn These represent the elements in the fitted six-dimensional matrix and the original, undisturbed matrix, respectively. For LOF2, only the number of components with true chemical significance is represented. and xijklmm These represent the elements in the analytically obtained six-dimensional matrix and the original matrix without interference, respectively.
[0128] e The value in square brackets represents the standard deviation (SD) of the result.
[0129] 3.1.2 Analysis of Six-Dimensional Simulation Array
[0130] First, a six-dimensional simulation matrix (σ = 1%) under moderate noise conditions was analyzed to compare and evaluate the performance of Six-way ACM and Six-way PARAFAC-ALS. The results are summarized in Table 3. As shown in the table, when the correct number of components O = 3 was selected, both algorithms successfully ran in 100 runs, obtaining good qualitative and quantitative results, demonstrating a "second-order advantage". The results for average recovery, RMSEP, and Cons were similar and quite satisfactory for both algorithms. The number of iterations and iteration time for Six-way ACM were 2.9 and 7.4 s, respectively, while those for Six-way PARAFAC-ALS were 17.2 and 24.0 s. Compared to Six-way PARAFAC-ALS, the Six-way ACM algorithm has a faster running speed.
[0131] Furthermore, this embodiment also explores the noise resistance of the method described in this embodiment. In this work, six-dimensional data arrays with three different noise levels (including 0.1%, 1%, and 5%) were simulated and analyzed using Six-way ACM and Six-way PARAFAC-ALS. The results are shown in Table 3. The results show that both algorithms have strong noise tolerance, with 0 failures in 100 runs. The results for average recovery rate, RMSEP, and Cons are similar. For six-dimensional data arrays with low, medium, and high noise levels, the LOF1 and LOF2 obtained by Six-way PARAFAC-ALS are approximately two orders of magnitude higher than the LOF2 obtained by Six-way ACM. The values in square brackets represent the standard deviation (SD) of the results; the SD value obtained by Six-way ACM is smaller than that obtained by Six-way PARAFAC-ALS. Furthermore, there are significant differences between Six-way ACM and Six-way PARAFAC-ALS in terms of the number of iterations and iteration time. Six-way ACM consistently runs faster than Six-way PARAFAC-ALS, and the results obtained by Six-way ACM are more accurate and stable. Because the objective function of Six-way PARAFAC-ALS uses a large matrix, it consumes a significant amount of memory, resulting in slow convergence and low computational efficiency. In contrast, the method in this embodiment uses AHLD based on an incompletely extended matrix form as the initial objective function, greatly improving the iteration speed.
[0132] In multidimensional correction methods, the selection of the number of components is crucial for the decomposition of multidimensional data arrays. Incorrect component selection often leads to erroneous qualitative and quantitative results. However, due to the complexity and diversity of real-world matrices, choosing the correct number of components is difficult. Developing algorithms insensitive to excessive component counts has significant practical value and demonstrates clear advantages in real-world applications. Therefore, sensitivity to component counts has become an important indicator for comparing algorithm performance. In this work, excessive component counts (O=4 and O=5) were selected to evaluate the performance of two methods when processing simulated data arrays. The results are shown in Table 4. When O=4, Six-way PARAFAC-ALS failed 59 out of 100 runs. When O=5, the error rate reached as high as 73%. The analytical plot for O=5 is shown in Table 4. Figure 5For Six-way ACM, it can run successfully and obtain similar decomposition results regardless of whether O=4 or O=5. The average iteration time and the number of iterations required gradually increase with the increase of the number of components. It is worth noting that only the correct decomposition results are used for the calculation of performance parameters. Furthermore, the RMSEP, SD of average recovery, Cons, and SD of LOF2 obtained from Six-way PARAFAC-ALS are all greater than those from Six-way ACM, indicating that the consistency of the decomposition results is not as good as that of Six-way ACM. Figure 5 As shown in (A1-F1), when too many scores are selected, the Six-way ACM algorithm can treat these scores as noise, thus not affecting the qualitative and quantitative information of the target analyte. However, this is not the case with Six-way PARAFAC-ALS. Figure 5 The (A2-F2) result clearly shows a two-factor degradation problem, producing unsatisfactory or even erroneous results. This also indicates that Six-way PARAFAC-ALS can only obtain correct results by selecting the correct number of components.
[0133] The results from the six-dimensional simulation matrix demonstrate that the method in this embodiment not only has strong noise resistance but also fast convergence speed and is insensitive to an excessive number of components. Furthermore, it fully demonstrates that the method in this embodiment has good practicality and stability.
[0134] Table 4 Performance parameters of the two algorithms for processing six-dimensional analog arrays with noise level σ = 1%.
[0135]
[0136] 3.2 Analysis of Real Chemical Six-Dimensional Data
[0137] To explore and analyze potentially high-order instrument data in the future, this embodiment employs excitation-emission matrix fluorescence combined with three auxiliary methods to measure different samples. The samples are then stacked along their dimensions to obtain a true six-dimensional chemical matrix of size 55×41×4×3×3×13. The aforementioned algorithm is used to analyze this matrix to explore the feasibility of qualitative and quantitative analysis of thiabendazole in environmental water samples. Furthermore, the performance of the Six-way ACM in processing the true six-dimensional matrix is investigated, and the results are compared with those of the Six-way PARAFAC-ALS. Finally, to demonstrate the "higher-order advantage," a three-dimensional calibration algorithm (ATLD algorithm) is used to analyze the three-dimensional matrix (55×41×13, excitation × emission × sample).
[0138] 3.2.1 Data Preprocessing and Spectral Characteristics
[0139] After measuring each sample with an F-7000 instrument, a data matrix of size 55×41 was obtained. During the analysis, Rayleigh scattering and Raman scattering disrupt the linear structure of the original data; therefore, interpolation and blank subtraction methods were used to eliminate the effects of Rayleigh and Raman scattering. The three-dimensional fluorescence spectra of thiabendazole, imidacloprid, and PR03-specified water samples after scattering subtraction are shown below. Figure 6 As shown. From Figure 6 As can be seen, thiabendazole exhibits fluorescence peaks at excitation wavelengths of 250 nm and 295 nm, and an emission wavelength of 340 nm. Imidacloprid shows fluorescence peaks at excitation wavelengths of 250 nm and 310 nm, and an emission wavelength of 380 nm. The real sample of PR03 shows fluorescence peaks at excitation wavelengths around 250 nm and 300 nm, and an emission wavelength around 370 nm. Clearly, there is severe fluorescence peak overlap and even peak embedding between the target analyte and known interfering substances. Without any pre-separation or pretreatment, traditional univariate analysis methods cannot obtain qualitative and quantitative results for the target analyte. Fortunately, the method in this embodiment has "second-order advantage" or even "higher-order advantage," allowing for easy and accurate extraction of qualitative and quantitative information of the target substance.
[0140] 3.2.2 Effects of pH, voltage and dilution ratio
[0141] To construct a true six-dimensional chemical data array, this embodiment selected three auxiliary methods affecting fluorescence intensity: pH buffer solution, dilution ratio, and detection voltage. Different pH buffer solutions have a significant impact on the fluorescence intensity of substances. This embodiment explored the effect of pH buffer solution on the fluorescence intensity of thiabendazole and imidacloprid. Within a suitable pH range (5.5-7.0), the fluorescence intensity of thiabendazole decreased with increasing pH, while the fluorescence intensity of imidacloprid gradually increased. Furthermore, the changes in fluorescence intensity of thiabendazole at different dilution ratios (1:1, 2:1, 3:1) and three detection voltages (600V, 620V, and 640V) were investigated. With increasing voltage, the fluorescence intensity of thiabendazole and imidacloprid increased, but so did the instrument noise. The trend of thiabendazole fluorescence intensity with pH, dilution ratio, and detection voltage is shown in the figure below. Figure 7 As shown (a3-a5, dashed lines). From Figure 7 It can be seen that there are serious collinearity problems in both the dilution ratio dimension and the detection voltage dimension, which places higher demands on the required data analysis methods.
[0142] 3.2.3 Qualitative and quantitative analysis of thiabendazole in river water
[0143] First, the CORCONDIA method was used to estimate the number of components to improve the computational speed and accuracy of the six-dimensional correction algorithm. The results showed that when O < 2, the CORCONDIA value was 100%, while when O = 3, the CORCONDIA value rapidly decreased to below 70%. The CORCONDIA value decreased further with increasing component numbers. Therefore, 2 was chosen as the correct component number, indicating the presence of cofactor interference from the target analyte thiamethoxam and the known interfering analyte imidacloprid, or analytes with similar fluorescence peaks. Then, the six-dimensional data array was decomposed using the method of this embodiment and the Six-way PARAFAC-ALS algorithm to obtain the corresponding qualitative and quantitative results. Furthermore, to demonstrate "higher-order advantage," a three-dimensional correction algorithm, represented by the ATLD algorithm, was used to analyze the three-dimensional data array (pH 5.5, dilution ratio 1:1, voltage 640). The Six-way ACM resolution spectrum is shown below. Figure 7 As shown.
[0144] All three algorithms successfully extracted the pure signal of thiabendazole even in the presence of interference. The normalized excitation, emission, pH, dilution ratio, and voltage curves obtained by Six-way ACM and Six-way PARAFAC-ALS were consistent with the normalized actual spectrum (dashed lines). The consistency between the normalized spectrum obtained by ATLD and the actual spectrum was not as good as that of the proposed Six-way ACM. By performing linear regression on the true concentration information and the resolved relative concentration information of the calibration set to establish corresponding calibration curves, the analyte concentration in the real sample and the real spiked predicted sample was predicted, yielding quantitative results. The specific quantitative results and performance parameters of the three algorithms are summarized in Table 5. Due to the relative uniqueness of hexagonal and trilinear decomposition, the correlation coefficient, average recovery rate, and RMSEP results obtained from repeated runs were almost identical. To avoid randomness, the average number of iterations and the average iteration time (out of a total of 100 runs) for effective runs were calculated, and the number of failures was also obtained. For Six-way ACM, the mean recovery and standard deviation of thiabendazole were 100 ± 0.26%, the correlation coefficient was 0.9997, and the RMSEP was 2.98 ng / mL. -1 The quantitative results were satisfactory. All 100 runs were successful, with an average of 16.0 iterations and an average iteration time of 2.78 s. Based on the Six-way PARAFAC-ALS algorithm, one run failed out of 100 runs, resulting in an average recovery rate of 127.8%, which is outside the reasonable range (80%-120%). In effective runs, the correlation coefficient, average recovery rate, RMSEP, average number of iterations, and average iteration time obtained by the Six-way PARAFAC-ALS algorithm were 0.9997, 100.6 ± 0.26%, and 3.00 ng / mL, respectively.-1 The time to recovery was 32.6 and 4.55 s, respectively. The Six-way PARAFAC-ALS algorithm suffers from error execution and therefore cannot guarantee 100% feasibility. The correlation coefficient, average recovery rate, RMSEP, average number of iterations, and average iteration time obtained by the ATLD algorithm were 0.9977, 99.4 ± 4.23%, and 12.3 ng / mL, respectively. -1 The results were 6.39 and 0.007s. Although the qualitative and quantitative results obtained by ATLD were within acceptable ranges, they were clearly not as good as the results of Six-way ACM.
[0145] Table 5. Quantitative results and performance parameters obtained by two algorithms when analyzing the six-dimensional data array of real excitation-emission-pH-dilution ratio-voltage-sample.
[0146]
[0147] a This indicates the average recovery rate and standard deviation of thiabendazole in samples PR01-PR03.
[0148] To further compare algorithm performance and explore "higher-order advantages," the quality factor (limit of detection (LOD), limit of quantitation (LOQ), sensitivity (SEN), and selectivity (SEL)) results of the three algorithms are summarized in Table 6. The quality factor results obtained by the two six-dimensional calibration algorithms are similar, with their LOD and LOQ values being lower than those of the three-dimensional calibration algorithm, enabling better detection of low-volume pesticides. Their SEN and SEL are higher than those of the three-dimensional calibration method, further demonstrating that the method in this embodiment possesses "higher-order advantages," namely higher selectivity and sensitivity, more stable and accurate results, and better handling of collinearity issues.
[0149] In summary, this invention proposes a novel six-linear decomposition method, namely the Six-way ACM (Six-dimensional Algorithm Combination Method), which can successfully analyze real six-dimensional data and a series of six-dimensional simulated arrays with varying noise levels. The six modes of the real data's six-dimensional array include excitation wavelength, emission wavelength, pH, dilution ratio, detection voltage, and sample. From the simulated and real data results, it can be concluded that the overall performance of the Six-way ACM in this embodiment is excellent. It not only has a fast convergence speed and is insensitive to excessive components and noise, but also obtains stable and accurate qualitative and quantitative results even in the presence of known interference, unknown interference, peak overlap, and strong collinearity. Compared with three-dimensional correction methods, the Six-way ACM also exhibits "higher-order advantages," extracting richer information, possessing strong anti-collinearity capabilities, and higher sensitivity and selectivity. In conclusion, this invention not only provides data analysis tools for future high-order instruments but also provides real data support and methodological reference for theoretical research in higher-order tensor algebra.
[0150] Table 6. Results of thiabendazole quality factors in environmental water samples obtained based on three algorithms.
[0151]
[0152] Example 2
[0153] Embodiment 2 of the present invention provides a terminal device corresponding to Embodiment 1 above. The terminal device can be a processing device for a client, such as a mobile phone, a laptop, a tablet computer, a desktop computer, etc., to execute the method of the above embodiments.
[0154] The terminal device in this embodiment includes a memory, a processor, and a computer program stored in the memory; the processor executes the computer program in the memory to implement the steps of the method in Embodiment 1 described above.
[0155] In some implementations, the memory may be high-speed random access memory (RAM), and may also include non-volatile memory, such as at least one disk storage device.
[0156] In other implementations, the processor can be any type of general-purpose processor, such as a central processing unit (CPU) or a digital signal processor (DSP), and there is no limitation here.
[0157] Example 3
[0158] Embodiment 3 of the present invention provides a computer-readable storage medium corresponding to Embodiment 1 above, on which a computer program / instructions are stored. When the computer program / instructions are executed by a processor, they implement the steps of the method of Embodiment 1 above.
[0159] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any combination thereof.
[0160] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented in various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0161] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0162] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0163] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0164] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A six-dimensional data decomposition method, characterized by, The method comprises the following steps: S1, construct a six-linear component model, obtain a six-dimensional number array with a size of I×J×K×L×M×N X I×J×K×L×M×N ; The six linear component model expression is as follows: i=1, 2, …, I; j=1, 2, …, J; k=1, 2, …, K; l=1, 2, …, L; m=1, 2, …, M; n=1, 2, …, N Where x ijklmn for X I×J×K×L×M×N The elements in, i.e. represents the n The response intensity of a sample in the i-th, j-th, k-th, l-th, and m-th channels; a io ,b jo ,c ko ,d lo ,e mo and f no These represent the elements in the first through sixth contour matrices A, B, C, D, E, and F, respectively; e ijklmn It is a residual array E The elements in the table are: O represents the number of components in the study system; I represents the number of excitation wavelength channels; J represents the number of emission wavelength channels; K represents the number of pH levels; L represents the number of dilution levels; M represents the number of voltage levels; and N represents the total number of samples. S2, randomly initialize profile matrices A, B, C, D, E; set X I×J×K×L×M×N respectively along different dimensions, six different three-dimensional number arrays are obtained in the form of incomplete expansion matrix X I×JKLM×N , X J×KLMN×I , X K×LMNI×J , X L×MNIJ×K , X M×NIJK×L , X N×IJKL×M ; extracting X I×JKLM×N the nth pre-slice matrix of the matrix, X J×KLMN×I the ith pre-slice matrix of the matrix, X K×LMNI×J the jth pre-slice matrix of the matrix, X L×MNIJ×K the kth pre-slice matrix of the matrix, X M×NIJK×L the lth pre-slice matrix of the matrix, X N×IJKL×M the mth pre-slice matrix of the matrix, respectively, S3, according to initialized A, B, C, D, E, using formula calculate F; according to the calculated F, initialized B, C, D, E, using formula calculate A; according to the calculated F, A, initialized C, D, E, using formula calculate B; according to the calculated F, A, B, initialized D, E, using formula calculate C; according to the calculated F, A, B, C, initialized E, using formula calculate D; according to the calculated F, A, B, C, D, using formula calculate E; wherein, diagm() represents extracting diagonal elements into column vector; + represents Moore-Penrose generalized inverse of matrix; ⊙ represents Khatri-Rao product; f (n) , a (i) , b (j) , c (k) , d (l) , e (m) respectively represent the nth row of F, the ith row of A, the jth row of B, the kth row of C, the lth row of D, and the mth row of E. S4, judging whether the first convergence condition is established: If yes, the optimized six profile matrices are obtained; otherwise, the profile matrix obtained in this iteration is taken as the initial value, and the step S3 is returned; wherein, SSR (m) represents the sum of squares of all elements in the residual matrix after m iterations, and ε1 is the first threshold value set. E 2. The six-dimensional data decomposition method of claim 1, wherein, When the first convergence condition is met, further comprising: S5, will X I×J×K×L×M×N Different two-dimensional matrices X are obtained respectively in the form of fully expanded matrices I×JKLMN , X J×KLMNI , X K×LMNIJ , X L×MNIJK , X M×NIJKL , X N×IJKLM ; S6, taking the optimized second to sixth profile matrices obtained in step S4 as initialization values; S7, substitute the initialization value into formula A=X I×JKLMN [(F⊙E⊙D⊙C⊙B) T ] + , obtain the first profile matrix A of current optimization; substitute the first profile matrix A of current optimization, the initial values of the third to sixth profile matrices into formula B=X J×KLMNI [(A⊙F⊙E⊙D⊙C) T ] + , obtain the second profile matrix B of current optimization; substitute the first profile matrix A of current optimization, the second profile matrix B of current optimization, the initial values of the fourth to sixth profile matrices into formula C=X K×LMNIJ [(B⊙A⊙F⊙E⊙D) T ] + , obtain the third profile matrix C of current optimization; substitute the first profile matrix A of current optimization, the second profile matrix B of current optimization, the third profile matrix C, the initial values of the fifth to sixth profile matrices into formula D=X L×MNIJK [(C⊙B⊙A⊙F⊙E) T ] + , obtain the fourth profile matrix D of current optimization; substitute the first profile matrix A of current optimization, the second profile matrix B of current optimization, the third profile matrix C of current optimization, the fourth profile matrix D of current optimization, the initial value of the sixth profile matrix into formula E=X M×NIJKL [(D⊙C⊙B⊙A⊙F) T ] + , obtain the fifth profile matrix E of current optimization; substitute the first profile matrix A of current optimization, the second profile matrix B of current optimization, the third profile matrix C of current optimization, the fourth profile matrix D of current optimization, the fifth profile matrix E of current optimization into formula F=X N×IJKLM [(E⊙D⊙C⊙B⊙A) T ] + , obtain the sixth profile matrix F of current optimization; S8, judging whether the second convergence condition is established: If yes, ending and outputting the final first to sixth profile matrices; otherwise, taking the second to sixth profile matrices obtained in the current iteration as the initialization values and returning to step S7; (n) denotes the square sum of all elements in the residual matrix after the nth iteration, and ε2 is the second threshold value set. E denotes the square sum of all elements in the residual matrix after the nth iteration, and ε2 is the second threshold value set.
3. The six-dimensional data decomposition method of claim 1 or 2, wherein, ε1 is set to 10 -3 .
4. The six-dimensional data decomposition method of claim 2, wherein, ε2 is set to 10 -6 .
5. A terminal device comprising a memory, a processor, and a computer program stored on the memory; characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1-4.
6. A computer readable storage medium having stored thereon computer programs / instructions; characterized in that, The computer program / instructions are executed by the processor to implement the steps of the method of any one of claims 1-4.
Citation Information
Patent Citations
Method for measuring content of carbendazim in fruit juice based on three-dimensional fluorescence
CN114460050A
Similarity index: a rapid classification method for multivariate data arrays
US20090306932A1