Classification Model Construction Method, Bronze Rust Component Analysis Method and System

The SVM-based classification model with optimized kernel functions and preprocessing techniques addresses the challenge of identifying bronze artifact corrosion components, achieving enhanced accuracy and robustness in distinguishing between different corrosion types.

CN116956166BActive Publication Date: 2025-07-15UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310931032.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-26
Publication Date
2025-07-15
Estimated Expiration
2043-07-26

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the components of bronze rust, especially the weak characteristic absorption peak signal and background noise, which leads to low accuracy of classification models.

Method used

The optimized Support Vector Machine (SVM) classification model is adopted, and the optimized kernel function and loss function are set, combined with preprocessing and principal component analysis, the importance of frequency points and the robustness of data are improved, and the classification accuracy of the model is improved.

Benefits of technology

It improves the accuracy of the analysis of the components of the bronze rust substance, enhances the ability to identify the demarcation line of the corrosion layer, reduces the operation complexity and noise influence, and improves the generalization performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116956166B_ABST
    Figure CN116956166B_ABST
Patent Text Reader

Abstract

Classification model construction method, bronze corrosion component analysis method and system. By setting the kernel function of the support vector machine classification model, this classification model construction method can map the bronze corrosion sample data to a sample space that is easier to divide by a hyperplane. At the same time, it improves the ability to rank the importance of frequency points, and is applicable to the classification of bronze corrosion with weak terahertz signals, small feature absorption peak intensities and difficult to separate from background noise or other feature absorption peaks, improving the accuracy of the SVM classification model after training for the analysis of the components of bronze corrosion products. In addition, by optimizing the loss function and combining it with the optimized kernel function, the accuracy and generalization ability of the support vector machine classification model after training can be effectively further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of detection of bronze corrosion components, and particularly to a method for constructing a classification model, a method and a system for analyzing bronze corrosion components. Background Art

[0002] Bronze wares are cultural relics that span thousands of years and are crucial for studying the development of human civilization. However, whether they are museum collections, handed-down items, or buried deep underground for a long time, different degrees of corrosion will appear on the surface of bronze wares. Moreover, although greatly affected by the compositional differences and environmental differences of the bronze wares themselves, bronze ware cultural relics are difficult to follow a universal corrosion mechanism, and various corrosion products will appear on their surfaces. Among these corrosion products, some will continuously expand and penetrate the corrosion areas of bronze wares, with greater destructiveness. For example, copper chloride, this kind of corrosion product is considered "harmful rust", while some corrosion products have much weaker destructiveness to bronze wares and are also considered "harmless rust". In the protection of cultural relics, non-destructively detecting the corrosion components on the surface of bronze wares has important guiding significance for formulating cultural relic protection strategies and methods.

[0003] Terahertz waves have various advantages in non-destructive detection of cultural relics. Terahertz waves not only have good penetration for many dielectric materials and non-polar substances, but also have a high lateral resolution due to their shorter wavelength compared to microwaves. At the same time, the photon energy of terahertz waves is only in the order of millielectron volts and will not cause harmful ionization reactions. In addition, the time resolution of terahertz pulse waves is in the picosecond order of magnitude, which can identify material stratifications at the micron level and obtain internal tomographic information. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for constructing a classification model. This construction method is based on the SVM (Support Vector Machine) classification model, uses the frequency-domain signals of bronze corrosion samples to train the SVM classification model, and when training, by setting an optimized kernel function, maps the corrosion samples to a sample space that is easier to divide the hyperplane for the support vector machine model, and at the same time strengthens the importance ranking of frequency points, thereby effectively improving the accuracy of the trained SVM classification model for analyzing the components of bronze corrosion products.

[0005] The above purpose is achieved by the following technical solutions:

[0006] A method for constructing a classification model, comprising the following steps:

[0007] Collect the time-domain signals of bronze corrosion samples and convert the time-domain signals into frequency-domain signals;

[0008] Train a support vector machine classification model based on the frequency-domain signals to obtain a trained classification model;

[0009] Among them, the kernel function of the support vector machine classification model is:

[0010] or

[0011] or

[0012]

[0013] In the formula, K(x (i) , x (j) ) is the kernel function required for machine learning; x (i) is the value corresponding to the i-th frequency point in the sample points; x (j) is the value corresponding to the j-th frequency point in the sample points; σ is the bandwidth of the Gaussian kernel; d is the degree of the polynomial.

[0014] In this technical solution, a terahertz time-domain spectrometer (TDS) is used to collect the time-domain signal of the bronze corrosion sample, and the time-domain signal is converted into a frequency-domain signal by performing a fast Fourier transform (FFT) on the time-domain signal to obtain the positions of the characteristic absorption peaks.

[0015] In this technical solution, the bronze corrosion sample can be a standard sample of corrosion powder with known composition. For example, common bronze corrosion powders such as atacamite, paratacamite, botallackite, and cuprous chloride, or a standard corrosion sample obtained by pressing a mixture of multiple corrosion powders in a certain proportion into a sheet; the bronze corrosion sample can also be the surface corrosion of a bronze cultural relic. Its frequency-domain signal is divided into a training set and a test set, and the model is trained using the frequency-domain signal in the training set. The frequency-domain signal input into the model is the corresponding frequency point and absorption peak intensity.

[0016] Since the bronze corrosion product absorbs the terahertz signal very strongly, resulting in a weak terahertz signal received by the terahertz time-domain spectrometer. At the same time, the absorption peak intensity of some corrosion products in the frequency-domain signal is small, and it is not easy to separate from the background noise or the absorption peaks of other corrosion products, making it more difficult to identify those very weak absorption peaks.

[0017] Several traditional kernel functions of support vector machines are mainly used to transform input data into another high-dimensional sample space through the kernel function formula, draw a decision hyperplane and two support vector planes on the left and right of the decision hyperplane, so as to perform classification tasks. However, since the characteristic absorption peak signals of some bronze rusts are very weak, they cannot be well separated from other types of rust products when reflected in the sample space, resulting in difficulty in accurately dividing the decision hyperplane and the other two support vector planes to make the accuracy of the model reach the best effect. In addition, the closer the frequency points near the support vector plane are to the decision hyperplane, that is, the more difficult they are to classify compared with the frequency points at other positions in the new sample space. These frequency points are usually some important frequency points with weak absorption peaks. If the importance ranking of such frequency points is missing or unclear, it will also affect the accuracy of model classification.

[0018] To solve the problems existing in the traditional kernel function when converting bronze rust sample data, in this technical solution, three optimized kernel functions are adopted. Among the three kernel functions, the first kernel function x (j) ) is optimized on the basis of the Gaussian function and the linear kernel function. The characteristic that the Gaussian function can map data to an infinite dimension is used to make up for the shortcoming that the linear kernel function can only solve linearly separable problems, so as to improve the classification accuracy of the model. Although its ability to rank the importance of frequency points is only applicable to those linearly separable sample points, and for samples with slightly more complex distributions, its defect in extracting important frequency points is more obvious, it can effectively improve the accuracy of the model compared with the traditional kernel function.

[0019] The second kernel function is optimized on the basis of the Gaussian function and the polynomial function, making up for the shortcomings of the polynomial function, increasing the accuracy of data classification, and improving the accuracy. Moreover, it integrates the advantages of the polynomial kernel function, that is, through the setting of the power number, a summary prediction can be realized, achieving a controllable and precise feature dimension elevation, improving the generalization ability of the model while improving the accuracy of the model, making the model reach an ideal state.

[0020] The third kernel function combines the polynomial kernel function and the Laplace kernel function, achieving the advantages of a good feature dimension elevation of the polynomial kernel function and the good generalization ability of the Laplace kernel function for unknown samples without overfitting phenomena, making the accuracy of the model's ranking of the importance of frequency points improve. However, due to the improvement of its generalization ability, the classification performance of the model does not reach the best.

[0021] In some embodiments, the cross-grid search method can be used to select a kernel function that better matches the loss function and hyperparameters. In one or more preferred embodiments, the second kernel function is preferably adopted.

[0022] By setting the kernel function of the support vector machine classification model, the bronze corrosion sample data can be mapped to a sample space where it is easier to divide the hyperplane. At the same time, the ability to rank the importance of frequency points is improved, which is applicable to the classification of bronze corrosion with weak terahertz signals, small feature absorption peak intensities, and difficulties in separating from background noise or other feature absorption peaks, and improves the accuracy of the SVM classification model trained for analyzing the components of bronze corrosion products.

[0023] Furthermore, the loss function of the support vector machine classification model is as follows:

[0024] loss = lg(1 + exp(-y i (w T x i + b))); or

[0025] loss = rnax(0, 1 + y i (w T x i + b)); or

[0026] loss = max(0, exp(-y i (w T x i + b)))

[0027] In the formula, x i is the value corresponding to the i-th sample point and is a column matrix; y i is the label value corresponding to the i-th sample point; w is the normal vector of the hyperplane, which determines the direction of the hyperplane; b is the displacement term, which determines the distance between the hyperplane and the origin.

[0028] Among the large amount of bronze corrosion data obtained in the early stage, there will inevitably be some abnormal data due to the influence of factors such as machines, humans, and the environment. These data are removed through traditional preprocessing and denoising operations. Therefore, in the new sample space after the action of the kernel function, the decision hyperplane and the support vector plane are interfered by abnormal points and are divided inaccurately or even misclassified, making it difficult to find the corresponding weak feature absorption peaks of various bronze corrosion products in the final classification model.

[0029] In this technical solution, through our optimized loss function, the positions of the decision hyperplane and the two support vector planes obtained by converting data to the new sample space by any of the aforementioned kernel functions can be finely adjusted. This further improves the accuracy and stability of the classification and the importance frequency point ranking performance of the finally trained model.

[0030] As a preferred embodiment of the present invention, the following steps are further included. Before the frequency-domain signal is input into the support vector machine classification model, a preprocessing function is used to transform the frequency-domain signal, and the preprocessing function is:

[0031]

[0032] In the formula, x is the frequency-domain signal before transformation, X is the frequency-domain signal after transformation, and x (i) is the value corresponding to the i-th frequency point in the sample point; x min is the minimum value corresponding to all frequency points in the sample point; x max is the maximum value corresponding to all frequency points in the sample point; n is the number of frequency points in the sample point.

[0033] In this technical solution, aiming at the problems that the terahertz signal of bronze corrosion products is weak, the intensity of some characteristic absorption peaks is low and close to noise, and the signals between different bronze corrosion products vary greatly, a preprocessing function is set to preprocess the frequency-domain signal input into the support vector machine classification model.

[0034] In this technical solution, the adopted preprocessing function can not only map all data to a new small interval, making the accuracy of the data and the connection between data closer, but also for the outliers in the data, due to the sample standard deviation function in the denominator, the preprocessed data is more robust and not easily affected by the outliers in the data. Moreover, the use of half of the sum of the minimum value and the maximum value in the numerator instead of the average value of the original data reduces a part of the computational complexity and greatly improves the preprocessing speed.

[0035] Further, the time-domain signal of the bronze corrosion sample is collected and converted into a first frequency-domain signal, the background signal is collected and converted into a background frequency-domain signal, and the background signal in the first frequency-domain signal is eliminated to obtain the frequency-domain signal.

[0036] In this technical solution, the time-domain signal of the collected bronze corrosion sample is converted into a first frequency signal E f青铜器 , the background signal is collected and converted into a background frequency-domain signal E f背景 , and then the first frequency-domain signal is processed to remove noise, that is, the processed frequency-domain signal is Through the above steps, most of the noise can be eliminated, as well as the influence of amplitude and phase drift due to the too long single measurement time of the terahertz time-domain spectrometer itself, ensuring the effectiveness of the frequency-domain signal, reducing the influence of background noise on the characteristic absorption peaks with small intensity, and further improving the accuracy of the final classification result.

[0037] Furthermore, the frequency domain signal is reduced in dimension by using principal component analysis, and the information content of the feature quantity of the frequency domain signal after the dimension reduction is more than 98% of the information content of the feature quantity of the frequency domain signal before the dimension reduction.

[0038] The spectrum of bronze data measured by terahertz time-domain spectrometer is very wide, and the number of feature quantities is also very large, which makes the training process of SVM classification model time-consuming, and will increase the degree of overfitting in the process of training the model, making the model not generalizable. Therefore, principal component analysis (PCA) can be used to reduce the dimension of the converted frequency domain signal.

[0039] In this technical solution, the existing principal component analysis method can be used to reduce the dimension of the frequency domain signal. However, in order to ensure that the characteristic absorption peak is weakly retained, the feature quantity after dimensionality reduction should have at least 98% of the information of the original feature quantity, so that most of the original information can be retained while improving the model training speed, thereby making the analysis results of the trained classification model more accurate and more generalizable.

[0040] Furthermore, the bronze rust sample includes a standard rust sample, and the preparation of the standard rust sample includes the following steps: rust powder with known composition is fully mixed with polyethylene powder and then pressed into tablets to obtain a rust standard sample, and the mass ratio of the rust powder to the polyethylene powder is 1:2 to 1:3.

[0041] Bronze rust samples have strong absorption of terahertz signals, which can easily cause weak signals during measurement. In order to improve the intensity of the collected signal, when making the standard sample, the rust powder is mixed with polyethylene powder, and then pressed into sheets to obtain the rust standard sample. However, too much polyethylene powder will affect the signal intensity of the rust standard sample, while a low content of polyethylene powder will easily cause the rust standard sample to be more sparse as a whole, and the signal intensity will also be affected. Therefore, preferably, the mass ratio of the rust powder to the polyethylene powder is 1:2 to 1:3, and more preferably, the mass ratio of the rust powder to the polyethylene powder is 1:3.

[0042] After the standard rust sample is prepared, it is left to stand for a period of time, and the frequency domain signal of the bronze rust sample is detected using a terahertz time-domain spectrometer, and the frequency domain signal is used as a type of data for training the SVM classification model.

[0043] Furthermore, the rust powder comprises a first rust powder and a second rust powder, the first rust powder and the second rust powder are mixed according to an initial ratio and pressed into polyethylene sheets for sampling, and on the premise that the total mass of the rust powder remains unchanged, the first rust powder is increased or decreased by a preset weight ratio based on the initial ratio to obtain the remaining mixing ratio and pressed into polyethylene sheets for sampling.

[0044] When different layers of rust form on the surface of bronze ware, the boundary between different rust layers usually has two or more types of rust. In order to further improve the recognition ability of the classification model for the boundary of rust layers, in this technical solution, when preparing the standard rust sample, two types of rust powders and polyethylene are pressed together to make a sample.

[0045] The first rust powder and the second rust powder can be mixed and pressed in a certain initial ratio to obtain the first sample. The initial ratio can be set according to the actual situation of bronze rust formed in nature. For example, the lower layer of atacamite is usually paratacamite, so the common ratio of the two, such as 35:65, is mixed to make a sample to obtain the first sample. Subsequently, with a preset weight ratio of 5%, atacamite and paratacamite are mixed in a ratio of 30:70 to make a sample to obtain the second sample, and so on. Finally, by increasing the number of standard rust samples, the trained classification model can more accurately identify the boundary mixing situation, and to a certain extent, can roughly judge the mass ratio of various rusts at a certain point on the bronze cultural relic based on the frequency domain signal.

[0046] Furthermore, the value of the penalty factor of the support vector machine classification model is 12.25, and the value of gamma is 1. When the penalty factor C is greater than 12.25, the accuracy of the training set will increase, but the accuracy of the test set will decrease, and overfitting will occur. When C is less than 12.25, the accuracy of both the training set and the test set will decrease, indicating that the model is in an underfitting state. Therefore, preferably, the value of the penalty factor is set to 12.25, and the value of gamma is 1. Combining the foregoing preprocessing function, loss function, and kernel function, the classification performance and generalization performance of the SVM model trained with different types of data of bronze rust products can be optimized.

[0047] The present invention also provides an analysis method for analyzing the rust components of bronze ware using a classification model constructed based on any of the foregoing classification model construction methods. The analysis method specifically includes the following steps:

[0048] Collect the time-domain signal of the rust on the surface of the bronze cultural relic, and convert the time-domain signal into a frequency-domain signal of the cultural relic;

[0049] Input the frequency-domain signal of the cultural relic into the trained classification model constructed by any of the foregoing classification model construction methods, and the trained classification model outputs the rust component information corresponding to the frequency-domain signal of the cultural relic.

[0050] The present invention also provides a bronze rust component analysis system, including:

[0051] A terahertz time-domain spectrometer is used to collect the time-domain signals of the rust on the surface of bronze cultural relics and convert the time-domain signals into frequency-domain signals of the cultural relics;

[0052] An analysis unit is used to analyze the frequency-domain signals of the cultural relics by using the trained classification model constructed by any of the foregoing classification model construction methods, and obtain the rust component information corresponding to the frequency-domain signals of the cultural relics.

[0053] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0054] 1. By setting the kernel function of the support vector machine classification model, the present invention can map the bronze rust sample data to a sample space that is easier to divide the hyperplane. At the same time, the ability to rank the importance of frequency points is improved, which is applicable to the classification of bronze rust with weak terahertz signals, small characteristic absorption peak intensities, and difficult to separate from background noise or other characteristic absorption peaks, and improves the accuracy of the trained SVM classification model for analyzing the components of bronze rust;

[0055] 2. By optimizing the loss function, the present invention can fine-tune the positions of the decision hyperplane and the two support vector planes obtained by converting the data to a new sample space by the kernel function, so that the accuracy and stability of the classification and the ranking performance of important frequency points of the finally trained model can be further improved;

[0056] 3. The preprocessing function adopted by the present invention can not only map all data to a new small interval, making the accuracy of the data and the connection between the data closer, but also for the outliers in the data, due to the sample standard deviation function in the denominator, the preprocessed data is more robust and not easily affected by the outliers in the data. Moreover, half of the sum of the minimum value and the maximum value is used in the numerator to replace the average value of the original data, reducing a part of the computational complexity and greatly improving the speed of preprocessing;

[0057] 4. By setting the values of the penalty factor and gamma, and combining the foregoing preprocessing function, loss function and kernel function, the present invention can optimize the classification performance and generalization performance of the SVM model trained with different types of data of bronze rust products;

[0058] 5. The present invention can further improve the recognition ability of the classification model for the boundary line of the rust layer by making samples of the standard sample of the mixed rust powder. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not constitute a limitation to the embodiments of the present invention. In the drawings:

[0060] Figure 1 It is a flowchart of the classification model construction method in a specific embodiment of the present invention;

[0061] Figure 2 It is a flowchart of the bronze corrosion component analysis method in a specific embodiment of the present invention.

[0062] Figure 3 It is a comparison chart of bronze corrosion component recognition between the training model constructed in a specific embodiment of the present invention and the training model constructed by the traditional method. Specific Embodiments

[0063] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the embodiments and the accompanying drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and do not limit the present invention.

[0064] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "front", "rear", "left", "right", "upper", "lower", "vertical", "horizontal", "high", "low", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the protection scope of the present invention.

[0065] Embodiment 1:

[0066] As Figure 1 shown, the classification model construction method includes the following steps:

[0067] Collect the time-domain signal of the bronze corrosion sample and convert the time-domain signal into a frequency-domain signal;

[0068] Train the support vector machine classification model based on the frequency-domain signal to obtain the trained classification model;

[0069] Among them, the kernel function of the support vector machine classification model is:

[0070] Or

[0071] Or

[0072]

[0073] In the formula, K(x (i) , x (j) ) is the kernel function required for machine learning; x (i) is the value corresponding to the i-th frequency point in the sample points; x (j)is the value corresponding to the j-th frequency point in the sample points; σ is the bandwidth of the Gaussian kernel; d is the degree of the polynomial.

[0074] In this embodiment, by setting the kernel function of the support vector machine classification model, the bronze rust sample data can be mapped to a sample space where the hyperplane is easier to divide, and the ability to rank the importance of frequency points is improved. It is applicable to the classification of bronze rust with weak terahertz signals, small characteristic absorption peak intensities, and difficulty in separating from background noise or other characteristic absorption peaks, and improves the accuracy of the SVM classification model for analyzing the composition of bronze rust after training.

[0075] In some preferred embodiments, the kernel function of the support vector machine classification model is preferably the second kernel function.

[0076] Embodiment 2:

[0077] On the basis of Embodiment 1, the loss function of the support vector machine classification model is:

[0078] loss = lg(1 + exp(-y i (w T x i + b))); or

[0079] loss = max(0, 1 + y i (w T x i + b)); or

[0080] loss = max(0, exp(-y i (w T x i + b)))

[0081] where x i is the value corresponding to the i-th sample point, which is a column matrix; y i is the label value corresponding to the i-th sample point; w is the normal vector of the hyperplane, which determines the direction of the hyperplane; b is the displacement term, which determines the distance between the hyperplane and the origin.

[0082] In some preferred embodiments, the loss function loss = max(0, exp(-y i (w T x i + b))) is preferably adopted. This loss function is rewritten based on the hinge function and the logarithmic function, and it not only includes the advantage of the relatively high robustness of the original hinge loss function, being insensitive to outliers and noise, but also has the very good representation of probability distribution of the logarithmic loss function. Therefore, using the modified loss function has a significant improvement in the training results of the model.

[0083] In some preferred embodiments, the value of the penalty factor of the support vector machine classification model is 12.25, and the value of gamma is 1. In this embodiment, by setting the value of the penalty factor to 12.25 and the value of gamma to 1, combined with the aforementioned preprocessing function, loss function, and kernel function, the classification performance and generalization performance of the SVM model trained with data of different types of bronze corrosion products can reach the optimal level.

[0084] Embodiment 3:

[0085] Based on the above embodiments, the following steps are further included. Before the frequency-domain signal is input into the support vector machine classification model, a preprocessing function is used to transform the frequency-domain signal. The preprocessing function is:

[0086]

[0087] In the formula, x is the frequency-domain signal before transformation, X is the frequency-domain signal after transformation, x (i) is the value corresponding to the i-th frequency point in the sample points; x min is the minimum value corresponding to all frequency points in this sample point; x max is the maximum value corresponding to all frequency points in this sample point; n is the number of frequency points in this sample point.

[0088] In this embodiment, the preprocessing function used can not only map all data to a new small interval, making the accuracy of the data and the connection between data closer, but also for the outliers in the data, due to the sample standard deviation function in the denominator, the preprocessed data is more robust and not easily affected by the outliers in the data. Moreover, using half of the sum of the minimum value and the maximum value instead of the average value of the original data in the numerator reduces a part of the computational complexity and greatly improves the speed of preprocessing.

[0089] In some preferred embodiments, the time-domain signal of the bronze corrosion sample is collected and converted into a first frequency-domain signal, the background signal is collected and converted into a background frequency-domain signal, and the background signal in the first frequency-domain signal is eliminated to obtain the frequency-domain signal. This embodiment can eliminate most of the noise and the influence of amplitude and phase drift caused by the too long single measurement time of the terahertz time-domain spectrometer itself, ensure the effectiveness of the frequency-domain signal, reduce the influence of background noise on the small-intensity characteristic absorption peaks, and further improve the accuracy of the final classification result.

[0090] In some preferred embodiments, the principal component analysis method is used to reduce the dimension of the frequency-domain signal, and the information content of the feature quantity of the frequency-domain signal after dimension reduction is more than 98% of the information content of the feature quantity of the frequency-domain signal before dimension reduction.

[0091] Example 4:

[0092] Based on the above embodiments, the bronze ware rust sample includes a standard rust sample, and the preparation of the standard rust sample includes the following steps: thoroughly mix rust powder with known composition and polyethylene powder and then press them into tablets to obtain a rust standard sample, and the mass ratio of the rust powder to the polyethylene powder is 1:2 to 1:3.

[0093] In some preferred embodiments, the rust powder includes a first rust powder and a second rust powder. The first rust powder and the second rust powder are mixed according to an initial ratio and pressed into tablets with polyethylene for sample preparation. On the premise that the total mass of the rust powder remains unchanged, on the basis of the initial ratio, the first rust powder is increased or decreased by a preset weight ratio to obtain other mixing ratios and pressed into tablets with polyethylene for sample preparation. This embodiment can enable the trained classification model to more accurately identify the situation of boundary mixing, and to a certain extent, can roughly judge the mass ratio of various rust substances at a certain point on the bronze cultural relic based on the frequency domain signal.

[0094] Example 5:

[0095] Based on the above embodiments, as Figure 2 shown in the bronze ware rust component analysis method, includes the following steps:

[0096] Collect the time-domain signal of the rust on the surface of the bronze ware cultural relic, and convert the time-domain signal into the frequency-domain signal of the cultural relic;

[0097] Input the frequency-domain signal of the cultural relic into the trained classification model constructed by any of the above classification model construction methods, and the trained classification model outputs the rust component information corresponding to the frequency-domain signal of the cultural relic.

[0098] In some preferred embodiments, the obtained frequency-domain signal of the cultural relic can also be preprocessed by data processing methods. For example, in one or more embodiments, a preprocessing function is used to process the frequency-domain signal of the cultural relic; in one or more embodiments, the background frequency-domain signal in the frequency-domain signal of the cultural relic is removed; in one or more embodiments, the principal component analysis method is used to perform dimensionality reduction processing on the data.

[0099] Example 6:

[0100] The bronze ware rust component analysis system includes:

[0101] A terahertz time-domain spectrometer, which is used to collect the time-domain signal of the rust on the surface of the bronze ware cultural relic and convert the time-domain signal into the frequency-domain signal of the cultural relic;

[0102] An analysis unit for analyzing the rust component information corresponding to the cultural relic frequency domain signal by using the trained classification model constructed by the classification model construction method in any of the foregoing embodiments.

[0103] In some embodiments, it further includes a robotic arm for moving the probe head of the terahertz time-domain spectrometer to achieve fixed-angle scanning of the surface of the bronze ware and obtain the signals of the rust layer on the entire surface of the bronze ware.

[0104] In some embodiments, the terahertz time-domain spectrometer mainly includes components such as a femtosecond laser, a lock-in amplifier, an electro-optic crystal, a photoconductive antenna, and an optical module. The electro-optic crystal emits femtosecond-level terahertz pulse waves under the excitation of the femtosecond laser. Through the focusing and guiding of the optical module, the object to be measured is detected in the form of transmission or reflection. The detected signals are converged onto the photoconductive antenna for sampling, thereby obtaining the terahertz signal with sample information.

[0105] Embodiment 7:

[0106] 350 total terahertz samples of bronze ware rust products are obtained. According to the training set: test set = 4:1, the types of bronze rust include atacamite, cuprous chloride, lead carbonate, copper oxide, tin oxide, stannous oxide, and basic copper phosphate, seven kinds in total.

[0107] Set the hyperparameters penalty factor C = 12.25 and gamma = 1, use the Gaussian function as the kernel function, and use the logarithmic loss function as the loss function to train the SVM model to obtain SVM model 1.

[0108] Set the hyperparameters penalty factor C = 12.25 and gamma = 1, and use the second kernel function as the kernel function The loss function is loss = max(0, exp(-y i (w T x i +b))), train the SVM model to obtain SVM model 2, and before inputting the data, use the preprocessing function to process the data.

[0109] The test results are as shown in Figure 3 (a) and (b). The accuracy rate of SVM model 1 for identifying bronze ware rust products is 70.00%, while the accuracy rate of SVM model 2 for identifying bronze ware rust products is as high as 94.28%.

[0110] In the present invention, the terms "first", "second", etc. (such as the first kernel function, the second kernel function, etc.) are only used to distinguish the corresponding components for the sake of clarity, and are not intended to limit any order or emphasize importance, etc. In addition, the term "connection" used in the present invention, without special explanation, can be directly connected or indirectly connected via other components.

[0111] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for constructing a classification model, characterized in that, It includes the following steps: Collect the time-domain signal of the bronze corrosion sample, and convert the time-domain signal into a frequency-domain signal; Train the support vector machine classification model based on the frequency-domain signal to obtain the trained classification model; Among them, the kernel function of the support vector machine classification model is: or or where x (i) is the value corresponding to the i-th frequency point in the sample points; x (i) is the value corresponding to the j-th frequency point in the sample points; σ is the bandwidth of the Gaussian kernel; d is the degree of the polynomial; The loss function of the support vector machine classification model is: loss = lg(1 + exp(-y i (w T x i + b))) ; or loss = max(0, 1 + y i (w T x i + b)); or loss=max(0,exp(-y i (w T x i +b))) where x i is the value corresponding to the i-th sample point and is a column matrix; y i is the label value corresponding to the i-th sample point; w is the normal vector of the hyperplane; b is the displacement term.

2. The classification model construction method according to claim 1, wherein It further includes the following steps. Before the frequency-domain signal is input into the support vector machine classification model, a preprocessing function is used to convert the frequency-domain signal, and the preprocessing function is: Where x is the frequency-domain signal before conversion, X is the frequency-domain signal after conversion, and x (i) is the value corresponding to the i-th frequency point among the sample points; x min is the minimum value corresponding to all frequency points in the sample point; x max is the maximum value corresponding to all frequency points in the sample point; n is the number of frequency points in the sample point.

3. The classification model construction method according to claim 1, wherein Collect the time-domain signal of the bronze corrosion sample and convert it into the first frequency-domain signal, collect the background signal and convert it into the background frequency-domain signal, and eliminate the background signal in the first frequency-domain signal to obtain the frequency-domain signal.

4. The classification model construction method according to claim 3, wherein Use the principal component analysis method to reduce the dimension of the frequency-domain signal, and the information volume of the feature quantity of the frequency-domain signal after dimension reduction is more than 98% of the information volume of the feature quantity of the frequency-domain signal before dimension reduction.

5. The classification model construction method according to claim 1, characterized in that The bronze corrosion sample includes a standard corrosion sample, and the preparation of the standard corrosion sample includes the following steps: fully mix the corrosion powder with known composition and polyethylene powder and then press it into a tablet to obtain a corrosion standard sample, and the mass ratio of the corrosion powder to the polyethylene powder is 1:2 to 1:

3.

6. The classification model construction method according to claim 5, wherein The corrosion powder includes the first corrosion powder and the second corrosion powder. The first corrosion powder and the second corrosion powder are mixed according to the initial ratio and pressed into a sample with polyethylene. On the premise that the total mass of the corrosion powder remains unchanged, the first corrosion powder is increased or decreased by a preset weight ratio on the basis of the initial ratio to obtain the remaining mixing ratios and then pressed into a sample with polyethylene.

7. The classification model construction method according to any one of claims 1 to 6, characterized in that The value of the penalty factor of the support vector machine classification model is 12.25, and the value of gamma is 1.

8. Method for analyzing components of bronze corrosion, characterized in that, It includes the following steps: Collect the cultural relic time-domain signal of the corrosion on the surface of the bronze cultural relic, and convert the time-domain signal into a cultural relic frequency-domain signal; Input the cultural relic frequency-domain signal into the trained classification model constructed by the classification model construction method described in any one of claims 1 to 7, and the trained classification model outputs the corrosion component information corresponding to the cultural relic frequency-domain signal.

9. A bronze corrosion component analysis system, characterized in that, It includes: A terahertz time-domain spectrometer for collecting the cultural relic time-domain signal of the corrosion on the surface of the bronze cultural relic and converting the time-domain signal into a cultural relic frequency-domain signal; An analysis unit for analyzing the cultural relic frequency-domain signal by using the trained classification model constructed by the classification model construction method described in any one of claims 1 to 7 to obtain the corrosion component information corresponding to the cultural relic frequency-domain signal.

Citation Information

Patent Citations

  • A performance fault detection method and device based on a support vector machine

    CN109902731A

  • Internal defect imaging method based on terahertz waves, electronic equipment and storage medium

    CN114689598A