Automatic extraction method of time-varying structure modal frequency based on image segmentation

CN121236385BActive Publication Date: 2026-08-21NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511386004.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-08-21
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

[0005]然而,现有模态识别方法仍存在明显不足

Benefits of technology

[0033] 1. The automatic extraction method for time-varying structural modal frequencies based on image segmentation disclosed in this invention transforms the modal frequency identification problem into a time-spectrum image segmentation problem by introducing an image segmentation neural network. The trained model can automatically and effectively segment the modal ridge region in the time-spectrum image, reducing manual intervention and enabling automated identification and tracking of structural modal frequencies under complex working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121236385B_ABST
    Figure CN121236385B_ABST
Patent Text Reader

Abstract

The application discloses a time-varying structure modal frequency automatic extraction method based on image segmentation, and particularly relates to the field of time-varying structure modal frequency extraction. A windowed Fourier transform of a localized window function is introduced to vibration response data containing noise, so as to obtain a noise-containing time-frequency spectrum image containing frequency component distribution of a time-varying structure response signal at different times and corresponding amplitude characteristics. The position of a modal ridge line is determined for a preprocessed time-frequency spectrum image, a labeled region is generated, and then a labeled data set is obtained and used as a data set of a model for training, so as to realize automatic segmentation of the ridge line region in the time-frequency spectrum image. A short-time Fourier transform is performed on a preprocessed vibration response signal under excitation, so as to obtain a real time-varying structure time-frequency spectrum. The preprocessed real time-varying structure time-frequency spectrum is input into the model, and then the model identifies the ridge line region of the real time-varying structure modal frequency and outputs a segmentation result. The segmentation result is post-processed, and finally the time-varying structure modal frequency is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of time-varying structural modal frequency extraction, and more specifically to an automatic method for extracting time-varying structural modal frequencies based on image segmentation. Background Technology

[0002] With the rapid development of aerospace technology, aircraft are gradually evolving towards hypersonic speeds, longer ranges, and multi-mission capabilities. Especially driven by next-generation aerospace warfare and deep space exploration missions, aircraft not only need higher speeds and maneuverability but also need to consider lightweight design and mission adaptability. This trend makes the structural dynamics of aircraft increasingly complex, significantly enhancing the coupling effects between aerodynamics, thermal, dynamics, and structure. Accurately understanding the dynamic characteristics of aircraft in complex service environments has become a core scientific issue in aircraft design and operational support.

[0003] During flight, aircraft exhibit typical time-varying characteristics of structural modal frequencies, a phenomenon particularly prominent in new aerospace vehicles such as hypersonic glide missiles and single-stage-to-orbit combined propulsion vehicles. The main reasons for this include: firstly, the widespread adoption of blended wing-body configurations and lightweight structures in these vehicles significantly enhances the aerodynamic-dynamic-elastic coupling effect, placing higher demands on the design of guidance and control systems; secondly, the complex mission profiles of these vehicles, requiring them to endure extreme service environments such as ultra-high speeds, unusual temperatures, and severe weather over extended periods, cause their structural modal frequencies to change dynamically over time, directly impacting flight control accuracy and flight safety.

[0004] Accurate acquisition and tracking of modal frequencies are crucial in aircraft design and operation. First, modal frequencies are fundamental parameters for structural dynamic design, determining the overall response characteristics of an aircraft under complex loads. Second, they are essential for load analysis, accurately assessing aerodynamic, thermodynamic, and inertial coupling effects. Third, in attitude control system design, the stability and effectiveness of control laws are closely related to modal frequencies; erroneous frequency estimation can lead to the excitation of structural modes, thus affecting flight stability. Finally, the evolution of modal frequencies provides a reliable basis for structural health monitoring and is a vital indicator for identifying structural damage and degradation. Therefore, studying the time-varying characteristics of modal frequencies is not only a key aspect of lightweight and high-performance aircraft design but also a necessary condition for ensuring the safe execution of missions.

[0005] However, existing modal identification methods still have significant shortcomings. Model-based methods typically rely on high-precision dynamic modeling, which is not only computationally complex but also computationally inefficient, making it difficult to meet the real-time requirements of flight missions. In complex environments with large-amplitude noise, these methods are susceptible to interference and loss of accuracy, especially under hypersonic flight conditions where nonlinear effects are significant, further reducing their applicability. These deficiencies limit the efficient and accurate acquisition of modal frequency time-varying characteristics, hindering the application of aircraft in real-world mission environments. Summary of the Invention

[0006] To address this, this invention discloses an automatic extraction method for time-varying structural modal frequencies based on image segmentation. This method obtains a time-spectrum image by performing a windowed Fourier transform on the vibration signal and transforms the frequency extraction problem into an image segmentation problem involving ridge recognition of the time-spectrum image, thereby achieving automatic tracking of modal frequencies. This invention can operate stably under complex and non-white noise load environments, effectively avoiding the generation of spurious modes, exhibiting high accuracy and strong robustness, significantly improving computational efficiency, and allowing for the evaluation of the reliability of the recognition results, thus possessing strong engineering applicability.

[0007] To achieve the above objectives, the present invention provides the following technical solution: an automatic extraction method for time-varying structural mode frequencies based on image segmentation, comprising:

[0008] A time-varying structural response signal is generated by numerical calculation and a database is constructed. A windowed Fourier transform with a localized window function is introduced into the noisy vibration response data to obtain a noisy time-spectrum image containing the frequency component distribution and corresponding amplitude characteristics of the time-varying structural response signal at different times.

[0009] The modal ridge positions are determined by interpolation based on system theory in the preprocessed time-spectrum image, and labeled regions are generated to obtain a labeled dataset. The labeled dataset is used as the training dataset for the image segmentation neural network model to achieve automatic segmentation of ridge regions in the time-spectrum image.

[0010] A random excitation is applied to the real time-varying structural signal, and the measured vibration response signal is then acquired and preprocessed until the vibration response signal meets the time-frequency resolution requirements of the short-time Fourier transform.

[0011] The vibration response signal under preprocessing excitation is subjected to short-time Fourier transform to obtain the real time-varying structure time spectrum. The preprocessed real time-varying structure time spectrum (including size normalization, coordinate axis deletion and amplitude normalization) is input into the trained image segmentation neural network model. The image segmentation neural network model then segments the real time-varying structure time spectrum, identifies the ridge region of the real time-varying structure modal frequency and outputs the segmentation result. The segmentation result is post-processed using the structural modal frequency extraction and tracking method to finally obtain the time-varying structure modal frequency.

[0012] Preferably, the time-varying structural response signal database is constructed as follows:

[0013] A finite element model of a four-degree-of-freedom time-varying spring-mass-damped system is established. While keeping the stiffness and damping parameters constant, the mass parameters are treated to decay over time. The mass size of each mass block changes with time as follows:

[0014] m i (t)=m i0 -q i t,i=1,2,3,4;

[0015] Where: m i0 q represents the initial mass of the i-th mass block. i Here, we define the rate of mass decay over time, ensuring that the mass of any mass block remains greater than 0 at the final moment.

[0016] Subsequently, random excitations were applied to each node, and the vibration response signal data x(t) of each node was solved using the Newmark-β method. Multiple sets of different rates of change were generated by Latin hypercube sampling, and time-varying structural response signal databases were constructed by solving them sequentially.

[0017] Preferably, the process of generating a noisy spectral image is as follows:

[0018] The windowed Fourier transform (WFT) of the response signal x(t) is expressed as:

[0019]

[0020] Where: G WFT (ω,t) represents the WFT transform result, τ represents the position of the window function on the time axis, and X(ω) and W(ω) are the Fourier transforms of the signal x(t) and the window function w(t), respectively. To eliminate the interference of the negative frequency part, equation (1) uses the positive frequency part x of the original signal. + (t) is used for calculation, x + (t) the analytic signal x(t) passing through the original signal x(t) a (t) is obtained, i.e., x + (t)=xa (t) / 2.

[0021] Preferably, the labeled dataset is constructed as follows:

[0022] Irrelevant elements, including coordinate axes, legends, and color bars, were removed from the noisy time-spectrum image, and the image size was uniformly adjusted to 256×256 pixels. Then, the theoretical modal frequencies of the system in different groups were mapped to the pixel coordinate system of the time-spectrum image through interpolation to determine the pixel position of the system modal ridge in the time-spectrum image. After that, by setting the ridge width, the upper and lower boundaries of the ridge were generated on the time axis of the time-spectrum image at fixed time intervals to form closed region markers. Finally, the standardized time-spectrum image and the corresponding ridge marker points were stored as paired text files to form a labeled dataset.

[0023] Preferably, the image segmentation neural network model includes a symmetrical encoder, decoder, and skip connections;

[0024] The encoder consists of an initial double convolutional module and four downsampling modules, comprising five feature extraction stages. The input is a 256×256×1 image, which, after passing through the initial double convolutional module, yields a 128×128×64 feature map. Subsequently, it undergoes four downsampling passes, with the feature map size decreasing sequentially to 64×64×128, 32×32×256, 16×16×512, and 8×8×1024, gradually extracting high-level semantic features.

[0025] The decoder is symmetrical to the encoder and consists of four upsampling modules and one output convolution module. The feature maps are progressively upsampled starting from the lowest layer of the encoder, and their sizes are restored to 16×16×512, 32×32×256, 64×64×128 and 128×128×64 respectively. They are then spliced ​​and fused with the features of the corresponding layers in the encoder through skip connections. After multi-level feature reconstruction, the final output layer is restored to 256×256×1, and pixel-level prediction results are achieved through 1×1 convolution.

[0026] In practical engineering applications, random excitations are applied to the real structure to be studied. Random excitations include force excitations, impact excitations, or environmental excitations. The modal frequency response signals of the real structure under the excitation are collected by the acquisition device and stored for subsequent time-frequency analysis and feature extraction. The acquisition device includes an accelerometer, a displacement sensor, and a strain sensor.

[0027] The preprocessing of vibration response signals under excitation includes filtering and resampling operations; filtering is used to remove high-frequency noise or low-frequency interference, and resampling is used to adjust the sampling rate so that the signal meets the time-frequency resolution requirements of short-time Fourier transform, thereby improving the accuracy of subsequent analysis.

[0028] The preferred method for structural modal frequency extraction and tracking is as follows:

[0029] The segmented image is binarized, and then the bone line extraction algorithm is used to extract the bone lines in the target region to obtain the frequency points at discrete time points; then the discrete points are connected to form local ridge lines.

[0030] For local ridge line discontinuities, the system determines whether to delete or connect them by judging whether the spacing between discontinuity points is within a set threshold.

[0031] For local ridge line overlap, the merging of local ridge lines is completed by judging whether the number of overlapping points and the spacing meet the threshold requirements;

[0032] Then, isolated points and short ridges are removed to obtain continuous ridges, and the time-varying modal frequencies of the structure are extracted.

[0033] 1. The automatic extraction method for time-varying structural modal frequencies based on image segmentation disclosed in this invention transforms the modal frequency identification problem into a time-spectrum image segmentation problem by introducing an image segmentation neural network. The trained model can automatically and effectively segment the modal ridge region in the time-spectrum image, reducing manual intervention and enabling automated identification and tracking of structural modal frequencies under complex working conditions.

[0034] 2. The automatic extraction method of time-varying structural modal frequencies based on image segmentation disclosed in this invention obtains discrete frequency points by performing image binarization and bone line extraction operations on the segmentation results; after connecting the frequency points, the discontinuity and overlap of local ridges are processed, isolated value points and short ridges are deleted to obtain continuous ridges, and finally the continuous modal frequencies of the whole time period are obtained, which effectively avoids false modal interference and improves the accuracy and continuity of the identification results.

[0035] 3. The automatic extraction method for time-varying structural modal frequencies based on image segmentation disclosed in this invention has good adaptability to high noise and large amplitude disturbance environments, and can work stably under non-white noise loads and significant nonlinear effects, ensuring the reliability of modal frequency identification.

[0036] 4. The automatic extraction method of time-varying structural modal frequencies based on image segmentation disclosed in this invention, combined with the identification of vibration response data under real working conditions, proves, through comparison with traditional calculation methods, that this method not only has high identification accuracy but also significantly improves computational efficiency. It can quickly complete data processing and analysis under limited computing resources and is more suitable for practical engineering. Attached Figure Description

[0037] Figure 1 A flowchart of the automatic extraction method for time-varying structural modal frequencies based on image segmentation provided by the present invention;

[0038] Figure 2A schematic diagram of the finite element model for creating the dataset provided by this invention;

[0039] Figure 3 This is a schematic diagram of the dataset annotation provided by the present invention;

[0040] Figure 4 The structure diagram of the neural network model for automatic extraction of structural modal frequencies provided by this invention;

[0041] Figure 5 The test set evaluation metrics and loss iteration curves provided by this invention;

[0042] Figure 6 The present invention provides a time-varying mass finite element model for a launch vehicle.

[0043] Figure 7 The rocket model provided by this invention has a time spectrum image;

[0044] Figure 8 The result of segmenting the time-spectrum image provided by the present invention using a neural network model;

[0045] Figure 9 The flowchart of the structural frequency extraction and tracking method provided by the present invention is shown below;

[0046] Figure 10 A comparison diagram of the segmentation results provided by this invention and the actual results. Detailed Implementation

[0047] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] like Figure 1 As shown, this embodiment proposes an automatic extraction method for time-varying structural mode frequencies based on image segmentation, including:

[0049] Step 1: Generate a time-varying structural response signal database using numerical calculations.

[0050] like Figure 2 As shown, a finite element model of a four-degree-of-freedom time-varying spring-mass-damped system is established. While keeping the stiffness and damping parameters constant, mass decay is considered to simulate the continuous fuel consumption process of an aircraft during flight. The mass size of each mass block changes with time as follows:

[0051] m i(t)=m i0 -q i t, i=1,2,3,4 (2)

[0052] Where: m i0 q represents the initial mass of the i-th mass block. i Here, we define the rate of mass decay over time, ensuring that the mass of any mass block remains greater than 0 at the final moment.

[0053] Within a certain range, 500 sets of mass decay rate coefficients are generated using the Latin hypercube sampling method. Gaussian noise excitation is applied to the model, and the vibration response of the four-degree-of-freedom system is calculated using the Newmark-β method, thus obtaining 500 sets of data samples.

[0054] Step 2: Generate a noisy time-spectrum image

[0055] Gaussian noise with a signal-to-noise ratio of 20dB is added to the vibration response dataset obtained in step 1 to simulate the response measurement noise. Then, a windowed Fourier transform is performed on the noise-added vibration response data to obtain the required noisy time-spectrum image dataset.

[0056] Step 3: Creating the labeled dataset

[0057] The coordinate axes, legend, and color bars of the time-spectrum image were removed, and the image size was uniformly adjusted to 256×256 pixels. Then, interpolation was used to map the theoretical modal frequencies of the system in different groups to the pixel coordinate system of the time-spectrum, determining the pixel positions of the system modal ridges in the time-spectrum image. Next, by pre-setting the ridge width (chosen as 10 pixels in this invention), the upper and lower boundaries of the ridges were generated at fixed time intervals on the time axis of the time-spectrum, forming closed region markers. Finally, the standardized time-spectrum image and the corresponding ridge marker points were stored as a paired document, forming an annotated dataset. A schematic diagram of the time-spectrum image annotation is shown below. Figure 3 As shown.

[0058] Step 4: Train the neural network image segmentation model

[0059] The neural network structure described in this invention is as follows: Figure 4 As shown, it consists of a symmetrical encoder, decoder, and skip connections. The encoder comprises an initial double convolutional module and four downsampling modules, containing a total of five feature extraction stages. The input is a 256×256×1 image, which, after passing through the initial double convolutional module, yields a 128×128×64 feature map; subsequently, it undergoes four downsampling passes, with the feature map size decreasing sequentially to 64×64×128, 32×32×256, 16×16×512, and 8×8×1024, progressively extracting high-level semantic features.

[0060] The decoder is symmetrical to the encoder and consists of four upsampling modules and one output convolutional module. Feature maps are progressively upsampled starting from the lowest layer of the encoder, with sizes successively restored to 16×16×512, 32×32×256, 64×64×128, and 128×128×64. These features are then concatenated and fused with features from corresponding layers in the encoder via skip connections. After multi-level feature reconstruction, the final output layer is restored to 256×256×1, achieving pixel-level prediction results through 1×1 convolutions.

[0061] The ratio of the training dataset to the validation dataset was set to 8:2, meaning 400 images were used for training and 100 images were used for validation. Then, amplitude normalization was performed on the dataset images to uniformly map the pixel values ​​to the range [0,1].

[0062] To improve the model's generalization ability, random scaling, random horizontal flipping, random vertical flipping, and random cropping data augmentation techniques are used to expand the dataset, effectively reducing the risk of overfitting and improving the model's prediction accuracy and generalization ability on unknown test data.

[0063] The nonlinear function ReLU is used as the activation function of the hidden layer. ReLU preserves the gradient in the non-negative interval, making the gradient update more stable during backpropagation, thus ensuring convergence speed and feature representation ability. Its expression is:

[0064]

[0065] The Dice Loss function is chosen as the loss function to measure the similarity between the image predicted by the image segmentation neural network model and the real image. The higher the overlap between the prediction and the real result, the smaller the Dice Loss value. The calculation formula is as follows:

[0066]

[0067] In the formula: Y pred Y represents the set of pixels representing the target region predicted by the model. ture This represents the set of pixels in the actual target area.

[0068] To objectively evaluate the model's performance in the temporal spectral ridge segmentation task, key metrics output from the training and validation processes are used to measure model performance. These metrics are as follows:

[0069] Global Correct (GC): Represents the proportion of pixels correctly classified by the model as a whole, defined as:

[0070]

[0071] Where: n cjThis represents the number of pixels with true class c and predicted class j. The numerator is the total number of pixels correctly classified, and the denominator is the total number of pixels.

[0072] Mean Intersection over Union (MIoU): Represents the average intersection-over-union ratio between the predicted and actual regions for each category.

[0073]

[0074] In the formula: the numerator is the number of pixels where the prediction and the actual values ​​of category c intersect, and the denominator is the number of pixels where they intersect.

[0075] The neural network model training environment was built on the PyTorch deep learning framework under Linux, with an Intel i7-13700K CPU and an NVIDIA RTX 4070 12G graphics card. The total number of training epochs was set to 200, the batch size to 8, and a stochastic gradient descent (SGD) optimizer was used. The initial learning rate was set to 0.01 to balance the initial convergence speed with later stability, the momentum was set to 0.9 to accelerate the consistency of gradient update directions, and the weight decay coefficient was set to 10⁻⁴. L2 regularization was used to suppress overfitting and ensure the model's generalization ability. To further optimize the training process, a learning rate scheduling strategy with warmup was introduced to dynamically adjust the learning rate.

[0076] Evaluation indicators during training, such as Figure 5 As shown, the training loss decreased rapidly after about 20 epochs and stabilized after 50 epochs. At the same time, the evaluation metrics remained stable at a high level, indicating that the model converged quickly in the early stage of training and maintained good performance in subsequent training. The overall training effect was quite ideal, indicating that the model can effectively learn the features in the data and has good performance.

[0077] Step 5: Acquire real time-varying structure signals

[0078] To verify the method of the present invention, a system was established as follows: Figure 6 The finite element model of the time-varying structure of the launch vehicle shown is used to simulate the rocket structure as its mass gradually decreases due to fuel consumption during flight.

[0079] The main structure of the launch vehicle is modeled using Bernoulli-Euler beam elements with a circular cross-section. The material properties are shown in Table 1. During modeling, the payload and fuel are equivalent to lumped masses attached to the nodes. The spatial coordinates of each node and the corresponding total lumped mass parameters are shown in Table 2.

[0080] Table 1 Material properties of the rocket model

[0081]

[0082] Table 2. Nodes and Lumped Mass of the Rocket Model

[0083]

[0084] To simulate the characteristic of rocket fuel depletion over time during flight, the concentrated mass of the simulated fuel was set to decrease linearly at a rate of 100 kg / s, meaning the fuel would be exhausted in 72 seconds. Uncorrelated random excitations were applied to all nodes of the model, and its lateral vibration response was calculated using the Newmark-β method.

[0085] Step 6: Filter and resample the real signal.

[0086] The vibration response signal was resampled at a frequency of 128 Hz. Since the rocket model has free-to-free boundary conditions, the response signal contains zero-frequency rigid body modes. To eliminate the influence of rigid body modes, the calculated vibration response was filtered using a Butterworth high-pass filter with a cutoff frequency of 1 Hz. Subsequently, Gaussian white noise with a signal-to-noise ratio of 40 dB was applied to the filtered response signal to simulate sensor measurement noise.

[0087] Step 7: Perform a windowed Fourier transform on the processed real signal to obtain the time spectrum.

[0088] The vibration response signal of the launch vehicle model, after filtering and resampling preprocessing, was subjected to a windowed Fourier transform to obtain the corresponding time-frequency spectrum. The time-frequency spectrum of the signal after the windowed Fourier transform is shown below. Figure 7 As shown, the overall trend of frequency component changes over time can be observed.

[0089] Step 8: Input the real-time spectrogram into the neural network model for identification.

[0090] The time-spectrum image obtained in step 7, after necessary preprocessing (including size normalization and amplitude normalization), is input into the trained neural network model for further processing. The neural network automatically identifies the ridge regions corresponding to the modal frequencies in the time-spectrum image and outputs the segmentation results. The results show that the neural network model can effectively extract the ridges corresponding to the first four bending modes of the launch vehicle model, but the extracted ridges still have a certain width and exhibit breaks in local areas. Therefore, post-processing is required to obtain continuous modal ridges.

[0091] Step 9: Frequency Extraction and Post-processing

[0092] like Figure 8 As shown, the proposed structural modal frequency extraction and tracking method is used for post-processing of the model segmentation results. Figure 9 As shown, firstly, the segmented image is binarized, and then a bone line extraction algorithm is used to extract bone lines within the target region, obtaining frequency points at discrete times. Then, the discrete points are connected into local ridges using an algorithm: for discontinuous local ridges, the algorithm determines whether to delete or connect them based on whether the distance between discontinuous points is within a set threshold; for overlapping local ridges, the algorithm merges them based on whether the number and spacing of overlapping points meet the threshold requirements. Subsequently, isolated points and short ridges are deleted, ultimately obtaining continuous ridges, and thus the time-varying modal frequencies of the structure are extracted.

[0093] Finally, a comparison of the results of this invention, the results of conventional methods, and the benchmark values ​​is provided. Figure 10 As shown. The reference value is calculated by introducing the proportional damping matrix at the initial moment, and the modal frequency reference value at each moment is obtained.

[0094] The results show that the method of this invention can effectively identify the first four modal frequencies of a launch vehicle model. These frequencies are located within the range of (0, 60] Hz and gradually increase over time. The identification results are highly consistent with the benchmark values, accurately tracking modal frequency changes caused by fuel consumption during flight. In contrast, traditional methods can only effectively identify the first two modal frequencies of the launch vehicle model; from the third frequency onwards, the identification results show obvious overlap and discontinuity, with the high-frequency identification results being very chaotic. Therefore, compared with existing technologies, the method disclosed in this invention has a significant advantage in identification accuracy.

[0095] Meanwhile, in terms of computational efficiency, the method of the present invention takes 0.62s to complete one frequency extraction, which is significantly better than the 1.22s of the traditional method, indicating that the present invention has higher computational efficiency while ensuring accuracy.

[0096] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. An automatic extraction method for time-varying structural mode frequencies based on image segmentation, characterized in that: include: A time-varying structural response signal is generated by numerical calculation and a database is constructed. A windowed Fourier transform with a localized window function is introduced into the noisy vibration response data to obtain a noisy time-spectrum image containing the frequency component distribution and corresponding amplitude characteristics of the time-varying structural response signal at different times. The modal ridge positions are determined by interpolation based on system theory in the preprocessed time-spectrum image, and labeled regions are generated to obtain the labeled dataset. Specifically: By interpolating the theoretical modal frequencies of the system in different groups to the pixel coordinate system of the time spectrum, the pixel positions of the system modal ridges in the time spectrum image are determined. Then, by presetting the ridge width, the upper and lower boundaries of the ridges are generated on the time axis of the time spectrum image at fixed time intervals to form closed region markers. Finally, the standardized time spectrum image and the corresponding ridge marker points are stored as paired text files to form a labeled dataset. The labeled dataset is used as the dataset for training the image segmentation neural network model to achieve automatic segmentation of ridge regions in the time spectrum image. A random excitation is applied to the real time-varying structure, and the measured vibration response signal is then acquired and preprocessed until the vibration response signal meets the time-frequency resolution requirement of the short-time Fourier transform. The vibration response signal under preprocessed excitation is subjected to short-time Fourier transform to obtain the real time-varying structure time spectrum. The preprocessed real time-varying structure time spectrum is input into the trained image segmentation neural network model. The image segmentation neural network model then segments the real time-varying structure time spectrum, identifies the ridge region of the real time-varying structure modal frequency and outputs the segmentation result. The segmentation result is post-processed using the structural modal frequency extraction and tracking method to finally obtain the time-varying structure modal frequency.

2. The automatic extraction method for time-varying structural mode frequencies based on image segmentation according to claim 1, characterized in that: The time-varying structural response signal database is constructed as follows: A finite element model of a four-degree-of-freedom time-varying spring-mass-damped system was established. While keeping the stiffness and damping parameters constant, the mass parameter was subjected to time-varying degradation. Then, random excitation was applied to each node, and the Newmark-β method was used to solve for the vibration response signal data of each node. Multiple sets of different rates of change are generated by Latin hypercube sampling, and time-varying structural response signal databases are constructed by solving them sequentially.

3. The automatic extraction method for time-varying structural mode frequencies based on image segmentation according to claim 1, characterized in that: The process of generating a noisy spectral image is as follows: Response signal The windowed Fourier transform (WFT) is represented as: In the formula: This represents the result of the WFT transformation. This represents the position of the window function on the time axis. and Signals Window function The Fourier transform; to eliminate interference from the negative frequency part, equation... Use the positive frequency portion of the original signal Perform calculations. Through the original signal Analyzed signal Obtain, that is .

4. The automatic extraction method for time-varying structural mode frequencies based on image segmentation according to claim 1, characterized in that: Image segmentation neural network models include symmetric encoders, decoders, and skip connections; The encoder consists of an initial double convolutional module and four downsampling modules, comprising five feature extraction stages. The input is a 256×256×1 image, which, after passing through the initial double convolutional module, yields a 128×128×64 feature map. Subsequently, it undergoes four downsampling passes, with the feature map size decreasing sequentially to 64×64×128, 32×32×256, 16×16×512, and 8×8×1024, gradually extracting high-level semantic features. The decoder is symmetrical to the encoder and consists of four upsampling modules and one output convolution module. The feature maps are progressively upsampled starting from the lowest layer of the encoder, and their sizes are restored to 16×16×512, 32×32×256, 64×64×128 and 128×128×64 respectively. They are then spliced ​​and fused with the features of the corresponding layers in the encoder through skip connections. After multi-level feature reconstruction, the final output layer is restored to 256×256×1, and pixel-level prediction results are achieved through 1×1 convolution.

5. The automatic extraction method for time-varying structural mode frequencies based on image segmentation according to claim 1, characterized in that: Random excitation includes force excitation, impact excitation, or environmental excitation. The modal frequency response signal of the real structure under excitation is acquired by a data acquisition device and stored for subsequent time-frequency analysis and feature extraction. The data acquisition device includes an accelerometer, a displacement sensor, and a strain sensor.

6. The automatic extraction method for time-varying structural mode frequencies based on image segmentation according to claim 1, characterized in that: The method for extracting and tracking structural modal frequencies is as follows: The segmented image is binarized, and then the bone line extraction algorithm is used to extract the bone lines in the target region to obtain the frequency points at discrete time points; then the discrete points are connected to form local ridge lines. For local ridge line discontinuities, the system determines whether to delete or connect them by judging whether the spacing between discontinuity points is within a set threshold. For local ridge line overlap, the merging of local ridge lines is completed by judging whether the number of overlapping points and the spacing meet the threshold requirements; Then, isolated points and short ridges are removed to obtain continuous ridges, and the time-varying modal frequencies of the structure are extracted.