Surface wave multi-modal dispersion curve probability inversion method based on modal self-matching
By using a modal self-matching method, Rayleigh wave multimodal dispersion data are separated, Bayesian equations are constructed, and an likelihoodless inference method is adopted to solve the problems of modal misidentification and computational time consumption, thus realizing efficient shear wave velocity profile inversion and uncertainty characterization.
Patent Information
- Application Number
- CN202510859574.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, Rayleigh wave multimodal dispersion curve inversion suffers from modal misidentification and computational time consumption, leading to uncertainty and inefficiency in the inversion results.
A modality self-matching method is adopted to separate apparent dispersion data points into base mode and higher mode datasets. A Bayesian equation is constructed and an likelihoodless inference method is used to quantify data similarity through distance measurement. The posterior distribution is solved using stochastic simulation methods such as Markov chain Monte Carlo, and the modality number is automatically matched.
It effectively avoids the problem of mode misidentification, improves the recognizability and computational efficiency of shear wave velocity profiles, and reasonably characterizes the uncertainty of inversion results.
Smart Images

Figure CN120911244A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of surface wave dispersion inversion technology, and particularly relates to a surface wave multi-modal dispersion curve probabilistic inversion method based on modal self-matching. BACKGROUND
[0002] Rayleigh wave propagation presents a multi-modal characteristic, that is, Rayleigh wave can propagate at different speeds (phase velocity) at a given frequency. According to the order of phase velocity from small to large, different modes of Rayleigh wave propagation are divided into a basic mode (0th), a first higher mode (1st), a second higher mode (2nd), and the like. A dispersion curve is composed of phase velocity points of different frequencies, and is referred to as a 0th dispersion curve, a 1st dispersion curve, and the like according to the propagation mode.
[0003] Inversion of the dispersion curve can estimate the shear wave velocity (v s ) profile of the near-surface rock-soil body, and is generally based on 0th dispersion curve data, because it is easier to obtain than higher mode dispersion curve data. However, 0th dispersion curve data has certain limitations, that is, the identified v s profile has a limited sensitive depth and low resolution. In order to alleviate this limitation, many studies attempt to jointly invert 0th and higher mode dispersion curve data, and prove that joint inversion of multi-modal dispersion curve data can increase the sensitive depth and longitudinal resolution of the identified v s profile, and thus can reduce the non-uniqueness of the v s profile identification result. Therefore, the present application mainly aims at joint inversion of multi-modal dispersion data.
[0004] Due to measurement errors, limited dispersion data and other factors, the v s profile identification result based on multi-modal dispersion data may have great uncertainty. In the study, the Bayesian inversion method is often used to quantify the uncertainty of the inversion result. In the Bayesian inversion problem of multi-modal dispersion data, it is crucial to evaluate the corresponding likelihood function. Because the apparent dispersion curve (that is, the set of all observed phase velocity points) obtained in practice may have a "modal jump" or "modal stacking" phenomenon of the dispersion curve, it is difficult to separate the data points belonging to different modes, and thus some data points may be incorrectly identified (that is, the modal misidentification problem). Because the evaluation of the likelihood function of multi-modal dispersion data involves forward calculation of the theoretical dispersion curve of different modes of a given v s profile. Therefore, the modal misidentification problem will cause the modal of the theoretical dispersion curve to be incorrectly set, resulting in incorrect calculation of the likelihood function of the multi-modal dispersion curve, and finally leading to errors related to modal misidentification in the inversion result.
[0005] In addition, because the likelihood function of the dispersion curve has high-dimensional nonlinear characteristics, the same is true for multi-modal dispersion curves, so the corresponding Bayesian equation usually does not have an analytical solution. In the literature, numerical solutions to such Bayesian equations are generally calculated based on stochastic simulation methods, such as Markov Chain Monte Carlo (MCMC), reversible-jump Markov Chain Monte Carlo (RJMCMC), Hamiltonian Monte Carlo (HMC), and the like. However, in the solving process, the above-mentioned stochastic simulation methods need to calculate the likelihood function a large number of times, resulting in a very time-consuming solving process.
[0006] In summary, there are two main problems in the current Rayleigh wave multi-modal dispersion curve probability inversion:
[0007] (1) The modal misidentification problem that may exist in the multi-modal data extraction process will lead to incorrect calculation of the likelihood function of the multi-modal dispersion curve;
[0008] (2) The posterior distribution related to the multi-modal dispersion curve may be very time-consuming in the literature using the random simulation method to solve it. SUMMARY
[0009] To solve the problems in the prior art, the present application provides a Rayleigh wave multi-modal dispersion curve probability inversion method based on modal self-matching, which solves the problems mentioned in the background art.
[0010] To achieve the above-mentioned purposes, the present application provides the following technical solutions: a Rayleigh wave multi-modal dispersion curve probability inversion method based on modal self-matching, comprising the following:
[0011] S1. When extracting the Rayleigh wave multi-modal dispersion curve, the apparent dispersion data points are divided into two categories, namely the base modal data set and the higher modal data set;
[0012] S2. Construct the Bayesian equation for inverting the multi-modal dispersion data under the framework of the likelihood-free inference method, i.e., specify the distribution type of the dispersion data and the model parameters, and establish the likelihood function;
[0013] S3. Establish a suitable distance measure to quantify the similarity of the dispersion data and its simulation samples, and specify the method that can be used to sample from the posterior distribution;
[0014] S4. Solve the posterior distribution using the likelihood-free inference method, and represent the uncertainty of the multi-modal dispersion data inversion result based on the approximate posterior samples of the model parameters.
[0015] Preferably, in step S1, modality numbers can be identified either manually by visual interpretation or by using a convolutional neural network (CNN) or Transformer architecture, with the time-frequency features of the dispersion curve as input, to train an automatic classification model for modality numbers. For example, by pre-training on a synthetic multimodal dispersion dataset through transfer learning and then fine-tuning it on actual test data, the classification accuracy in complex modality mixed scenarios can be improved.
[0016] Preferably, in step S1, since there may be modal skipping / stacking problems, it is generally impossible to correctly identify the modality number of all data points. Therefore, this patent proposes a strategy of dividing the data into two parts and processing them with different methods, that is, classifying the identifiable basic modal data into one category (basic modal dataset) and classifying the remaining data points into another category (higher modal dataset).
[0017] Preferably, step S2 specifically includes the following:
[0018] S21. The underground structure of the test site is simplified into a horizontally layered model. This model can be parameterized by the number of layers N and model parameters, including the layer thickness H and shear wave velocity v of each layer. s Compression wave speed v p and density ρ;
[0019] S22, Let v p Given ρ, only H and v of each layer are inverted. s ,use θ Represents all layers H and v s The set of parameters for the shear wave velocity profile;
[0020] S23, Given apparent dispersion curve data d Under the Bayesian framework θ The uncertainty is expressed by its posterior distribution p( θ | d Quantization. In this patent, bolded underlined letters are used to represent vector data;
[0021] S24. Within the framework of the likelihoodless inference method, the solution to p( θ | d It was transformed into a solution θ Simulated samples of apparent dispersion curve data joint posterior distribution The formula is expressed as follows:
[0022]
[0023] Where p( θ )for θ Prior distribution, For a given θ hour The prior distribution, i.e. d Distribution; It is the likelihood function, denoted as indicator function I. S(ε) Where S(ε) represents the set L represents quantization. d and Distance measure of similarity between; when I S(ε) (y) = 1, otherwise I S(ε) (y) = 0; P ε ( d ) is the normalization constant.
[0024] Preferably, in step S2, selecting different forms of prior data distributions allows for the reconstruction of the actual statistical characteristics of the data with varying degrees of complexity. This patented method is applicable to any form of prior data distribution and provides several commonly used options: a multivariate Gaussian distribution considering independent and identically distributed data errors, a Student-T distribution considering the variance of the integral data, and a multivariate Gaussian distribution considering the correlation of data variances, etc.
[0025] Preferably, in step S2, the selection of the prior distribution affects the identifiability of the model parameters. This patent method proposes several informative prior distributions. s The prior distribution can be set as a uniform distribution obtained from the typical range of shear wave velocities of the site material, or it can be set as a uniform distribution plus the prior v. s The Gaussian distribution of the profile (obtained directly from the data). The minimum thickness of H and the maximum identifiable depth of the profile can both be estimated from the data. With these as constraints, the prior distribution of H can be set as a uniform distribution or a Dirichlet distribution.
[0026] Preferably, in step S3, for higher modality datasets, a distance metric L with automatically matching modality numbers is selected; the modality information constraint inversion results of the base modality data are added to the distance metric, and a method based on stochastic simulation is used to solve the problem.
[0027] Preferably, the method based on stochastic simulation includes Markov chain Monte Carlo (MCMC) and subset simulation + MCMC methods.
[0028] Because modal misidentification can lead to incorrect assumptions about the modality of data points, it can cause... d and Inconsistent modal number eventually leads to errors or biases in the inversion results of multimodal dispersion curves (referred to as "modal misidentification error"). In order to avoid this error, the distance measure L of the modal number needs to be selected to automatically match (mainly for higher modal data sets). Such distance measures can be based on the frequency dispersion function values of each data point, or the corresponding theoretical apparent phase velocity. In addition, in order to avoid the fitting of the fundamental modal data points with the theoretical dispersion curve of the higher modal, the modal information of the fundamental modal data can be added to the distance measure to constrain the inversion results. Finally, due to the complexity and high dimensionality of the multimodal dispersion data and the model parameters, Generally, there is no analytical solution. A random simulation-based method can be used, such as Markov Chain Monte Carlo (MCMC), subset simulation + MCMC method, etc.
[0029] The beneficial effects of the present application are: the method of the present application inverts the Rayleigh wave multimodal dispersion curve to obtain the shear wave velocity profile, and the posterior samples can reasonably represent the uncertainty of the identified shear profile; the method uses a likelihood-free inference method to estimate the posterior distribution related to the multimodal dispersion curve, and automatically matches the modal number of the data points, thus avoiding the error inversion results caused by the modal misidentification problem; since there is no need to calculate the likelihood function, compared with the traditional random simulation-based method, the proposed method can more efficiently sample from the posterior distribution; the proposed method reasonably utilizes the multimodal dispersion curve data, and can improve the identifiability of the inversion shear wave velocity profile. BRIEF DESCRIPTION OF DRAWINGS
[0030] Figure 1 A flowchart of a multimodal dispersion curve probability inversion method based on modal self-matching in an embodiment;
[0031] Figure 2 A schematic diagram of the synthetic multimodal (including fundamental and 1st modal) dispersion data of the virtual site in the embodiment;
[0032] Figure 3 A schematic diagram of the parameter posterior samples and probability histogram obtained by the method proposed in the embodiment, and the inversion results based on the fundamental or multimodal data when the modal number of the data points is given;
[0033] Figure 4 A real v s profile and the v s profile obtained by the method proposed in the embodiment;
[0034] Figure 5 A synthetic dispersion curve and the fundamental and 1st modal theoretical dispersion curves calculated from the MAP of the v s profile in the embodiment. DETAILED DESCRIPTION
[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] Example 1
[0037] This embodiment provides a probabilistic inversion method for surface wave multimode dispersion curves based on modal self-matching, such as... Figure 1 As shown, it includes the following steps:
[0038] (1) To address the problem of mode misidentification when extracting Rayleigh wave multimodal dispersion curves, the apparent dispersion data points are divided into two categories: the basic mode dataset and the higher mode dataset.
[0039] Suppose a place has two layers of v s The virtual site profile and site material parameters are shown in Table 1. A seismic wave simulation software based on the finite difference method (i.e., "Tesseral") was used to simulate the propagation process of Rayleigh waves at the site, recording multiple seismic signals. Apparent dispersion curves were extracted from the multiple seismic records using the frequency (f)-wavenumber conversion method. d ),like Figure 2 As shown. In addition, to simulate observation errors, this embodiment adds a Gaussian random error with a standard deviation of 1% of the apparent phase velocity and a mean of 0 to each apparent phase velocity data point. Figure 2 The apparent dispersion curve data consists of dispersion curve data points of the fundamental mode (solid dots) and the 1st mode (circles), and there is a mode "jump" phenomenon between the dispersion curves of the two modes. In practice, this phenomenon can easily lead to mode misidentification, that is, all data points are identified as the fundamental mode by their mode numbers. This invention aims to solve this problem by inversion. Figure 2 The multimodal dispersion curve data shown demonstrates the proposed method.
[0040] Table 1 Material Parameters for Virtual Site
[0041]
[0042]
[0043] (2) Construct Bayes' equation for inverting multimodal dispersed data within the framework of the likelihood-free inference method, that is, specify the distribution type of dispersed data and model parameters, and establish the likelihood function;
[0044] In step (2), a horizontal stratified model is used to approximate the underground structure of the virtual site. The number of layers in the horizontal stratified model is N, which can be specified based on other site survey information. In this embodiment, it is assumed that N is known, i.e., N = 3; the model parameters to be inverted are the layer thickness H and v of each layer. s Therefore, the shear wave velocity v can be determined. s Section; and each layer v p Both ρ and ρ are assumed to be known and equal to their respective true values (see Table 1). In this embodiment, vectors are used respectively. H 2 = [H1, Inf] and Represents all layers H and v s A set, where the subscript order is denoted by . Given an apparent dispersion curve d (like Figure 2 ), θ The uncertainty is represented by the distribution p( θ | d ) Quantization. Within the Bayesian framework, p( θ | d )equal
[0045]
[0046] Where p( θ )for θ Prior distribution, p( d | θ )yes d Likelihood function, P( d ) is the normalization constant. Based on the conditional probability formula and assuming... H 2 and independent, Within the framework of likelihood-free inference methods, from the posterior distribution p( d | θ Sampling from model parameters can be transformed into sampling from model parameters. θ and d Simulated samples (i.e.) The joint posterior distribution p) ε ( θ N , y | d obs Sampling from )
[0047]
[0048] in For a given θ hour The distribution, and with d Identical distribution; It is a likelihood function, which is assumed to be an indicator function I in this embodiment. S(ε) Where S(ε) represents the set
[0049] (21) p( H 2) Assuming a distribution similar to a uniform Dirichlet distribution.
[0050]
[0051] Where N = 2, H max It is v s The maximum recognizable depth of the interface, H min It is the minimum identifiable layer thickness. In this embodiment, H is... max and H min Set them to 20m and 1m respectively.
[0052] (22) For Assuming the shear wave velocity v in each layer s Independent and subject to uniform distribution:
[0053]
[0054] Where N = 2, v smax and v smin It is each layer v s The maximum and minimum possible values are set to 700m and 100m, respectively.
[0055] (23) For the likelihood function p( d | θ In this calculation, it is assumed that the observed phase velocity d varies at different frequencies. m The errors ω (where the subscript m represents the numbering of the data points arranged in ascending order of dispersion) are independent and identically distributed, and ω follows a pattern with a mean of zero and a variance of σ. 2 The Gaussian distribution of d. Because d m =c R,m +ω m , where c R,m Represents the theoretical phase velocity and d m They have the same modal number, so d It follows a multivariate Gaussian distribution. Assume σ -2 Follows a gamma distribution, Γ(σ) -2 )~Gam(a,b), and integrate p( d |σ 2 , θ σ in ) 2 achievable
[0056]
[0057] Where n = 2a, M is the total number of data points; in this embodiment, M = 86. Therefore, p( d |θ )obeys a student-T distribution in M dimensions, i.e. d ~ T c R (0, δ 2 I m , n).
[0058] 3) Establish a proper distance measure to quantify the similarity between the dispersion data and its simulated samples, and specify a method that can be used to sample from the posterior distribution in step (2);
[0059] (31) Distance function is defined as
[0060]
[0061] where G is the dispersion curve characteristic equation, and G' represents the first derivative of G with respect to the phase velocity; represents the projection of the value of the characteristic equation G on the phase velocity axis, which is used to estimate the difference between the observed phase velocity and its theoretical value; is the observed error simulation value of d m , and
[0062] (32) In this embodiment, subset simulation is used to solve the Bayesian equation as shown in formula (1.5). The nested threshold sequence ε 1 = +∞ > ε 2 > … > ε J = ε region is selected as the intermediate region corresponding to the increasingly close approximation of the observed dispersion data d in the observation space, i.e. The Markov Chain Monte Carlo method is used to sample from these intermediate regions, thereby efficiently solving the posterior distribution At the same time, the approximate posterior samples of p θ | θ d ) are also obtained. Note that in this embodiment, ε 1 is set to infinity and ε = 2.1.
[0063] (4) The posterior distribution of step (2) is solved by using the likelihood-free inference method, and the approximate posterior samples of the model parameters are used to represent the uncertainty of the inversion result (i.e. the shear wave velocity profile) of the multimodal dispersion data.
[0064] Figure 3 The inversion result of the multimodal dispersion data obtained by the proposed method is shown. The blue scatter points in each subgraph under the diagonal line in the figure represent the posterior samples of different model parameters (i.e. H1, v s1 and v s2 ). The probability histogram of each parameter is obtained by counting the posterior samples, and is represented as Figure 3Histograms of each subplot on the diagonal. To compare the results with those obtained by the proposed method, this embodiment, given the modality numbering information of the data points, inverted the basic modal data and multimodal data (e.g., Figure 2 The Bayesian equation shown in formula (1.1) under different data conditions was directly solved using the subset simulation sampling method. Figure 3 The dashed and solid lines in the diagram represent the posterior probability density function (PDF) images of each model parameter obtained statistically from the inversion results of the base mode data and multimodal data, respectively. First, the results obtained by the proposed method (histogram) are basically consistent with the results based on multimodal data (solid line), demonstrating the accuracy and rationality of the proposed method. Then, the v obtained statistically from the inversion results of the base mode data... s2 The PDF image (dashed line) has a significantly wider distribution range than the results obtained based on multimodal modal data (solid line) and the results of the proposed method. This indicates that the results obtained based on multimodal modal data have a wider distribution range. s2 The uncertainty is reduced. The above results demonstrate that the proposed method can reduce v s The identification results are uncertain or not unique.
[0065] Figure 4 This demonstrates the determination of v based on posterior samples of the model parameters obtained using the proposed method. s The posterior most probable value (MAP) of the profile and the true v of the virtual site s The cross-sections are basically the same; Figure 5 Demonstrates v-based s The theoretical dispersion curves and apparent dispersion curves of the fundamental mode (solid line) and the 1st mode (dashed line) obtained from the profile MAP calculation are shown. The dispersion curves of the two modes are fitted to the data points of their respective modes. Although the proposed method does not assume mode numbering for the data, Figure 5 The modes of the theoretical dispersion curves perfectly match the observed data points, demonstrating that the proposed method can automatically match the mode numbers of the data points. These results further prove the effectiveness of the proposed method. In summary, the virtual site demonstrates that the proposed method can reasonably quantify the uncertainty of multimodal dispersion data inversion results. The proposed method does not require the calculation of the likelihood function, thus avoiding mode misidentification and offering advantages in computational efficiency.
[0066] It should be noted that, in the present document, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a", "comprising", or "comprises" does not, without further qualification, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0067] The terminology used in the description herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used in the description of the embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0068] It should be understood that the term "and / or" as used herein merely describes associated objects in a manner that one or more of the associated objects can be present, unless the context clearly indicates otherwise. In addition, the character " / " as used herein generally represents an "or" relationship between the front and rear associated objects.
[0069] Depending on the context, the word "if" as used herein can be interpreted to mean "when" or "while" or "in response to determining" or "in response to detecting." Similarly, the phrase "if it is determined" or "if [a stated condition or event] is detected" can be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [the stated condition or event]" or "in response to detecting [the stated condition or event]."
[0070] The "first\second" mentioned in the embodiments are only to distinguish similar objects, and do not represent a specific order of the objects. Understandably, the "first\second" can be interchanged in a specific order or sequence as allowed. It should be understood that the objects distinguished by "first\second" can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0071] Although the present application has been described in detail with reference to the foregoing embodiments, the technical solutions recorded in the foregoing embodiments can be modified by those skilled in the art, or some technical features can be replaced by equivalent ones, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A modal self-matching based surface wave multi-modal dispersion curve probabilistic inversion method, characterized in that, Comprise the following steps: S1, when extracting Rayleigh wave multi-modal dispersion curve, apparent dispersion data points are divided into two categories, namely, base mode data set and higher mode data set; S2, under the framework of likelihood-free inference method, Bayesian equation for inverting multi-modal dispersion data is constructed, i.e. distribution type of dispersion data and model parameters is specified, and likelihood function is established; S3, suitable distance measure is established to quantify the similarity of dispersion data and its simulation sample, and the method available for sampling from posterior distribution is specified; S4, posterior distribution is solved by using likelihood-free inference method, and the uncertainty of inversion result of multi-modal dispersion data is represented based on approximate posterior sample of model parameters.
2. The modal self-matching based probabilistic inversion method of surface wave multi-modal dispersion curves according to claim 1, characterized in that: In step S1, the identifiable base mode data is classified into one category, i.e. base mode data set; the remaining data points are classified into another category, i.e. higher mode data set.
3. The modal self-matching based probabilistic inversion method of surface wave multi-modal dispersion curves according to claim 1, characterized in that: In step S2, the following steps are specifically included: S21, simplifying the underground structure of the test site into a horizontal layered model, which can be parameterized as the number of layers N of the model and model parameters including the thickness H of each layer, the shear wave speed v s , the compressional wave speed v p , and the density p; S22, Let v p Given ρ, only H and v of each layer are inverted. s ,use θ Represents all layers H and v s The set of parameters for the shear wave velocity profile; S23, given apparent dispersion curve data d Under the Bayesian framework θ The uncertainty in the parameters is quantified by their posterior distribution p( θ | d ) S24. Within the framework of the likelihoodless inference method, the solution to p( θ | d It was transformed into a solution θ Simulated samples of apparent dispersion curve data joint posterior distribution The formula is expressed as follows: where p θ ) is θ the prior distribution, is the likelihood function, θ is the prior distribution of , i.e., the distribution of d ; is the likelihood function, set as the indicator function I S(ε) , where S(ε) represents the set L represents the distance measure quantifying the similarity between d and ; when I S(ε) (y) = 1, otherwise I S(ε) (y) = 0; P ε ( d ) is the normalization constant.
4. The modal self-matching based probabilistic inversion method of surface wave multi-modal dispersion curves according to claim 1, characterized in that: In step S2, selecting different forms of data prior distribution allows to restore the actual statistical characteristics of data with different complexity; any one of the following is selected as the data prior distribution: multivariate Gaussian distribution considering independent and identically distributed data error, Student-T distribution integrating data variance, or multivariate Gaussian distribution considering data variance correlation.
5. The modal self-matching based probabilistic inversion method of surface wave multi-modal dispersion curves according to claim 1, characterized in that: In step S3, a distance measure L of the automatically matched modality number is selected for the higher modality data set; the modality information constraint inversion result of the base modality data is added in the distance measure, and a method based on random simulation is used for solving 6. The modal self-matching based probabilistic inversion method of surface wave multi-modal dispersion curves according to claim 4, characterized in that: The method based on random simulation comprises Markov chain Monte Carlo (MCMC) and subset simulation+MCMC method.
Citation Information
Cited By
Shear wave velocity two-dimensional profile probability inversion method based on multi-source data fusion
CN121808195A