Seismic facies analysis method and system based on random projection tree
By applying the SOM clustering method based on a random projection tree based on SOM based on a SOM analysis, the problem of utilizing a small amount of drilling data in a deposition phase analysis is solved, and efficient and accurate deposition phase analysis is achieved.
Patent Information
- Application Number
- CN202311546045.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-20
- Publication Date
- 2025-05-20
AI Technical Summary
In the analysis of depositional phases, it is difficult for the prior art to effectively utilize a small amount of drilling data to accurately grasp the plane spread characteristics of the deposited phases, and the traditional methods rely on a large amount of drilling data and manual interpretation, which is inefficient.
The seismic layer attributes are dimensionality reduction and clustered by using a large data nonlinear dimensionality reduction method based on random projection trees and a PSO-optimized SOM clustering method, and visual results are generated to explain the sedimentary facies.
This method can efficiently perform deposition phase analysis under a small amount of drilling data, improving the accuracy and efficiency of the analysis and reducing the dependence on manual interpretation.
Smart Images

Figure CN120020602A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of exploration geophysics technology, and in particular to a seismic phase analysis method and system based on random projection tree. Background Technology
[0002] Sedimentary phase characteristics are extremely important in oil, gas and coalbed methane seismic exploration and development. Therefore, sedimentary phase research is of great significance. In the process of seismic exploration and development, there are great differences in research means and methods between the sedimentary phase research of the target layer buried deep underground and the sedimentary phase research of the outcrop.
[0003] The purpose of sedimentary phase analysis is to describe and interpret the extracted seismic reflection parameters, including but not limited to geometry, continuity, amplitude, frequency, velocity, coherence, etc. For three-dimensional seismic data, due to its large coverage area, relying solely on manual sedimentary phase description and conclusion requires a lot of time and experience of interpreters, but the relevant experience of interpreters is also very necessary in the process of sedimentary phase description. Even when using user-friendly sedimentary phase processing software, subjective results are still required through the experience of interpreters. In recent decades, sedimentary phase analysis technology has continued to develop, and the use of pattern recognition methods has gradually emerged. By clustering similar seismic signals, each clustered category represents a sedimentary feature. This is also an important way of sedimentary phase research at present.
[0004] In the process of sedimentary phase analysis, only through rock data can the target sedimentary phase markers be observed. However, since drilling coring is generally not carried out continuously, and the coring rate of the entire well area of an exploration well is usually between a few percent and more than ten percent, this has caused great difficulties in sedimentary phase research. In order to make a continuous sedimentary phase interpretation of the entire well, the well logging phase analysis can be carried out through the electrical side well data, but the phase interpretation in this way is highly multi-solution. Therefore, in addition to processing the above two types of data, it is also urgent to obtain more information from other data to improve the accuracy of sedimentary phase interpretation.
[0005] More importantly, even if the data of single well phase analysis is sufficient, the traditional research method can only obtain part of the information, and important information such as stratigraphic superposition pattern and sedimentary body shape are not used. Furthermore, even if the interpretation is completely correct, it is only a "one-hole view" after all. In order to further understand the planar distribution characteristics of sedimentary phases, a large number of sufficiently dense boreholes must be available, which is difficult to meet in the exploration stage. Therefore, a new means and method that can better grasp the planar change characteristics of sedimentary phases with only a small number of boreholes is urgently needed.
[0006] The method of sedimentary facies analysis is to identify the unique seismic reflection wave group characteristics and their morphological combinations within each sequence, endow them with certain geological meanings, and then interpret the sedimentary facies. This process is called sedimentary facies analysis.
[0007] There are two methods for sedimentary facies analysis and identification. The first method is to observe the seismic reflection characteristics manually and compare them with the established standard sedimentary facies characteristics to determine which sedimentary facies it belongs to. This method is generally applied to the interpretation and analysis of local seismic data, and the accuracy of interpretation and identification is relatively low. The other method is to analyze and calculate the seismic data volume or seismic attribute data by applying seismic data processing technology, computer technology and certain mathematical methods to extract the sedimentary facies that can reflect the changes of sedimentary facies. This is an efficient, advanced and quantitative sedimentary facies identification method.
[0008] Seismic waveform is the basic property of seismic data. It contains all qualitative and quantitative information, such as reflection mode, phase, frequency and amplitude information. It is the overall characteristic of seismic information. Its dynamic changes contain rich internal information and can truly reflect the characteristics of the underground structure. The waveform classification method is the most commonly used sedimentary facies analysis method. By classifying the seismic signal waveform, the division of sedimentary facies can be realized. Summary of the Invention
[0009] In view of the above problems, the present invention is proposed to provide a seismic facies analysis method and system based on a random projection tree that overcomes the above problems or at least partially solves the above problems.
[0010] According to one aspect of the present invention, a seismic facies analysis method based on a random projection tree is provided. The analysis method includes:
[0011] Step 1: Extract the seismic data to be processed according to the seismic data volume and the horizon file;
[0012] Step 2: Select seismic layer attributes according to the seismic data;
[0013] Step 3: Use a big data non-linear dimensionality reduction method based on a random projection tree to perform dimensionality reduction on the seismic layer attributes to obtain the dimensionality-reduced seismic layer attributes;
[0014] Step 4: Use a PSO-optimized SOM to perform clustering according to the dimensionality-reduced seismic layer attributes to obtain a clustering result;
[0015] Step 5: Output a visualization result according to the clustering result.
[0016] Optionally, the specific content of Step 1: Extract the seismic data to be processed according to the seismic data volume and the horizon file includes: processing the original seismic data, including extracting the seismic data between the horizons to be processed.
[0017] Optionally, the step 3: using a big data nonlinear dimensionality reduction method based on random projection tree to reduce the dimension of the seismic layer attributes, and obtaining the seismic layer attributes after dimensionality reduction specifically includes:
[0018] Use the random projection tree to obtain a space partition and obtain the partition space;
[0019] Get the k-nearest neighbors of each point in the partitioned space to obtain a preliminary KNN graph;
[0020] Based on the preliminary KNN graph, a neighborhood search algorithm is used to obtain potential neighbors to obtain a KNN graph;
[0021] For the weights of the edges in the k-nearest neighbor graph, they are set in a way similar to the conditional probability between two points.
[0022] Optionally, the step of obtaining a space partition by using a random projection tree, wherein obtaining the partitioned space specifically includes:
[0023] The conditional probability between different seismic trace attributes is:
[0024]
[0025] The conditional probability between the attributes of the same seismic trace layer is:
[0026] P i|i =0
[0027] In the conditional probability formula between different seismic trace attributes, by setting the perplexity σ i To set the parameters, where the conditional distribution P .|i The value of is equal to the perplexity log u , set and The weight between them.
[0028] Optional, the settings and The weights between include:
[0029] The probability weights of different seismic trace layer attribute conditions are:
[0030]
[0031] Use a probabilistic model for visualization, which preserves the similarity between vectors as much as possible in a low-dimensional space:
[0032] Given a set of vectors (r i ,v j ), the probability calculation formula of the existence of an edge between two seismic trace layer attributes is:
[0033] In the above formula, is the representation of v i in the low-dimensional space; f(·) is the probability function distance between vertices y i and y j ;
[0034] When the distance between y i and y j in the low-dimensional space is less than the first threshold, this binary edge is observed;
[0035] When the distance between y i and y j in the low-dimensional space is greater than the second threshold, it is observed that there is no such binary edge;
[0036] The formula defines the observation of a binary edge, which is further extended to a weighted edge; the possibility of observing a weighted edge is defined.
[0037] Optionally, the probability calculation formula for observing a binary edge in the low-dimensional space is:
[0038]
[0039] Using the skip-gram idea of word2vec, the method of randomly selecting some negative edges to replace all negative edges is used to optimize the model;
[0040] For each vertex i, some vertices are randomly sampled according to the noise distribution, and (i, j) is regarded as a negative edge;
[0041] Under the condition that M negative edges are sampled for each positive edge, the objective function is:
[0042]
[0043] Optionally, the specific steps of step 4: clustering the seismic layer attributes after dimensionality reduction by using SOM optimized by PSO to obtain the clustering result include:
[0044] Initializing the processed data;
[0045] Defining the fitness function and calculating and solving the individual extreme value and the global optimal solution;
[0046] Updating the velocity and position;
[0047] Setting the termination condition of the iteration process.
[0048] Optionally, the specific initialization of the processed data includes:
[0049] Set the maximum number of iterations, the number of independent variables of the objective function, the maximum velocity of the particles, and the position information as the entire search space;
[0050] Randomly initialize the velocity and position within the velocity interval and the search space, set the particle swarm size to M, and randomly initialize a flying velocity for each particle.
[0051] Optionally, the defining the fitness function and calculating and solving the individual extreme value and the global optimal solution specifically include:
[0052] Define the fitness function, and the individual extreme value is the optimal solution found by each particle;
[0053] Find a global value from the optimal solution to obtain the current global optimal solution;
[0054] Compare the current global optimal solution with the historical global optimal solution and update it.
[0055] Optionally, the updating the velocity and position specifically includes:
[0056] The formulas for updating the velocity and position
[0057] V id = ωV id + C 1 random(0, 1)(P id - X ia )
[0058] + C 2 random(0, 1)(P gd - X id )
[0059] X id = X id + V id
[0060] Among them, ω is called the inertia factor; by adjusting ω, the global optimization performance and the local optimization performance are adjusted;
[0061] C 1 and C 2 are called acceleration constants. The former is the individual learning factor of each ion, and the latter is the social learning factor of each particle. When C 1 and C 2 are constants, better solutions can be obtained. Usually, C 1 = C 2 = 2, and take C 1 = C 2 ∈[0, 4];
[0062] random(0, 1) represents a random number in the interval [0, 1], Pid represents the d-th dimension of the individual extreme value of the i-th variable, P gd represents the d-th dimension of the global optimal solution.
[0063] Optionally, the termination condition for setting the iterative process specifically includes: setting that the difference between the set number of iterations and the number of generations satisfies a minimum bound.
[0064] The present invention also provides a seismic facies analysis system based on a random projection tree, which applies the above-mentioned seismic facies analysis method based on a random projection tree. The analysis system includes:
[0065] A seismic data extraction module, configured to extract seismic data to be processed according to a seismic data volume and a horizon file;
[0066] A seismic layer attribute selection module, configured to select seismic layer attributes according to the seismic data;
[0067] A seismic layer attribute dimensionality reduction module, configured to perform dimensionality reduction on the seismic layer attributes by using a big data non-linear dimensionality reduction method based on a random projection tree to obtain the dimensionality-reduced seismic layer attributes;
[0068] An optimization clustering module, configured to perform clustering according to the dimensionality-reduced seismic layer attributes by using a PSO-optimized SOM to obtain a clustering result;
[0069] A result output module, configured to output a visualization result according to the clustering result.
[0070] A seismic facies analysis method based on a random projection tree provided by the present invention, the analysis method includes: Step 1: Extract seismic data to be processed according to a seismic data volume and a horizon file; Step 2: Select seismic layer attributes according to the seismic data; Step 3: Perform dimensionality reduction on the seismic layer attributes by using a big data non-linear dimensionality reduction method based on a random projection tree to obtain the dimensionality-reduced seismic layer attributes; Step 4: Perform clustering according to the dimensionality-reduced seismic layer attributes by using a PSO-optimized SOM to obtain a clustering result; Step 5: Output a visualization result according to the clustering result. Using the combination of domain space search and probability model to perform non-linear dimensionality reduction on seismic layer attributes is more suitable for the non-linear characteristics of seismic layer attribute data and can provide a scientific basis for the subsequent reservoir exploration process.
[0071] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically enumerates the specific embodiments of the present invention. Brief Description of the Drawings
[0072] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0073] Figure 1 The flowchart of a seismic facies analysis method based on a random projection tree provided by an embodiment of the present invention;
[0074] Figure 2 It is the result graph of the present invention when using actual seismic data. Specific embodiments
[0075] The following will describe the exemplary embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0076] The terms "including" and "having" and any variations thereof in the description of the embodiments, claims, and drawings of the present invention are intended to cover non-exclusive inclusion. For example, a series of steps or units are included.
[0077] The following will further describe the technical solutions of the present invention in detail with reference to the accompanying drawings and embodiments.
[0078] As Figure 1 shown, the present invention provides a seismic facies analysis method based on a random projection tree and self-organizing clustering, including the steps of:
[0079] Process the original seismic data, mainly extract the seismic data between the layers to be processed, to facilitate subsequent operations and reduce the corresponding computational amount.
[0080] Select and extract the seismic layer attributes that may contain seismic facies information according to the experience of petroleum exploration practitioners.
[0081] Design a non-linear dimensionality reduction method for seismic layer attributes based on a random projection tree and applicable to big data, and then reduce the dimensionality of the seismic layer attributes through this method.
[0082] Use the SOM clustering method improved by the particle swarm optimization method with the addition of the heterogeneity index to cluster the dimensionality-reduced seismic layer attributes.
[0083] Visualize the clustering results.
[0084] According to an actual example of the present invention, the non-linear dimensionality reduction algorithm based on a random projection tree in the above step 3 includes the following steps:
[0085] (1) Use a random projection tree to obtain a space partition. On this basis, find the k-nearest neighbors of each point to obtain a preliminary KNN (K-Nearest Neighbor) graph; on this basis, use a neighborhood search algorithm to find potential neighbors. Finally, a more accurate KNN graph is obtained. In this way, a KNN graph can be constructed efficiently. For the weights of the edges in the k-nearest neighbor graph, set them in a way similar to the conditional probability between two points.
[0086] The conditional probability between different seismic trace layer attributes is:
[0087]
[0088] The conditional probability between the same seismic trace layer attributes is:
[0089] P i|i = 0
[0090] In the conditional probability formula between different seismic trace layer attributes, set the parameter by setting the perplexity σ i where the value of the conditional distribution P ·|i is equal to the perplexity log u , and then set the weights between and . The conditional probability weights of different seismic trace layer attributes are:
[0091]
[0092] At the same time, use a probability model for visualization, which preserves the similarity between vectors as much as possible in the low-dimensional space. Given a set of vectors (v i , v j ), the probability calculation formula for the existence of an edge between two seismic trace layer attributes is:
[0093]
[0094] In the above formula, is the representation of v i in the low-dimensional space. f(·) is the probability function distance between the vertices y i and y j . When y i and y j are relatively close in the low-dimensional space, there is a greater probability of observing this larger binary edge. When y i and y jWhen the distance is relatively far in the low-dimensional space, the probability of observing this binary edge is relatively small. The above formula defines the observation of a binary edge and further extends it to a weighted edge, defining the possibility of observing a weighted edge. The probability calculation formula for observing a binary edge in the low-dimensional space is:
[0095]
[0096] Since the number of negative edges in the above weighted graph is approximately equal to the square of the number of nodes, it is costly to directly optimize the likelihood of this weighted graph. Using the skip-gram idea of word2vec, a method of randomly selecting some negative edges to replace all negative edges is adopted to optimize the model. For each vertex i, some vertices are randomly sampled according to the noise distribution, and (i, j) is regarded as a negative edge. Under the condition that M negative edges are sampled for each positive edge, the objective function is:
[0097]
[0098] By maximizing the above, the seismic layer attribute data points with low similarity are separated far from each other, meeting the purpose of dimensionality reduction.
[0099] According to the actual example of the present invention, the SOM clustering method in step 4 includes the following steps:
[0100] (1) Initialization: Each node randomly initializes its own parameters, and the number of parameters of each node is the same as the dimension of Input.
[0101] (2) For each input data, find the node that best matches it. Assume the input is D-dimensional, that is, X = {x 1 , …, x i , ……, x D}, then the discriminant function can be the Euclidean distance:
[0102] (3) After finding the activated node I(x), update the nodes adjacent to it. Let S i,j represent the distance between nodes i and j. For the nodes adjacent to I(x), assign them an update weight. The adjacent nodes are updated to different degrees according to the distance.
[0103]
[0104] (4) Update the node parameters according to the gradient descent method.
[0105] Δw ji = η(t)·T j,I(x) (t)·(x i - w ji )
[0106] (5) Iterate until convergence
[0107] According to the actual example of the present invention, the particle swarm optimization method in the above step 4 includes the following steps:
[0108] (1) Initialization
[0109] First, we set the maximum number of iterations, the number of independent variables of the objective function, the maximum speed of the particles, the position information as the entire search space. We randomly initialize the speed and position in the speed interval and the search space, set the particle swarm size as M, and randomly initialize a flying speed for each particle.
[0110] (2) Personal best and global best solution
[0111] Define the fitness function. The personal best is the optimal solution found by each particle. Find a global value from these optimal solutions, which is called the current global best solution. Compare it with the historical global best and update.
[0112] (3) Formulas for updating speed and position
[0113] V id = ωV id + C 1 random(0, 1)(P id - X id )
[0114] + C 2 random(0, 1)(P gd - X id )
[0115] X id = X id + V id
[0116] Among them, ω is called the inertia factor. When its value is negative and relatively large, the global optimization ability is strong and the local optimization ability is weak; when it is relatively small, the global optimization ability is weak and the local optimization ability is strong. By adjusting the size of ω, the global optimization performance and local optimization performance are adjusted. C 1 and C 2 are called acceleration constants. The former is the individual learning factor of each ion, and the latter is the social learning factor of each particle. When C 1 and C 2 are constants, better solutions can be obtained. Usually, C 1 = C 2 = 2, but it doesn't necessarily have to be equal to 2. Generally, C 1 = C 2 ∈[0, 4]. random(0, 1) represents a random number in the interval [0, 1], Pid represents the d-th dimension of the individual extreme value of the i-th variable, P gd represents the d-th dimension of the global optimal solution.
[0117] (4) Termination condition: Set that the difference between the set number of iterations and the number of generations satisfies the minimum bound.
[0118] The object of the present invention is, under the conditions of a given N-channel seismic data set D and the seismic horizon ranges L1 and L2 to be intercepted, to divide the seismic data into K (K≤N) partitions, where S k represents the data geometry of each partition after partitioning. A partition is called a cluster, and each cluster represents a sedimentary facies.
[0119] First, for the post-stack seismic data volume in the (time domain or depth domain) of the target layer, select the N-channel seismic data set D to be used for classification therefrom. According to the pre-given interpreted horizons, extract the seismic data contained between the horizons as classification data.
[0120] Suppose it is necessary to divide N seismic traces into K categories according to their P-dimensional attribute X, then the data can be expressed as X ij , where i = 1, 2,..., N represents the data point serial number, and j = 1, 2,..., P represents the P attributes extracted on each seismic trace.
[0121] Since too many seismic attributes will lead to data redundancy problems, it is necessary to process the extracted seismic layer attributes. Since the seismic layer attributes are high-dimensional and non-linear, using the traditional PCA principal component analysis method may not achieve good results. The present invention creatively proposes a non-linear dimensionality reduction method based on random projection trees and self-organizing clustering to reduce the dimensionality of the seismic layer attribute data. In this method, the local neighborhood relationship in the low-dimensional space is the same as that in the original nested space, which is more suitable for solving the non-linear feature dimensionality reduction problem of seismic data. Finally, the above data X ij is processed into X 限 , where i = 1, 2,..., N represents the data point serial number, k = 1, 2,..., Q, Q < P, representing the Q attributes obtained after dimensionality reduction of the P attributes extracted on each seismic trace.
[0122] As described above, the object of the present invention is to classify the seismic trace waveform shapes according to the layer attributes of the seismic trace waveforms in a certain section of the target layer, and form discrete "sedimentary facies" according to the classification results, so as to characterize and understand the lateral variation of seismic signals. Since this lateral variation of seismic signals is actually reflected in some seismic attribute slices, when the seismic attribute slices at the center of the waveform classification time window are displayed to the user on the computer, the user can see the lateral variation of seismic signals.
[0123] The above process does not require the user to manually classify seismic wave type data based on other information they possess. The present invention will automatically cluster seismic traces according to the seismic layer attributes extracted from the above seismic trace data in the form of SOM clustering. SOM clustering is similar to the KMeans algorithm, both of which divide sets of individuals with small distances into the same category and sets of individuals with large distances into different categories. It adopts a competitive learning strategy, where each neuron in the competitive layer competes with each other to vie for the right to update the weights, that is, the so-called "winner takes all" concept. More specifically, for the current input sample x, only the neuron most similar to it will be updated.
[0124] According to its basic idea, only one neuron is updated each time. For the winning neuron, it should also have a certain impact on the nearby neurons, and this impact should be from far to near and decrease with time. More specifically, a neighborhood N(w) with a radius of R is drawn centered on the winning neuron, and the weights falling within the neighborhood receive a weight update that changes according to the distance.
[0125] Regarding the problem of the coarseness of the SOM clustering results, such as the fact that there is less sample data corresponding to some categories in the clustering results, which is caused by its unsupervised learning mechanism and is not well shown in the clustering results. To avoid the above problems, the present invention proposes a technique of further processing the SOM clustering samples using an improved particle swarm optimization algorithm PSO.
[0126] After reducing the dimension of the seismic layer attribute data through a dimension reduction method based on a random projection tree, the seismic traces are clustered using the PSO particle swarm optimization method and the SOM clustering method according to the layer data corresponding to each seismic trace after dimension reduction. Finally, the sedimentary facies clustering result of the seismic data is obtained.
[0127] See Figure 2 , this figure is the result graph of clustering the actual seismic data using the above process. Therefore, the present invention has a good effect on sedimentary facies analysis.
[0128] Advantageous effects: The present invention proposes a seismic facies analysis method based on a random projection tree and self-organizing clustering from a new perspective. This method combines domain space search and probability models to perform non-linear dimension reduction on seismic layer attributes, which is more suitable for the non-linear characteristics of seismic layer attribute data and provides a scientific basis for the oil reservoir exploration process.
[0129] The above specific implementation manners have further detailed the purpose, technical solution, and advantageous effects of the present invention. It should be understood that the above is only the specific implementation manners of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A seismic phase analysis method based on random projection tree, characterized in that: The analysis method comprises: Step 1: Extract the seismic data to be processed according to the seismic data volume and layer file; Step 2: Selecting seismic layer attributes according to the seismic data; Step 3: Using a large data nonlinear dimensionality reduction method based on random projection tree to reduce the dimension of the seismic layer attributes, and obtain the seismic layer attributes after dimensionality reduction; Step 4: clustering the seismic layer attributes after dimension reduction using the SOM optimized by PSO to obtain clustering results; Step 5: Output visualization results based on the clustering results.
2. A seismic phase analysis method based on random projection tree according to claim 1, characterized in that: The step 1: extracting the seismic data to be processed according to the seismic data volume and the layer file specifically includes: processing the original seismic data, including extracting the seismic data between the layers to be processed.
3. A seismic phase analysis method based on random projection tree according to claim 1, characterized in that: The step 3: using a large data nonlinear dimensionality reduction method based on random projection tree to reduce the dimension of the seismic layer attributes, and obtaining the seismic layer attributes after dimensionality reduction specifically includes: A space partition is obtained by using a random projection tree to obtain a partitioned space; Obtaining the k-nearest neighbors of each point in the partitioned space to obtain a preliminary KNN graph; Based on the preliminary KNN graph, a neighborhood search algorithm is used to obtain potential neighbors to obtain a KNN graph; The weights of the edges in the k-nearest neighbor graph are set in a way similar to the conditional probability between two points.
4. A seismic phase analysis method based on random projection tree according to claim 3, characterized in that: The method of obtaining a space partition by using a random projection tree specifically includes: The conditional probability between different seismic trace attributes is: The conditional probability between the attributes of the same seismic trace layer is: P i|i =0 In the conditional probability formula between different seismic trace attributes, by setting the perplexity σ i To set the parameters, where the conditional distribution P ·|i The value is equal to the perplexity log u ,set up and The weight between .
5. A seismic phase analysis method based on random projection tree according to claim 4, characterized in that: The settings and The weights between them include: The conditional probability weights of different seismic trace layer attributes are: Use a probabilistic model for visualization, which preserves the similarity between vectors as much as possible in a low-dimensional space; Given a set of vectors (v i ,v j ), the probability calculation formula of the existence of an edge between two seismic trace layer attributes is: In the above formula, Yes i Representation in low-dimensional space; f(·) is the vertex y i and j Probability function distance between them; When i and j When the distance in the low-dimensional space is less than the first threshold, this binary edge is observed; When i and j When the distance in the low-dimensional space is greater than the second threshold, it is observed that there is no such binary edge; The formula defines the observation of a binary edge and further extends it to weighted edges; it defines the probability of observing a weighted edge.
6. A seismic phase analysis method based on random projection tree according to claim 5, characterized in that: The probability of observing a binary edge in the low-dimensional space is calculated as: Using the skip-gram idea of word2vec, the model is optimized by randomly selecting some negative edges to replace all negative edges. For each vertex i, randomly sample some vertices according to the noise distribution and treat (i, j) as a negative edge; Under the condition that M negative edges are obtained by sampling each positive edge, the objective function is:
7. The seismic phase analysis method based on random projection tree according to claim 1, characterized in that: The step 4: using the SOM optimized by PSO to cluster the attributes of the seismic layer after dimension reduction, and obtaining the clustering results specifically includes: Initialize processing data; Define the fitness function and calculate the individual extreme value and the global optimal solution; Update speed and position; Set the termination condition of the iteration process.
8. A seismic phase analysis method based on random projection tree according to claim 7, characterized in that: The initialization processing data specifically includes: Set the maximum number of iterations, the number of independent variables of the objective function, the maximum speed of the particle, and the position information for the entire search space; The speed and position are randomly initialized in the speed interval and search space, the particle swarm size is set to M, and each particle is randomly initialized with a flying speed.
9. A seismic phase analysis method based on random projection tree according to claim 7, characterized in that: Defining the fitness function and calculating and solving the individual extreme value and the global optimal solution specifically include: Define the fitness function, where the individual extreme value is the optimal solution found for each particle; Find a global value from the optimal solution to obtain the current global optimal solution; The current global optimal solution is compared with the historical global optimal solution and updated.
10. A seismic phase analysis method based on random projection tree according to claim 7, characterized in that: The update speed and position specifically include: Formula for updating velocity and position V id =ωV id +C1random(0,1)(P id -X id )+C2random(0,1)(P gd -X id ) X id =X id +V id Among them, ω is called the inertia factor; by adjusting ω, the global optimization performance and local optimization performance are adjusted; C1 and C2 are called acceleration constants. The former is the individual learning factor of each ion, and the latter is the social learning factor of each particle. When C1 and C2 are constants, a better solution is obtained. Usually, C1=C2=2 is set, and C1=C2∈[0,4] is taken; random(0,1) represents a random number in the interval [0,1], P id represents the dth dimension of the individual extreme value of the ith variable, P gd represents the dth dimension of the global optimal solution.
11. A seismic phase analysis method based on random projection tree according to claim 7, characterized in that: The termination condition of the iterative process is specifically set to include: reaching a set number of iterations and the difference between generations meeting a minimum limit.
12. A seismic phase analysis system based on random projection tree, using a seismic phase analysis method based on random projection tree as described in any one of claims 1 to 11, characterized in that: The analysis system comprises: A seismic data extraction module is used to extract seismic data to be processed based on the seismic data volume and the layer file; A seismic layer attribute selection module, used for selecting seismic layer attributes according to the seismic data; A seismic layer attribute dimensionality reduction module is used to perform dimensionality reduction of the seismic layer attributes by using a large data nonlinear dimensionality reduction method based on a random projection tree to obtain the seismic layer attributes after dimensionality reduction; An optimization clustering module is used to cluster the seismic layer attributes after dimension reduction using a PSO-optimized SOM to obtain a clustering result; The result output module is used to output a visualization result according to the clustering result.