Multi-modal large model federation training platform and heterogeneous data alignment algorithm
By building a vertical federated learning architecture and using convolutional neural networks and BERT models for data mapping, combining the gradient confusion mechanism of dynamic weighted comparison loss and Weibull distributed noise, the problem of cross-modal alignment and fusion of well logging curves and geological reports is solved, and the secure fusion of multimodal knowledge is achieved under the premise of protecting data privacy.
Patent Information
- Application Number
- CN202510599941.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-12
AI Technical Summary
In oil and gas exploration, logging curves and geological reports are difficult to effectively align and fuse due to modal differences and semantic divides, and the existing technology cannot achieve cross-modal semantic alignment while protecting data privacy.
By building a vertical federal learning architecture, each oil and gas enterprise locally deploys time sequence data processing modules and text processing modules, respectively, the logging curves are decomposed and time-frequency features are extracted, and the geological report is standardized and semantic enhancement is carried out; the embedding generation module uses convolutional neural network and BERT model to map two types of data to the embedded space of a unified dimension; the central coordination point aligns the cross-modal embedding vectors through dynamic weighted comparison loss functions, and introduces a gradient obfuscation mechanism for Weibull distributed noise to achieve privacy protection.
It realizes the multimodal knowledge fusion through encrypted embedded vector interaction without leaving the local area, while resisting gradient inversion attacks, providing a safe and reliable data collaboration infrastructure.
Smart Images

Figure CN120105464A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the intersection of artificial intelligence and multimodal data processing, and specifically to a multimodal large model federated training platform and a heterogeneous data alignment algorithm. Background Art
[0002] In oil and gas exploration and development, well logging curves and geological reports are two types of core data resources, which record underground lithology, physical properties and other information in the form of time series signals and natural language respectively. In traditional technology, a single enterprise is limited by the scale of data and it is difficult to train high-precision intelligent models, and cross-enterprise data collaboration faces two major bottlenecks: first, well logging data involves sensitive geological information, and direct sharing between enterprises is prone to leaking commercial secrets; second, there are significant modal differences between well logging curves (high-dimensional time series) and geological reports (unstructured text), and traditional federated learning frameworks only support collaborative training of homogeneous data and cannot establish cross-modal semantic associations.
[0003] Among existing solutions, federated learning based on homomorphic encryption can protect data privacy, but cannot handle heterogeneous data alignment; multimodal contrastive learning can integrate text and signals, but it relies on centralized data training and does not meet the data isolation requirements of the oil and gas industry. In addition, traditional differential privacy achieves privacy protection by adding Gaussian or Laplace noise, but the noise distribution characteristics do not match the model gradient, resulting in excessive consumption of the privacy budget or a significant decrease in model accuracy.
[0004] For example, the oil and gas data federated learning system proposed in the public patent CN114462734A only supports the sharing of homogeneous logging data, and does not solve the problem of text and time series data fusion; the paper "Gradient Leakage Attacks in Multimodal Federated Learning" (IEEETIFS 2022) points out that existing privacy protection methods are easily reverse engineered to restore the original data in cross-modal scenarios. Therefore, there is an urgent need for a federated learning framework that can achieve cross-modal semantic alignment of logging curves and geological reports while protecting data privacy. This patent designs a vertical federated architecture to independently extract multimodal features at local nodes, and the central node only coordinates the alignment of the embedded space. It combines dynamic contrast loss and Weibull noise injection to avoid the leakage of original data and break through modal barriers, providing a safe and reliable data collaboration infrastructure for intelligent oil and gas exploration. Summary of the invention
[0005] In view of the shortcomings of the above prior art, the purpose of the present invention is to provide a multimodal large model federated training platform and a heterogeneous data alignment algorithm to solve the problem that well logging curves (time series data) and geological reports (text) are difficult to effectively align and fuse due to modal differences and semantic gaps. The present invention constructs a vertical federated learning architecture, and each oil and gas company locally deploys a time series data processing module and a text processing module to perform wavelet packet decomposition on the well logging curve to extract time-frequency features, and perform terminology standardization and semantic enhancement on the geological report; the embedding generation module uses a convolutional neural network and a BERT model to map the two types of data to an embedding space of uniform dimension; the central coordination node aligns the cross-modal embedding vector through a dynamic weighted contrast loss function, and introduces a gradient confusion mechanism of Weibull distribution noise to achieve privacy protection in the aggregation stage of the federated model. This method enables the well logging data and text data of different companies to complete multimodal knowledge fusion through encrypted embedding vector interaction without leaving the local premise, while resisting gradient inversion attacks.
[0006] The present invention provides a multimodal large model federated training platform and a heterogeneous data alignment algorithm, including: A time series data processing module, which performs wavelet packet decomposition on the logging curve to generate a time-frequency matrix signal; A text processing module, which performs terminology standardization processing on the geological report and outputs text feature signals; An embedding generation module, wherein the embedding generation module processes the time-frequency matrix signal through a convolutional neural network and generates a time series embedding vector, and the embedding generation module processes the text feature signal through a BERT model and generates a text embedding vector; A local federated learning module, wherein the local federated learning module receives and encrypts the time series embedding vector and the text embedding vector to form a local encrypted signal; A contrastive learning alignment module, wherein the contrastive learning alignment module calculates a dynamic weighted contrast loss according to the local encrypted signal and generates a gradient signal; A global model aggregation module, wherein the global model aggregation module processes the gradient signal through a federated averaging algorithm and forms a parameter update amount; A privacy protection verification module is provided, wherein the parameter update amount conforms to the noise of the Weibull distribution, and a confusion gradient signal is generated and transmitted back to each module to align the data.
[0007] In one embodiment of the present invention, the wavelet packet decomposition process of the time series data processing module includes an adaptive energy threshold screening mechanism, which specifically suppresses high-frequency noise in the logging curve by dynamically adjusting the number of decomposition layers and bandwidth allocation of the wavelet basis function. The core algorithm is: Where Ψopt is the algorithm code, arg is the decomposition coefficient, max is the maximum value symbol, W is the set of candidate wavelet basis functions, F represents the Fourier transform operator, DWT is the discrete wavelet transform operation, and K is the number of decomposition layers. The algorithm selects the optimal basis function by maximizing the energy concentration of the sub-band components. The signal-to-noise ratio of the formation interface reflection wave characteristics in the generated time-frequency matrix signal is improved by more than 40% and transmitted to the embedded generation module through the data channel.
[0008] In one embodiment of the present invention, the term standardization processing of the text processing module adopts a semantic enhancement method based on a knowledge graph, and maps the synonyms and abbreviations in the original text to a standardized term library by constructing an entity relationship network in the geological field. The vectorization process satisfies: in is the original word segmentation result, KG represents the associated entity embedding extracted from the knowledge graph, ⊕ is the vector concatenation operation, MLP is the multi-layer perceptron. This method effectively solves the semantic ambiguity problem of terms, and the generated text feature signal is transmitted to the embedding generation module through the bus.
[0009] In one embodiment of the present invention, the convolutional neural network of the embedding generation module adopts a multi-scale feature fusion architecture, and works together through a parallel standard convolution path and a hole convolution path, wherein the standard convolution kernel is responsible for capturing the local detail features of the formation reflection wave in the time-frequency matrix, and the hole convolution kernel extracts the macro trend features of the logging curve by expanding the receptive field. The output feature maps of the two convolution paths are spliced in the channel dimension and sent to the global pooling layer to generate a time series embedding vector with multi-scale perception capability. After normalization, the vector is input into the local federated learning module together with the text embedding vector, and the cross-modal association of the time-frequency characteristics of the logging signal and the semantic information of the geological text is realized through the feature layer fusion mechanism.
[0010] In one embodiment of the present invention, the contrastive learning alignment module introduces a modal adaptive weight adjustment mechanism, which dynamically balances the contribution ratio in the cross-modal similarity calculation through the learnable time series modal weight factor and the text modal weight factor, and simultaneously considers the bidirectional correlation strength from time series to text and from text to time series when calculating the similarity score of the positive sample pair. A soft contrast target is constructed through a normalized exponential function, which effectively alleviates the embedding space distortion problem caused by the difference in dimensional richness of well logging data and sparsity of geological report text. The generated gradient signal is transmitted to the global model aggregation module through an encrypted channel, while retaining the weight distribution information between modalities to guide the collaborative optimization of the federated model.
[0011] In one embodiment of the present invention, the global model aggregation module integrates a momentum acceleration mechanism when implementing the federated averaging algorithm, smoothes random disturbances in the local node gradient update process by maintaining the parameter update direction in the form of a sliding average, and combines the exponentially decaying average of the historical update vector when aggregating the obfuscated gradients uploaded by each participating node. This mechanism can not only accelerate the convergence speed of the cross-enterprise joint training process, but also effectively suppress parameter fluctuations caused by privacy protection noise injection. The global model parameters that have been optimized by momentum must pass an integrity check before being distributed to each participating node to ensure the validity of the model update.
[0012] In one embodiment of the present invention, the privacy protection verification module adopts a composite noise injection strategy to simultaneously apply distributed noise with heavy-tail characteristics and Gaussian noise that obeys the exponential decay law in the gradient obfuscation stage, wherein the heavy-tail noise provides a basic privacy protection layer against gradient inversion attacks, and Gaussian noise is used to fill the correlation loopholes between feature dimensions. The two noises are coupled in the form of a Hadamard product and applied to the model parameter update. The composite noise mechanism establishes a multi-level privacy protection system while ensuring the availability of the model. The processed obfuscated gradient signal carries the noise distribution feature metadata when it is transmitted back to each participating node through a security protocol for local model recovery.
[0013] In one embodiment of the present invention, the local federated learning module implements an encrypted transmission mechanism based on random linear transformation. Before uploading the time series embedding vector and the text embedding vector to the central coordination node, a dynamically generated random orthogonal matrix is used to perform a reversible linear transformation on the embedding vector, and a random bias vector that conforms to a specific statistical distribution is superimposed. The encrypted transmission mechanism hides the numerical features of the original embedding vector while maintaining the relative distance relationship between the vectors, ensuring that the comparative learning alignment module can correctly calculate the cross-modal sample similarity. The central coordination node can only operate the encrypted feature space and cannot parse the semantic information of the original geological data.
[0014] In one embodiment of the present invention, the text processing module integrates an adversarial training framework in the terminology standardization stage, and improves the robustness of the model by constructing an adversarial game mechanism between a terminology discriminator and a noise generator. The noise generator is responsible for constructing confusing words that are similar in morphology to real geological terms, and the terminology discriminator learns to distinguish real terms from generated noise. In this process, the text encoder is forced to strengthen the semantic representation ability of lithology description keywords, so that the generated text feature vector can effectively filter out expression ambiguity and annotation noise in the logging interpretation report, thereby improving the mapping accuracy of professional terms and geophysical features in the cross-modal alignment process.
[0015] In one embodiment of the present invention, the BERT model of the embedding generation module adopts a context-aware dynamic mask pre-training strategy, dynamically adjusts the mask probability according to the part-of-speech tagging and document frequency of the terms in the model pre-training stage, implements high-probability masking on the core terms characterizing the lithology and reservoir properties to enhance the contextual reasoning ability of the model, and captures the dependency relationship between the mask position and the surrounding descriptive statements through the self-attention mechanism when processing the geological report text. The generated text embedding vector carries implicit professional semantic association information, and can establish a fine-grained mapping relationship between the physical characteristics of the logging curve and the professional description of the geological text when cross-modally aligned with the time series embedding vector. The multimodal large-model federated training platform and heterogeneous data alignment algorithm provided by the present invention build a vertical federated learning architecture. Each oil and gas enterprise locally deploys a time series data processing module and a text processing module to extract time-frequency features from the logging curve by wavelet packet decomposition, and standardizes the terminology and enhances the semantics of the geological report; the embedding generation module uses a convolutional neural network and a BERT model to map the two types of data to an embedding space of uniform dimension; the central coordination node aligns the cross-modal embedding vector through a dynamic weighted contrast loss function, and introduces a gradient confusion mechanism of Weibull distribution noise to achieve privacy protection in the aggregation stage of the federated model. This method enables the logging data and text data of different companies to complete multimodal knowledge fusion through encrypted embedding vector interaction without leaving the local area, while resisting gradient inversion attacks. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0017] Figure 1 This is the system architecture diagram of the multimodal large model federated training center and heterogeneous data alignment algorithm. DETAILED DESCRIPTION
[0018] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0019] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.
[0020] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0021] See also Figure 1 , which shows the multimodal large model federated training middle station and heterogeneous data alignment algorithm of the present invention. The multimodal large model federated training middle station and heterogeneous data alignment algorithm of the present invention include a time series data processing module, a text processing module, an embedding generation module, a local federated learning module, a contrastive learning alignment module, a global model aggregation module and a privacy protection verification module. The time series data processing module performs wavelet packet decomposition on the logging curve to generate a time-frequency matrix signal; the text processing module performs terminology standardization processing on the geological report and outputs a text feature signal; the embedding generation module processes the time-frequency matrix signal through a convolutional neural network and generates a time series embedding vector, and the embedding generation module processes the text feature signal through the BERT model (an open source machine learning framework) and generates a text embedding vector; the local federated learning module receives and encrypts the time series embedding vector and the text embedding vector to form a local encrypted signal; the contrastive learning alignment module calculates the dynamic weighted contrast loss according to the local encrypted signal and generates a gradient signal; the global model aggregation module processes the gradient signal through a federated average algorithm and forms a parameter update amount; the privacy protection verification module generates a confusion gradient signal for the parameter update amount that conforms to the Weibull distribution noise and returns it to each module to align the data.
[0022] like Figure 1As shown, the present invention relates to a multimodal large model federated training platform and a heterogeneous data alignment algorithm, which is characterized by realizing cross-enterprise privacy security and multimodal data fusion in the field of oil and gas exploration through a modular architecture design, and specifically includes the following core modules and their collaborative operation mechanisms: First, the time series data processing module is responsible for multi-scale feature extraction of the logging curves stored locally by the oil and gas enterprises, and uses adaptive wavelet packet decomposition technology to perform time-frequency analysis on the original logging signals. By dynamically selecting the optimal wavelet basis function, the signal is multi-layered and decomposed to generate a time-frequency matrix signal containing formation interface reflection waves and lithology change characteristics. The signal is transmitted to the embedding generation module through a data channel; the text processing module is responsible for geological The unstructured text data in the report is standardized, and the semantics of professional terms such as "high-porosity sandstone" and "fault development zone" in the report are enhanced based on the geological ontology library and knowledge graph. The text context dependency is captured using a bidirectional long short-term memory network, and a text feature signal carrying geological semantic features is output. The signal is transmitted to the embedding generation module through the bus; the embedding generation module includes a parallel time encoder and a text encoder. The time encoder uses a multi-scale hole convolutional neural network to process the time-frequency matrix signal, and generates a time embedding vector with the ability to characterize the physical properties of the formation by fusing the output features of the local detail convolution path and the global trend convolution path. The text encoder uses a dynamic mask pre- The trained lightweight BERT model processes text feature signals and extracts semantic embedding vectors of key geological concepts through the self-attention mechanism. The two embedding vectors are input into the local federated learning module after dimension alignment and normalization. The local federated learning module implements an encrypted transmission mechanism based on random linear transformation, performs reversible matrix transformation and random bias superposition on the time series embedding vector and the text embedding vector, generates encrypted local signals and uploads them to the central coordination node to ensure that the original logging data and geological report content cannot be parsed during the transmission process. After receiving the encrypted signals uploaded by each participating node, the contrastive learning alignment module in the central coordination node calculates the cross-modal embedding space through a dynamic weighted contrast loss function. Similarity measurement, introduces a learnable modal weight factor into the loss function, balances the contribution ratio of the time-frequency characteristics of the logging curve and the semantic characteristics of the geological text, constructs a soft comparison target through a normalized exponential function, forces the time series and text embedding vectors of the same sample to be close to each other in the projection space, and pushes away the embedding vectors of different samples. The generated gradient signal is transmitted to the global model aggregation module through a secure link; the global model aggregation module adopts a federated averaging algorithm with an integrated momentum acceleration mechanism to perform a sliding average calculation on the gradient update uploaded by each participating node, and smoothes the parameter fluctuations caused by noise disturbances by accumulating historical gradient directions. The aggregated global model parameters are distributed to the privacy protection verification module after verification;The privacy protection verification module implements a composite noise injection strategy, applies heavy-tailed noise and exponentially decaying Gaussian noise that conform to the Weibull distribution during the model parameter update phase, and couples the statistical characteristics of the two noises through the Hadamard product, while maintaining the model convergence stability while defending against gradient inversion attacks. The processed obfuscated gradient signal is transmitted back to each participating node through a security protocol to complete the privacy protection update of the local model. End-to-end collaboration is achieved between modules through hierarchical data flows: the time series data processing module and the text processing module respectively complete the feature extraction of the original data, the embedding generation module establishes a cross-modal vector mapping relationship, the local federated learning module ensures data transmission security, and the central coordination node coordinates global comparative learning optimization and noise protection, forming a closed-loop federated training system with "data not out of the domain and knowledge shareable", and finally achieving semantic alignment and knowledge fusion of logging curves and geological reports under the premise of privacy protection. When the system is running, the logging curve data is always stored in the local server of the oil and gas enterprise. The central coordination node only processes the encrypted embedding vector and the model parameters with added noise. Through the combined effect of the shape parameter regulation of Weibull noise and the dynamic weighted contrast loss, the problem of the difficult balance between privacy protection and model accuracy in traditional methods is solved, providing a safe and efficient multi-modal collaborative basic platform for cross-enterprise oil and gas exploration intelligent analysis. ;
[0023] Furthermore, the technical implementation and coordination mechanism of the core modules of the multimodal large model federated training platform and heterogeneous data alignment algorithm involved in the present invention are specifically developed as follows: In the time series data processing module, for the high-dimensional time series data of the well logging curve in oil and gas exploration, the adaptive wavelet packet decomposition technology is used to extract multi-scale features, and the spectrum energy distribution characteristics of the signal are dynamically analyzed to construct a multi-layer decomposition tree structure and iteratively select the optimal wavelet basis function. For example, when the high-frequency formation interface reflection wave component appears in the logging signal, the module automatically switches to the Daubechies series wavelet basis to enhance the transient feature capture capability; and for the low-frequency lithology gradual change signal, the Symlets wavelet basis is preferably used to improve the trend tracking accuracy. During the decomposition process, the signal subband components are separated layer by layer through a three-level filtering operation, and finally spliced to form a time-frequency matrix. The row dimension of the matrix corresponds to the number of sampling points, and the column dimension is determined by the number of decomposition layers. It can completely retain the time-frequency localization characteristics of the well logging curve, and transmit it to the input end of the embedded generation module in real time through a high-speed data channel, providing a high signal-to-noise ratio time-frequency feature base for subsequent cross-modal alignment.
[0024] In one embodiment of the present invention, the text processing module designs a set of term enhancement processes based on domain knowledge graphs for the unstructured text data of geological reports. First, a geological ontology library containing core concepts such as lithology classification, structural characteristics, and reservoir parameters is constructed, and a bidirectional attention mechanism is used to identify professional entities in the report, such as mapping "high porosity and permeability sandstone" to the standardized term "high porosity and permeability sandstone reservoir". Subsequently, the enhanced text sequence is contextually semantically modeled through a bidirectional long short-term memory network. The forward layer of the network captures the grammatical dependencies of the terms, and the backward layer extracts the reverse semantic associations. After the outputs of the two are vector-concatenated at each time step, a global text feature vector is generated through average pooling in the time dimension. This vector not only contains the semantic information of key geological concepts such as "fault direction" and "pore structure", but also retains the contextual logical relationship of the description sentence, significantly improving the representation quality of the text features. After normalization, it is input into the text encoder of the embedding generation module.
[0025] like Figure 1 As shown in the figure, the embedding generation module realizes the unified mapping of cross-modal feature space through heterogeneous neural network architecture. The time series encoder uses a multi-scale hole convolution network to process the time-frequency matrix, in which the 3×3 standard convolution kernel is responsible for extracting the local detail features of the formation reflection wave, and the 5×5 hole convolution kernel expands the receptive field with an expansion rate of 2 to capture the macroscopic morphological changes of the logging curve. After the output feature maps of the two convolution paths are spliced in the channel dimension, they are sent to the channel attention mechanism for feature recalibration, and the time series embedding vector with multi-scale perception is generated through global average pooling. The text encoder uses a lightweight BERT model pre-trained with dynamic masking. In the fine-tuning stage, the masking probability is dynamically adjusted according to the part-of-speech tagging of the terms, and high-probability masking is implemented for core parameters such as "oil saturation" and "fracture development index", forcing the model to strengthen the contextual reasoning ability. The final output text embedding vector carries implicit geological semantic association information. The two embedding vectors are dimensionally aligned after L2 normalization, and the random linear transformation encryption mechanism of the local federated learning module is used to form an irreversible encrypted signal uploaded to the central coordination node.
[0026] like Figure 1As shown in the figure, the contrastive learning alignment module innovatively introduces a modal adaptive weight adjustment mechanism. When calculating cross-modal similarity, the contribution ratio of the two is dynamically balanced through the learnable time series modal weight factor and the text modal weight factor. Specifically, when constructing the similarity score of the positive sample pair, the bidirectional correlation strength from time series to text and text to time series is calculated synchronously, and the normalized exponential function is used to construct the soft contrast target. This mechanism effectively alleviates the problem of embedding space distortion caused by the dimensional richness of logging data and the sparsity of geological text. For example, in the semantic alignment process of "acoustic time difference anomaly" and "carbonate dissolution pore development", the text modal weight is dynamically increased to compensate for the abstractness of professional terms. When the generated gradient signal is transmitted to the global model aggregation module through an encrypted channel, it carries the weight distribution information between modalities to guide the collaborative optimization of the federated model.
[0027] Furthermore, the global model aggregation module integrates a momentum acceleration mechanism when implementing the federated averaging algorithm. By maintaining the exponentially decaying average of the historical gradient update vector, the parameter fluctuations caused by the privacy protection noise injection are smoothed. In the specific implementation, the obfuscated gradients of each participating node are first weighted and fused with the historical gradients in the momentum buffer. The weight coefficient is dynamically adjusted according to the model training stage. A higher momentum factor is set at the beginning of training to accelerate convergence, and then gradually reduced in the later stage to improve parameter stability. The aggregated global model parameters need to be hashed to ensure integrity to prevent man-in-the-middle attacks from tampering with the model update amount. This mechanism enables the joint training process to increase the model convergence speed by about 35% while ensuring privacy security, especially when dealing with large differences in data distribution across enterprises, effectively avoiding the problem of local model divergence. After the updated global parameters are processed by the composite noise injection of the privacy protection verification module, they are transmitted back to each participating node through the security protocol to complete the federated learning closed loop.
[0028] like Figure 1As shown, the multimodal large model federated training middle station and heterogeneous data alignment algorithm involved in the present invention are further refined as follows: In the privacy protection verification module, a composite noise injection strategy is used to build a multi-level protection system, and a noise generation mechanism of a mixture of Weibull distribution and Gaussian distribution is designed for the diversity characteristics of gradient inversion attacks in federated learning scenarios. Weibull noise presents a heavy-tailed characteristic through the regulation of shape parameters, imposes small perturbations in most gradient dimensions to ensure model convergence, and randomly generates large noise in key feature dimensions to destroy the possibility of attackers reconstructing data; Gaussian noise adopts an exponential decay variance design, and gradually reduces the noise intensity as the number of training rounds increases, providing strong privacy protection in the early stage of training, and reducing the impact on model accuracy in the later stage. The two types of noise are nonlinearly coupled through the Hadamard product, so that the noise distribution retains the defensive advantages of Weibull noise and has the smooth characteristics of Gaussian noise. For example, when processing gradient parameters related to the spectral characteristics of logging curves, Weibull noise focuses on interfering with the gradient values corresponding to high-frequency components, while Gaussian noise uniformly covers all dimensions. Before the obfuscated gradient signal is transmitted back to the participating nodes, the noise distribution metadata is attached for the local model to compensate for the noise, ensuring the consistency of the global model update. This module forms a closed-loop collaboration with the global model aggregation module, dynamically adjusts the noise injection ratio in each round of federated training, and balances security and availability through the privacy budget consumption monitoring mechanism.
[0029] Furthermore, the encrypted transmission mechanism of the local federated learning module adopts a design that combines random linear transformation with dynamic key management, and implements multi-level obfuscation for time series embedding vectors and text embedding vectors. Before uploading the embedding vector, the module generates a random orthogonal transformation matrix, rotates and projects the embedding vector through matrix multiplication, and superimposes a bias vector that conforms to the uniform distribution. The strict orthogonality between the column vectors of the orthogonal matrix ensures that the transformed vector space maintains the original relative distance, so that the comparative learning calculation of the central coordination node is not affected by encryption. For example, the error between the cosine similarity of the encrypted embedding vectors of two similar logging samples and the original value does not exceed the set threshold. The transformation matrix and the bias vector are dynamically updated through the key distribution service, and each training batch uses an independent key. Even if a single batch of keys is leaked, the historical communication data cannot be cracked. This mechanism forms a complementary protection with the noise injection of the privacy protection verification module: the former protects the intermediate features of the data transmission process, and the latter protects against parameter leakage in the model update phase, and jointly builds an end-to-end privacy barrier. The adversarial training framework integrated in the text processing module improves the robustness of terminology standardization through the generative adversarial network architecture, including the dynamic game process between the term generator and the discriminator. The generator receives the latent space noise vector and contextual semantic features, and outputs realistically shaped confusing terms, such as mutating "muddy siltstone" into "silty mudstone"; the discriminator analyzes the grammatical rationality and semantic consistency of the terms based on the bidirectional attention mechanism, and identifies subtle anomalies in the generated terms. In this adversarial process, the text encoder is forced to learn more discriminative feature representations, such as establishing sensitivity to the differences in mineral composition between "breccia" and "conglomerate". At the same time, the adversarial samples generated by the generator are used for data enhancement to expand the boundary cases in the training set, so that the model can still accurately extract semantic features when dealing with handwriting recognition errors or ambiguous terms in actual exploration reports. This mechanism works in conjunction with the knowledge graph enhancement module, the former improves the model's anti-interference ability, and the latter ensures the standardized mapping of professional terms, jointly optimizing the information density and noise resistance of the text feature vector.
[0030] like Figure 1As shown in the figure, the dynamic mask pre-training strategy of the BERT model in the embedding generation module implements differentiated processing according to the linguistic characteristics and domain importance of geological terms. In the pre-training stage, the module analyzes the part-of-speech tagging results and term frequency-inverse document frequency weights of the input text in real time, and implements a masking operation with increasing probability for noun-based professional terms (such as "porosity" and "permeability"), while maintaining a low masking rate for conjunctions and conventional description sentences. For example, when processing "reservoir porosity is between 15% and 20%", the system implements a high-probability mask on "porosity", forcing the model to infer the masked content through the contextual numerical range and unit. At the same time, the partial masking technology is used to randomly mask some characters instead of the complete vocabulary for long terms (such as "carbonate rock dissolution pores"), strengthening the model's understanding of the internal structure of the term. The BERT model optimized by dynamic masking shows stronger geological semantic reasoning ability in the fine-tuning stage. The generated text embedding vector can establish cross-modal association patterns such as "high resistivity" and "tight sandstone", "low acoustic time difference" and "fracture development", and significantly improve the efficiency of comparative learning alignment. The data flow and control flow between modules are finely coordinated through the intelligent scheduling engine. Before injecting noise, the privacy protection verification module obtains the convergence status indicator of the current training stage from the global model aggregation module, and dynamically adjusts the shape parameter of the Weibull noise and the variance decay rate of the Gaussian noise. For example, the noise intensity is reduced to accelerate convergence in the stage where the model loss decreases rapidly, and the noise ratio is increased in the loss plateau period to strengthen privacy protection. The key management service of the local federated learning module establishes a secure handshake protocol with the contrastive learning module of the central coordination node to ensure the orthogonality check of the encryption matrix and the synchronous update of the key. After each round of iteration, the adversarial training framework of the text processing module feeds back the adversarial sample features output by the generator to the embedding generation module for updating the dynamic masking strategy of the BERT model, forming a positive cycle of adversarial data enhancement and feature learning. When the system is running, the time series data processing module, text processing module and embedding generation module of each participating node form a local feature extraction pipeline. The central coordination node aggregates cross-node knowledge through the federated averaging algorithm, and then forms global knowledge feedback after processing by the privacy protection verification module. This operating paradigm of "decentralized feature extraction-centralized knowledge fusion-secure knowledge distribution" enables multimodal semantic association and joint model optimization of logging curves and geological reports of different companies in the oil and gas exploration field in a completely physically isolated environment, thus overcoming the industry difficulties of coexistence of data silos and privacy risks in traditional methods.
[0031] The multimodal large-model federated training platform and heterogeneous data alignment algorithm of the present invention construct a vertical federated learning architecture. Each oil and gas enterprise locally deploys a time series data processing module and a text processing module to perform wavelet packet decomposition on the logging curve to extract time-frequency features, and perform terminology standardization and semantic enhancement on the geological report; the embedding generation module uses a convolutional neural network and a BERT model to map the two types of data to an embedding space of uniform dimension; the central coordination node aligns the cross-modal embedding vector through a dynamic weighted contrast loss function, and introduces a gradient confusion mechanism of Weibull distribution noise to achieve privacy protection in the aggregation stage of the federated model. This method enables the logging data and text data of different companies to complete multimodal knowledge fusion through encrypted embedding vector interaction without leaving the local area, while resisting gradient inversion attacks.
[0032] Therefore, the multimodal large model federated training platform and heterogeneous data alignment algorithm of the present invention can solve the problem that well logging curves and geological reports are difficult to effectively align and fuse due to modal differences and semantic gaps.
[0033] The above embodiments are merely illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Anyone familiar with the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by a person of ordinary skill in the art without departing from the spirit and technical concept disclosed by the present invention shall still be covered by the claims of the present invention.
Claims
1. Multimodal large model federated training platform and heterogeneous data alignment algorithm, characterized by: include: A time series data processing module, which performs wavelet packet decomposition on the logging curve to generate a time-frequency matrix signal; A text processing module, which performs terminology standardization processing on the geological report and outputs text feature signals; An embedding generation module, wherein the embedding generation module processes the time-frequency matrix signal through a convolutional neural network and generates a time series embedding vector, and the embedding generation module processes the text feature signal through a BERT model and generates a text embedding vector; A local federated learning module, wherein the local federated learning module receives and encrypts the time series embedding vector and the text embedding vector to form a local encrypted signal; A contrastive learning alignment module, wherein the contrastive learning alignment module calculates a dynamic weighted contrast loss according to the local encrypted signal and generates a gradient signal; A global model aggregation module, wherein the global model aggregation module processes the gradient signal through a federated averaging algorithm and forms a parameter update amount; A privacy protection verification module is provided, wherein the parameter update amount conforms to the noise of the Weibull distribution, and a confusion gradient signal is generated and transmitted back to each module to align the data.
2. The multimodal large model federated training platform and heterogeneous data alignment algorithm according to claim 1 is characterized in that: The wavelet packet decomposition process of the time series data processing module includes an adaptive energy threshold screening mechanism, which suppresses the high-frequency noise in the logging curve by dynamically adjusting the decomposition layer number and bandwidth allocation of the wavelet basis function. The core algorithm is: Where W is the set of candidate wavelet basis functions, F represents the Fourier transform operator, DWT is the discrete wavelet transform operation, and K is the number of decomposition layers. The algorithm selects the optimal basis function by maximizing the energy concentration of the sub-band components. The signal-to-noise ratio of the formation interface reflection wave characteristics in the generated time-frequency matrix signal is improved by more than 40% and transmitted to the embedding generation module through the data channel.
3. The multimodal large model federated training platform and heterogeneous data alignment algorithm according to claim 1 is characterized in that: The terminology standardization processing of the text processing module adopts a semantic enhancement method based on knowledge graph. By constructing an entity relationship network in the geological field, the synonyms and abbreviations in the original text are mapped to a standardized terminology library. The vectorization process satisfies: in is the original word segmentation result, KG represents the associated entity embedding extracted from the knowledge graph, ⊕ is the vector concatenation operation, MLP is the multi-layer perceptron. This method effectively solves the semantic ambiguity problem of terms, and the generated text feature signal is transmitted to the embedding generation module through the bus.
4. The multimodal large model federated training platform and heterogeneous data alignment algorithm according to claim 1 is characterized in that: The convolutional neural network of the embedding generation module adopts a multi-scale feature fusion architecture, and works together through a parallel standard convolution path and a hole convolution path. The standard convolution kernel is responsible for capturing the local detail features of the formation reflection wave in the time-frequency matrix, and the hole convolution kernel extracts the macro trend features of the logging curve by expanding the receptive field. The output feature maps of the two convolution paths are spliced in the channel dimension and sent to the global pooling layer to generate a time series embedding vector with multi-scale perception capabilities. After normalization, the vector is input into the local federated learning module together with the text embedding vector, and the cross-modal association of the time-frequency characteristics of the logging signal and the semantic information of the geological text is realized through the feature layer fusion mechanism.
5. The multimodal large model federated training platform and heterogeneous data alignment algorithm according to claim 1 is characterized in that: The contrastive learning alignment module introduces a modal adaptive weight adjustment mechanism, which dynamically balances the contribution ratio in the cross-modal similarity calculation through the learnable time series modal weight factor and the text modal weight factor, and simultaneously considers the bidirectional correlation strength from time series to text and from text to time series when calculating the similarity score of the positive sample pair. A soft contrast target is constructed through a normalized exponential function, which effectively alleviates the embedding space distortion problem caused by the difference in dimensional richness of logging data and sparsity of geological report text. When the generated gradient signal is transmitted to the global model aggregation module through an encrypted channel, the weight distribution information between modalities is retained to guide the collaborative optimization of the federated model.
6. The multimodal large model federated training platform and heterogeneous data alignment algorithm according to claim 1 is characterized in that: The global model aggregation module integrates a momentum acceleration mechanism when implementing the federated averaging algorithm, smoothes random disturbances in the local node gradient update process by maintaining the parameter update direction in the form of a sliding average, and combines the exponentially decaying average of the historical update vector when aggregating the obfuscated gradients uploaded by each participating node. This mechanism can not only accelerate the convergence speed of the cross-enterprise joint training process, but also effectively suppress parameter fluctuations caused by privacy protection noise injection. The global model parameters that have been optimized by momentum must pass an integrity check before being distributed to each participating node to ensure the effectiveness of the model update.
7. The multimodal large model federated training platform and heterogeneous data alignment algorithm according to claim 6 is characterized in that: The privacy protection verification module adopts a composite noise injection strategy, and simultaneously applies distributed noise with heavy-tail characteristics and Gaussian noise that obeys the exponential decay law in the gradient confusion stage. The heavy-tail noise provides a basic privacy protection layer against gradient inversion attacks, and Gaussian noise is used to fill the correlation loopholes between feature dimensions. The two noises are coupled in the form of a Hadamard product and applied to the model parameter update. The composite noise mechanism establishes a multi-level privacy protection system while ensuring the availability of the model. The processed confused gradient signal carries the noise distribution feature metadata when it is transmitted back to each participating node through a security protocol for local model recovery.
8. The multimodal large model federated training platform and heterogeneous data alignment algorithm according to claim 1 is characterized in that: The local federated learning module implements an encrypted transmission mechanism based on random linear transformation. Before uploading the time series embedding vector and the text embedding vector to the central coordination node, a dynamically generated random orthogonal matrix is used to perform a reversible linear transformation on the embedding vector, and a random bias vector that conforms to a specific statistical distribution is superimposed. The encrypted transmission mechanism hides the numerical features of the original embedding vector while maintaining the relative distance relationship between the vectors, ensuring that the contrastive learning alignment module can correctly calculate the cross-modal sample similarity. The central coordination node can only operate the encrypted feature space and cannot parse the semantic information of the original geological data.
9. The multimodal large model federated training platform and heterogeneous data alignment algorithm according to claim 1, characterized in that: The text processing module integrates an adversarial training framework in the terminology standardization stage, and improves the robustness of the model by constructing an adversarial game mechanism between a terminology discriminator and a noise generator. The noise generator is responsible for constructing confusing words that are similar in morphology to real geological terms, and the terminology discriminator learns to distinguish real terms from generated noise. In this process, the text encoder is forced to strengthen the semantic representation ability of lithology description keywords, so that the generated text feature vector can effectively filter out expression ambiguity and annotation noise in the logging interpretation report, and improve the mapping accuracy of professional terms and geophysical features in the cross-modal alignment process.
10. The multimodal large model federated training platform and heterogeneous data alignment algorithm according to claim 1, characterized in that: The BERT model of the embedding generation module adopts a context-aware dynamic mask pre-training strategy. In the model pre-training stage, the mask probability is dynamically adjusted according to the part-of-speech tagging and document frequency of the terms, and high-probability masking is implemented on the core terms representing the lithology and reservoir properties to enhance the contextual reasoning ability of the model. When processing geological report texts, the dependency relationship between the mask position and the surrounding description sentences is captured through the self-attention mechanism. The generated text embedding vector carries implicit professional semantic association information, and can establish a fine-grained mapping relationship between the physical characteristics of the logging curve and the professional description of the geological text when cross-modally aligned with the time series embedding vector.
Citation Information
Patent Citations
Safe and efficient vehicle voice recognition method and system based on federal learning
CN116564339A
RPA service processing method and system based on artificial intelligence
CN118887044A
Multi-source heterogeneous data fusion and processing method based on big data
CN119783037A
Federated large model adaptive learning system
US20250103952A1
Cited By
AI and password technology fused identity authentication method
CN120301605A
Safety evaluation method, device and equipment for multi-modal large model and storage medium
CN120321041A
Federal cross-modal retrieval method and system based on interaction prompt
CN120723920A
A federated cross-modal retrieval method and system based on interactive prompts
CN120723920B
Power big data privacy protection method and system based on federated learning
CN120822242A