Microscopic residual oil intelligent identification and quantification method

CN122414499BActive Publication Date: 2026-09-15SOUTHWEST PETROLEUM UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610873500.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-15
Estimated Expiration
2046-06-17

AI Technical Summary

Technical Problem

[0007]针对现有技术中的上述不足,本发明提供的一种微观剩余油智能识别及量化方法,解决了微观剩余油分析中“识别-建模-模拟-决策”流程割裂、自动化程度低、结果不可量化的问题,通过多模态感知、智能体协同与物理模拟决策的端到端闭环,实现了从数据到优化策略的全自动、定量化智能分析

Benefits of technology

本发明首先通过领域知识增强的多模态金字塔池化视觉Transformer(MPP-ViT)模型,对岩心微观图像与关联地质文本进行深度跨模态融合与理解,通过生成语义分割图,实现对剩余油赋存状态的“认知级”语义化识别;然后,将这一富含语义的识别结果,作为总控指令,动态调度一个由多个专业化智能体构成的协同工作流,这些智能体各司其职,依次完成高精度定量计算与物理过程模拟,并最终通过优化决策模块,生成针对性的提高采收率策略建议,整个过程实现了从原始数据到决策知识的端到端自动化流转与优化,本发明至少具有以下有益效果:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122414499B_ABST
    Figure CN122414499B_ABST
Patent Text Reader

Abstract

The application provides a micro residual oil intelligent identification and quantification method, and belongs to the field of micro residual oil analysis. The method comprises the following steps: preprocessing a micro CT image of a visual mode and a geological description text of a text mode, and inputting the micro CT image and the geological description text into a multi-modal pyramid pooling visual model to perform deep cross-modal fusion and semantic identification, so as to generate a high-precision residual oil semantic segmentation graph; inputting the semantic segmentation graph into a task scheduler, performing seepage simulation and potential prediction through dynamic scheduling and in cooperation with an intelligent agent, obtaining a dynamic simulation result, and based on a current residual oil state, utilizing a decision intelligent agent to perform closed-loop optimization to obtain an optimal strategy and a verification simulation result; and based on the semantic segmentation graph, the dynamic simulation result, the optimal strategy and the verification simulation result, generating a structured report. The application solves the problems of a split identification-modeling-simulation-decision process, low automation degree and unquantifiable results in micro residual oil analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of microscopic residual oil analysis, and particularly relates to a method for intelligent identification and quantification of microscopic residual oil. Background Technology

[0002] In the mid-to-late stages of oil and gas field development, especially during the high water-cut period, crude oil in underground reservoirs is often retained as dispersed "residual oil" within complex rock pore networks. Accurately identifying the microscopic state, quantity, and spatial distribution of this residual oil is crucial geological evidence for assessing reservoir potential and developing precise enhanced oil recovery strategies. With the widespread adoption of high-precision imaging technologies such as micro-CT and focused ion beam scanning electron microscopy, digital core analysis has become a core tool for pore-scale research. However, how to automatically, accurately, and intelligently extract residual oil information with clear geological significance from the massive amounts of complex microscopic images acquired and directly use it to support development decisions remains a long-standing technical challenge in this field.

[0003] Currently, the technical approach commonly used in the industry follows a linear process of "imaging, processing, and simulation," relying heavily on traditional digital image processing methods. For example, widely used thresholding methods (such as the Otsu method) distinguish rock skeletons, pore spaces, and fluids by setting grayscale thresholds. However, in microscopic images, the grayscale values ​​of oil and water phases often overlap significantly and are unevenly distributed due to mineral composition interference and imaging limitations, making it difficult to accurately segment with a single threshold, especially in reservoirs with complex wettability. While edge detection and watershed algorithms have made some improvements, they are extremely sensitive to image noise and easily produce a large number of artifacts or missegmentation in complex pore structures. The fundamental limitation of these traditional methods is that they only perform mathematical operations based on pixel grayscale or simple textures, representing a shallow "what you see is what you get" visual processing that completely fails to understand the geological and physical meaning behind the image content. The results are highly dependent on the operator's experience in parameter adjustment, resulting in low automation and reproducibility.

[0004] To improve recognition accuracy, deep learning techniques, especially convolutional neural networks, have been introduced into this field. Semantic segmentation models, such as U-Net, can automatically learn features from large amounts of labeled data to achieve end-to-end pore and fluid segmentation, with significantly improved accuracy compared to traditional methods. However, existing deep learning-based methods are essentially still purely "visual models," with input limited to image pixel data. During training and inference, these models cannot acquire and utilize geological descriptive textual information closely related to the core sample (such as key semantics like "strongly water-wet sandstone" and "post-polymer flooding"). This makes the model a "black box" lacking domain knowledge, and its generalization ability is limited by the coverage of the training set, leading to performance instability when encountering new lithologies or development conditions. More importantly, its output is only pixel labels of "category A" or "category B," failing to provide interpretable geological origin inferences such as "this is corner oil retained due to wettability," resulting in a disconnect between "recognition" and "geological understanding."

[0005] After obtaining the static distribution of fluids, researchers typically need to use microscopic flow simulation techniques such as pore network models or lattice Boltzmann methods to predict the dynamic utilization potential of remaining oil under different development measures. However, in existing technology systems, the image recognition and segmentation module and the physical numerical simulation module operate independently and sequentially. The entire process is unidirectional and open-loop: manual or AI models complete the segmentation, and the result file is imported into simulation software for calculation. Uncertainties or errors discovered during the simulation process cannot be automatically fed back to guide the correction and optimization of the front-end model. Furthermore, current simulations mainly perform "computation" and "prediction" functions, that is, given a preset development plan (such as injecting a certain chemical agent), calculate its effect. However, existing technologies lack a built-in, adaptive, and efficient decision-making mechanism to automatically and intelligently search for the optimal development strategy for the current specific remaining oil distribution from thousands of possible combinations of plans. They still rely heavily on researchers' experience and repetitive manual trial and error, resulting in an efficiency bottleneck in the chain from microscopic analysis to macroscopic decision-making.

[0006] In recent years, multimodal models have made groundbreaking progress in cross-modal understanding and generation tasks, demonstrating a powerful ability to integrate visual and textual information. Meanwhile, intelligent agent technology has provided a new paradigm for achieving autonomous planning and executing complex task sequences. However, how to deeply integrate these cutting-edge artificial intelligence technologies with microscopic seepage physics and petroleum geology to construct an intelligent system capable of mimicking the complete cognitive process of domain experts—from "observation-understanding-analysis-decision-making"—and thus provide a one-stop solution to the entire process from intelligent perception of microscopic images to development strategy optimization suggestions, remains a challenge that has not yet been publicly and effectively implemented in existing technical solutions. Summary of the Invention

[0007] To address the aforementioned shortcomings in existing technologies, this invention provides a method for intelligent identification and quantification of microscopic residual oil. This method solves the problems of fragmented "identification-modeling-simulation-decision" processes, low automation, and unquantifiable results in microscopic residual oil analysis. Through an end-to-end closed loop of multimodal perception, agent collaboration, and physical simulation decision-making, it achieves fully automated and quantitative intelligent analysis from data to optimization strategies.

[0008] To achieve the above objectives, the technical solution adopted by this invention is: a method for intelligent identification and quantification of microscopic residual oil, comprising the following steps: S1. Obtain visual and textual modalities, preprocess them, and generate semantic segmentation maps based on the preprocessed image-text pair data using a multimodal pyramid pooling visual model that performs cross-modal fusion and semantic recognition. S2. Input the semantic segmentation graph into the task scheduler, and perform seepage simulation and potential prediction by dynamically scheduling and coordinating the intelligent agents to obtain dynamic simulation results. Based on the current remaining oil state, use the decision intelligent agent to find the optimal strategy and verify the simulation results. S3. Based on semantic segmentation graphs, dynamic simulation results, optimal strategies, and their verification simulation results, a structured report is generated to complete the intelligent identification and quantification of microscopic residual oil.

[0009] Further, S1 includes the following steps: The visual modality and the corresponding unstructured geological description text are obtained; whereby the unstructured geological description text is defined as the original text string. ; Based on the acquired visual modality, the image Preprocessing is performed, in which, The first microscopic image representing the target rock core i Zhang two-dimensional slice image; From the original text string Key geological parameters are extracted to form a structured representation. and structured representation Transform into a coherent text description ; Based on image Preprocessing structure and text description To obtain image-text pair data ,in, This represents the image after preprocessing. Based on image-text pair data A semantic segmentation map is generated by using a multimodal pyramid pooling visual model for cross-modal fusion and semantic recognition. .

[0010] Furthermore, the multimodal pyramid pooling visual model includes: Visual encoder Used to transfer images The image is divided into N image patches, linearly projected and positionally encoded, and then output as a visual feature sequence through an L-layer encoder. Based on visual feature sequences By utilizing the pyramid pooling module to perform pooling and fusion at multiple scales in parallel, visual features containing multi-scale context are output. ; Text encoder Used for text-based descriptions The global text semantic vector is obtained using a text encoder. ; Cross-modal attention fusion module Used for visual feature-based and text features The text-guided attention is calculated, and the attention calculation result is fused with the original visual features to output deep fusion features. ; The output layer is used to incorporate deep fused features during the inference phase. The data is reconstructed into a spatial feature map, which is then input into the segmentation decoder. By calculating the probability map of each pixel belonging to each semantic category, the decoder outputs a pixel-level semantic segmentation map. And the confidence score of each pixel belonging to each category.

[0011] Furthermore, S2 includes the following steps: Obtain semantic segmentation map With images And based on the semantic segmentation graph The set of categories defined in For intelligent agents Generate task list ,in, Indicates the target category. Indicates origin from semantic segmentation graph Category coarse binary mask, This represents the preprocessed image. Indicates the first One task; For each task Apply a coarse binary mask to the category Alternatively, bounding boxes can be used as spatial cues input into the visual base model. ; Based on the visual fundamental model Using spatial cues in images Pixel-level retouching is performed within the context to obtain a binary mask. ,in, This represents the number of vertical pixels in an image with a fixed size. This indicates the number of pixels horizontally in an image with a fixed size. Based on the refined binary mask set Using intelligent agents Automatic calculation of structured quantization parameter set Output a refined segmentation image. ; Based on refined segmentation map Reconstructing three-dimensional binary digital cores for simulation ; Given initial conditions And boundary conditions, using a lightweight proxy simulator Predict the evolution of the saturation field, where, Indicates the initial saturation distribution. Indicates time; Based on predicted saturation field evolution and reconstructed 3D binary digital core Lightweight agent simulator Perform scenario simulation and output the curve of average remaining oil saturation changing over time. Based on the curve of average residual oil saturation over time... The degree of extraction was calculated. ; Using intelligent agents Output dynamic simulation results Among them, dynamic simulation results Including extraction degree Curve of average residual oil saturation over time And the remaining oil distribution map at the final moment; From the structured quantization parameter set Extracting a low-dimensional vector that can characterize the current remaining oil. Define the state, where, Represents the set of real numbers. Represents the dimension of the state vector; Adjustable injection strategy parameters Defined as an action, where, Represents the action space; Given the current state and injection strategy Based on the environmental transition to a new state And generate a reward, where the reward function is... The expression is as follows:

[0012] in, Indicates the incremental level of extraction at each stage. Indicates the residual oil saturation at the end of the stage. Indicates the cost of the measures, , and All represent weights; Based on state Injection strategy And rewards, to build a reinforcement learning framework; intelligent agents As an intelligent agent In an interactive environment within a reinforcement learning framework, rewards are maximized through interactive trial and error. To train the policy network ,in, This represents the mathematical expectation of the cumulative reward. Represents the trajectory. Indicates the time step index; For the new core condition Utilizing the trained policy network Generate the optimal strategy ,in, Indicates the injection of policy variables; The generated optimal strategy Feedback to the intelligent agent Perform a final high-fidelity verification simulation to obtain the verification simulation results. .

[0013] Furthermore, the expression for the evolution of the saturation field is as follows:

[0014] in, Represents the remaining oil saturation field. Represents a pressure field. Indicates the proxy model parameters. express The remaining oil saturation field at any given time, express Constant pressure field express The remaining oil saturation field at any given time, express Constant pressure field Indicates the time step. This indicates the injection strategy.

[0015] Furthermore, the degree of extraction The expression is as follows:

[0016]

[0017] in, This represents the average remaining oil saturation at the initial moment. This represents the average remaining oil saturation at the final moment. Represents the spatial domain of porous media. x Represents spatial coordinates. Indicates spatial location x ,time The remaining oil saturation at that location.

[0018] Furthermore, step S3 includes the following steps: Based on semantic segmentation graph Confidence plots for each category, and structured quantization parameter sets. Refined Segmentation Diagram Dynamic simulation results Optimal Strategy and its verification simulation results To integrate information from multiple sources; The report generation engine has a built-in domain knowledge rule base. Among them, the domain knowledge rule base Used to automatically perform correlation analysis and generate conclusions; Based on integrated multi-source information, and utilizing a domain knowledge rule base Generate association rule-driven knowledge and combine it with verification simulation results. Conduct an assessment; Based on the evaluation results, according to predefined document modules The system automatically assembles and generates a final structured report, completing the intelligent identification and quantification of microscopic residual oil.

[0019] The beneficial effects of this invention are: This invention first utilizes a domain-knowledge-enhanced multimodal pyramid pooled visual Transformer (MPP-ViT) model to perform deep cross-modal fusion and understanding of core microscopic images and associated geological text. By generating a semantic segmentation map, it achieves "cognitive-level" semantic recognition of the remaining oil occurrence status. Then, this semantically rich recognition result is used as a central control command to dynamically schedule a collaborative workflow composed of multiple specialized intelligent agents. These agents each perform their respective functions, sequentially completing high-precision quantitative calculations and physical process simulations. Finally, through an optimized decision-making module, targeted suggestions for improving oil recovery strategies are generated. The entire process achieves end-to-end automated flow and optimization from raw data to decision knowledge. This invention has at least the following beneficial effects: Improving recognition accuracy and interpretability: Multimodal fusion enables the model to distinguish visually similar fluids based on geological semantics and output labels with geological origins, solving the problems of high misjudgment rate and uninterpretable results of traditional methods; Achieve end-to-end automation and efficiency leap: Through intelligent agent collaboration, identification, quantification, simulation, and optimization are linked into an automated pipeline, and intelligent agents are introduced to realize a closed loop from analysis to autonomous optimization, which greatly improves efficiency; Provides direct and quantitative optimal decision support: Through reinforcement learning closed-loop automatic optimization, it generates targeted development strategies and quantitative benefit assessments, directly transforming micro-analysis into the basis for engineering decisions. Attached Figure Description

[0020] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0021] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0022] To address the aforementioned technical problems, this invention proposes a novel and systematic technical solution. Its core lies in constructing a microscopic residual oil analysis paradigm characterized by "multimodal intelligent perception-driven, multi-agent collaborative execution." This solution breaks away from the traditional linear and isolated process, forming an intelligent closed-loop system with perception, cognition, decision-making, and feedback capabilities. Based on the above overall concept, the complete technical solution of this invention is specifically implemented through the following methods, such as... Figure 1 As shown, this invention provides a method for intelligent identification and quantification of microscopic residual oil, the implementation of which is as follows: S1. Obtain the visual and textual modalities, preprocess them, and based on the preprocessed image-text pair data, use a multimodal pyramid pooling visual model to perform cross-modal fusion and semantic recognition to generate a semantic segmentation map. The implementation method is as follows: The visual modality and the corresponding unstructured geological description text are obtained; whereby the unstructured geological description text is defined as the original text string. ; Based on the acquired visual modality, the image Preprocessing is performed, in which, The first microscopic image representing the target rock core i Zhang two-dimensional slice image; From the original text string Key geological parameters are extracted to form a structured representation. and structured representation Transform into a coherent text description ; Based on image Preprocessing structure and text description To obtain image-text pair data ,in, This represents the image after preprocessing. Based on image-text pair data A semantic segmentation map is generated by using a multimodal pyramid pooling visual model for cross-modal fusion and semantic recognition. ; Multimodal pyramid pooling visual models include: The multimodal pyramid pooling visual model includes: Visual encoder Used to transfer images The image is divided into N image patches, linearly projected and positionally encoded, and then output as a visual feature sequence through an L-layer encoder. Based on visual feature sequences By utilizing the pyramid pooling module to perform pooling and fusion at multiple scales in parallel, visual features containing multi-scale context are output. ; Text encoder Used for text-based descriptions The global text semantic vector is obtained using a text encoder. ; Cross-modal attention fusion module Used for visual feature-based and text features The text-guided attention is calculated, and the attention calculation result is fused with the original visual features to output deep fusion features. ; The output layer is used to incorporate deep fused features during the inference phase. The data is reconstructed into a spatial feature map, which is then input into the segmentation decoder. By calculating the probability map of each pixel belonging to each semantic category, the decoder outputs a pixel-level semantic segmentation map. And the confidence score of each pixel belonging to each category; S2. Input the semantic segmentation map into the task scheduler. Through dynamic scheduling and collaborative intelligent agents, perform seepage simulation and potential prediction to obtain dynamic simulation results. Based on the current remaining oil state, use the decision intelligent agent to perform closed-loop optimization to obtain the optimal strategy and verify the simulation results. The implementation method is as follows: Obtain semantic segmentation map With images And based on the semantic segmentation graph The set of categories defined in For intelligent agents Generate task list ,in, Indicates the target category. Indicates origin from semantic segmentation graph Category coarse binary mask, This represents the preprocessed image. Indicates the first One task; For each task Apply a coarse binary mask to the category Alternatively, bounding boxes can be used as spatial cues input into the visual base model. ; Based on the visual fundamental model Using spatial cues in images Pixel-level retouching is performed within the context to obtain a binary mask. ,in, Indicates the number of pixels vertically in the image. Indicates the number of pixels horizontally in the image; Based on the refined binary mask set Using intelligent agents Automatic calculation of structured quantization parameter set Output a refined segmentation image. ; Based on refined segmentation map Reconstructing three-dimensional binary digital cores for simulation ; Given initial conditions And boundary conditions, using a lightweight proxy simulator Predict the evolution of the saturation field, where, Indicates the initial saturation distribution. Indicates time; Based on predicted saturation field evolution and reconstructed 3D binary digital core Lightweight agent simulator Perform scenario simulation and output the curve of average remaining oil saturation changing over time. Based on the curve of average residual oil saturation over time... The degree of extraction was calculated. ; Using intelligent agents Output dynamic simulation results Among them, dynamic simulation results Including extraction degree Curve of average residual oil saturation over time And the remaining oil distribution map at the final moment; From the structured quantization parameter set Extracting a low-dimensional vector that can characterize the current remaining oil. Define the state, where, Represents the set of real numbers. Represents the dimension of the state vector; Adjustable injection strategy parameters Defined as an action, where, Represents the action space; Given the current state and injection strategy Based on the environmental transition to a new state And generate rewards; Based on state Injection strategy And rewards, to build a reinforcement learning framework; intelligent agents As an intelligent agent In an interactive environment within a reinforcement learning framework, rewards are maximized through interactive trial and error. To train the policy network ,in, This represents the mathematical expectation of the cumulative reward. Represents the trajectory. Indicates the time step index; For the new core condition Utilizing the trained policy network Generate the optimal strategy ,in, Indicates the injection of policy variables; The generated optimal strategy Feedback to the intelligent agent Perform a final high-fidelity verification simulation to obtain the verification simulation results. ; S3. Based on semantic segmentation graphs, dynamic simulation results, optimal strategies, and their verification simulation results, a structured report is generated to complete the intelligent identification and quantification of microscopic residual oil. The implementation method is as follows: Based on semantic segmentation graph Confidence plots for each category, and structured quantization parameter sets. Refined Segmentation Diagram Dynamic simulation results Optimal Strategy and its verification simulation results To integrate information from multiple sources; The report generation engine has a built-in domain knowledge rule base. Among them, the domain knowledge rule base Used to automatically perform correlation analysis and generate conclusions; Based on integrated multi-source information, and utilizing a domain knowledge rule base Generate association rule-driven knowledge and combine it with verification simulation results. Conduct an assessment; Based on the evaluation results, according to predefined document modules The system automatically assembles and generates a final structured report, completing the intelligent identification and quantification of microscopic residual oil.

[0023] In this embodiment, multimodal data acquisition and fusion perception are performed in S1, aiming to address the fundamental deficiency of traditional "pure vision models" in their inability to understand and utilize domain knowledge when processing geological images. This step innovatively proposes a "Multimodal Pyramid Pooling VisionTransformer (MPP-ViT)" architecture, designed to perform deep, structured cross-modal fusion understanding of microscopic core images and associated geological descriptive text, and output a high-precision "semantic map" of remaining oil segmentation with geological semantics, as detailed below: 1. Acquisition and Standardization Preprocessing of Multi-Source Heterogeneous Data 1) Data Acquisition: Visual Modal V: Acquiring three-dimensional microscopic CT scan data of the target core sample. Defined as or representative two-dimensional slice sequences extracted from them. ,in, Indicates the first Slices are typically 1 (grayscale) or 3 (RGB). These represent the image height and width, respectively. Represents the spatial dimensions of three-dimensional volume data.

[0024] Text modality T: Obtaining unstructured geological description text that strictly corresponds to visual samples. The text information should cover key geophysical and experimental conditions, including but not limited to: rock physical properties (such as lithology, porosity φ, permeability K), fluid properties, and system state (such as wettability, water saturation). (and development history, etc.) 2) Data standardization preprocessing: Image preprocessing: Noise reduction: for images The preprocessed image is obtained by applying algorithms such as nonlocal mean filtering. .

[0025] Gray-level normalization: linearly maps image gray-level values ​​to the [0,1] interval. .

[0026] Uniform size: Scaling or cropping all images to a fixed size. , This represents the number of vertical pixels in an image with a fixed size. This represents the number of horizontal pixels in an image with a fixed size. This represents the normalized image.

[0027] Text preprocessing: Through named entity recognition or information extraction models From the original text string Key geological parameters are extracted to form a structured representation. , to structure representation Transformed into a coherent text description of the model input The format is: "Lithotype: [Lithotype]; Porosity: [φ]%; Permeability: [K] mD; Wettability: [Wettability Description];..." This step converts unstructured text into machine-readable semantic vectors to prepare the input.

[0028] 3) Preprocessed data pairs: The final result is a standardized image-text pair dataset. ,in, For preprocessed images (or from) Representative slices extracted.

[0029] 2. Deep cross-modal fusion and semantic segmentation based on a multimodal pyramid pooling visual model The core of this invention lies in the proposed multimodal pyramid pooling visual model MPP-ViT. This multimodal pyramid pooling visual model is specifically designed for microscopic residual oil segmentation, integrating global modeling of the visual Transformer, multi-scale feature extraction via pyramid pooling, and a cross-modal attention fusion mechanism. Its architecture and workflow are as follows: 1) Overall Model Architecture: The multimodal pyramid pooling visual model adopts an encoder-decoder architecture, which specifically includes: The visual feature extraction branch consists of a Vision Transformer (ViT) encoder and a Pyramid Pooling Module (PPM).

[0030] Text feature extraction branch: consists of a pre-trained Transformer text encoder.

[0031] Cross-modal attention fusion module: Enables semantic-level interaction between visual and textual features.

[0032] Feature decoding and segmentation module: adopts a U-Net-style segmentation decoder to output pixel-level segmentation maps.

[0033] 2) Visual feature extraction and multi-scale context coding: ViT visual encoder : The preprocessed image Divided into N Image blocks , Indicates the first i Image blocks, Indicates the first N Image blocks, Represents an image patch. This represents the number of vertical pixels in an image with a fixed size. This represents the number of horizontal pixels in an image with a fixed size. C The number of channels is represented by a linear projection and positional encoding, which, after passing through an L-layer Transformer encoder, outputs a visual feature sequence. ,in, Represents the dimensions of visual features. Indicates visual encoder operation, It represents the set of real numbers.

[0034] Pyramid Pooling Module (PPM): To capture multi-scale features, it pools intermediate layer features (or reshaped spatial feature maps) from the visual encoder. Input pyramid pooling module. This module performs pooling and fusion at multiple scales in parallel:

[0035] in, Feature pooling is The grid, s This indicates the grid size for the pooling operation. This represents the number of vertical pixels in an image with a fixed size. This represents the number of horizontal pixels in an image with a fixed size. d Represents the dimensions of visual features. for convolution, The features are upsampled back to their original size, and finally multi-scale information is fused. :

[0036] in, , , and These represent the pooling scales respectively. The processed multi-scale feature map.

[0037] After further fusion by convolutional layers, the output is visual features rich in multi-scale context. ,in, This indicates the dimensions of the fused visual features. Represents the set of real numbers. This represents the number of vertical pixels in an image with a fixed size. This indicates the number of horizontal pixels in an image with a fixed size.

[0038] 3) Text Feature Encoder : Describe the text Input a pre-trained Transformer text encoder Through its Transformer architecture, it obtains text feature sequences. Typically, features from the sequence start marker are extracted or processed to obtain the global text semantic vector. ,in, Indicates the sequence length. Represents the dimension of text features. This indicates text encoder operation.

[0039] 4) Cross-modal attention fusion module This module achieves deep integration of visual and textual elements. It employs a cross-attention mechanism, guided by text, to enhance visual features.

[0040] visual features Flatten into a sequence .

[0041] With global text semantic vectors For query Visual sequence For key Sum Calculate text-guided attention:

[0042] Output attention By fusing the original visual features with gated residual connections, deep fused features are obtained. :

[0043] in, Indicates learnable parameters, Indicates splicing, This indicates element-wise multiplication.

[0044] 5) Semantic segmentation, decoding, and output: U-Net decoder: Deeply fuses features The data is reconstructed into a spatial feature map and input into a U-Net-style segmentation decoder. This segmentation decoder progressively recovers high-resolution spatial details by passing the data through a series of upsampling layers and making skip connections with the hierarchical features of the ViT encoder.

[0045] Segmentation Head: A segmentation head is connected to the end of the decoder ( (Convolution + Softmax) outputs a probability map of each pixel belonging to each semantic category. .

[0046] Final output: The model output is a pixel-level semantic segmentation map. The semantic category set C = {background, skeleton, pore water, film-like residual oil, pore corner residual oil, cluster-like residual oil, droplet-like residual oil, columnar residual oil} can also output confidence maps for each category.

[0047] 6) Model training mechanism: The multimodal pyramid pooling visual model uses a multi-task loss function that combines segmentation accuracy and modality alignment. Perform end-to-end training:

[0048]

[0049]

[0050]

[0051] in, This represents the standard cross-entropy loss, used to optimize pixel classification. This represents Dice loss, which alleviates class imbalance. This represents a cross-modal contrastive loss (such as InfoNCE) that aligns image-text pairs in the feature space. Represents cross-modal contrast loss The weight, This represents the total number of pixels in the image. Indicates the total number of categories. Indicates the first The pixel belongs to the first The true label of the class, The model predicts the first... The pixel belongs to the first The probability value of the class. Indicates the predicted first Class feature map, The first one representing the true label Class feature map, Represents cosine similarity. Indicates the temperature coefficient. Indicates batch size, Indicates the sample index in the batch. Indicates the first Visual features of an image Indicates the first Textual features of a text Indicates the first in the batch Textual features of a text.

[0052] 7) Output: The final output is a pixel-level semantic segmentation map. This result deeply integrates visual evidence with geological text semantics, forming a high-quality "semantic blueprint" that drives subsequent collaborative decision-making by intelligent agents.

[0053] In this embodiment, dynamic guidance and collaborative reasoning of geological text semantics for microscopic image segmentation tasks are realized, which significantly improves the accuracy and interpretability of complex residual oil morphology identification.

[0054] In this embodiment, S2 is the agent-based collaborative scheduling and specialized task execution. This step is the core execution and computation hub of the invention, aiming to solve the bottlenecks in traditional analysis processes, namely, the separation of identification and quantization simulation, reliance on manual serial operations, and lack of closed-loop feedback. This step uses a task scheduler (Dispatcher) to take the output of S1 as input, dynamically scheduling and coordinating a series of highly specialized agents to sequentially complete tasks such as high-precision quantization, physical simulation, and policy optimization in a pipelined and automated manner. Its core idea is to construct a collaborative computing workflow driven by data flow and executed by agents, as detailed below: 1. Image analysis agents perform refined quantitative analysis. This intelligent agent The goal is to obtain accurate and computable geometric and physical parameters based on S1 semantic understanding.

[0055] 1) Input and Task Analysis: Input: Receive the semantic segmentation graph from S1 and original image .

[0056] Task parsing: The task scheduler uses semantic segmentation graphs... The semantic category set defined in For intelligent agents Generate task list ,in, Indicate the target category (e.g., "film-like residual oil"), Indicates origin from semantic segmentation graph Category coarse binary mask, This represents the preprocessed image. Indicates the first One task.

[0057] 2) Cue-based fine-grained segmentation: intelligent agent At its core is a visual foundation model that supports Promptable engineering. (e.g., a domain-adjusted version of the Segment Anything Model). For each task , category coarse mask Or its bounding box serves as a spatial prompt, inputting into the visual base model. Visual basic model The model utilizes spatial cues in the original image. Pixel-level refinement within the context:

[0058] This step outputs a high-precision binary mask. This enabled boundary optimization for various types of residual oil areas. Represents the basic visual model operations. This indicates a spatial prompt message.

[0059] 3) Automatic quantization calculation of multiple parameters: Based on refined mask set intelligent agent Automatically calculate a comprehensive set of quantitative parameters : a. Saturation parameter: Calculates the pixel volume ratio of various types of residual oil, i.e., digital saturation. :

[0060] in, Represents the set of all categories related to pore space. Indicates the first time after refinement In a pixel-like mask, the image located at the [missing information] line, number The pixel values ​​in the column represent the category index variable.

[0061] b. Morphological parameters: For each connected component (oil cluster), calculate its equivalent diameter. Count the number of items such as aspect ratio and roundness. and size distribution histogram.

[0062] c. Distribution parameters: Calculate the fractal dimension of the spatial distribution. (Using box counting) to assess the complexity of the remaining oil distribution; estimate the contact angle distribution of the oil-water-rock three phases (based on local image gradient and geometric information).

[0063] d. Topology parameters: Analyze the connectivity between the remaining oil clusters and the pore network, and calculate their coordination number (the number of connected pore throats).

[0064] 4) Output: Intelligent agent Output structured quantization parameter set And refined segmentation image This provides accurate initial conditions for subsequent physics simulations.

[0065] 2. Physical simulation of intelligent agents to perform seepage simulation and potential prediction This intelligent agent The goal is to reproduce the microscopic seepage physics process in the digital world and predict the dynamic response of remaining oil under different development measures.

[0066] 1) Input and Digital Core Reconstruction: Input: Received from the agent Refined segmentation image and structured quantization parameter set .

[0067] Model Reconstruction: Based on Refined Segmentation Map The two-dimensional semantic labels of each slice are stacked in voxels according to their spatial coordinates (slice index in the Z direction) to construct an initial three-dimensional semantic volume, where the value of each voxel is its corresponding semantic category label; to solve the problem of the initial three-dimensional semantic volume To address the label discontinuity issue caused by independent slice classification in the Z-direction, a three-dimensional isotropic diffusion process is introduced to calculate the probability distribution of each semantic category in three-dimensional space. This probability distribution is obtained by solving a three-dimensional diffusion equation from the initial labels. The diffusion coefficient is inversely proportional to the local gradient of the image to protect the true boundaries. Finally, each voxel... The semantic labels are based on the probability of each category. The maximum a posteriori probability is redistributed to generate a final three-dimensional digital core with spatial continuity. In this model, each voxel is uniquely labeled as skeleton (0), pore water (1), or various types of residual oil (2,3,...).

[0068] 2) Lightweight proxy simulator: intelligent agent At its core is a lightweight agent simulator based on Physics-Informed Neural Networks (PINN). .

[0069] The lightweight agent simulator By training on a large amount of high-fidelity simulation data (from a Lattice Boltzmann Method (LBM) pore network model (PNM)), the physical laws governing multiphase flow (such as Darcy's law and capillary pressure equation) were learned. Given initial conditions... and boundary conditions (injection strategy) It can predict the evolution of the saturation field in milliseconds:

[0070] in, Represents the remaining oil saturation field. Represents a pressure field. Indicates the proxy model parameters. express The remaining oil saturation field at any given time, express Constant pressure field express The remaining oil saturation field at any given time, express Constant pressure field Indicates the time step. This represents the injection strategy. Its training objective is to minimize the error between the output of a high-fidelity simulator (such as LBM) and the residuals of the physical equations.

[0071] 3) Scenario simulation and potential prediction: Given a development scenario (i.e., injection strategy) Such as injection type, concentration Flow rate ), drive a lightweight agent simulator Perform a fast forward simulation. The simulation output includes: Curve of average residual oil saturation over time:

[0072] Final recovery factor (RF):

[0073] in, This represents the average remaining oil saturation at the initial moment. This represents the average remaining oil saturation at the final moment. Represents the spatial domain of porous media. x Represents spatial coordinates. Indicates spatial location x ,time The remaining oil saturation at that location.

[0074] Contribution of utilization of residual oil under different occurrence types.

[0075] 4) Output: intelligent agent Output dynamic simulation results Dynamic simulation results Including extraction degree Curve of average residual oil saturation over time Together with the remaining oil distribution map at the final moment, this constitutes a set of "virtual experiment" data.

[0076] 3. The decision-making optimization agent achieves closed-loop optimization. This intelligent agent The goal is to automatically find the optimal development strategy based on the current remaining oil status. It will involve an intelligent agent. As its interactive environment, it forms a closed loop of reinforcement learning (RL).

[0077] 1) Construction of reinforcement learning framework: State: Defined from the structured quantization parameter set The low-dimensional vector extracted that can characterize the remaining oil. ,For example ,in, Represents the set of real numbers. Represents the dimension of the state vector. Indicates the characteristic parameters of the first type of reservoir. Indicates the characteristic parameters of the second type of reservoir. Indicates the crack development index. Indicates the total number of injection and production wells. This represents the equivalent average porosity, and T represents the transpose symbol. Action: Defined as a controllable injection strategy parameter. ,For example ,in, Represents the action space. Indicates the type of injected fluid. Indicates the concentration of the injected chemical agent. Indicates the injection rate; Environment: that is, the physical simulation of the intelligent agent. Given the current state and injection strategy The environment shifts to a new state And generate rewards .

[0078] Reward function: Carefully designed to quantify development goals, for example:

[0079] in, Indicates the incremental level of extraction at each stage. Indicates the residual oil saturation at the end of the stage. Indicates the cost of the measures, , and All represent weights.

[0080] 2) Strategy learning and optimization: intelligent agent Deep reinforcement learning algorithms (such as Proximal Policy Optimization, PPO) are used to learn an optimal policy. The optimal strategy is a parameter with Neural networks. Through extensive interaction with intelligent agents. Interactive trial and error in the environment to maximize cumulative rewards Thus, the policy network is trained. ,in, This represents the mathematical expectation of the cumulative reward. Represents the trajectory. Indicates the time step index.

[0081] 3) Strategy generation and verification: Reasoning phase: For a new core condition The trained policy network The optimal injection strategy can be generated directly:

[0082] in, Indicates the injection of policy variables; Closed-loop verification: the generated optimal strategy It will immediately provide feedback to the intelligent agent. A final high-fidelity verification simulation is performed to ensure its effectiveness and robustness under actual physical constraints. The verification results can be further used for fine-tuning the policy network, forming a self-improving closed loop.

[0083] In this embodiment, the semantic blueprint-driven multi-agent collaborative architecture proposed in this invention designs an automated workflow and data interface standard that dynamically schedules specialized agents such as image analysis, physical simulation, and decision optimization to execute in an orderly manner, using semantic recognition results as a unified task input. The microscopic percolation rapid simulation based on a lightweight agent model proposed in this invention integrates a lightweight agent simulator based on a Physical Information Neural Network (PINN) into the physical simulation agent. This invention enables rapid, physically constrained prediction of seepage processes within seconds, supporting closed-loop optimization. Furthermore, the proposed microscale reinforcement learning decision-making closed-loop construction utilizes a reinforcement learning framework with remaining oil quantity parameters as states, development measures as actions, and a physically simulated intelligent agent as the environment, achieving autonomous online optimization and verification of development strategies.

[0084] In this embodiment, S3 is the result synthesis and report generation. This step is the value realization and output terminal of the present invention, aiming to intelligently integrate, correlate, and present the multi-dimensional and multi-scale information generated by S1 and S2, generating a comprehensive quantitative and potential assessment report that can directly serve geological research, reservoir engineering, and development decisions. This step achieves automated transformation from raw analysis data to structured decision-making knowledge through preset knowledge templates, correlation rules, and a visualization engine, as detailed below: 1. Multi-source information fusion and structured analysis This stage is the central hub for information integration and knowledge extraction, responsible for gathering, associating, and deeply interpreting all the outputs from the preceding steps.

[0085] 1) Multi-source information integration: The system automatically collects and time-aligns the following multi-modal data streams to form a comprehensive database. : Perceptual data: from semantic segmentation graphs And confidence plots for each category.

[0086] Quantization data: from structured quantization parameter sets With refined segmentation image .

[0087] Simulation data: from dynamic simulation results Dynamic simulation results Including average residual oil saturation-time series Extraction degree and the evolution of pressure fields, among which, Indicates time step The average remaining oil saturation at time T represents the total number of time steps. i Indicates the time step index; Optimized data (if enabled): from the optimal strategy and its verification simulation results .

[0088] 2) Knowledge generation driven by association rules: The report generation engine has a built-in domain knowledge rule base. It is used to automatically perform association analysis and generate advanced conclusions. The association rules adopt "if-then" logic or graph-based reasoning. Key analyses include: a. Structure-attribution association analysis: Example of association rule: IF aperture AND wettability is hydrophilic; THEN's main occurrence type tends to be "film-like".

[0089] Execution: System call to structured quantization parameter set Pore ​​size distribution and semantic segmentation map Distribution of remaining oil types in the data, and calculation of conditional probability. And generate quantitative conclusions such as "the probability of film-like residual oil occurrence is as high as 85% in hydrophilic pores with a pore size of less than 5 μm".

[0090] b. Potential Comparison and Attribution Analysis: Based on dynamic simulation results Different development scenarios The simulation results are automatically sorted and the potential difference is calculated:

[0091] in, Describing a scenario Relative to the context The difference in recovery rate (RF), Indicate development scenario The final extraction degree, Indicate development scenario The final extraction level.

[0092] Attributing potential differences by combining geological semantics from semantic segmentation maps. For example, if surfactant-driven... Polymer drive incremental recovery of production at different stages 8% higher, system-related structured quantitative parameter set If a high percentage of film-like oil is found in the core, the conclusion is: "Because film-like oil dominates (accounting for XX%), surfactant flooding that reduces interfacial tension is expected to increase the recovery rate by 8% more than polymer flooding that increases viscosity." c. Strategy effectiveness evaluation: For the optimal strategy Calculate its expected benefit indicators, such as the incremental input-output ratio (iROI):

[0093] Combined with verification simulation results To assess the robustness and risks of the strategy.

[0094] 2. Automated generation and output of structured reports Based on the above analysis results, and according to the predefined document template The system automatically assembles and generates the final report.

[0095] 1) Report Content Module: The generated report This is a structured document, containing the following core modules: Module 1: Sample and Task Summary, showcasing core images, key geological parameter tables (derived from multi-source heterogeneous data), and analysis objectives.

[0096] Module 2: Fine characterization of residual oil.

[0097] Visualization: Overlay display of semantic segmentation results and spatial distribution cloud maps of various types of remaining oil.

[0098] Data table: A detailed, structured table of quantitative parameters (saturation, quantity, size, distribution fractal dimension, etc.).

[0099] Module 3: Development Potential Prediction.

[0100] Visualization: Residual oil saturation versus time curves under different scenarios Ultimate recovery rate comparison bar chart and remaining oil utilization potential radar chart.

[0101] Text analysis: Automatically generated potential ranking and key control factor analysis.

[0102] Module 4: Optimization Strategy Recommendations.

[0103] Strategy Description: Clearly list the recommended measures and the optimal strategy. Specific parameters (type, concentration, rate, etc.).

[0104] Expected outcome: To demonstrate and validate the extraction rate of the simulation. The curves provide the expected increase in recovery rate, incremental return on investment (iROI), and other metrics.

[0105] Key points and risk warnings for implementation: precautions for operation based on rule base generation.

[0106] Module 5: Uncertainty Description. Quantitatively or qualitatively describe the main sources of uncertainty in the analysis, such as model confidence, simulation parameter range, and geological parameter errors.

[0107] 2) Report format and output: The report can generate interactive web pages (HTML), standard documents (PDF / DOCX), and machine-readable structured data files (JSON / XML).

Claims

1. A method for intelligent identification and quantification of microscopic residual oil, characterized in that, Includes the following steps: S1. Obtain visual and textual modalities, preprocess them, and generate semantic segmentation maps based on the preprocessed image-text pair data using a multimodal pyramid pooling visual model that performs cross-modal fusion and semantic recognition. Multimodal pyramid pooling visual models include: Visual encoder Used to transfer images Divided into N Each image patch, after linear projection and positional encoding, is passed through an L-layer encoder to output a visual feature sequence. Based on visual feature sequences By utilizing the pyramid pooling module to perform pooling and fusion at multiple scales in parallel, visual features containing multi-scale context are output. ; Text Encoder Used for text-based descriptions The global text semantic vector is obtained using a text encoder. ; Cross-modal attention fusion module Used for visual feature-based and text features The text-guided attention is calculated, and the attention calculation result is fused with the original visual features to output deep fusion features. ; The output layer is used to incorporate deep fused features during the inference phase. The data is reconstructed into a spatial feature map, which is then input into the segmentation decoder. By calculating the probability map of each pixel belonging to each semantic category, the decoder outputs a pixel-level semantic segmentation map. And the confidence score of each pixel belonging to each category; S2. Input the semantic segmentation graph into the task scheduler, and perform seepage simulation and potential prediction by dynamically scheduling and coordinating the intelligent agents to obtain dynamic simulation results. Based on the current remaining oil state, use the decision intelligent agent to find the optimal strategy and verify the simulation results. S3. Based on semantic segmentation graphs, dynamic simulation results, optimal strategies, and their verification simulation results, a structured report is generated to complete the intelligent identification and quantification of microscopic residual oil.

2. The method for intelligent identification and quantification of microscopic residual oil according to claim 1, characterized in that, S1 includes the following steps: The visual modality and the corresponding unstructured geological description text are obtained; whereby the unstructured geological description text is defined as the original text string. ; Based on the acquired visual modality, the image Preprocessing is performed, among which, The first microscopic image representing the target rock core i Zhang two-dimensional slice image; From the original text string Key geological parameters are extracted to form a structured representation. and structured representation Transform into a coherent text description ; Based on image Preprocessing structure and text description To obtain image-text pair data ,in, This represents the image after preprocessing. Based on image-text pair data A semantic segmentation map is generated by using a multimodal pyramid pooling visual model for cross-modal fusion and semantic recognition. .

3. The intelligent identification and quantification method for microscopic residual oil according to claim 1, characterized in that, S2 includes the following steps: Obtain semantic segmentation map With images And based on the semantic segmentation graph The set of categories defined in For intelligent agents Generate task list ,in, Indicates the target category, Indicates origin from semantic segmentation graph The category of coarse binary mask, This represents the image after preprocessing. Indicates the first One task; For each task Apply a coarse binary mask to the category Alternatively, bounding boxes can be used as spatial cues input into the visual base model. ; Based on the visual fundamental model Using spatial cues in images Pixel-level retouching is performed within the context to obtain a binary mask. ,in, This represents the number of vertical pixels in an image with a fixed size. This indicates the number of horizontal pixels in an image with a fixed size. Based on the refined binary mask set Using intelligent agents Automatic calculation of structured quantization parameter set Output a refined segmentation image. ; Based on refined segmentation map Reconstructing three-dimensional binary digital cores for simulation ; Given initial conditions And boundary conditions, using a lightweight proxy simulator Predict the evolution of the saturation field, where, Indicates the initial saturation distribution. Indicates time; Based on predicted saturation field evolution and reconstructed 3D binary digital core Lightweight agent simulator Perform scenario simulation and output the curve of average remaining oil saturation changing over time. Based on the curve of average residual oil saturation over time... The degree of extraction was calculated. ; Using intelligent agents Output dynamic simulation results Among them, dynamic simulation results Including extraction degree Curve of average residual oil saturation over time And the remaining oil distribution map at the final moment; From the structured quantization parameter set Extracting a low-dimensional vector that can characterize the current remaining oil. Define the state, where, Represents the set of real numbers. Represents the dimension of the state vector; Adjustable injection strategy parameters Defined as an action, where, Represents the action space; Given the current state and injection strategy Based on the environmental transition to a new state And generate a reward, where the reward function is... The expression is as follows: in, Indicates the incremental level of extraction at each stage. Indicates the residual oil saturation at the end of the stage. Indicates the cost of the measures, , and All represent weights; Based on state Injection strategy And rewards, to build a reinforcement learning framework; intelligent agents As an intelligent agent In an interactive environment within a reinforcement learning framework, rewards are maximized through interactive trial and error. To train the policy network ,in, This represents the mathematical expectation of the cumulative reward. Represents the trajectory. Indicates the time step index; For the new core condition Utilizing the trained policy network Generate the optimal strategy ,in, Indicates the injection of policy variables; The generated optimal strategy Feedback to the intelligent agent Perform a final high-fidelity verification simulation to obtain the verification simulation results. .

4. The intelligent identification and quantification method for microscopic residual oil according to claim 3, characterized in that, The expression for the evolution of the saturation field is as follows: in, Represents the remaining oil saturation field. Represents a pressure field. Indicates the proxy model parameters. express The remaining oil saturation field at any given time, express Constant pressure field, express The remaining oil saturation field at any given time, express Constant pressure field Indicates the time step. This indicates the injection strategy.

5. The intelligent identification and quantification method for microscopic residual oil according to claim 3, characterized in that, The degree of extraction The expression is as follows: in, This represents the average remaining oil saturation at the initial moment. This represents the average remaining oil saturation at the final moment. Represents the spatial domain of porous media. x Represents spatial coordinates. Indicates spatial location x ,time The remaining oil saturation at that location.

6. The method for intelligent identification and quantification of microscopic residual oil according to claim 1, characterized in that, S3 includes the following steps: Based on semantic segmentation graph Confidence plots for each category, and structured quantization parameter sets. Refined Segmentation Diagram Dynamic simulation results Optimal Strategy and its verification simulation results To integrate information from multiple sources; The report generation engine has a built-in domain knowledge rule base. Among them, the domain knowledge rule base Used to automatically perform correlation analysis and generate conclusions; Based on integrated multi-source information, and utilizing a domain knowledge rule base Generate association rule-driven knowledge and combine it with verification simulation results. Conduct an assessment; Based on the evaluation results, according to predefined document modules The system automatically assembles and generates a final structured report, completing the intelligent identification and quantification of microscopic residual oil.

Citation Information

Patent Citations

  • Melanoma lesion area segmentation method based on CLIP multi-mode fusion network

    CN120807924A

  • Multi-agent cooperation method based on spatial perception and multi-modal AI fusion

    CN122113983A