A three-dimensional mesh reconstruction method based on large model adaptive feature extraction

By constructing a 3D mesh reconstruction method based on a large model and optimizing mesh deformation using robust feature extraction and multi-scale reconstruction modules, the problem of low accuracy in complex organ reconstruction in existing technologies is solved, and efficient adaptive 3D reconstruction of multiple organs is achieved.

CN120782998BActive Publication Date: 2025-11-07HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511262964.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2025-11-07
Estimated Expiration
2045-09-05

AI Technical Summary

Technical Problem

Existing 3D reconstruction algorithms have low reconstruction accuracy when dealing with organ tissues in complex abdominal environments, and deep learning methods have high computational complexity and long training and inference times, making it difficult to achieve efficient adaptive reconstruction of complex structures of multiple organs.

Method used

A 3D mesh reconstruction method based on a large model is constructed, including robust feature extraction, dynamic semantic feature selection, voxel-point cloud mapping, and multi-scale reconstruction modules. Feature extraction and fusion are performed through a pre-trained large model and a language model, and mesh deformation is optimized by combining the multi-scale reconstruction module to achieve high-precision 3D reconstruction.

Benefits of technology

While maintaining global structural consistency, the method gradually corrects local deformation errors, improves reconstruction accuracy in complex scenarios, solves the mesh distortion problem of traditional methods, and achieves efficient multi-organ adaptive reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120782998B_ABST
    Figure CN120782998B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional mesh reconstruction method based on large model adaptive feature extraction, compared with the prior art, the application constructs a multi-scale feature guiding mechanism through a medical input image and a target medical task, under the premise of keeping the consistency of the global structure, gradually corrects the local deformation error by using an iterative optimization strategy, and finally converges the predicted mesh to the real anatomical form in the topological structure and geometric precision, effectively solving the mesh distortion problem of the traditional method in the large deformation scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer vision, and particularly relates to a three-dimensional mesh reconstruction method based on large model adaptive feature extraction. BACKGROUND

[0002] With the development of artificial intelligence technology and the upgrading of medical imaging equipment, preoperative three-dimensional reconstruction has brought revolutionary progress in improving the accuracy of liver surgery and reducing surgical risk. Preoperative three-dimensional reconstruction is to import patient CT, MRI raw data into three-dimensional reconstruction software, then perform three-dimensional modeling to form a three-dimensional visual model, which provides doctors with more abundant and intuitive lesion information and helps doctors diagnose more accurately. However, existing three-dimensional reconstruction algorithms are mostly based on template registration, which directly deforms a high-precision template network using image features. Although this method can achieve good results for simple structures with small deformations, it is limited by the performance of the deformation module and the smoothing regularity. Generally, only a single structure can be reconstructed, and the reconstruction accuracy is often limited in such complex situations. Therefore, it is necessary to develop a high-precision adaptive reconstruction algorithm for multiple complex organ structures to enable doctors to clearly understand the surgical process before surgery, which is beneficial for doctors to assess surgical risks and plan surgical designs in advance.

[0003] To solve the problem of low reconstruction accuracy of traditional methods in processing complex medical images, Quan et al. used a Transformer architecture for three-dimensional reconstruction to improve reconstruction accuracy and detail performance using its powerful long-distance dependency processing capability. Buttongkum et al. introduced fracture representation learning to extract fracture features using deep learning technology and combined it with dual-plane X-ray images for three-dimensional reconstruction. However, deep learning methods require a large amount of labeled data for training, and image labeling can be complex and time-consuming. The generalization ability of the model may be affected by the diversity of the training data. In addition, the model has high computational complexity, which may result in long training and inference times.

[0004] Therefore, how to design an efficient three-dimensional reconstruction method for adaptive reconstruction of multiple complex organ structures is a technical problem to be solved. SUMMARY

[0005] Therefore, it is necessary to provide a three-dimensional mesh reconstruction method based on large model adaptive feature extraction to solve the problems of the prior art, which comprises the following steps:

[0006] S1: constructing a three-dimensional mesh reconstruction model, wherein the three-dimensional mesh reconstruction model comprises a large model robust feature extraction module, a dynamic semantic feature selection module, a voxel-point cloud mapping module, and a multi-scale reconstruction module;

[0007] S2: obtaining a medical input image, performing feature extraction on the medical input image by the large model robust feature extraction module to obtain image extraction features;

[0008] S3: confirming a target medical task, constructing a semantic prompt word based on the medical input image and the target medical task, encoding the semantic prompt word by a pre-trained language large model to obtain image semantic features;

[0009] S4: performing fusion processing on the image extraction features and the image semantic features based on the dynamic semantic feature selection module to obtain dynamic semantic features of the medical input image;

[0010] S5: performing spatial mapping of the dynamic semantic features from voxels to point clouds based on the voxel-point cloud mapping module to obtain multi-scale point cloud image spatial features;

[0011] S6: the multi-scale reconstruction module deforms an initial mesh template based on the point cloud image spatial features to obtain a target three-dimensional mesh corresponding to the medical input image.

[0012] Preferably, step S2 comprises:

[0013] S21: obtaining a medical input image and preprocessing the medical input image;

[0014] S22: inputting the preprocessed medical input image into a pre-trained visual large model SAM encoder to extract global features and local features of the medical input image;

[0015] S23: performing global average pooling, upsampling and convolution processing on the global features and local features to obtain image extraction features.

[0016] Preferably, step S3 comprises:

[0017] S31: confirming a target medical task based on the medical input image;

[0018] S32: constructing a semantic prompt word based on the medical input image and the target medical task;

[0019] S33: segmenting the semantic prompt word into multiple discrete sequences by a word segmenter;

[0020] S34: mapping the multiple discrete sequences into high-dimensional continuous vectors by an embedding layer;

[0021] S35: processing the high-dimensional continuous vectors by a dimension reduction processing module and performing feature pooling to obtain image semantic features.

[0022] Preferably, the dynamic semantic feature extraction module comprises a plurality of trainable convolutional layers, and step S4 comprises:

[0023] S41: point-multiplying the image extraction feature and the image semantic feature to obtain a point-multiplying result;

[0024] S42: performing feature fusion on the point-multiplying result to obtain the dynamic semantic feature of the medical input image.

[0025] Preferably, step S5 comprises:

[0026] S51: performing three-line difference value on each voxel feature of the dynamic semantic feature and mapping to a vertex of an initial mesh template;

[0027] S52: performing dynamic semantic feature learning based on all vertices mapped to the initial mesh template to generate a multi-scale point cloud image space feature.

[0028] Preferably, the multi-scale reconstruction module comprises a plurality of GCN up-sampling modules, each GCN up-sampling module being composed of a cascaded up-sampling layer, a convolutional layer and an activation layer, and step S6 comprises:

[0029] S61: sequentially inputting the multi-scale point cloud image space feature into the plurality of GCN up-sampling modules, and each GCN up-sampling module expanding the number of grid vertices of the input point cloud image space feature to twice the original number;

[0030] S62: after each GCN up-sampling module fuses and performs trilinear interpolation on the point cloud image space feature after the number of grids is expanded, obtaining a target three-dimensional grid corresponding to the medical input image.

[0031] Preferably, the grid loss function of any one grid of the point cloud image space feature is:

[0032] (1);

[0033] wherein, pred and true represent the predicted value and the true value of the grid respectively, is a geometric consistency loss, is a regularization loss.

[0034] Preferably, the geometric consistency loss is represented by the following formula:

[0035] (2);

[0036] wherein, is a predicted grid of the Sth stage, is a true value grid, and is a weight coefficient of the correspondence loss, is a curvature weighted chamfer loss, is represented by the following formula:

[0037] (3);

[0038] wherein, and are point sets of the real mesh and the predicted mesh respectively, , is the curvature of the point u is the nearest neighbor of the point v in the point set is the nearest neighbor of u in is a point in is a point in is a point in

[0039] is a normal consistency loss between meshes, and the calculation formula is as follows:

[0040] (4);

[0041] wherein, is a normal, S is the number of deformation stages, and C is the number of meshes.

[0042] Preferably, the regularization loss is represented by the following formula:

[0043] (5);

[0044] wherein, is a Laplacian smoothing loss of displacement distance, is an intra-single mesh normal consistency loss, is a side length loss, is a predicted displacement vector, , and are weight coefficients of the correspondence loss respectively;

[0045] The Laplacian smoothing loss of displacement distance is represented by the following formula:

[0046] (6);

[0047] The intra-single mesh normal consistency loss is represented by the following formula:

[0048] ​​ (7);

[0049] edge length loss is expressed by the following formula:

[0050] (8);

[0051] L is a Laplacian operator of the mesh, is a vertex set of the predicted mesh, and is an adjacent face, is an edge set of the predicted mesh, and are two vertices of the edge .

[0052] Compared with the prior art, the present application has the following beneficial effects: the present application constructs a multi-scale feature guiding mechanism through a medical input image and a target medical task, gradually corrects local deformation errors by using an iterative optimization strategy on the premise of maintaining global structural consistency, and finally converges the predicted mesh to a real anatomical morphology in terms of topological structure and geometric precision, effectively solving the mesh distortion problem of traditional methods in a large deformation scene. BRIEF DESCRIPTION OF DRAWINGS

[0053] The exemplary embodiments of this application can be more fully understood with reference to the following description when taken in connection with the following drawings, in which like reference numerals represent analogous elements or steps. The drawings provide further assistance in understanding the application. They constitute a part of this specification and are included to further provide specific embodiments of the application, and to explain the application together with the embodiments of the application, and do not constitute a limitation of the application. In the drawings, the same reference numerals represent the same components or steps.

[0054] Figure 1 A three-dimensional mesh reconstruction method based on large model adaptive feature extraction is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0055] Exemplary embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided so that the present disclosure can be more thoroughly understood, and the scope of the present disclosure can be accurately conveyed to those skilled in the art.

[0056] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0057] In addition, the technical features involved in the different embodiments of the application described below can be combined with each other as long as there is no conflict.

[0058] With reference to Figure 1 The embodiment discloses a three-dimensional grid reconstruction method based on large model adaptive feature extraction, comprising the following steps:

[0059] S1: constructing a three-dimensional grid reconstruction model, the three-dimensional grid reconstruction model comprising a large model robust feature extraction module, a dynamic semantic feature selection module, a voxel-point cloud mapping module and a multi-scale reconstruction module;

[0060] S2: obtaining a medical input image, extracting features of the medical input image through the large model robust feature extraction module to obtain image extraction features;

[0061] Specifically, step S2 comprises:

[0062] S21: obtaining a medical input image and preprocessing the medical input image;

[0063] S22: inputting the preprocessed medical input image into a pre-trained visual large model SAM encoder to extract global features and local features of the medical input image;

[0064] S23: performing global average pooling, upsampling and convolution processing on the global features and the local features to obtain image extraction features.

[0065] In the embodiment, the CT / MR image is input into the image encoder of the pre-trained visual large model SAM, the purpose is to extract large model robust features of the image, and the role is to convert the original image into semantic features with a size S2 comprises the following detailed process:

[0066] Firstly, the input CT / MR image is uniformly adjusted to 128x128 size by an interpolation method such as bilinear interpolation or nearest neighbor interpolation, and normalized to regulate the image pixel value range. Then the preprocessed image is input into the pre-trained visual large model SAM encoder based on the Vision Transformer (ViT) architecture, which is constructed by the MAE (Masked Autoencoders) self-supervised pre-training method. The normalized image is divided into 16x16 pixel image blocks and converted into embedding vectors, combined with position encoding to retain spatial position information, and then passed through a multi-layer Transformer network to capture global and local features in the image.

[0067] Further, the extracted features are globally averaged to obtain global feature representation, and the generated global feature representation is mapped to 128x128 size image semantic features through upsampling and convolution operations .

[0068] Wherein, the pre-training process of the image encoder of the visual large model SAM is completed by the original development team of the visual large model SAM based on the MAE (Masked Autoencoder) self-supervised learning method on millions of general image datasets, and the training details include: using Vision Transformer (ViT) architecture, through the pre-training task of randomly masking image blocks and reconstructing pixel values, the model learns the general visual feature representation capability. In this application, the trained model parameters are directly loaded, and the encoder is not updated or fine-tuned, but only used as a fixed feature extractor to convert the input CT / MR medical image into a feature representation with semantic information . Because existing studies have shown that a pre-trained large model extracts features plus a small number of trainable convolution layers to form an adapter (Adaptor) sufficient to extract effective features, the operation of the adapter can be seen in S5. This way of directly using pre-trained parameters not only ensures the stability of feature extraction, but also avoids additional training costs.

[0069] S3: Confirm the target medical task, construct semantic prompt words based on the medical input image and the target medical task, encode the semantic prompt words through the pre-trained language large model to obtain image semantic features;

[0070] Specifically, step S3 includes:

[0071] S31: Based on the medical input image, confirm the target medical task;

[0072] S32: Based on the medical input image and the target medical task, construct semantic prompt words;

[0073] S33: segmenting the semantic prompt word into a plurality of discrete sequences by a tokenizer;

[0074] S34: mapping the plurality of discrete sequences into high-dimensional continuous vectors by an embedding layer;

[0075] S35: processing the high-dimensional continuous vectors by a dimension reduction processing module and performing feature pooling to obtain image semantic features.

[0076] Specifically, according to the current input image modality (such as CT) and the specific anatomical structure reconstruction task (such as liver reconstruction), a text prompt word with clear semantic direction is constructed (e.g. "Reconstruct Liver Mesh From Abdomen CT Image"), which is encoded by the text encoder of the pre-trained language large model LLM, and converted into semantic features with the same size as the image extraction features .

[0077] Specifically, first, the input semantic prompt text string Input is segmented into a discrete Token sequence by a tokenizer, denoted as Each "token" after tokenization is mapped to a unique number , forming an input ID sequence that the model can process ; then the discrete ID sequence is mapped to a high-dimensional continuous vector representation by an embedding layer (symbols in natural language (such as words) are essentially discrete and cannot be directly operated mathematically. The embedding layer maps each symbol to a high-dimensional vector (e.g. 300-dimensional or 768-dimensional), so that it has continuous numerical features.) to form the input feature of the Transformer encoder; the encoder is a neural network module based on self-attention mechanism (Self-Attention), which is used to convert the input sequence (text prompt word Input) into a feature representation with context awareness, composed of multiple layers of self-attention mechanism and feedforward neural network module, which captures long-distance dependencies and context information in the text, and generates a feature matrix with rich semantic information; then a dimension reduction processing module is used to perform principal component analysis and dimension compression on the high-dimensional features output by the encoder, and finally a text semantic feature vector with fixed dimension is extracted through a feature pooling operation.

[0078] S4: performing fusion processing on the image extraction features and the image semantic features based on a dynamic semantic feature selection module to obtain dynamic semantic features of the medical input image;

[0079] In this embodiment, the dynamic semantic feature extraction module includes a small number of layers (such as convolutional layers, fully connected layers) for fine-tuning the features output by the large model to adapt to downstream tasks. By multiplying the image extraction features and the image semantic features, the dot product result is input into a small number of layers, which can achieve the effect of fusion.

[0080] Step S4 includes:

[0081] S41: Dot product the image extraction features and the image semantic features to obtain a dot product result;

[0082] S42: Perform feature fusion on the dot product result to obtain the dynamic semantic features of the medical input image.

[0083] Current research has shown that the features extracted by the pre-trained large model plus a small number of trainable convolutional layers to form an adapter (dynamic semantic feature extraction module) are sufficient to extract effective features, so the semantic features F SAM extracted by the image encoder of the above SAM and the semantic features R converted by the text encoder of the LLM are multiplied and merged for feature extraction; then the features extracted by the pre-trained model are further adjusted, optimized and fused by the convolutional layer as the adapter layer. Then the adapter is fine-tuned, most of the parameters of the pre-trained image encoder of SAM and the text encoder of LLM are frozen, and only the parameters of the adapter layer are updated. This not only reduces the demand for computing resources, but also avoids excessive adjustment of the pre-trained model, thereby constructing a dynamic semantic feature extraction module.

[0084] S5: Perform voxel-to-point cloud spatial mapping on the dynamic semantic features based on the voxel-point cloud mapping module to obtain multi-scale point cloud image spatial features;

[0085] Step S5 includes:

[0086] S51: Perform three-line difference value for each voxel feature of the dynamic semantic features and map to the vertices of the initial mesh template;

[0087] S52: Based on all the vertices mapped to the initial mesh template, learn the dynamic semantic features to generate multi-scale point cloud image spatial features.

[0088] Specifically, the semantic features extracted by the dynamic semantic feature extraction module and the voxels are Euclidean data, while the multi-scale spatial point cloud data features used for coarse-to-fine multi-scale reconstruction are non-Euclidean data, so voxel-point cloud spatial mapping must be performed. The discrete voxel features on the feature map are mapped to the vertices of the initial mesh template through a three-linear interpolation operation, thereby obtaining the spatial domain features of the mesh.

[0089] Specifically, the initial mesh template is the basic method of current three-dimensional mesh reconstruction, according to human experience, a corresponding initial mesh template is put in, and through step-by-step learning, it is slowly deformed into the target mesh. In three-dimensional reconstruction, the initial mesh template is a basic structure for representing the geometric shape of the object surface, usually composed of a grid of polygons (mostly triangles), including vertices, edges and faces. The dynamic semantic feature has the characteristics of the target mesh, and the dynamic semantic feature is mapped to the initial mesh template through trilinear interpolation, so that the initial mesh template learns the dynamic semantic feature to gradually become the target mesh.

[0090] For each mesh vertex that needs to be interpolated Through trilinear interpolation, based on the principle of spatial local linear approximation, that is, assuming that the feature value changes linearly within a region, the feature values of the surrounding 8 discrete vertices are weighted and averaged to estimate the feature value of the target vertex. Specifically, assuming that the coordinates of the surrounding 8 existing vertices are , , , , , , , , , , .

[0091] First, for the vertex to be interpolated , , , , , ,

[0092] , (9);

[0093] (10);

[0094] Along the y-axis direction, linear interpolation is performed on the intermediate coordinates , , , ,

[0095] , (11);

[0096] The trilinear interpolation can fuse the extracted voxel features with the features of the template mesh vertices, realize the function of feature mapping, and the core function is to generate a smooth and geometrically consistent feature representation for any point (mesh vertex) in the continuous space through the known feature value in the voxel, which plays a role in feature continuity and geometric precision improvement for multi-scale mesh reconstruction in S6.

[0097] Specifically, the mesh vertices are initially from the initial template mesh, after adding the dynamic semantic features through trilinear interpolation, they are deformed into more detailed mesh vertices through upsampling, and then the dynamic semantic features through trilinear interpolation are added again, and the mesh vertices are deformed into more detailed mesh vertices through upsampling, and the process is repeated three to four times, and the mesh formed each time is different in detail, that is, from coarse to fine multi-scale spatial features.

[0098] S6: The multi-scale reconstruction module deforms the initial mesh template based on the point cloud image spatial features to obtain a target three-dimensional mesh corresponding to the medical input image.

[0099] Specifically, the multi-scale reconstruction module includes a plurality of GCN upsampling modules, and each GCN upsampling module is composed of a cascaded upsampling layer, a convolution layer and an activation layer, and step S6 includes:

[0100] S61: sequentially input the multi-scale point cloud image spatial features into the plurality of GCN upsampling modules, and each level of GCN upsampling module expands the number of mesh vertices of the input point cloud image spatial features to twice the original number;

[0101] S62: each level of GCN upsampling module fuses and trilinearly interpolates the point cloud image spatial features after the number of meshes is expanded, to obtain a target three-dimensional mesh corresponding to the medical input image.

[0102] Specifically, a multi-scale reconstruction framework from coarse to fine is constructed by using a plurality of GCN upsampling modules, and the voxel features are accurately mapped to the mesh vertex space through the trilinear interpolation technology, to realize high-precision three-dimensional reconstruction under a large deformation scene. Specifically, the system first maps the coarse-grained voxel features to the initial template mesh vertices through trilinear interpolation at the lowest resolution level, to form an initial feature representation containing global shape prior, as the input of the first level GCN module; then the mesh resolution is gradually improved through the cascaded GCN upsampling module, wherein the number of mesh vertices is expanded to twice the original number at each upsampling stage, and the higher resolution voxel features extracted through trilinear interpolation are fused synchronously, to realize progressive enhancement of geometric details.

[0103] The GCN upsampling module consists of upsampling layers, convolutional layers, and activation layers. Through a cascaded structure, it progressively increases the mesh resolution to twice the original size. A dedicated mesh deformation module is configured at each scale level. This module, composed of three graph convolutional layers, optimizes the mesh topology and feature representation under the guidance of spatial features. The specific deformation process is achieved through displacement vector prediction, mathematically expressed as follows:

[0104] (12);

[0105] in, It is the vertex position of the current step. It is the predicted displacement vector. This refers to the vertex position after deformation. This framework, through a multi-scale feature-guided mechanism, gradually corrects local deformation errors using an iterative optimization strategy while maintaining global structural consistency. Ultimately, it enables the predicted mesh to converge to the true anatomical shape in terms of topological structure and geometric accuracy, effectively solving the mesh distortion problem in large deformation scenarios using traditional methods.

[0106] Specifically, the grid loss function for any grid in the spatial features of a point cloud image is:

[0107] (1);

[0108] in, These represent the predicted and true values ​​of the grid, respectively. For geometrically consistent loss, This is the loss due to regularization.

[0109] Among them, geometrically consistent loss This can be expressed by the following formula:

[0110] (2);

[0111] in, For the prediction grid of stage S, For truth grids, and These are the weighting coefficients corresponding to the loss. Curvature-weighted chamfer loss, This can be expressed by the following formula:

[0112] (3);

[0113] in, and These are the point sets of the real grid and the predicted grid, respectively. , It is a point u curvature, is the nearest neighbor of point v in the point set , is the nearest neighbor of point u in the point set , is a point in the point set ;

[0114] Grid-to-grid normal consistency loss The calculation formula is as follows:

[0115] (4);

[0116] wherein, is the normal, S is the number of deformation stages, and C is the number of meshes.

[0117] wherein, the regularization loss is represented by the following formula:

[0118] (5);

[0119] wherein, is the Laplacian smoothing loss of displacement distance, is the normal consistency loss within a single mesh, is the edge length loss, is the predicted displacement vector, , and are weight coefficients of the corresponding loss, respectively.

[0120] Laplacian smoothing loss of displacement distance is represented by the following formula:

[0121] (6);

[0122] Normal consistency loss within a single mesh is represented by the following formula:

[0123] (7);

[0124] Edge length loss is represented by the following formula:

[0125] (8);

[0126] L is the Laplace operator of the mesh, is the vertex set of the predicted mesh, and are adjacent faces, is the edge set of the predicted mesh, and is the side of the two vertices.

[0127] It should be noted that the flowchart and block diagrams in the drawings show the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagrams can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations thereof, can be implemented by a dedicated hardware-based system that performs the specified functions or acts or combinations thereof, or can be implemented by a combination of dedicated hardware and computer instructions.

[0128] Those skilled in the art can clearly understand that, for the convenience and brevity, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0129] In several embodiments provided in the present application, it should be understood that the disclosed apparatus and method can be implemented by other ways. The apparatus embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there can be another division manner in actual implementation, and for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some communication interfaces, and can be electrical, mechanical or other forms.

[0130] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0131] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit.

[0132] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0133] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and they should be covered in the scope of the claims and the specification of the present application.

Claims

1. A three-dimensional mesh reconstruction method based on large model adaptive feature extraction, characterized in that, The method comprises the following steps: S1: constructing a three-dimensional grid reconstruction model, the three-dimensional grid reconstruction model comprising a large model robust feature extraction module, a dynamic semantic feature selection module, a voxel-point cloud mapping module and a multi-scale reconstruction module; S2: obtaining a medical input image, extracting features of the medical input image through the large model robust feature extraction module to obtain image extraction features; S3: confirming a target medical task, constructing a semantic prompt word based on the medical input image and the target medical task, encoding the semantic prompt word through a pre-trained language large model to obtain image semantic features; S4: fusing the image extraction features and the image semantic features based on the dynamic semantic feature selection module to obtain dynamic semantic features of the medical input image; S5: performing spatial mapping of the dynamic semantic features from voxel to point cloud based on the voxel-point cloud mapping module to obtain multi-scale point cloud image spatial features; S6: deforming an initial grid template based on the point cloud image spatial features through the multi-scale reconstruction module to obtain a target three-dimensional grid corresponding to the medical input image; The multi-scale reconstruction module comprises a plurality of GCN up-sampling modules, each GCN up-sampling module comprising a cascaded up-sampling layer, a convolution layer and an activation layer, and step S6 comprises: S61: sequentially inputting the multi-scale point cloud image spatial features into the plurality of GCN up-sampling modules, and expanding the grid vertex number of the input point cloud image spatial features to twice the original number at each level of the GCN up-sampling module; S62: fusing and tri-linearly interpolating the point cloud image spatial features after the grid number expansion to obtain the target three-dimensional grid corresponding to the medical input image.

2. The method of claim 1, wherein, Step S2 comprises: S21: obtaining a medical input image and preprocessing the medical input image; S22: inputting the preprocessed medical input image into a pre-trained visual large model SAM encoder to extract global features and local features of the medical input image; S23: performing global average pooling, up-sampling and convolution processing on the global features and local features to obtain image extraction features.

3. The method of claim 2, wherein, Step S3 comprises: S31: confirming a target medical task based on the medical input image; S32: constructing a semantic prompt word based on the medical input image and the target medical task; S33: segmenting the semantic prompt word into a plurality of discrete sequences through a word segmenter; S34: mapping the plurality of discrete sequences into high-dimensional continuous vectors through an embedding layer; S35: processing the high-dimensional continuous vectors through a dimension reduction processing module and performing feature pooling to obtain image semantic features.

4. The method of claim 3, wherein, The dynamic semantic feature extraction module comprises a plurality of trainable convolution layers, and step S4 comprises: S41: point multiplying the image extraction features and the image semantic features to obtain a point multiplication result; S42: fusing the point multiplication result to obtain dynamic semantic features of the medical input image.

5. The method of claim 4, wherein, Step S5 comprises: S51: Perform trilinear interpolation for each voxel feature of the dynamic semantic feature, and map to the vertices of the initial mesh template; S52: Perform dynamic semantic feature learning based on all vertices mapped to the initial mesh template, and generate a multi-scale point cloud image space feature.

6. The method of claim 5, wherein, The mesh loss function of any one mesh of the point cloud image space feature is: (1); where, pred and true respectively represent the predicted and true values of the mesh, is the geometry consistency loss, is the regularization loss.

7. The method of claim 6, wherein, geometric consistency loss is expressed by the following equation: (2); wherein, is the predicted grid for the Sth stage, is the ground truth grid, and is the weight coefficient for the corresponding loss, is the curvature weighted chamfer loss, is represented by the following equation: (3); wherein, and are respectively the point sets of the real grid and the predicted grid, , is the curvature of the point u , is the nearest neighbor of the point v in the point set , is the nearest neighbor of u in the point set , is a point in the point set ; For the inter-grid normal consistency loss, the calculation formula is as follows: (4); wherein, is the normal, S is the number of deformation stages, and C is the number of meshes.

8. The method of claim 7, wherein, regularization loss is expressed by the following equation: (5); wherein, is a Laplacian smoothing loss for displacement distance, is a normal consistency loss within a single mesh, is a side length loss, is a predicted displacement vector, , and are weight coefficients for the corresponding losses, respectively. laplacian smoothing loss of displacement distance is expressed by the following equation: (6); Single grid-in method normal consistency loss is expressed by the following equation: (7); side length loss is expressed by the following equation: (8); L is the Laplacian of the mesh, is a set of vertices of the prediction mesh, and is an adjacent face, is a set of edges of the prediction mesh, and are two vertices of an edge of the prediction mesh.

Citation Information

Patent Citations

  • Surface grid reconstruction method for medical image

    CN118470253A

  • Three-dimensional grid reconstruction method based on adaptive template

    CN118470254A