Multi-modal industrial anomaly detection method fusing image masks and hybrid experts
Through the multimodal industrial anomaly detection method that combines image masks and mixed expert models, the problem of difficult modeling of modal missing and deep semantic associations in the prior art is solved, and more efficient and robust industrial anomaly detection performance is achieved.
Patent Information
- Application Number
- CN202510469130.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing multimodal industrial anomaly detection model is difficult to effectively deal with the problem of modal missing, and traditional methods rely on precise sensor calibration, making it difficult to model deep semantic correlations between modals.
Using a multimodal industrial anomaly detection method that integrates image masks and hybrid expert models, the image mask training strategy uses the image mask training strategy to reconstruct the CLIP visual features of the mask image as the pre-training goal, and build a multimodal industrial anomaly detection large model, including the linear projection layer, the embedding layer, the ViT encoder module and the MLP output head of 3D data.
This method can effectively deal with the problem of modal missing without relying on precise sensor calibration, improve the accuracy and robustness of abnormal detection, and enhance the migration performance of the model in downstream tasks.
Smart Images

Figure CN119992234A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of industrial anomaly detection, and in particular to a multimodal industrial anomaly detection method integrating image mask and hybrid experts. Background Art
[0002] In recent years, large models can be effectively transferred to various downstream tasks and maintain good performance by pre-training on large amounts of data. Masked image modeling (MIM) performs masking operations on the input image, and through training, the model can reconstruct the masked part based on the visible image blocks. However, there are obvious limitations in using original pixels as the reconstruction target. Since natural original images are information-sparse, simple pixel-level restoration tasks can often only capture low-level geometric and structural information, while ignoring high-level semantic features. Therefore, it is difficult to maintain good performance under a billion-scale model size. The pre-training target is to reconstruct the image-text aligned visual features of the masked image blocks. This pre-training task can simultaneously capture the geometric structure information and high-level semantic abstract features of natural images, covering the information required for most visual tasks, and can achieve good performance in a wide range of downstream tasks.
[0003] As the parameters of large models become larger and their structures become more complex, the parameters expand to billions or even trillions of sizes, and how to maintain efficiency and accuracy becomes a challenge. Traditional models activate all layers and neurons in the network for each input, which often leads to huge computational costs and memory consumption. The mixture of experts (MoE) model is an advanced network architecture that dynamically selects and activates some experts to process different inputs during training and inference, which can reduce computational overhead while maintaining model performance. MoE includes a gating network and multiple expert networks. The gating network acts as a selector, sparsely activating a small number of experts for each input to process data.
[0004] Traditional multimodal methods usually achieve 2D and 3D fusion through projection or feature stitching, such as projecting the point cloud onto the image plane to generate a depth map, or combining the features of the two modalities through early / late fusion strategies. However, these methods often rely on accurate sensor calibration and have difficulty modeling deep semantic associations between modalities. In addition, multimodal data in industrial scenarios often have noise, occlusion, and sparsity problems (such as missing part of the structure of the point cloud), which further increases the difficulty of fusion. At the same time, most current multimodal anomaly detection models cannot effectively handle the problem of missing modalities, or the model performance degrades severely when the modality is missing. Summary of the invention
[0005] The purpose of this invention is to fuse image mask and hybrid expert model, and propose a novel industrial multimodal large model pre-training method. This method uses image mask training strategy, takes CLIP visual features of reconstructed mask image as pre-training target, does not require expensive supervised training but is based on self-supervised training, expands the scale of pre-trained large model of the model, can lead to qualitative change of learning performance in migration, and enables large model to be better migrated to downstream tasks.
[0006] The object of the present invention is achieved through the following technical solution: a multimodal industrial anomaly detection method integrating image mask and hybrid expert, comprising the following steps:
[0007] S1, collect high-resolution aligned 2D images and 3D point cloud data of industrial components and perform preprocessing;
[0008] S2, projecting the 3D point cloud data onto the 2D image plane to construct a depth map, and projecting the depth map to a dimension aligned with the 2D image;
[0009] S3, dividing and randomly masking the 2D image and the projected 3D image;
[0010] S4. Construct a large multimodal industrial anomaly detection model, wherein the large multimodal industrial anomaly detection model includes a linear projection layer of 3D data, an embedding layer, a ViT encoder module including a mixed expert layer, and an MLP output head; and pre-train the CLIP features of the masked image blocks as reconstruction targets;
[0011] S5. Freeze the linear projection layer of the 3D data, fine-tune the model, and use the pre-trained and fine-tuned model for industrial anomaly detection.
[0012] Furthermore, the acquisition of high-resolution aligned 2D images and 3D point cloud data of industrial components specifically includes: adding multimodal data of abnormal industrial components with various abnormal manifestations based on artificial simulation of defects, and acquiring paired 2D images and 3D point cloud data of the same object.
[0013] Furthermore, the preprocessing includes: performing dedistortion and normalization operations on the 2D image, correcting the lens distortion according to the camera internal parameters, removing outliers based on the DBSCAN algorithm for the 3D point cloud data, and adding Gaussian noise.
[0014] Furthermore, the projecting of the 3D point cloud data onto the 2D image plane to construct a depth map includes: using the camera intrinsic parameter K and the extrinsic parameter (R, t) to project the point cloud Projection onto a plane:
[0015]
[0016] According to the projection coordinates and depth value Generate Depth Map .
[0017] Furthermore, dividing and randomly masking the 2D image and the projected 3D image includes: inputting each modality into an image Divide into image blocks, where is the image resolution, is the number of channels, and the size of each image block is ;
[0018] The obtained multimodal image blocks are randomly sampled, and the input image blocks are destroyed with mask marks, the average mask rate of the two-modal image blocks is kept at 40%, and the unmasked image blocks are spliced.
[0019] Furthermore, the reconstruction target is specifically: splicing the divided 2D image and the projected 3D image block, using a CLIP encoder to extract image features, and the features are used as mask reconstruction targets for subsequent pre-training tasks.
[0020] Furthermore, the multimodal industrial anomaly detection large model specifically includes: a linear projection layer of 3D data, an embedding layer, several ViT encoder layers and an MLP output head, wherein the unmasked visible image blocks are mapped to the D-dimensional space through the linear projection layer and the position embedding is added, and input to the ViT encoder layer, wherein the VIT encoder module includes a multi-head attention layer and a hybrid expert layer, each hybrid expert layer includes a gating network and several expert layers, and the obtained output is projected to the same dimension as the CLIP visual feature through the MLP output head.
[0021] Furthermore, the pre-training specifically includes: taking CLIP visual features as reconstruction targets, calculating reconstruction losses using negative cosine similarity for back propagation, and the loss function is expressed as:
[0022]
[0023] in, is the visual feature output by the CLIP encoder, is the output of the multimodal industrial anomaly detection model, and N is the total number of image patches.
[0024] On the other hand, a multimodal industrial anomaly detection device that fuses image masks and mixed experts is also provided, including a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, the multimodal industrial anomaly detection method that fuses image masks and mixed experts is implemented.
[0025] On the other hand, a computer-readable storage medium is also provided, on which a program is stored. When the program is executed by a processor, the multimodal industrial anomaly detection method that integrates image masks and mixed experts is implemented.
[0026] Beneficial effects of the present invention: Compared with the single-modal anomaly detection model, the multi-modal model can effectively utilize the characteristics of different modalities to make up for the shortcomings of other modalities. For example, the 2D image modality is more sensitive to the color attributes of the object, and can capture all the information on the surface of the object, but lacks the depth attribute; similarly, the 3D point cloud can effectively express the depth information of the object, but cannot perceive the internal information of the object. The multi-modal industrial anomaly detection large model described in the present invention will have a better anomaly detection effect than the single-modal model. At the same time, in actual detection, the data processed by the model described in the present invention will not be limited to single-modal or multi-modal input data, can effectively deal with the problem of missing modalities, and has rich application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A flow chart of a multi-modal industrial large model anomaly detection method integrating image mask and hybrid experts implemented by the present invention;
[0028] Figure 2 This is a flow chart of the multimodal model pre-training method based on image mask reconstruction according to the present invention;
[0029] Figure 3 This is a schematic diagram of the structure of the multi-modal industrial anomaly detection large model of the present invention;
[0030] Figure 4 is a schematic diagram of the hybrid expert layer structure in the present invention;
[0031] Figure 5 A multimodal industrial anomaly detection device that integrates image masks and mixed experts is provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0032] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.
[0033] A multimodal industrial anomaly detection method integrating image mask and mixed experts, such as Figure 1 As shown, the specific steps are as follows:
[0034] (1) Based on high-definition industrial cameras, high-precision millimeter-wave radars, lidars and other equipment, high-resolution aligned 2D images and 3D point cloud data of industrial components are collected, and 2D and 3D multimodal data of abnormal industrial components are added by artificially simulating defects. Data preprocessing: clean the obtained industrial multimodal data set, remove noise and low-quality data, and normalize and standardize the data to facilitate subsequent model processing.
[0035] (2) 3D point cloud processing: The input is a pair of aligned 2D images and 3D point cloud data. The calibrated camera and radar parameters are used to project the data into a depth map. After a linear projection layer, the depth map obtained by projecting the 3D point cloud is reshaped to be consistent with the 2D image, and the two modal data are aligned in the same coordinate system.
[0036] (3) Image segmentation and masking: The 2D image and the projected 3D image are divided into image blocks of size 16×16. Some image blocks are randomly masked in proportion, that is, the content of the corresponding image blocks is destroyed. The mask reconstruction pre-training task aims to reconstruct the CLIP features of the masked image blocks.
[0037] (4) Model pre-training. The present invention uses image mask reconstruction as a pre-training task, and the reconstruction target is the CLIP visual feature.
[0038] (5) Model fine-tuning and anomaly detection task application. The multimodal industrial anomaly model proposed in this invention can process single-modal data sets, such as pure image data or pure 3D point cloud data, and can also process multimodal data.
[0039] The multimodal data obtained in step (1) needs to be paired 2D images and 3D point cloud data based on the same object. The data set is required to meet the ratio of normal samples to abnormal samples of no less than 10:1. The multimodal data of abnormal industrial components added by artificial simulation of defects, etc., must include various forms of abnormalities such as scratches, dents, cracks, and surface damage.
[0040] In step (1), the data preprocessing operation is to dedistort and normalize the 2D image, and correct the lens distortion according to the camera intrinsic parameters. For the 3D point cloud data, outliers are removed based on the DBSCAN algorithm and appropriate Gaussian noise is added.
[0041] In step (2), the projection operation of the 3D point cloud data is performed by using the camera internal parameters (K) and external parameters (R, t) to transform the point cloud Projection onto a plane:
[0042]
[0043] According to the projection coordinates and depth value Generate Depth Map .
[0044] In step (2), a linear projection layer is used to transform the depth map Projected to a dimension aligned with the 2D image, the same image segmentation method and Transformer encoder can be used to extract features of the two modalities.
[0045] In step (3), for the aligned multimodal data, i.e., the 3D point cloud image block and the 2D image block, each modality is input to the image Divide into image blocks, where is the image resolution, is the number of channels, and the size of each image block is The divided multimodal image blocks are spliced and a CLIP encoder is used to extract image features, which are used as the mask reconstruction target for subsequent pre-training tasks.
[0046] Regarding the CLIP encoder described in step (3), CLIP (Contrastive Language-Image Pre-Training) is contrastive text-image pre-training. The model is pre-trained with paired text-image data as input. It includes an encoder for image training and an encoder for text training. Both use the Transformer encoder structure. The model learns the matching relationship between text-image pairs during the training phase. CLIP can effectively extract image-text alignment visual features, and the extracted image-text alignment visual features are the target of image mask reconstruction. The image features output by the CLIP encoder are ,in Image encoder for CLIP.
[0047] Regarding the image segmentation operation in step (3), the obtained multimodal image blocks are randomly sampled, and the input image blocks are destroyed with mask marks, keeping the average mask rate of the two-modal image blocks at 40%, and the masked image blocks are spliced.
[0048] The model pre-training task process in step (4) is as follows: Figure 2 As shown, the left side of the figure shows data preprocessing and image segmentation operations. The pre-training process will continuously train the weights of the linear projection layer for processing 3D data. The multimodal image blocks obtained after image segmentation and image masking operations are extracted by the large industrial anomaly detection model described in the present invention, and then a linear projection layer is used to project the multimodal features to the same dimension as the CLIP visual features. The goal of pre-training is to reconstruct the multimodal features into CLIP visual features, and use negative cosine similarity as the mask reconstruction loss function. This training process can effectively improve the model's ability to extract deep-level image features. During the pre-training process, the parameters related to the frozen CLIP encoder will be frozen. Based on the excellent feature extraction ability of CLIP, the parameters of the linear projection layer of 3D data, the large anomaly detection model, and the terminal linear projection are guided to update.
[0049] Regarding the multimodal industrial anomaly detection model mentioned in step (4), its structure is as follows: Figure 3As shown, it includes a linear projection layer of 3D data, an embedding layer, N Vision Transformer (ViT) encoder modules, and an MLP output head.
[0050] Regarding the image embedding module in step (4), the image block is mapped to the D-dimensional space through the linear projection layer, and the position embedding is added to obtain the sequence ,in is the vector of the i-th image block, E is the projection matrix, Embed information for the position of the i-th image block, is the modality information of the i-th image block, which is input to the ViT encoder.
[0051] The output features of the visible image block after being processed by the ViT encoder. Among the N-layer encoders, the shallow encoders focus on extracting local features of the image data and are sensitive to position information. The high-level encoders focus more on extracting global features of the image and have better robustness to local disturbances (such as noise and occlusion in the image). The multi-layer encoder design enables the model to locate anomalies more accurately. After being projected to the same dimension as the CLIP visual feature through the linear projection layer, ,in is the output feature after encoder processing, is the projection matrix.
[0052] The CLIP visual feature is used as the reconstruction target, and the reconstruction loss is calculated using negative cosine similarity, which is expressed as:
[0053]
[0054] The structure of the Mixture of Experts (MoE) in the ViT encoder structure in the industrial multimodal anomaly detection model in step (4) is as follows: Figure 4 As shown in the figure, MoE is a machine learning model that can activate only part of the network when processing input data, so as to speed up training and reasoning. In the encoder architecture of ViT, the input sequence will pass through the multi-head attention layer and FFN layer of each layer, and MoE is introduced to replace the FFN layer. MoE includes 1 gated network and n expert networks.
[0055] The hybrid expert neural network used in the present invention specifically includes the following expert networks:
[0056] 1 Fully Connected Expert: It includes two fully connected layers, and its mathematical expression is:
[0057]
[0058] Among them, the input of each expert network is x, , and , are the weight matrices and biases of the two fully connected layers.
[0059] 2Convolutional Experts: includes one or more convolutional layers, and its mathematical expression is:
[0060]
[0061] Among them, Conv is the convolution function, K is the convolution kernel size, and b is the convolution layer bias.
[0062] 3Zero Expert: Directly discard the input and output a zero vector. Its mathematical expression is:
[0063]
[0064] 4 Copy Expert: Skip the current expert and return to the input directly. Its mathematical expression is:
[0065]
[0066] 5 Constant Expert: Perform constant correction and replacement on input. Its mathematical expression is:
[0067] or
[0068] The gating network of the MoE layer is responsible for selecting a sparse expert combination for each input token. For input x, the output of the MoE layer is ,in is the output of the gating network, is the output of the i-th expert network. The gated network is a network with softmax. ,in is the weight matrix of the gating network for each expert, and the obtained The scores of each expert.
[0069] The selection of experts is based on Top-K gating, which selects the k experts with the highest scores and sets the others to 0.
[0070]
[0071] Based on this, , only the first k experts with the highest scores are retained to process the input data, and the remaining experts do not perform calculations, which can effectively reduce computing resources.
[0072] Regarding the model fine-tuning mentioned in step (5), after the pre-training task is completed, the multimodal anomaly model of the present invention has good generalization ability to process multimodal industrial data. When processing different industrial data sets in actual applications, only fine-tuning is required to achieve good performance. During model fine-tuning, the parameters of the image embedding module, the MoE layer in the ViT encoder, and the forward propagation network will be updated, while the 3D data projection module will remain frozen.
[0073] Compared with single-modal anomaly detection models, multimodal models can effectively utilize the characteristics of different modalities and make up for the shortcomings of other modalities. For example, the 2D image modality is more sensitive to the color attributes of objects and can capture all the information on the surface of objects, but lacks depth attributes; similarly, 3D point clouds can effectively express the depth information of objects, but cannot perceive the internal information of objects.
[0074] Based on the multimodal industrial anomaly detection model described in the present invention, the backbone structure for multimodal data feature extraction is the same ViT encoder, so it can cope with the problem of missing modalities. In actual detection, the processed data will not be limited to single-modal or multi-modal input data, and the application scenarios are rich.
[0075] Corresponding to the aforementioned embodiment of a multimodal industrial anomaly detection method integrating image masks and mixed experts, the present invention also provides an embodiment of a multimodal industrial anomaly detection device integrating image masks and mixed experts.
[0076] See also Figure 5 A multimodal industrial anomaly detection device that integrates image masks and mixed experts provided in an embodiment of the present invention includes a memory and one or more processors. The memory stores executable codes. When the processor executes the executable codes, it is used to implement a multimodal industrial anomaly detection method that integrates image masks and mixed experts in the above-mentioned embodiment.
[0077] An embodiment of a multimodal industrial anomaly detection device that integrates image masks and hybrid experts provided by the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From a hardware perspective, if Figure 5As shown, a hardware structure diagram of a multi-modal industrial anomaly detection device integrating image mask and hybrid expert provided by the present invention is provided for any device with data processing capability, except Figure 5 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0078] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0079] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0080] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, a multimodal industrial anomaly detection method integrating image mask and mixed expert in the above embodiment is implemented.
[0081] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capability. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.
[0082] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the multimodal industrial anomaly detection method that integrates image masks and mixed experts.
[0083] Those skilled in the art will readily appreciate other embodiments of the present application after considering the description and practicing the contents disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The description and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the claims.
[0084] It should be understood that the above general description and the detailed description below are only exemplary and explanatory and cannot limit the present application. The present application is not limited to the precise structure described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is limited only by the attached claims.
Claims
1. A multimodal industrial anomaly detection method integrating image mask and hybrid experts, characterized in that: The steps include: S1, collect high-resolution aligned 2D images and 3D point cloud data of industrial components and perform preprocessing; S2, projecting the 3D point cloud data onto the 2D image plane to construct a depth map, and projecting the depth map to a dimension aligned with the 2D image; S3, dividing and randomly masking the 2D image and the projected 3D image; S4. Construct a large multimodal industrial anomaly detection model, wherein the large multimodal industrial anomaly detection model includes a linear projection layer of 3D data, an embedding layer, a ViT encoder module including a mixed expert layer, and an MLP output head; and pre-train the CLIP features of the masked image blocks as reconstruction targets; S5. Freeze the linear projection layer of the 3D data, fine-tune the model, and use the pre-trained and fine-tuned model for industrial anomaly detection.
2. A multimodal industrial anomaly detection method integrating image mask and hybrid experts according to claim 1, characterized in that: The acquisition of high-resolution aligned 2D images and 3D point cloud data of industrial components specifically includes: adding multimodal data of abnormal industrial components with various abnormal manifestations based on artificial simulation of defects, and acquiring paired 2D images and 3D point cloud data for the same object.
3. The multimodal industrial anomaly detection method integrating image mask and hybrid experts according to claim 1, characterized in that: The preprocessing includes: performing dedistortion and normalization operations on the 2D image, correcting the lens distortion according to the camera internal parameters, removing outliers based on the DBSCAN algorithm and adding Gaussian noise to the 3D point cloud data.
4. The multimodal industrial anomaly detection method integrating image mask and hybrid experts according to claim 1, characterized in that: The projecting of 3D point cloud data onto a 2D image plane to construct a depth map includes: projecting the point cloud onto the plane using camera intrinsic parameters and extrinsic parameters, and generating a depth map according to projection coordinates and depth values of the point cloud.
5. The multimodal industrial anomaly detection method integrating image mask and hybrid experts according to claim 1, characterized in that: The 2D image and the projected 3D image are divided and randomly masked, including: dividing each modality input image into a number of image blocks according to the size of the image block, randomly sampling the obtained multimodal image blocks, destroying the input image blocks with mask marks, maintaining the average mask rate of the two modal image blocks at 40%, and splicing the unmasked image blocks.
6. The multimodal industrial anomaly detection method integrating image mask and hybrid experts according to claim 1, characterized in that: The reconstruction goal is specifically to splice the divided 2D image and the projected 3D image block, and use a CLIP encoder to extract image features, which are used as mask reconstruction targets for subsequent pre-training tasks.
7. The multimodal industrial anomaly detection method integrating image mask and hybrid experts according to claim 1, characterized in that: The multimodal industrial anomaly detection large model specifically includes: a linear projection layer for 3D data, an embedding layer, several ViT encoder layers and an MLP output head, wherein the unmasked visible image blocks are mapped to a D-dimensional space through the linear projection layer and position embedding is added, and then input to the ViT encoder layer, wherein the VIT encoder module includes a multi-head attention layer and a hybrid expert layer, each hybrid expert layer includes a gating network and several expert layers, and the obtained output is projected to the same dimension as the CLIP visual feature through the MLP output head.
8. The multimodal industrial anomaly detection method integrating image mask and hybrid experts according to claim 1, characterized in that: The pre-training specifically includes: taking CLIP visual features as reconstruction targets, calculating reconstruction losses with negative cosine similarity according to visual features output by the CLIP encoder and output of a large multimodal industrial anomaly detection model, and performing back propagation.
9. A multimodal industrial anomaly detection device integrating image mask and hybrid expert, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, a multimodal industrial anomaly detection method integrating image mask and hybrid expert as described in any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, a multimodal industrial anomaly detection method integrating image mask and mixed experts as described in any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Detection method for opening and closing state of disconnecting link, medium and system
CN112837262A
Video base model acquisition method based on non-mask alignment
CN116310995A
Pre-training method, device and equipment of image-text understanding model and storage medium
CN116796287A
Hybrid expert target detection system and method
CN118675030A
Large-format remote sensing image semantic segmentation method for super-long context modeling
CN119478403A
Cited By
Vehicle condition estimation method, device and equipment based on multi-modal data fusion, storage medium and product
CN121010770A
Model calculation path dynamic reconstruction method and system based on visual prompt
CN121724165A
Multi-modal anomaly prediction method and system for logistics supply chain based on hybrid expert model
CN122333268A
City downstream task prediction method, device and equipment
CN122452880A