Image data processing method and device, equipment and medium

By using feature extraction methods based on the splicing order of the target sequence and the encoding order sequence, and combining them with attention matrix for MRI image sequence feature fusion, the problem of repeated extraction is solved, and the feature fusion capability and efficiency are improved.

CN121330433APending Publication Date: 2026-01-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410924406.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-07-10
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing medical image feature extraction models repeatedly extract image features when stitching together MRI image sequences, resulting in wasted computational resources and reduced feature fusion capabilities.

Method used

By obtaining the concatenation order and encoding order of the target sequence, we perform first feature extraction and multi-business dimension feature extraction, and combine the attention matrix to perform feature fusion to generate potential sequence features.

Benefits of technology

It effectively reduces the waste of computing resources, enhances feature fusion capabilities, and improves the efficiency and accuracy of MRI image sequence feature extraction and fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330433A_ABST
    Figure CN121330433A_ABST
Patent Text Reader

Abstract

The invention provides an image data processing method and apparatus, a device and a medium. The method comprises the steps of obtaining N target input image sequences and a target spliced image sequence; obtaining a target coding sequence, and performing first feature extraction processing on each target input image sequence in the N target input image sequences to obtain a target input image feature of each target input image sequence; performing second feature extraction processing on the target spliced image sequence on multiple service dimensions, and performing feature fusion to obtain target fusion sequence features; and determining a target attention matrix based on the target input image features of each target input image sequence, and determining target potential sequence features associated with the N target input image sequences based on the target attention matrix and the target fusion sequence features. By adopting the method and the device, the waste of computing resources can be reduced, effective fusion of image sequence features can be ensured, and the feature fusion capability is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to an image data processing method and device, equipment and medium. BACKGROUND

[0002] Magnetic resonance imaging (MRI) is a non-invasive medical imaging technology, which has the advantages of high resolution, good contrast, no radiation damage, etc. At present, different MRI image sequences can be obtained through this medical imaging technology, and then different MRI sequence features can be extracted from these MRI image sequences through artificial intelligence technology, so as to provide certain diagnostic auxiliary information for subsequent medical diagnosis through the extracted MRI sequence features.

[0003] However, the inventors have found in practice that, in the case of obtaining MRI image sequences of different modalities, the obtained MRI image sequences of different modalities can be input into a medical image feature extraction model in a sequence connection manner, so that the medical image feature extraction model can further perform image feature extraction on the concatenated image sequence in the image dimension. However, since the different MRI image sequences participating in sequence concatenation are all from the same object, in the process of performing image feature extraction on the concatenated image sequence by the existing medical image feature extraction model, there is a phenomenon that the image features of a large number of similar regions in the concatenated image sequence are repeatedly extracted. Therefore, when performing feature fusion on the concatenated image sequence, the repeatedly extracted image features are also fused without distinction, which not only causes waste of computing resources to some extent, but also reduces the feature fusion capability of multi-sequence MRI images. SUMMARY

[0004] The embodiments of the present application provide an image data processing method, device, equipment and medium, which can not only reduce the waste of computing resources, but also ensure the effective fusion of image sequence features, thereby improving the feature fusion capability.

[0005] The embodiments of the present application provide an image data processing method, device, equipment and medium, which can not only reduce the waste of computing resources, but also ensure the effective fusion of image sequence features, thereby improving the feature fusion capability.

[0006] obtaining N target input image sequences and target concatenated image sequences corresponding to the N target input image sequences; the target concatenated image sequence refers to an image sequence obtained by sequentially concatenating the N target input image sequences according to a target sequence concatenation order; N is a positive integer greater than 1;

[0007] obtain a target coding sequence corresponding to the target sequence splicing order, perform first feature extraction processing on each of the N target input image sequences based on the target coding sequence, and obtain target input image features of each of the target input image sequences;

[0008] perform second feature extraction processing on the target spliced image sequence in multiple business dimensions, obtain target business dimension features of the target spliced image sequence in the multiple business dimensions, perform feature fusion on the target business dimension features in the multiple business dimensions, and obtain a target fusion sequence feature of the target spliced image sequence;

[0009] determine a target attention matrix used for attention calculation based on the target input image features of each of the target input image sequences, and determine target latent sequence features associated with the N target input image sequences based on the target attention matrix and the target fusion sequence feature.

[0010] Embodiments of the present application provide an image data processing device, comprising:

[0011] a sequence acquisition module configured to acquire N target input image sequences and target spliced image sequences corresponding to the N target input image sequences; the target spliced image sequence refers to an image sequence obtained by sequentially splicing the N target input image sequences according to a target sequence splicing order; N is a positive integer greater than 1;

[0012] a first extraction module configured to obtain a target coding sequence corresponding to the target sequence splicing order, perform first feature extraction processing on each of the N target input image sequences based on the target coding sequence, and obtain target input image features of each of the target input image sequences;

[0013] a second extraction module configured to perform second feature extraction processing on the target spliced image sequence in multiple business dimensions, obtain target business dimension features of the target spliced image sequence in the multiple business dimensions, perform feature fusion on the target business dimension features in the multiple business dimensions, and obtain a target fusion sequence feature of the target spliced image sequence;

[0014] a feature determination module configured to determine a target attention matrix used for attention calculation based on the target input image features of each of the target input image sequences, and determine target latent sequence features associated with the N target input image sequences based on the target attention matrix and the target fusion sequence feature.

[0015] Optionally, the device further comprises a sequence generation module.

[0016] The sequence generation module is configured to generate a target missing image sequence based on a target latent sequence feature if the target coding order sequence indicates that there is a target missing image sequence in the N target input image sequences; the target missing image sequence refers to an image sequence in which an input modality is missing in the N input modalities corresponding to the N target input image sequences; and one target input image sequence corresponds to one input modality.

[0017] The first extraction module includes an encoding processing unit and a feature extraction unit.

[0018] The encoding processing unit is configured to obtain a sequence splicing order, and perform encoding processing on the N target input image sequences based on the sequence splicing order by the binary coding component to obtain a target coding order sequence corresponding to the sequence splicing order.

[0019] The feature extraction unit is configured to input the target coding order sequence and each of the N target input image sequences into a target image processing model, and perform first feature extraction processing on each of the target input image sequences based on the target coding order sequence by the target image processing model to obtain target input image features of each of the target input image sequences.

[0020] The target image processing model includes a first feature extraction component.

[0021] The feature extraction unit includes an encoding matching subunit and a second feature extraction subunit.

[0022] The encoding matching subunit is configured to obtain N binary encodings from the target coding order sequence, and perform binary encoding matching processing on each of the N target input image sequences based on the target sequence splicing order to obtain a matching encoding corresponding to each of the N target input image sequences.

[0023] The target feature extraction subunit is configured to input each of the N target input image sequences and the matching encoding corresponding to each of the target input image sequences into the first feature extraction component, and perform first feature extraction processing on each of the target input image sequences based on the matching encoding by the first feature extraction component to obtain target input image features of each of the target input image sequences.

[0024] The N target input image sequences include a first input image sequence and a second input image sequence, the matching encoding corresponding to the first input image sequence is a first matching encoding, and the matching encoding corresponding to the second input image sequence is a second matching encoding.

[0025] The target feature extraction subunit is also specifically configured to input the first input image sequence and the first matching code into a first feature extraction component, perform first feature extraction processing on the first input image sequence and the first matching code by the first feature extraction component, and obtain first extracted features corresponding to the first input image sequence and a first code vector corresponding to the first matching code;

[0026] The target feature extraction subunit is also specifically configured to perform feature fusion on the first extracted features and the first code vector to obtain first target extracted features corresponding to the first input image sequence.

[0027] The target feature extraction subunit is also specifically configured to input the second input image sequence and the second matching code into the first feature extraction component, perform first feature extraction processing on the second input image sequence and the second matching code by the first feature extraction component, and obtain second extracted features corresponding to the second input image sequence and a second code vector corresponding to the second matching code.

[0028] The target feature extraction subunit is also specifically configured to perform feature fusion on the second extracted features and the second code vector to obtain second target extracted features corresponding to the second input image sequence.

[0029] The target feature extraction subunit is also specifically configured to use the first target extracted features and the second target extracted features as target input image features of each target input image sequence.

[0030] The second extraction module comprises a multi-dimensional feature extraction unit, a target dimension determination unit, a target feature determination unit, and a feature fusion unit.

[0031] The multi-dimensional feature extraction unit is configured to input the target spliced image sequence into a target image processing model, perform dimension feature extraction processing on the target spliced image sequence in a plurality of business dimensions by a second feature extraction component in the target image processing model, and obtain business dimension features of the target spliced image sequence in the plurality of business dimensions.

[0032] The target dimension determination unit is configured to determine a target business dimension from the plurality of business dimensions, and obtain target business dimension features of the target spliced image sequence in the target business dimension from the business dimension features of the target spliced image sequence in the plurality of business dimensions.

[0033] The target feature determination unit is configured to obtain the target business dimension features of the target spliced image sequence in the plurality of business dimensions when each business dimension in the plurality of business dimensions is determined as the target business dimension.

[0034] The feature fusion unit is configured to perform feature fusion on the target business dimension features of each target input image sequence in the plurality of business dimensions to obtain target fused sequence features of the target spliced image sequence.

[0035] wherein the plurality of service dimensions comprise a channel dimension, a height dimension, and a width dimension;

[0036] The target dimension determination unit comprises a target dimension determination subunit, a first feature determination subunit, a pooling processing subunit, a second feature determination subunit, and a weighting processing subunit.

[0037] The target dimension determination subunit is configured to, if the channel dimension is determined as the target service dimension among the plurality of service dimensions, determine the service dimensions other than the channel dimension among the plurality of service dimensions as the to-be-processed service dimensions; a first service dimension among the to-be-processed service dimensions is the height dimension, and a second service dimension among the to-be-processed service dimensions is the width dimension.

[0038] The first feature determination subunit is configured to acquire a first service dimension feature of the target spliced image sequence in the first service dimension and a second service dimension feature of the target spliced image sequence in the second service dimension from the service dimension features of the target spliced image sequence in the plurality of service dimensions.

[0039] The pooling processing subunit is configured to perform block processing on the service feature map constituted by the first service dimension feature in the first service dimension and the second service dimension feature in the second service dimension to obtain a plurality of block feature maps corresponding to the service feature map, perform pooling processing on the plurality of block feature maps to obtain a pooling feature of the plurality of block feature maps, and take the pooling feature of the plurality of block feature maps as a feature weight of the target spliced image sequence in the target service dimension.

[0040] The second feature determination subunit is configured to acquire a third service dimension feature of the target spliced image sequence in the target service dimension from the service dimension features of the target spliced image sequence in the plurality of service dimensions.

[0041] The weighting processing subunit is configured to perform weighting processing on the third service dimension feature in the target service dimension based on the feature weight in the target service dimension to obtain a target service dimension feature of the target spliced image sequence in the target service dimension.

[0042] The feature determination module comprises a first matrix processing unit, a second matrix processing unit, a third matrix processing unit, and a latent sequence feature determination unit.

[0043] The first matrix processing unit is configured to acquire a first weight matrix, perform dot product operation processing on the target input image feature of each target input image sequence and the first weight matrix to obtain a first target processing matrix corresponding to each target input image sequence.

[0044] The second matrix processing unit is configured to obtain a second weight matrix, and perform dot product operation processing on the target input image features of each target input image sequence and the second weight matrix to obtain a second target processing matrix corresponding to each target input image sequence.

[0045] The third matrix processing unit is configured to obtain a third weight matrix, and perform dot product operation processing on the target fusion sequence features and the third weight matrix to obtain a third target processing matrix.

[0046] The latent sequence feature determination unit is configured to determine a target attention matrix based on the second target processing matrix and the third target processing matrix, and determine target latent sequence features associated with the N target input image sequences based on the target attention matrix and the target fusion sequence features.

[0047] The latent sequence feature determination unit comprises a target matrix determination subunit, an associated sequence feature determination subunit, and a feature splicing processing subunit.

[0048] The target matrix determination subunit is configured to perform dot product operation processing on the second target processing matrix and the third target processing matrix to obtain the target attention matrix.

[0049] The associated sequence feature determination subunit is configured to perform dot product operation processing on the target attention matrix and the first target processing matrix corresponding to each target input image sequence to obtain target associated sequence features associated with the N target input image sequences.

[0050] The feature splicing processing subunit is configured to perform feature splicing processing on the target associated sequence features and the target fusion sequence features to obtain the target latent sequence features associated with the N target input image sequences.

[0051] Optionally, the device further comprises a model training module.

[0052] The model training module comprises a training sequence acquisition unit, a sample sequence processing unit, a sample feature extraction unit, and a target model training unit.

[0053] The training sequence acquisition unit is configured to obtain first-type input image sequences and second-type input image sequences for training the initial image processing model; the first-type input image sequences refer to image sequences that completely contain N input modalities, and the second-type input image sequences refer to image sequences in which there is a missing input modality among the N input modalities.

[0054] The sample sequence processing unit is configured to input the first type of input image sequence and the second type of input image sequence as a sample input image sequence, input a first type of spliced image sequence corresponding to the first type of input image sequence and a second type of spliced image sequence corresponding to the second type of input image sequence as a sample spliced image sequence, and determine a sample coding sequence of the sample spliced image sequence based on a sample sequence splicing order of the sample spliced image sequence.

[0055] The sample feature extraction unit is configured to input the sample input image sequence, the sample spliced image sequence, and the sample coding sequence to the initial image processing model, perform feature extraction processing on the sample input image sequence, the sample spliced image sequence, and the sample coding sequence by the initial image processing model, and obtain sample latent sequence features associated with the sample input image sequence.

[0056] The target model training unit is configured to determine predicted label feature information associated with the sample input image sequence based on the sample latent sequence features, perform model training on the initial image processing model based on the predicted label feature information and sample label feature information of the sample input image sequence, and obtain a target image processing model. The target image processing model is configured to predict and output target latent sequence features of a target missing image sequence with missing input modalities from N target input image sequences.

[0057] The training sequence acquisition unit includes a sequence extraction subunit, a sequence division subunit, a modal missing processing subunit, and a training sequence determination subunit.

[0058] The sequence extraction subunit is configured to acquire sample candidate spliced image sequences associated with N input modalities, extract sampling spliced image sequences associated with the N input modalities from the sample candidate spliced image sequences based on sampling proportions associated with the N input modalities. One sampling spliced image sequence includes N sampling image sequences corresponding to the N input modalities.

[0059] The sequence division subunit is configured to divide the sampling spliced image sequence into a first sampling spliced image sequence and a second sampling spliced image sequence based on an input missing proportion associated with the N input modalities.

[0060] The modal missing processing subunit is configured to perform input modal missing processing on N sampling image sequences corresponding to the N input modalities in the second sampling spliced image sequence according to modal missing information associated with the N input modalities, and obtain a sampling spliced missing image sequence corresponding to the second sampling spliced image sequence. The number of sequence of the sampling spliced missing image sequence is consistent with the number of sequence of the second sampling spliced image sequence.

[0061] The training sequence determining subunit is configured to splice the first sample image sequence as a first type of input image sequence for training the initial image processing model, and splice the sample splicing missing image sequence as a second type of input image sequence for training the initial image processing model.

[0062] The sample feature extraction unit comprises a first sample extraction subunit, a second sample extraction subunit, and a sample feature determining subunit.

[0063] The first sample extraction subunit is configured to input the sample input image sequence, the sample splicing image sequence, and the sample coding sequence to the initial image processing model, and perform first feature extraction processing on each sample input image sequence in the sample input image sequence based on the sample coding sequence by the initial image processing model to obtain sample input image features of each sample input image sequence.

[0064] The second sample extraction subunit is configured to perform second feature extraction processing on the sample splicing image sequence in multiple business dimensions to obtain sample business dimension features of the sample splicing image sequence in the multiple business dimensions, and perform feature fusion on the sample business dimension features in the multiple business dimensions to obtain sample fusion sequence features of the sample splicing image sequence.

[0065] The sample feature determining subunit is configured to determine a sample attention matrix for attention calculation based on the sample input image features of each sample input image sequence, and determine sample latent sequence features associated with the sample input image sequence based on the sample attention matrix and the sample fusion sequence features.

[0066] In an aspect, the present application provides a computer device, comprising a memory and a processor, the memory being connected to the processor, the memory being configured to store a computer program, and the processor being configured to call the computer program to enable the computer device to perform the method provided in the above aspect.

[0067] In an aspect, the present application provides a computer readable storage medium, the computer readable storage medium storing a computer program, the computer program being adapted to be loaded and executed by a processor to enable a computer device having the processor to perform the method provided in the above aspect.

[0068] According to an aspect of the present application, a computer program product or computer program is provided, which comprises computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the method provided in the above aspect.

[0069] The embodiment of the present application can obtain N target input image sequences and a target spliced image sequence corresponding to the N target input image sequences. The target spliced image sequence here refers to an image sequence obtained by sequentially splicing the N target input image sequences according to a target sequence splicing order, and N is a positive integer greater than 1. Further, a target coding order sequence corresponding to the target sequence splicing order can be obtained, and a first feature extraction process can be performed on each of the N target input image sequences based on the target coding order sequence, thereby obtaining a target input image feature of each target input image sequence. It should be understood that by performing the first feature extraction on each target input image sequence based on the target coding order sequence, the local information (or private feature) of each target input image sequence can be more effectively extracted.

[0070] In addition, a second feature extraction process can be performed on the target spliced image sequence in multiple business dimensions, so that when the target business dimension features in the multiple business dimensions of the target spliced image sequence are obtained, the target business dimension features in the multiple business dimensions are fused, and then the target fusion sequence features of the target spliced image sequence are obtained. It can be understood that by performing the second feature extraction process in multiple business dimensions, the global information of the sequence can be effectively utilized. Further, a target attention matrix used for attention calculation can be determined based on the target input image feature of each target input image sequence, and then a target latent sequence feature associated with the N target input image sequences is determined based on the target attention matrix and the target fusion sequence feature. It should be understood that by introducing the attention matrix, the features that need to be focused on can be determined, thereby saving computing resources and improving efficiency. In addition, each target input image sequence can learn the information of the other target input image sequence, thereby fusing to generate a deeper target latent sequence feature and improving the ability of feature fusion. BRIEF DESCRIPTION OF DRAWINGS

[0071] In order to more clearly illustrate the technical solutions of the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0072] Figure 1 is a structural schematic diagram of a network architecture provided by an embodiment of the present application;

[0073] Figure 2 is a method schematic diagram for training a target image processing model provided by an embodiment of the present application;

[0074] Figure 3is a schematic diagram of image processing by a target image processing model provided by an embodiment of the present application;

[0075] Figure 4 is a schematic diagram of an image data processing process provided by an embodiment of the present application;

[0076] Figure 5 is a schematic diagram of an image data processing method provided by an embodiment of the present application;

[0077] Figure 6 is a schematic diagram of a first feature extraction process on an input image sequence provided by an embodiment of the present application;

[0078] Figure 7 is another schematic diagram of a first feature extraction process on an input image sequence provided by an embodiment of the present application;

[0079] Figure 8 is a schematic diagram of multi-dimensional feature extraction and fusion on a stitched image sequence provided by an embodiment of the present application to obtain fused features;

[0080] Figure 9 is a schematic diagram of obtaining dimensional features of a stitched image sequence in a business dimension provided by an embodiment of the present application;

[0081] Figure 10 is a schematic diagram of obtaining target latent sequence features based on a cross-attention mechanism provided by an embodiment of the present application;

[0082] Figure 11 is a schematic diagram of image processing to generate a target image sequence provided by an embodiment of the present application;

[0083] Figure 12 is a schematic diagram of a structure of an image data processing apparatus provided by an embodiment of the present application;

[0084] Figure 13 is a schematic diagram of a structure of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0085] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0086] Before the image data processing method in the embodiments of the present application is described in detail, the technical terms involved in the present application are explained.

[0087] CNN (Convolutional Neural Network): A deep learning model primarily used for processing and analyzing data with grid structures, such as images and videos. It extracts features from input data by applying convolution and pooling operations at different levels, and uses these features for classification, recognition, or regression tasks. In the embodiments of the present application, CNN can be used for feature extraction when extracting features from each of the N target input image sequences to obtain target input image features of each target input image sequence.

[0088] ViT (Vision Transformer) model: The ViT model is a deep learning model based on the Transformer architecture, specifically designed for computer vision tasks. It was proposed by Dosovitskiy et al. in 2020. The architecture of the ViT model includes multiple Transformer encoder layers, each composed of multi-head self-attention mechanisms and feedforward neural networks. During training, the ViT model learns high-level feature representations in images by pre-training on large-scale image data. These pre-trained features can then be fine-tuned for downstream tasks such as image classification, object detection, and image segmentation. In the embodiments of the present application, the VIT model can be used for feature extraction from each of the N target input image sequences to obtain target input image features of each target input image sequence.

[0089] Transformer: Transformer is a deep learning model for natural language processing (NLP) tasks, first proposed by Vaswani et al. in 2017. Its English name is "Transformer", and its Chinese name is "Transformer". Compared with traditional recurrent neural networks (RNN) or convolutional neural networks (CNN), Transformer adopts a completely new architecture that captures long-distance dependencies in input sequences through attention mechanisms. In the embodiments of the present application, the VIT model used for feature extraction from each of the N target input image sequences includes multiple Transformer encoder layers, which can then be used for feature extraction.

[0090] MRI: Magnetic Resonance Imaging, is a medical imaging technique that uses magnetic fields and radio waves to generate detailed images of the internal structures and tissues of the human body. MRI uses the principle of interaction between water molecules in the human body and strong magnetic fields to perform imaging. In a magnetic field, the atomic nuclei of water molecules will align in a certain direction, and when a radio wave pulse is added, the atomic nuclei of water molecules will rearrange and emit signals of a specific frequency after the pulse is stopped. By detecting the intensity and spatial distribution of these signals, images with high contrast and spatial resolution can be generated. MRI can provide various different modalities of image sequences, such as T1 modality corresponding to T1 weighted image sequence, T2 modality corresponding to T2 weighted image sequence, T1Gd modality corresponding to T1Gd image sequence, and FLAIR modality corresponding to FLAIR image sequence, etc. The parameter settings and sequence types of each sequence can be adjusted according to the needs to provide detailed anatomical information, lesion areas, and tissue characteristics, etc. Among them, in the embodiments of the present application, the N target input image sequences can include but are not limited to T1 weighted image sequence, T2 weighted image sequence, T1Gd image sequence and FLAIR image sequence.

[0091] Cross-attention mechanism: an attention mechanism in neural network models, usually used for models processing two or more different input sequences. In cross-attention mechanism, the model can learn how to establish a connection between two sequences and use this connection to perform tasks such as machine translation, image description generation, etc. It can better capture the association between different parts when processing sequence data, so as to more effectively integrate information. Among them, in the embodiments of the present application, the target input image features of the N target input image sequences and the target fusion features of the target splicing image sequence can be effectively fused through the cross-attention mechanism, so as to avoid ignoring the features of any sequence.

[0092] Please refer to Figure 1 , Figure 1 is a structural diagram of a network architecture provided by an embodiment of the present application. As Figure 1 shown, the network architecture can include a server 2000 and a terminal device cluster, which can include one or more terminal devices, and the number of terminal devices is not limited here. As Figure 1 shown, the terminal device cluster can specifically include terminal device 10a, terminal device 10b, and terminal device 10c,..., terminal device 10n, etc. Among them, each terminal device in the terminal device cluster can include: a smart phone, a tablet computer, a notebook computer, a palm computer, a personal computer, a smart television, a smart watch, a vehicle-mounted device, a wearable device, etc. smart terminal, not limited here. It should be understood that, as Figure 1Each terminal device in the terminal device cluster shown can be installed with an application client related to image data processing, which can respectively interact with the above-mentioned Figure 1 server 2000 when running in each terminal device. Among them, the application client can run on the terminal device in the form of a browser, or run on the terminal device in the form of an independent application (APP), etc. The specific form of the client is not limited here.

[0093] Among them, the server 2000 can be a standalone physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and basic cloud computing services such as big data and artificial intelligence platforms. Among them, as shown, the terminal device 10a, the terminal device 10b, and the terminal device 10c, …, the terminal device 10n can be respectively connected to the above-mentioned server 2000 to facilitate each terminal device to interact with the server 2000 through the network connection. Among them, the network connection here does not limit the connection method, which can be directly or indirectly connected through wired communication, or directly or indirectly connected through wireless communication, or through other methods, which are not limited by the present application. Figure 1

[0094] Among them, the scheme provided by the present application can be independently completed by any one of the terminal devices (for example, terminal device 10a) in the terminal device cluster, or independently completed by the server (for example, server 2000), or completed by the terminal device and the server. For example, the server 2000 can use the image data processing method of the present application to train the initial image processing model according to the obtained training data set, obtain the target image processing model, and use the target image processing model to extract features from the input image sequence and perform multi-dimensional feature extraction on the spliced image sequence, thereby obtaining the final latent sequence feature, and using the obtained latent sequence feature for subsequent research of the server 2000 or sending it to the terminal device (for example, terminal device 10a) for the user of the terminal device to view and analyze. For this, the present application does not make specific limitations.

[0095] It can be understood that, Figure 1 only exemplarily represents the possible network architecture of the technical scheme of the present application, and does not limit the specific architecture of the technical scheme of the present application, that is, the technical scheme of the present application can also provide other forms of network architecture.

[0096] ​In the embodiments of the present application, feature extraction is performed on MRI image sequences. It should be understood that magnetic resonance imaging (MRI) is a non-invasive medical imaging technology with high resolution, good contrast, and no radiation damage. At present, different MRI image sequences can be obtained through this medical imaging technology, and different MRI sequence features can be extracted from these MRI image sequences through artificial intelligence technology, so as to provide certain diagnostic auxiliary information for subsequent medical diagnosis through the extracted MRI sequence features. However, in actual application, feature extraction is performed on MRI images, and there is a phenomenon of repeated feature extraction, which leads to waste of computing resources. In addition, when feature fusion is performed, the complexity of feature fusion is increased, and the feature fusion capability of multi-sequence MRI images is reduced.

[0097] Therefore, the embodiments of the present application provide an image data processing method, which can effectively extract MRI image sequences through an image processing model (i.e., a target image processing model described below) and effectively fuse the extracted MRI image sequence features. In this way, not only the waste of computing resources can be reduced, but also the feature fusion capability can be improved.

[0098] It should be understood that the image data processing method provided by the embodiments of the present application needs to be processed using a corresponding image processing model. In the embodiments of the present application, the image processing model is collectively referred to as a target image processing model. Next, the process of training the target image processing model will be described in detail. Figure 2 is a schematic diagram of a method for training a target image processing model provided by the embodiments of the present application. It should be understood that the steps of training the target image processing model can include the following steps S101-S104:

[0099] Step S101, obtaining a first type of input image sequence and a second type of input image sequence for training an initial image processing model.

[0100] It should be understood that in the embodiments of the present application, the first type of input image sequence and the second type of input image sequence are used to train the initial image processing model. Figure 2 In the respective steps of the embodiments corresponding to the above, the nouns with "sample" indicate that they are applied to the application training stage, i.e. Figure 2 The training stage of the target image processing model in the embodiments corresponding to the above.

[0101] The initial image processing model refers to a model used to train the target image processing model. The initial image processing model can be used to train the target image processing model through preprocessing, feature extraction, and other processing procedures on the first type of input image sequence and the second type of input image sequence.

[0102] wherein, the first type of input image sequence refers to the image sequence which contains all N input modalities, and the second type of input image sequence refers to the image sequence which has missing input modality in the N input modalities.

[0103] wherein, it should be understood that the N input modalities can include but are not limited to the above-mentioned T1 modality, T2 modality, T1 Gd modality, FLAIR modality, etc., and then the MRI image sequences corresponding to the above-mentioned input modalities can be referred to as T1-weighted image sequence, T2-weighted image sequence, T1 Gd image sequence and FLAIR image sequence. Next, the above-mentioned sequences will be briefly introduced.

[0104] (1) T1-weighted image sequence: T1-weighted imaging is one of the most commonly used sequences in MRI. In this sequence, the brain parenchyma is shown as gray and white bands, and the fluid region is shown in a preset color (e.g., black or dark gray). T1-weighted image sequence is very sensitive to detecting certain lesions such as cystic lesions, and is commonly used for diagnosing brain, bone and abdominal diseases, etc.

[0105] (2) T2-weighted image sequence: T2-weighted imaging sequence can show high signal of water inside the human body, and is mainly used for detecting edema and water content changes around the lesion. This sequence is more sensitive than T1 image sequence, and is commonly used for diagnosing tumors, white matter lesions and meningiomas, etc.

[0106] (3) T1 Gd image sequence: T1 Gd image sequence is a special type of T1-weighted image sequence, where Gd represents a contrast agent (Gadolinium). In MRI scanning, T1-weighted image sequence has high sensitivity to the signal intensity of tissues, and can usually show the anatomical structure of brain tissue. However, for some types of lesions such as tumors or vascular lesions, T1-weighted image sequence may not provide sufficient contrast. In order to improve the visualization of these lesions, medical personnel (e.g., doctors) may inject contrast agents (such as Gd-DTPA) into patients before MRI scanning. This contrast agent can enter the brain vascular system and accumulate in the defective or abnormal tissues of the vascular wall. When using Gd-DTPA, the T1-weighted image sequence becomes a T1 Gd image sequence.

[0107] (4) FLAIR image sequence: FLAIR (fluid attenuated inversion recovery) sequence is an important sequence with high sensitivity. In this sequence, liquid regions appear in a preset color (e.g., black), which can accurately display the lesions inside the brain. The FLAIR image sequence is often used to detect brain diseases, and is particularly important for the diagnosis of diseases such as white matter lesions and brain monocytosis.

[0108] Wherein, for the specific process of obtaining the first and second input image sequences, first, sample candidate splicing image sequences associated with N input modalities are obtained, and sample splicing image sequences associated with N input modalities are extracted from the sample candidate splicing image sequences based on the sampling ratio associated with N input modalities. It should be understood that a sample splicing image sequence includes N sample image sequences corresponding to N input modalities.

[0109] Wherein, it should be noted that in the embodiments of the present application, brain MRI images will be described, assuming that 3000 brain MRI images of subjects are used in the embodiments of the present application, and each MRI image of a subject includes four aligned sequences: TI weighted image sequence, T2 weighted sequence, T1Gd image sequence and FLAIR image sequence. It should be understood that the T1 weighted image sequence corresponds to the T1 modality, the T2 weighted image sequence corresponds to the T2 modality, the T1Gd image sequence corresponds to the T1Gd modality, and the FLAIR image sequence corresponds to the FLAIR modality. Therefore, it should be understood that the 3000 MRI image sequences each include four aligned sequences: TI weighted image sequence, T2 weighted sequence, T1Gd image sequence and FLAIR image sequence, i.e. 3000 TI weighted image sequences, 3000 T2 weighted sequences, 3000 T1Gd image sequences and 3000 FLAIR image sequences. It should be understood that the 3000 MRI image sequences are the above-mentioned sample candidate splicing sequences.

[0110] Wherein, it should be understood that the above-mentioned sample candidate splicing sequences will be subjected to image intensity normalization, i.e. the image intensity is normalized to [-1, 1]. The image intensity refers to the gray value of each pixel in the image, and the purpose of normalization is to adjust the scale of MRI image data, avoid saturation problems, facilitate comparison and visualization, and be conducive to subsequent training efficiency. In addition, the normalized sample candidate splicing sequence will also be subjected to cropping, i.e. the central region is cropped to 128*192*192. The main reason for cropping is that the sample candidate splicing sequence may have a large background color (e.g. black) that affects the subsequent training process, and cropping can highlight the image content.

[0111] In the training process, for each subject, a certain number of image sequences are randomly selected, i.e., some of the four aligned sequences are selected. It is assumed that N of the N input modalities here is 3, i.e., three input modality-associated sequences are selected from the four aligned sequences, such as TI-weighted image sequences, T2-weighted sequences, and T1Gd image sequences. It should be noted that a certain number of image sequences are randomly selected, and the selected sequence modality categories will not be limited in the embodiments of the present application.

[0112] Further, the above sampling ratio refers to the ratio for sampling the above 3000 TI-weighted image sequences, 3000 T2-weighted image sequences, and 3000 T1Gd image sequences. For example, the sampling ratio here can be 1:1:1, for example, 1000 sequences are uniformly extracted for each input modality. In other words, 1000 are extracted from 3000 T1-weighted image sequences, 1000 are extracted from 3000 T2-weighted image sequences, and 1000 are extracted from 3000 T1Gd image sequences. Then, the number of the sampled and spliced image sequences obtained by the above extraction is 3000, and each sampled and spliced image sequence includes three sampled image sequences corresponding to three input modalities (i.e., T1-weighted image sequences, T2-weighted image sequences, and T1Gd image sequences).

[0113] Further, after obtaining the above sampled and spliced image sequences, the sampled and spliced image sequences can be divided into first sampled and spliced image sequences and second sampled and spliced image sequences based on an input missing ratio associated with the N input modalities. The input missing ratio can be 30%, which will not be limited in the embodiments of the present application. It should be understood that the above is described taking 3000 as an example, the number of first sampled and spliced sequences is 3000*(100%-30%) = 2100, and the second sampled and spliced sequences are 3000*30% = 900.

[0114] Further, according to the modality missing information associated with the N input modalities, the N input modality corresponding N sampled image sequences in the second sampled and spliced image sequences are subjected to input modality missing processing to obtain a sampled and spliced missing image sequence corresponding to the second sampled and spliced image sequence; the sequence number of the sampled and spliced missing image sequence is consistent with the sequence number of the second sampled and spliced image sequence.

[0115] For example, the missing modal information is randomly missing one or more input modal information, for example, the T1 modal information can be missing, the T2 modal information can also be missing, and the T1 and T2 modal information can be missing. It should be understood that the above-mentioned second sampling spliced image sequence includes 900, that is, 900 second sampling spliced image sequences are subjected to missing modal processing, for example, 300 second sampling spliced image sequences are subjected to missing T1 modal processing, 300 second sampling spliced image sequences are subjected to missing T2 modal processing, and 300 second sampling spliced image sequences are subjected to missing T1 and T2 modal processing, and finally a sampling spliced missing image sequence is obtained. It should be understood that the sampling spliced missing image sequence includes 300 second sampling spliced image sequences missing the T1 modal, 300 second sampling spliced image sequences missing the T2 modal, and 300 second sampling spliced image sequences missing the T1 modal and the T2 modal.

[0116] Finally, the above-mentioned first sampling spliced image sequence can be used as a first type of input image sequence for training the initial image processing model, and the sampling spliced missing image sequence can be used as a second type of input image sequence for training the initial image processing model.

[0117] It should be understood that, taking the above-mentioned N as 3 as an example, the first type of input image sequence refers to an image sequence that completely contains three input modalities (i.e., the T1 modal, the T2 modal, and the T1Gd modal), and the second type of input image sequence refers to an image sequence in which there is missing input modal in the three input modalities.

[0118] It should be understood that, by using the above-mentioned input missing ratio associated with the N input modalities, the sampling spliced image sequence is divided into the first sampling spliced image sequence and the second sampling spliced image sequence, which can train the target image processing model to have the ability to realize modal transfer even in the presence of missing sequences, thereby improving the generalization of the target image processing model.

[0119] In step S102, the first type of input image sequence and the second type of input image sequence are used as sample input image sequences, the first type of spliced image sequence corresponding to the first type of input image sequence and the second type of spliced image sequence corresponding to the second type of input image sequence are used as sample spliced image sequences, and a sample coding sequence of the sample spliced image sequence is determined based on a sample sequence splicing order of the sample spliced image sequence.

[0120] The first type of stitched image sequence refers to the image sequence obtained by stitching together the first type of image sequences according to the stitching order of the first sample sequence. The second type of stitched image sequence refers to the image sequence obtained by stitching together the second type of image sequences according to the stitching order of the second sample sequence. Both the stitching order of the first and second sample sequences refer to the stitching order during the model training phase.

[0121] The process of determining the sample encoding order sequence based on the sample stitched image sequence in step S102 can be found in the following... Figure 5 The process of obtaining the target encoding sequence in step S202 of the corresponding embodiment is described, and will not be repeated here.

[0122] Step S103: Input the sample input image sequence, the sample spliced ​​image sequence, and the sample encoding sequence into the initial image processing model. The initial image processing model performs feature extraction processing on the sample input image sequence, the sample spliced ​​image sequence, and the sample encoding sequence to obtain the sample latent sequence features associated with the sample input image sequence.

[0123] The aforementioned latent sequence features refer to the sequence features obtained by the initial image processing model during the model training phase.

[0124] Specifically, the initial image processing model can input the sample input image sequence, the sample stitched image sequence, and the sample encoding order sequence into an initial image processing model. The initial image processing model then performs a first feature extraction process on each sample input image sequence based on the sample encoding order sequence, thereby obtaining the sample input image features for each sample input image sequence. Further, a second feature extraction process can be performed on the sample stitched image sequence across multiple business dimensions, obtaining sample business dimension features of the sample stitched image sequence in multiple business dimensions. These sample business dimension features are then fused to obtain the sample fused sequence features of the sample stitched image sequence. Finally, a sample attention matrix for attention calculation can be determined based on the sample input image features of each sample input image sequence. Based on the sample attention matrix and the sample fused sequence features, the sample latent sequence features associated with the sample input image sequence are determined.

[0125] The specific process of step S103 can be found in the following sections. Figure 5 The descriptions of steps S203 to S204 in the corresponding embodiments will not be elaborated here.

[0126] Step S104: Based on the latent sequence features of the samples, determine the predicted label feature information associated with the sample input image sequence. Based on the predicted label feature information and the sample label feature information of the sample input image sequence, train the initial image processing model to obtain the target image processing model.

[0127] It should be understood that the aforementioned predicted label feature information refers to the associated label feature information predicted based on the latent sequence features of the samples and the sample input image sequence. This information is task-related. For example, if the task is to reconstruct an image sequence with a missing modality, then the aforementioned predicted label feature information is the label feature information associated with the image sequence with a missing input modality. Furthermore, the aforementioned sample label feature information refers to the true labels or label-related feature information used to label the aforementioned sample input image sequence during model training. Corresponding to the aforementioned predicted label feature information, this information can be used as training targets to train the initial image processing model.

[0128] During model training, predicted label feature information and sample label feature information can be used, combined with sample input image sequence and sample latent sequence features, to train the initial image processing model. This involves supervised learning methods, which continuously adjust the model parameters by minimizing the difference between predicted label feature information and sample label feature information, thereby obtaining the target image processing model.

[0129] The target image processing model is used to complete the corresponding image processing task. For example, the task can be to predict the target latent sequence features of the target missing image sequence with missing input modalities from N target input image sequences.

[0130] For further details, please see Figure 3 , Figure 3 This is a schematic diagram illustrating image processing using a target image processing model provided in an embodiment of this application. It should be understood that, as... Figure 3 The image processing model 30A shown can correspond to the above. Figure 2 In the corresponding embodiment, the target image processing model is obtained by training the initial image processing model. Specifically, the image processing model 30A (i.e., the target image processing model) includes a first feature extraction component, a second feature extraction component, and a cross-scale communication component.

[0131] Next, based on Figure 3 The image data processing method provided in the embodiments of this application will be described. For example... Figure 3As shown, firstly, sequences 30a, 30b, and 30c are input into image processing model 30A. Next, on one hand, these sequences are fed to the first feature extraction component in image processing model 30A, where the first feature extraction component extracts the features corresponding to each sequence (hereinafter referred to as target input image features). On the other hand, these sequences are concatenated in a certain order (hereinafter referred to as target sequence concatenation order) to obtain a concatenated sequence (hereinafter referred to as target concatenated image sequence), which is then fed to the second feature extraction component in image processing model 30A, where the second feature extraction component extracts the features corresponding to the concatenated sequence (hereinafter referred to as target fusion sequence features).

[0132] Furthermore, these aforementioned features are then fed into the cross-scale communication component in the image processing model 30A. The cross-scale communication component employs a cross-attention mechanism to further process these features, ultimately outputting a result such as... Figure 3 Feature 30d is shown.

[0133] It should be understood that, in this embodiment, feature 30d can be used for subsequent tasks, and can subsequently be connected to a generation network that indicates various downstream tasks, and this generation network can generate the final result based on feature 30d. For a detailed description of the downstream tasks and the subsequent generation network, please refer to the following sections. Figure 5 The description of step S204 in the corresponding embodiment will not be repeated here.

[0134] For further details, please see Figure 4 , Figure 4 This is a schematic diagram of an image data processing procedure provided in an embodiment of this application. For example... Figure 4 As shown, the embodiments of this application can be executed by a terminal device, for example, Figure 1 In the corresponding embodiment, any terminal device in the terminal device cluster (e.g., terminal device 10a) can also be executed by a server, which can be the aforementioned Figure 1 The server 2000 in the corresponding embodiment.

[0135] like Figure 4 As shown, the input image sequence includes sequence 400a, sequence 400b, ..., sequence 400n. It should be understood that each input image sequence corresponds to an input modality. For example, sequence 400a can be a T1 modality, in which case sequence 400a can be called a T1-weighted image sequence; similarly, sequence 400b can be a T2 modality, in which case sequence 400b can be called a T2-weighted image sequence. It should be understood that in this embodiment, sequences 400a, 400b, ..., 400n can be collectively referred to as the target input image sequence.

[0136] Furthermore, for the aforementioned target input image sequences (i.e., sequences 400a, 400b, ..., 400n), sequence concatenation can be performed to obtain the concatenated image sequence, as shown below. Figure 4 The splicing sequence shown can be collectively referred to as the target spliced ​​image sequence in this embodiment of the application.

[0137] The aforementioned target image sequence is stitched together in a specific order (e.g., ...). Figure 4 The sequence is obtained by splicing the given sequence in the specified order. Furthermore, during splicing, it can also be done as follows: Figure 4 As shown, each input image in the target input image sequence is encoded by a binary encoding component, thereby obtaining an encoding order sequence corresponding to the target stitched image sequence.

[0138] On the one hand, the input image sequence is input into the image processing model 40A, wherein the image processing model 40A can correspond to the above. Figure 2 The target image processing model in the corresponding embodiment. Specifically, the first feature extraction component of image processing extracts features from each of the above-mentioned input image sequences based on the encoded order sequence, thereby obtaining the input image features of each input image (i.e., such as...). Figure 4 The input image feature set shown includes feature 401a, feature 401b, ..., feature 401n. It should be understood that feature 401a is the image feature corresponding to sequence 400a, feature 401b is the image feature corresponding to sequence 400b, and so on, feature 401n is the image feature corresponding to sequence 400n.

[0139] On the other hand, the second feature extraction component in the image processing model 40A extracts features from the spliced ​​sequence to obtain the features corresponding to the encoded sequence (i.e., feature 403a). In this embodiment, feature 403a can be collectively referred to as the target fusion sequence feature.

[0140] Furthermore, such as Figure 4 The cross-scale communication components shown include matrices 410a, 410b, and 410c, which are obtained based on the aforementioned input image feature set and feature 403a. For details, please refer to the subsequent sections. Figure 10 The description of the corresponding embodiments will not be repeated here. It should be understood that matrix 420a can be obtained from matrices 410b and 410c through the cross-scale communication component. In the embodiments of this application, matrix 420a can be referred to as the target attention matrix.

[0141] Finally, based on matrices 410a and 420a, the final feature 430a can be determined. In this embodiment, feature 430a can be referred to as the target latent sequence feature.

[0142] It should be understood that, on the one hand, the first feature extraction component can extract features from each image sequence in the input image sequence, and the second feature extraction component can extract features from the encoded sequence. Finally, by using a cross-scale communication component to simultaneously fuse features from multiple sequences, features from any single sequence can be avoided, thus obtaining feature 430a (i.e., the target latent sequence feature). It should be understood that after obtaining feature 430a, it can be integrated into a generative network to achieve various tasks. For details on how to subsequently integrate a generative network to achieve various tasks, please refer to the following sections. Figure 5 The description of step S204 in the corresponding embodiment will not be repeated here.

[0143] It should be noted that, in the specific embodiments of this application, when data related to MRI image sequences (e.g., input image sequences) is involved, a prompt interface or pop-up window needs to be displayed on the corresponding terminal device (e.g., terminal device 10a). This prompt interface or pop-up window is used to inform the user that data such as MRI image sequences is currently being collected. The data acquisition steps only begin after the user confirms the prompt interface or pop-up window; otherwise, the process ends. Furthermore, when the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions.

[0144] For further details, please see Figure 5 , Figure 5 This is a schematic flowchart of an image data processing method provided in an embodiment of this application. Figure 5 As shown, this method can be executed by a computer device, which can be... Figure 1 This method is executed by any one of the terminal devices in the shown terminal device cluster (e.g., terminal device 10a). The method may include at least the following steps S201-S204:

[0145] Step S201: Obtain N target input image sequences and the target stitched image sequence corresponding to the N target input image sequences.

[0146] It should be understood that, in Figure 5 In the various steps of the corresponding embodiments, the terms containing "target" indicate that they are applied in the practical application stage, i.e., as described above. Figure 2 The practical stage of the target image processing model in the corresponding embodiment.

[0147] The N target input image sequences are obtained through different MRI scanning modes, and can be referred to the above. Figure 4 The corresponding embodiments include sequences 400a, 400b, etc. It should be understood that the N target input image sequences here refer to the image sequences input during the application stage of the aforementioned target image processing model. These N target input image sequences correspond to N input modalities, with each target input image sequence corresponding to one input modality. The N input modalities may include, but are not limited to, the aforementioned T1 modality, T2 modality, T1Gd modality, and FLAIR modality, etc. For example, the N target input images may include the T1-weighted image sequence corresponding to the T1 modality, the T2-weighted image sequence corresponding to the T2 modality, the T1Gd image sequence corresponding to the T1Gd modality, and the FLAIR image sequence corresponding to the FLAIR modality, etc.

[0148] The target stitched image sequence refers to the image sequence obtained by stitching together N target input image sequences according to the target sequence stitching order; N is a positive integer greater than 1. It should be noted that the target sequence stitching order refers to the order in which the N target input image sequences are stitched together. For example, assuming the N target input image sequences include sequence A1, sequence A2, and sequence A3, then the image sequence obtained by stitching together the N target sequence sequences (i.e., the target stitched image sequence) according to the target sequence stitching order can be sequence A1, sequence A2, and sequence A3, or it can be sequence A2, sequence A1, and sequence A3. In this embodiment, the target sequence stitching order is not limited.

[0149] Furthermore, since the aforementioned N target input images were obtained through different MRI scanning modes, the sequence of N target input images may be spatially misaligned due to different scanning parameters, spatial resolution, respiratory factors, etc. Therefore, it is necessary to align the aforementioned N target input images. For example, methods for aligning the aforementioned N target input images can include obtaining initial alignment, image registration, and parameter adjustment. Obtaining initial alignment can involve determining the approximate initial position of the images by marking them, and then manually aligning some marker points or anatomical structures in the N target input image sequence. Image registration refers to using an image registration algorithm to align the aforementioned N target input image sequences to the same spatial coordinate system. Parameter adjustment refers to adjusting some parameters during registration to obtain the optimal alignment result. These parameters may include rotation, translation, scaling, etc.

[0150] It should be understood that since the N target input image sequences correspond to different input modalities and provide different contrasts for tissues and lesions, the alignment operation described above makes the N target input image sequences easier to use for tasks such as feature extraction, image segmentation, and machine learning, which is beneficial to improving the accuracy and reliability of subsequent target image processing models when performing image processing.

[0151] Step S202: Obtain the target encoding order sequence corresponding to the target sequence splicing order, and perform the first feature extraction process on each of the N target input image sequences based on the target encoding order sequence to obtain the target input image features of each target input image sequence.

[0152] It should be understood that the target encoding order sequence refers to the sequence obtained by encoding N target input image sequences according to the sequence concatenation order. This encoding step is completed by the binary encoding component. First, the sequence concatenation order is obtained, and then the binary encoding component encodes the N target input image sequences based on the sequence concatenation order, thereby obtaining the target encoding order sequence corresponding to the sequence concatenation order.

[0153] The binary encoding component encodes N target input image sequences. The specific encoding rule is as follows: missing sequences are encoded as 0, and complete sequences are encoded as 1. Taking the aforementioned N target input image sequences A1, A2, and A3 as an example, assuming the target sequence concatenation order is A1, A2, and A3, where A1, A2, and A3 are all complete sequences, it should be understood that the binary encoding component encodes sequence A1 as 1 and sequences A2 and A3 as 1, meaning the target encoding order sequence is [1, 1, 1].

[0154] In this embodiment, the binary encoding component can also be referred to as an input missing sequence guidance component. It should be understood that, by encoding the missing sequence as 0 and the complete sequence as 1 according to the encoding rules of the binary encoding component, the target image processing model can effectively distinguish whether a missing sequence exists among the N target input image sequences in subsequent processing, and then perform relevant processing when a missing sequence exists among the N target input image sequences. This relevant processing may include reconstructing the missing sequence.

[0155] Furthermore, the specific steps of performing the first feature extraction process on each of the N target input image sequences to obtain the target input image features of each target input image sequence may include: obtaining N binary codes from the target encoding sequence, and performing binary code matching processing on each of the N target input image sequences based on the target sequence splicing order, thereby obtaining the matching code corresponding to each of the N target input image sequences.

[0156] Here, N binary codes refer to the codes obtained by encoding N target input image sequences one-to-one. Binary code matching processing refers to matching the N target input image sequences with their corresponding binary codes. For example, taking the aforementioned N target input image sequences A1, A2, and A3 as an example, the target code sequence is [1, 1, 1]. Then, the N binary codes are binary codes 1, 1, 1. In this case, binary code matching processing will match sequence A1 with binary code 1, sequence A2 with binary code 1, and sequence A3 with binary code 1.

[0157] Furthermore, the specific steps of performing the first feature extraction process on each of the N target input image sequences to obtain the target input image features of each target input image sequence further include: inputting each target input image sequence and the matching code corresponding to each target input image sequence into the first feature extraction component, and then having the first feature extraction component perform the first feature extraction process on each target input image sequence based on the matching code, thereby obtaining the target input image features of each target input image sequence.

[0158] It should be understood that the process of training the target image processing model can be referred to the above. Figure 2 The descriptions in the corresponding embodiments will not be repeated here. Furthermore, the target image processing model includes a first feature extraction component, which may correspond to the aforementioned... Figure 3 The first feature extraction component in the corresponding embodiment. It should be understood that the first feature extraction component refers to a component used to extract the features of the target input image from the target input image sequence. The first feature extraction component can refer to the VIT model, which stands for Vision Transformer model. For a detailed description of the VIT model, please refer to the foregoing detailed explanation of the VIT model; it will not be repeated here.

[0159] Continuing with the example of the aforementioned N target input image sequences, including sequences A1, A2, and A3, the following explanation is provided. It should be understood that sequence A1 and its corresponding matching code 1 are input into the first feature extraction component (e.g., the VIT model). On one hand, the first feature extraction component performs a first feature extraction process on sequence A1, obtaining a feature vector of sequence A1 (e.g., feature vector a11). On the other hand, the first feature extraction component also performs a first feature extraction on matching code 1, extracting a feature vector of matching code 1 (e.g., feature vector a12). The feature extraction on matching code 1 is to ensure that the dimension of matching code 1 is aligned with that of feature vector a11, i.e., the dimension of feature vector a12 is the same as that of feature vector a11. Further, feature vectors a11 and a12 are fused to obtain a fused feature (e.g., feature a13), which can be referred to as the target input image feature of sequence A1. It should be understood that the process of extracting features from sequences A2 and A3 based on matching coding to obtain the target input image processing features of sequences A2 and A3 can be found in the explanation of sequence A1, and will not be repeated here.

[0160] The target input image features of each target input image sequence can refer to the private features of each target input image sequence. Therefore, in this embodiment, the first feature extraction component can also be called a private feature extraction component. Private features refer to the abstract representation (e.g., feature a11) extracted from each target input image sequence by the first feature extraction component (e.g., the VIT model), and then fused with the matched encoded feature vector (e.g., feature vector a12). Private features capture the unique organizational information and structural features in each target input image sequence, such as the shape, edges, texture, and possible lesions and tissue differentiation information in the target input image sequence (e.g., a11). For example, suppose the above N (e.g., 3) input modes include T1 mode, T2 mode and T1Gd mode, the input mode of the aforementioned sequence A1 is T1 mode, that is, sequence A1 is a TI-weighted image sequence; the input mode of sequence A2 is T2 mode, that is, sequence A2 is a T2-weighted image sequence; the input mode of sequence A3 is T3 mode, that is, sequence A3 is a TIGd image sequence. Therefore, the target input image features of sequence A1 (i.e., T1-weighted image sequence) can refer to the features extracted from the T1-weighted image sequence under the VIT model, such as structural features (e.g., gray matter, white matter, etc.), morphological features (e.g., cortical thickness), texture, etc.; the target input image sequence features of sequence A2 (i.e., T2-weighted image sequence) can refer to the features extracted from the T2-weighted image sequence under the VIT model, which usually involves fluid signals, water distribution, and other features related to tissue type; the target input image features of sequence A3 (i.e., T1Gd image sequence) can refer to the features extracted from the T1Gd image sequence under the VIT model, such as the features of contrast-enhanced areas and lesion areas, etc.

[0161] For further information, please see [link / reference]. Figure 6 , Figure 6 This is a schematic diagram illustrating a process for performing a first feature extraction on an input image sequence, as provided in an embodiment of this application. Wherein, as... Figure 6 As shown, this includes sequences 60a, 60b, and 60c. It should be understood that sequences 60a, 60b, and 60c can be referred to as the N target input image sequences in step S202 above. Among them, as... Figure 6 As shown, sequences 60a, 60b, and 60c are all complete sequences.

[0162] like Figure 6As shown, according to the encoding rules of the aforementioned binary encoding component, the binary encoding component encodes sequence 60a as 1, sequence 60b as 1, and sequence 60c as 1, thus obtaining the encoded sequence [1,1,1]. Further, it should be understood that encoding 1 (i.e., the matching encoding) is matched with sequence 60a, matched with sequence 60b, and matched with sequence 60c. Further, these are each input into the VIT model (i.e., the aforementioned first feature extraction component).

[0163] Specifically, code 1 and sequence 60a are input into the VIT model, which extracts features from sequence 60a and code 1 respectively. Then, the extracted features are fused to obtain the following result: Figure 6 Feature 601a is shown. Similarly, inputting code 1 and sequence 60b into the VIT model yields feature 601b; inputting code 1 and sequence 60c into the VIT model yields feature 601c.

[0164] It should be understood that features 601a, 601b, and 601c can correspond to the target input image features of the above N (e.g., 3) target input image sequences.

[0165] In one feasible embodiment, it is assumed that the above-mentioned N target input image sequences include sequence B1, sequence B2, and sequence B3, and the target sequence concatenation order is sequence B1, sequence B2, and sequence B3. Here, sequence B1 is a missing sequence, i.e., sequence B1 is a completely black image containing no modal information. Furthermore, sequences B2 and B3 are complete sequences. It should be understood that, at this time, the binary encoding component encodes sequence B1 as 0 and sequences B2 and B3 as 1, i.e., the target encoding order sequence is [0, 1, 1].

[0166] Furthermore, the specific steps for performing the first feature extraction process on sequences B1, B2, and B3 respectively to obtain their corresponding target input image features can be found in the above-described steps for performing the first feature extraction process on sequences A1, A2, and A3 respectively to obtain their corresponding target input image features, which will not be repeated here.

[0167] For further information, please see [link / reference]. Figure 7 , Figure 7 This is a schematic diagram illustrating another process for performing a first feature extraction on an input image sequence, as provided in an embodiment of this application. Wherein, as... Figure 7 As shown, this includes sequences 70a, 70b, and 70c. It should be understood that sequences 70a, 70b, and 70c can be referred to as the N target input image sequences in step S202 above. Among them, as... Figure 7As shown, sequence 70a is a completely black sequence, meaning it is a missing sequence, while sequences 70b and 70c are both complete sequences.

[0168] like Figure 7 As shown, according to the encoding rules of the aforementioned binary encoding component, the binary encoding component encodes sequence 70a (i.e., the missing sequence) as 0, sequence 70b as 1, and sequence 70c as 1, thus obtaining the encoded sequence [0,1,1]. Further, it should be understood that encoding 0 (i.e., the matching encoding) is matched with sequence 70a, encoding 1 (i.e., the matching encoding) is matched with sequence 70b, and encoding 1 (i.e., the matching encoding) is matched with sequence 70c. Further, they are each input into the VIT model (i.e., the aforementioned first feature extraction component).

[0169] Specifically, the code 0 and sequence 70a are input into the VIT model, which extracts features from sequence 70a and code 0 respectively. Then, the extracted features are fused to obtain the following result: Figure 7 Feature 701a is shown. Similarly, inputting code 1 and sequence 70b into the VIT model yields feature 701b; inputting code 1 and sequence 70c into the VIT model yields feature 701c.

[0170] It should be understood that features 701a, 701b, and 701c can correspond to the target input image features of the above N (e.g., 3) target input image sequences.

[0171] It should be understood that, such as Figure 6 and Figure 7 The description of the corresponding embodiments specifically explains the encoding rules of the binary encoding component and its function: when a missing sequence exists in the input image sequence, it provides guidance on identifying the missing sequence, enabling the target image processing model to distinguish the missing sequence and perform relevant processing during subsequent feature extraction. For example, when the target image processing model identifies missing sequences in N target input image sequences, it can disregard the information of the missing sequence in subsequent processing, focusing instead on fusing information from other complete sequences. Alternatively, when the target image processing model identifies missing sequences in N target input image sequences, it can reconstruct the image based on the extracted features, thereby restoring the sequence containing the missing modality.

[0172] In an optional embodiment, the aforementioned first feature extraction component can be either a VIT model or a traditional CNN (Convolutional Neural Network), i.e., using CNNs for feature extraction, such as ResNet. However, in this embodiment, the first feature extraction component is not limited.

[0173] Step S203: Perform second feature extraction processing on the target stitched image sequence in multiple business dimensions to obtain target business dimension features of the target stitched image sequence in multiple business dimensions, and fuse the target business dimension features in multiple business dimensions to obtain the target fused sequence features of the target stitched image sequence.

[0174] In image processing and analysis, business dimensions refer to the feature dimensions of an image in different aspects or perspectives. Multiple business dimensions can include channel dimensions, height dimensions, and width dimensions. Each business dimension has a specific meaning and role in processing and analyzing images.

[0175] In this context, channel dimension refers to the amount of information present at each pixel location in an image. For color images, there are typically three channels: RGB. For medical images, different sequences or modalities may constitute different channels. Channel dimension reflects different types or sources of information in the image, such as color, spectral information at different wavelengths, or different medical image sequences. In deep learning, each channel can be considered an independent dimension in the input data, and the model (e.g., the target image processing model in this embodiment) can extract specific features from each channel to aid in classification, detection, or segmentation tasks.

[0176] In this context, the height dimension refers to the number of pixels in the vertical direction of an image, i.e., the number of rows in the image. In convolutional neural networks (CNNs) or visual attention models (such as the VIT model), the height dimension is used to capture spatial information in the vertical direction of an image. Through convolutional operations or attention mechanisms, the model can learn the feature relationships between different rows in an image, such as texture, edges, or structure.

[0177] In this context, the width dimension refers to the number of pixels in the horizontal direction of the image, i.e., the number of columns. Similar to the height dimension, the width dimension is also used to capture spatial information in the horizontal direction of the image. Through convolutional operations or self-attention mechanisms, the model can learn the feature relationships between different columns in the image, thereby understanding the structure, shape, or other local information in the image.

[0178] Specifically, step S203 is processed by the aforementioned target image processing model, which includes a second feature extraction component. It should be understood that the second feature extraction component can be a component within the target image processing model, specifically responsible for extracting features from the image sequence across multiple business dimensions (i.e., channel dimension, height dimension, and width dimension). Furthermore, since step S203 involves feature fusion, in this embodiment, the second feature extraction component can also be referred to as a fused feature extraction component.

[0179] It should be noted that the specific processing procedure for obtaining the target business dimension features of the target stitched image sequence in multiple business dimensions can be as follows: input the target stitched image sequence into the target image processing model, and the second feature extraction component in the target image processing model performs dimensional feature extraction processing on the target stitched image sequence in multiple business dimensions to obtain the business dimension features of the target stitched image sequence in multiple business dimensions.

[0180] Among them, the business dimension features of the target stitched image sequence in multiple business dimensions refer to the dimensional features of the target stitched image sequence in the channel dimension, the dimensional features in the height dimension, and the dimensional features in the width dimension.

[0181] Taking the aforementioned target stitched image, which includes a first input image sequence, a second input image sequence, and a third input image sequence, as an example, the first input image sequence is referred to as sequence H1, the second input image sequence as sequence H2, and the third input image sequence as sequence H3. The business dimension features of the target stitched image sequence across multiple business dimensions include dimensional feature h11 in the channel dimension, dimensional feature h12 in the height dimension, and dimensional feature h13 in the width dimension.

[0182] Furthermore, after obtaining the target business dimension features of the target stitched image sequence across multiple business dimensions, the target business dimension can be determined from these multiple business dimensions. Then, from the business dimension features of the target stitched image sequence across multiple business dimensions, the target business dimension features of the target stitched image sequence across the target business dimension can be obtained. Finally, when each of the multiple business dimensions is determined to be the target business dimension, the target business dimension features of the target stitched image sequence across multiple business dimensions are obtained.

[0183] For example, if the channel dimension is determined as the target business dimension among multiple business dimensions, then the business dimensions other than the channel dimension among the multiple business dimensions are determined as business dimensions to be processed; where the business dimensions to be processed refer to the dimensions that need to be enriched for features, the first business dimension among the business dimensions to be processed is the height dimension, and the second business dimension among the business dimensions to be processed is the width dimension.

[0184] Furthermore, the first business dimension feature of the target stitched image sequence in the first business dimension can be obtained from the business dimension features of the target stitched image sequence in multiple business dimensions, and the second business dimension feature of the target stitched image sequence in the second business dimension can also be obtained. It should be understood that, taking the above-mentioned target stitched image sequence including a first input image sequence, a second input image sequence, and a third input image sequence as an example, assuming that the target business dimension is the channel dimension, then the first business dimension feature of the target stitched image sequence in the first business dimension is dimension feature h12, and the second business dimension feature in the second business dimension is dimension feature h13.

[0185] Furthermore, the business feature map, composed of the first business dimension features on the first business dimension and the second business dimension features on the second business dimension, is divided into blocks to obtain multiple block feature maps corresponding to the business feature map. The multiple block feature maps are then pooled to obtain pooled features of the multiple block feature maps. The pooled features of the multiple block feature maps are used as the feature weights of the target stitched image sequence on the target business dimension.

[0186] Furthermore, from the business dimension features of the target stitched image sequence across multiple business dimensions, the third business dimension feature of the target stitched image sequence in the target business dimension is obtained. Finally, based on the feature weights in the target business dimension, the third business dimension feature in the target business dimension is weighted to obtain the target business dimension feature of the target stitched image sequence in the target business dimension.

[0187] The business feature map refers to the feature map obtained by concatenating the first and second business dimension features of the target stitched image sequence. It should be understood that, continuing with the example of the target stitched image sequence including the first input image sequence (sequence H1), the second input image sequence (sequence H2), and the third input image sequence (sequence H3), the business feature map consists of dimension features h12 and h13. The block feature map refers to the feature map obtained by dividing the business feature map into blocks. For example, taking the business feature map composed of dimension features h12 and h13 as an example, pooling dimension features h12 and h13 yields the feature weights (e.g., weight C1) of the target stitched image sequence in the target business dimension (i.e., channel dimension). This reduces the dimensionality of the block feature map and extracts important feature information. Furthermore, convolving dimension feature h11 with the feature weights (i.e., C1) yields the target business dimension features (e.g., feature D1) of the target stitched image sequence in the target business dimension (i.e., channel dimension).

[0188] Furthermore, it should be understood that the height dimension and the width dimension will take turns as the target business dimension. Therefore, the process of obtaining the height dimension features (e.g., feature D2) of the target stitched image sequence in the height dimension and the width dimension features (e.g., feature D3) of the target stitched image sequence in the width dimension can be found in the explanation of obtaining the target business dimension features of the target stitched image sequence in the channel dimension, which will not be repeated here.

[0189] It should be understood that the pooling method used in the above pooling processing can be average pooling, random pooling, or max pooling. The specific pooling method is chosen based on the task to be performed by the target image processing model or the characteristics of the input data (e.g., a sequence of N target input images). For example, max pooling can be chosen if more detailed information needs to be preserved; average pooling can be chosen if feature maps need to be smoothed and noise reduced; and random pooling can be chosen if the randomness and robustness of the target image processing model need to be increased. However, in this embodiment, the pooling method is not limited.

[0190] Finally, it should be understood that after obtaining the features (e.g., features D1, D2, and D3) of the target stitched image sequence across multiple business dimensions (i.e., channel dimension, height dimension, and width dimension), linear mapping is required to obtain the feature weights of the target stitched image sequence across multiple business dimensions. Linear mapping refers to transforming high-dimensional features into lower-dimensional features. It should be understood that high-dimensional features may contain a large number of redundant and repetitive features. Linear mapping reduces redundant features and retains the key features, thereby reducing the complexity of data processing and the waste of computational resources in subsequent task stages. For example, linear mapping is performed on the aforementioned features (e.g., features D1, D2, and D3) to obtain the mapped feature E1 for feature D1, the mapped feature E2 for feature D2, and the mapped feature E3 for feature D3. Among them, mapping feature E1 can represent the feature weight of the target stitched image in the channel dimension, mapping feature E2 can represent the feature weight of the target stitched image in the height dimension, and mapping feature E3 can represent the feature weight of the target stitched image in the width dimension. It should be understood that, furthermore, these feature weights need to be superimposed on the original image sequence (i.e., the target image stitched sequence). Through this weight superposition, the important structures and important parts of the original image sequence can be effectively highlighted. Finally, the superimposed features are fused to obtain the final feature map, which can be called the target fusion sequence feature of the target stitched image sequence in the aforementioned step S203.

[0191] In an alternative implementation, step S203 can also achieve the same effect through convolution followed by deconvolution. In other words, a convolutional neural network is used to extract features from the target stitched image sequence across multiple business dimensions using different convolutional kernels, with each kernel handling features from different business dimensions. Furthermore, a deconvolution kernel is used to restore the extracted features to the original image size (i.e., the size of the target stitched image sequence). It should be understood that convolution and deconvolution effectively meet the requirement of extracting features across multiple business dimensions.

[0192] For further information, please see [link / reference]. Figure 8 , Figure 8 This is a schematic diagram illustrating a method for extracting and fusing multi-dimensional features from a spliced ​​image sequence to obtain fused features, as provided in an embodiment of this application. Figure 8 The splicing sequence 80A shown can correspond to the aforementioned target spliced ​​image sequence, wherein the splicing sequence 80A includes sequence 80a, sequence 80b and sequence 80c.

[0193] It should be understood that this processing is handled by the second feature extraction component in the target image processing model. Furthermore, when performing feature extraction across multiple business dimensions, the processing is performed on the entire target stitched image stitching sequence 80A. In other words, sequences 80a, 80b, and 80c are treated as a whole and processed globally to enable information exchange between multiple sequences.

[0194] First, the height dimension is taken as the aforementioned target business dimension, such as Figure 8 As shown, feature extraction and pooling are performed on the spliced ​​sequence 80A to obtain feature 811a in the height dimension of the spliced ​​sequence 80A. This feature 811a is the height dimension feature of the spliced ​​sequence 80A (i.e., the target spliced ​​image sequence) in the height dimension.

[0195] Secondly, the width dimension is used as the aforementioned target business dimension, such as... Figure 8 As shown, feature extraction and pooling are performed on the spliced ​​sequence 80A to obtain feature 821a in the width dimension of the spliced ​​sequence 80A. Among them, feature 821a is the width dimension feature of the spliced ​​sequence 80A (i.e., the target spliced ​​image sequence).

[0196] Furthermore, the channel dimension is taken as the aforementioned target business dimension, such as... Figure 8 As shown, feature extraction and pooling are performed on the spliced ​​sequence 80A to obtain feature 831a in the channel dimension of the spliced ​​sequence 80A. It should be understood that feature 831a is the channel dimension feature of the spliced ​​sequence 80A (i.e., the target spliced ​​image sequence) in the channel dimension.

[0197] Furthermore, such as Figure 8 As shown, feature 811a (i.e., the height dimension feature) is linearly mapped to obtain feature 812a, where feature 812a is the feature weight generated by the concatenated sequence 80A in the height dimension. Similarly, feature 821a (i.e., the width dimension feature) is linearly mapped to obtain feature 822a, where feature 822a is the feature weight generated by the concatenated sequence 80A in the width dimension. Likewise, feature 831a (i.e., the channel dimension feature) is linearly mapped to obtain feature 822a, where feature 822a is the feature weight generated by the concatenated sequence 80A in the channel dimension.

[0198] Finally, the spliced ​​sequence 80A is weighted with feature 812a (i.e., the feature weight of spliced ​​sequence 80A in the height dimension), the spliced ​​sequence 80A is weighted with feature 822a (i.e., the feature weight of spliced ​​sequence 80A in the width dimension), and the spliced ​​sequence 80A is superimposed with feature 832a (i.e., the feature weight of spliced ​​sequence 80A in the channel dimension). The superimposed features are then fused to obtain the following result: Figure 8 Feature map 840a is shown. It should be understood that feature map 840a may correspond to the target fusion sequence features in step S203 above.

[0199] For further information, please see [link / reference]. Figure 9 , Figure 9 This is a schematic diagram illustrating how to obtain the dimensional features of a stitched image sequence in the business dimension, as provided in an embodiment of this application. The stitched sequence 90A can correspond to the above-mentioned... Figure 8 The splicing sequence 80A in the corresponding embodiment.

[0200] Specifically, the second feature extraction component extracts features from the spliced ​​sequence 90A across multiple business dimensions. First, it extracts features from the spliced ​​sequence 90A along the channel dimension, obtaining feature 901a. Second, it extracts features from the spliced ​​sequence 90A along the width dimension, obtaining feature 901b. Finally, it extracts features from the spliced ​​sequence 90A along the height dimension, obtaining feature 901c.

[0201] It should be noted that feature 901c described above is different from the above. Figure 8 The corresponding features 811a and 901a require further processing to obtain feature 811a. This processing may refer to the pooling process described above. Similarly, feature 901a differs from the aforementioned... Figure 8The corresponding feature 831a, and the aforementioned feature 901b are different from the above. Figure 8 The corresponding feature is 821a.

[0202] Specifically, features 901a and 901b are pooled to obtain feature 901d. Finally, feature 901d is fused with feature 901c to obtain feature 920a. It should be understood that feature 920a is the feature extracted from sequence 90a in the height dimension, and feature 920a corresponds to the aforementioned... Figure 8 Feature 811a in the corresponding embodiment.

[0203] Understandable, Figure 8 In the corresponding embodiment, the process of splicing sequences 80A to obtain their features in each business dimension can be found in [reference needed]. Figure 9 The description of the corresponding embodiments will not be repeated here.

[0204] It should be understood that by performing feature extraction, pooling, and feature fusion on the target stitched image sequence across multiple business dimensions, information interaction between and within each target input image sequence can be achieved. In other words, cross-sequence information fusion of each target input image sequence is realized. It can be understood that the target fused sequence features can also be called global information or a global information graph, which refers to the fusion information between and within each target input image sequence in the target stitched image sequence.

[0205] Step S204: Determine the target attention matrix for attention calculation based on the target input image features of each target input image sequence; and determine the target potential sequence features associated with the N target input image sequences based on the target attention matrix and the target fusion sequence features.

[0206] It should be understood that the target attention matrix is ​​a matrix calculated based on the target input image features obtained in step S202, targeting the target potential sequence features.

[0207] In this embodiment, step S204 can be implemented by a cross-scale communication component in the target image processing model. This cross-scale communication component can obtain the aforementioned target latent sequence features based on a cross-attention mechanism.

[0208] In this embodiment, the process of the cross-scale communication component acquiring the target attention matrix specifically includes: First, acquiring a first weight matrix, and performing a dot product operation between the target input image features of each target input image sequence and the first weight matrix to obtain a first target processing matrix corresponding to each target input image sequence. Further, acquiring a second weight matrix, and then performing a dot product operation between the target input image features of each target input image sequence and the second weight matrix to obtain a second target processing matrix corresponding to each target input image sequence. Additionally, acquiring a third weight matrix, and then performing a dot product operation between the target fusion sequence features and the third weight matrix to obtain a third target processing matrix. Finally, the target attention matrix can be determined based on the second and third target processing matrices.

[0209] The first, second, and third weight matrices refer to the parameter matrices commonly used in attention mechanisms. The first weight matrix can be called the value weight matrix (Wv), the second weight matrix can be called the key weight matrix (Wk), and the third weight matrix can be called the query weight matrix (Wq).

[0210] It should be noted that the value vector matrix of each target input image sequence can be obtained by calculating the target input image features and the value weight matrix (Wv) for each target input image sequence. For example, if there are N target input image sequences, and the target input image features of the N target input image sequences include the first input image feature_1, the second input image feature_2, ..., the Nth input image feature_n, then multiplying the value weight matrix (Wv) with the target input image features (i.e., feature_1, feature_2, ..., feature_3) of each target input image sequence yields n value vector (V) matrices (i.e., V1, V2, ..., Vn), for example, V1 = Wv * feature_1. It should be understood that these N value vector (V) matrices can correspond to the first target processing matrix for each target input image sequence mentioned above.

[0211] Similarly, by calculating the key vector matrix of each target input image sequence using the target input image features and the key weight matrix (Wk), the key vector matrix of each target input image sequence can be obtained. For example, if there are N target input image sequences, and the target input image features of the N target input image sequences include the first input image feature_1, the second input image feature_2, ..., the Nth input image feature_n, then multiplying the key weight matrix (Wk) with the target input image features (i.e., feature_1, feature_2, ..., feature_3) of each target input image sequence yields N key vector (K) matrices (i.e., K1, K2, ..., Kn), for example, K1 = Wk * feature_1. It should be understood that these N key vector (K) matrices correspond to the second target processing matrix for each target input image sequence mentioned above.

[0212] Furthermore, the query vector can be obtained by calculating the target fused sequence features and the query weight matrix (Wq) obtained in step S203 above. For example, if the target fused sequence feature is Integrated_feature, multiplying the query weight matrix (Wq) and Integrated_feature yields the query vector matrix (Q), such as Q = Wq * Integrated_feature. It should be understood that the query vector matrix (Q) here corresponds to the third target processing matrix mentioned above.

[0213] It should be noted that the value weight matrix (Wv), key weight matrix (Wk), and query weight matrix (Wq) mentioned above are learned through the training process (e.g., the training process of the aforementioned target image processing model), and the weights are adjusted through continuous training. Optionally, the value weight matrix (Wv), key weight matrix (Wk), and query weight matrix (Wq) can also be determined manually in practice. Their main function is to dynamically calculate the importance of each input in a specific task based on the input feature vector, thereby improving the model's (e.g., the target image processing model's) ability to model complex data relationships.

[0214] In the cross-attention mechanism, the dot product operation between the second target processing matrix (i.e., key vector matrices K1, K2, ..., Kn) and the third target processing matrix (i.e., query vector matrix Q) generates N attention score matrices (W1, W2, ..., Wn), which are the aforementioned target attention matrices. These target attention matrices characterize the correlation or similarity between the query vector matrix Q and each key vector matrix, indicating the importance of each position in the input sequence (i.e., each target input image sequence) given a query.

[0215] Furthermore, the cross-scale communication component needs to perform softmax normalization on the aforementioned N attention score matrices (W1, W2, ..., Wn) to convert the raw scores into attention weights. These weights determine which parts of the value vector (V1, V2, ..., Vn) should be focused and weighted to produce the final output. Specifically, further processing is performed: a dot product operation is performed on the target attention matrix and the first target processing matrix corresponding to each target input image sequence to obtain the target association sequence features associated with the N target input image sequences.

[0216] Taking the target attention matrix (W1, W2, ..., Wn) and the first target processing matrix (V1, V2, ..., Vn) corresponding to each target input image sequence as an example, each target attention matrix is ​​multiplied by its corresponding first target processing matrix using a dot product operation. For instance, the dot product of W1 and V1 yields U1, W2 and V2 yields U2, and Wn and Vn yields Un. Adding U1, U2, ..., Un together yields the final context vector (i.e., target-related sequence features). This context vector focuses on the most relevant parts of the input sequence (i.e., each target input image sequence) to the query, enabling more effective information processing and prediction in subsequent tasks.

[0217] Furthermore, the cross-scale communication component concatenates the target-associated sequence features and the target-fused sequence features to obtain the target latent sequence features associated with the N target input image sequences. In other words, the aforementioned context vector (i.e., target-associated sequence features) and the target-fused sequence features are concatenated or added, and the network layer outputs the target latent sequence features.

[0218] Among them, the target latent sequence features refer to the hidden features learned by the target image processing model from the aforementioned N target input image sequences. These target latent sequence features can be used by the generators in subsequent connections to generate the correct target sequences based on different sequence prediction tasks.

[0219] It should be understood that the aforementioned cross-attention mechanism generates a query vector matrix Q using the target fusion sequence features and generates key vector (K) matrices (i.e., K1, K2, ..., Kn) and value vector (V) matrices (i.e., V1, V2, ..., Vn) using the target input image features of each input image sequence. This mechanism can effectively fuse the target fusion sequence features and the target input image features of each input image sequence. By calculating the attention between the query vector matrix Q and each key vector (K) matrix, the target image processing model can capture the multi-scale information interaction between different sequences (i.e., N target input image sequences) and avoid ignoring the features of any sequence.

[0220] For further information, please see [link / reference]. Figure 10 , Figure 10 This is a schematic diagram illustrating a method for obtaining target latent sequence features based on a cross-attention mechanism, as provided in an embodiment of this application. Wherein, Figure 10 This can be processed by the aforementioned cross-scale communication components, such as... Figure 10 As shown, the input image 101A includes features 1001a, 1001b, and 1001c. It should be understood that features 1001a, 1001b, and 1001c can correspond to the target input image features of the aforementioned N target input image sequences, where N is 3.

[0221] It should be noted that matrix Wv can correspond to the aforementioned first weight matrix (i.e., value weight matrix (Wv)), and matrix Wk can correspond to the aforementioned second weight matrix (i.e., key weight matrix (Wk)). Further, each feature in the input image feature 101A (i.e., feature 1001a, feature 1001b, and feature 1001c) is processed by a dot product operation with matrix Wv to obtain the following... Figure 10 Matrix V1, matrix V2, and matrix V3 shown are collectively referred to as matrix set V (i.e., the first target processing matrix mentioned above). Similarly, each feature in the input image feature 101A (i.e., feature 1001a, feature 1001b, and feature 1001c) is processed by a dot product operation with matrix Wk to obtain the following... Figure 10 The matrices K1, K2, and K3 shown are collectively referred to as matrix set K (i.e., the second target processing matrix mentioned above).

[0222] It should be noted that, such as Figure 10 Feature 1001a can correspond to the aforementioned target fusion sequence feature, and matrix Wq can correspond to the aforementioned third weight matrix (i.e., the query weight matrix (Wq)). Then, performing a dot product operation between feature 1001a (i.e., the target fusion sequence feature) and matrix Wq yields the following result: Figure 10 The matrix Q shown is the third target processing matrix mentioned above.

[0223] Furthermore, each matrix (i.e., matrix K1, matrix K2, and matrix K3) in the matrix set K (i.e., the second target processing matrix) is subjected to a dot product operation with matrix Q (i.e., the third target processing matrix mentioned above), resulting in the following... Figure 10 The matrices W1, W2, and W3 shown are collectively referred to as the matrix set W. It should be understood that this matrix set W can correspond to the aforementioned target attention matrix.

[0224] It should be understood that the matrix set W (i.e., the target attention matrix) and the matrix set V (i.e., the first target processing matrix) are processed by performing a dot product operation to obtain the context vector, which is then fused with feature 1001d to obtain the following... Figure 10 The feature Z shown is referred to as the target latent sequence feature in this embodiment of the application.

[0225] It should be understood that traditional feature fusion methods (such as simple concatenation or summation) cannot fully utilize the correlation information from different data sources. However, in this embodiment, the cross-attention mechanism adopted in the cross-scale communication module can effectively calculate the attention weight of each data source to other data sources, i.e., the attention weight between the target input image features of the N target input image sequences and the target fused sequence features of the target concatenated image sequence. Furthermore, based on the needs of the actual task, it can focus on the more important features among these features, strengthen relevant information and suppress irrelevant information, thereby ensuring effective feature fusion.

[0226] Furthermore, efficient utilization of computing resources is crucial when processing large-scale data. In this embodiment, the cross-attention mechanism avoids unnecessary computation and redundant feature representations during feature fusion. By calculating attention weights (i.e., the target attention matrix), it is possible to accurately determine which features are more important, thereby reducing the waste of computing resources.

[0227] In one alternative implementation, if the target encoding order sequence indicates that there is a target missing image sequence among the N target input image sequences, then the target missing image sequence can be generated based on the target latent sequence features.

[0228] Here, a target missing image sequence refers to an image sequence in which one of the N input modalities corresponding to N target input image sequences is missing. One target input image sequence corresponds to one input modality.

[0229] In step S202, the binary encoding component encodes N target input image sequences to obtain a target encoding order sequence. During the encoding process, the presence of a missing sequence (i.e., a target missing image sequence) can be determined by the existence of a sequence encoded as 0. It should be understood that if a sequence encoded as 0 exists among the N target input image sequences during the encoding process, the target image processing model can determine that the current task is to generate the target missing image sequence. Therefore, the generative network connected in the target image processing model at this time generates the target missing image sequence based on the target latent sequence features. It should be understood that, according to the aforementioned... Figure 2 In the corresponding embodiment, the target image processing model has been continuously transferred and trained through the initial image processing model, enabling it to reconstruct the missing image sequence of the target based on the target latent sequence features.

[0230] Therefore, in the case of modality loss in a multimodal input image sequence (i.e., N target input image sequences), this application embodiment can also ensure the effective fusion of sequence features of the multimodal input image sequence, and then accurately restore the missing sequence corresponding to the missing modality through the potential sequence features obtained by effective fusion.

[0231] Optionally, if during the encoding process, there is no sequence encoded as 0 in the target encoding sequence, the target image processing model can determine that the current task is to generate a fused image sequence. Therefore, the generative network connected in the target image processing model at this time generates a fused image sequence based on the target latent sequence features. This fused image sequence refers to the fusion of features from the aforementioned N target input image sequences and the image features from the target stitched image sequence. This fused image sequence can represent the lesion information after MRI scanning at a deeper level and more intuitively. It should be understood that, according to the aforementioned... Figure 2 In the corresponding embodiment, the target image processing model has been continuously transferred and trained through the initial image processing model, enabling it to generate fused image sequences based on the target latent sequence features.

[0232] Optionally, if, during the encoding process, there is no sequence with a code of 0 in the target encoding sequence, the target image processing model can still determine that the current task is a classification task. Therefore, the generative network connected in the target image processing model at this time classifies the disease based on the features of the aforementioned N target input image sequences and the image features of the target stitched image sequence, according to the target latent sequence features. Through this classification task, the current disease type can be more clearly identified, allowing for targeted treatment. It should be understood that, based on the aforementioned... Figure 2In the corresponding embodiment, the target image processing model has been continuously transferred and trained through the initial image processing model, enabling it to classify diseases based on the target latent sequence features.

[0233] Optionally, if there are no sequences encoded as 0 in the target encoding sequence, the target image processing model can also determine that the current task is a segmentation task. In this case, the target latent sequence features can be features such as edges, textures, and colors from the aforementioned N target input image sequences that help distinguish lesions from normal tissue. Therefore, the generative network connected in the target image processing model can be based on image segmentation algorithms, such as pixel-based segmentation, region-based segmentation, and edge detection, to identify and mark the boundaries of lesions. The segmentation results can then be post-processed, such as removing small regions and filling holes, to obtain more accurate and smoother lesion boundaries. Finally, the segmentation results are compared and analyzed with the original image to generate a visual output, helping medical personnel (e.g., doctors) to analyze and diagnose lesions.

[0234] In this embodiment of the application, the downstream tasks that the generator network of the target image processing model needs to complete after obtaining the target latent sequence features are not limited.

[0235] For further information, please see [link / reference]. Figure 11 , Figure 11 This is a schematic diagram illustrating an image processing method for generating a target image sequence according to an embodiment of this application. Wherein, as... Figure 11 As shown, sequences 110a, 110b, and 110c can correspond to the aforementioned Figure 7 In the corresponding embodiments, sequences 70a, 70b, and 70c are used. The spliced ​​sequence 110d is a sequence obtained by splicing sequences 110a, 110b, and 110d together.

[0236] Furthermore, the first feature extraction component and the second feature extraction component can be as described above. Figure 3 The corresponding embodiment refers to the first feature extraction component and the second feature extraction component. The input image feature set includes features 1101a, 1101b, and 1101c, which can correspond to the aforementioned... Figure 7 The corresponding embodiments refer to features 701a, 701b, and 701c. The first feature extraction component extracts each feature from the input image feature set, which can be seen in [reference needed]. Figure 7 The processing steps in the corresponding embodiments will not be described in detail here. Furthermore, the process by which the second feature extraction component extracts feature 1102a can be found in [reference needed]. Figure 8The description of extracting feature map 840a from spliced ​​sequence 80A in the corresponding embodiment will not be repeated here.

[0237] Furthermore, feature 1102a and the input image feature set are input into the cross-scale communication component to obtain feature 1110a. This process can be found in [reference needed]. Figure 10 The process of obtaining feature Z in the corresponding embodiment will not be described in detail here.

[0238] Furthermore, feature 1110a is input into the generator network to generate sequence 110e. It should be understood that, as... Figure 11 The sequence 110a shown is a completely black image, i.e., a missing sequence. Therefore, the generator network can determine that the task is to restore and reconstruct this missing sequence, and sequence 110e is the original sequence 110a containing modal information.

[0239] Optional, and should be understood, if Figure 11 The sequence 110a shown is not a completely black image, that is, the sequence 110a is a complete sequence containing modal information. Therefore, the generator network can determine that the task is to generate a fused image sequence. The sequence 110e is an image sequence that fuses the information of the aforementioned sequences 110a, 110b, 110c and the spliced ​​sequence 110d.

[0240] Furthermore, in an optional implementation, the image data processing method provided in this application embodiment is applicable not only to brain MRI images but also to other body parts, such as the knee, shoulder, etc. Therefore, the image data processing method provided in this application embodiment is not limited to brain MRI images.

[0241] This application embodiment can obtain N target input image sequences and corresponding target stitched image sequences. Here, the target stitched image sequence refers to the image sequence obtained by stitching the N target input image sequences according to the target sequence stitching order, where N is a positive integer greater than 1. Further, the target encoding order sequence corresponding to the target sequence stitching order can be obtained, and a first feature extraction process can be performed on each of the N target input image sequences based on the target encoding order sequence to obtain the target input image features of each target input image sequence. It should be understood that by performing the first feature extraction on each target input image sequence based on the target encoding order sequence, the local information (or private features) of each target input image sequence can be extracted more effectively.

[0242] Furthermore, a second feature extraction process can be performed on the target stitched image sequence across multiple business dimensions. This allows for the fusion of these multiple business dimension features to obtain the target fused sequence features of the target stitched image sequence. It is understandable that performing a second feature extraction process across multiple business dimensions effectively utilizes the global information of the sequence. Further, a target attention matrix for attention calculation can be determined based on the target input image features of each target input image sequence. Then, based on the target attention matrix and the target fused sequence features, the target latent sequence features associated with the N target input image sequences are determined. It should be understood that introducing an attention matrix allows for the identification of key features, thereby saving computational resources and improving efficiency. In addition, each target input image sequence can learn from each other, thereby fusing and generating deeper target latent sequence features and enhancing feature fusion capabilities.

[0243] Furthermore, if the target encoding sequence indicates that there is a target missing image sequence among the N target input image sequences, then the target missing image sequence can be generated based on the target latent sequence features. Therefore, this embodiment can also ensure the effective fusion of sequence features of the multimodal input image sequences (i.e., N target input image sequences) when modality is missing, and then accurately reconstruct the missing sequence corresponding to the missing modality through the latent sequence features obtained by effective fusion.

[0244] For further information, please see [link / reference]. Figure 12 , Figure 12 This is a schematic diagram of the structure of an image data processing device provided in an embodiment of this application. Figure 12 As shown, the image data processing device can be a computer program (including program code) running on a computer device. For example, the image data processing device can be application software, and the computer device can be a terminal device, such as... Figure 1 The terminal device 10a shown can be used to perform corresponding steps in the image data processing method provided in the embodiments of this application. Figure 12 As shown, the image data processing device may include: a sequence acquisition module 11, a first extraction module 12, a second extraction module 13, and a feature determination module 14;

[0245] The sequence acquisition module 11 is used to acquire N target input image sequences and the corresponding target stitched image sequences; the target stitched image sequence refers to the image sequence obtained by stitching the N target input image sequences according to the target sequence stitching order; N is a positive integer greater than 1;

[0246] The first extraction module 12 is used to obtain the target encoding order sequence corresponding to the splicing order of the target sequence, and perform a first feature extraction process on each of the N target input image sequences based on the target encoding order sequence to obtain the target input image features of each target input image sequence.

[0247] The second extraction module 13 is used to perform second feature extraction processing on the target stitched image sequence in multiple business dimensions to obtain the target business dimension features of the target stitched image sequence in multiple business dimensions, and to fuse the target business dimension features in multiple business dimensions to obtain the target fused sequence features of the target stitched image sequence.

[0248] The feature determination module 14 is used to determine the target attention matrix for attention calculation based on the target input image features of each target input image sequence, and to determine the target potential sequence features associated with the N target input image sequences based on the target attention matrix and the target fusion sequence features.

[0249] The specific implementation methods of the sequence acquisition module 11, the first extraction module 12, the second extraction module 13, the feature determination module 14, and the sequence generation module 15 can be found in the above description. Figure 5 The descriptions of steps S201-S204 in the corresponding embodiments will not be repeated here.

[0250] Optionally, the apparatus further includes: a sequence generation module 15;

[0251] The sequence generation module 15 is used to generate a target missing image sequence based on the target latent sequence features if the target encoding order sequence indicates that there is a target missing image sequence among the N target input image sequences; the target missing image sequence refers to an image sequence in which an input mode is missing among the N input modes corresponding to the N target input image sequences; one target input image sequence corresponds to one input mode.

[0252] For the specific implementation of the sequence generation module 15, please refer to the above. Figure 11 The description of the corresponding embodiments will not be repeated here.

[0253] The first extraction module 12 includes: an encoding processing unit 121 and a feature extraction unit 122;

[0254] The encoding processing unit 121 is used to obtain the sequence splicing order, and the binary encoding component encodes the N target input image sequences based on the sequence splicing order to obtain the target encoding order sequence corresponding to the sequence splicing order;

[0255] The feature extraction unit 122 is used to input the target encoding sequence and each target input image sequence from the N target input image sequences into the target image processing model, and the target image processing model performs a first feature extraction process on each target input image sequence based on the target encoding sequence to obtain the target input image features of each target input image sequence.

[0256] The specific implementation methods of the encoding processing unit 121 and the feature extraction unit 122 can be found in the above description. Figure 5 The description of step S202 in the corresponding embodiments will not be repeated here.

[0257] The target image processing model includes a first feature extraction component;

[0258] Feature extraction unit 122 includes: encoding matching subunit and target feature extraction subunit;

[0259] The encoding matching extraction subunit is used to input the target encoding sequence into the first feature extraction component in the target image processing model. The first feature extraction component performs first feature extraction processing on the target encoding sequence to obtain the target encoding sequence features corresponding to the target encoding sequence.

[0260] The target feature extraction subunit is used to input the sequence feature extraction subunit, which is used to input each of the N target input image sequences into the first feature extraction component. The first feature extraction component performs a first feature extraction process on each target input image sequence and obtains the target input image features of each target input image sequence based on the target encoded sequence features.

[0261] Among them, the N target input image sequences include a first input image sequence and a second input image sequence, the matching code corresponding to the first input image sequence is the first matching code, and the matching code corresponding to the second input image sequence is the second matching code;

[0262] The target feature extraction subunit is also specifically used to input the first input image sequence and the first matching code into the first feature extraction component, and the first feature extraction component performs first feature extraction processing on the first input image sequence and the first matching code to obtain the first extracted features corresponding to the first input image sequence and the first encoding vector corresponding to the first matching code;

[0263] The target feature extraction subunit is also specifically used to fuse the first extracted features and the first encoding vector to obtain the first target extracted features corresponding to the first input image sequence;

[0264] The target feature extraction subunit is also specifically used to input the second input image sequence and the second matching code into the first feature extraction component, and the first feature extraction component performs first feature extraction processing on the second input image sequence and the second matching code to obtain the second extracted features corresponding to the second input image sequence and the second encoding vector corresponding to the second matching code;

[0265] The target feature extraction subunit is also specifically used to fuse the second extracted features and the second encoding vector to obtain the second target extracted features corresponding to the second input image sequence;

[0266] The target feature extraction subunit is also specifically used to use the first target extraction feature and the second target extraction feature as the target input image features for each target input image sequence.

[0267] The specific implementation methods of the first feature extraction subunit and the second feature extraction subunit can be found in the above description. Figure 6 or Figure 7 The descriptions in the corresponding embodiments will not be repeated here.

[0268] The second extraction module 13 includes: a multi-dimensional feature extraction unit 131, a target dimension determination unit 132, a target feature determination unit 133, and a feature fusion unit 134;

[0269] The multi-dimensional feature extraction unit 131 is used to input the target stitched image sequence into the target image processing model, and the second feature extraction component in the target image processing model performs dimensional feature extraction processing on the target stitched image sequence in multiple business dimensions to obtain the business dimension features of the target stitched image sequence in multiple business dimensions.

[0270] The target dimension determination unit 132 is used to determine the target business dimension from multiple business dimensions, and to obtain the target business dimension features of the target stitched image sequence in the target business dimension from the business dimension features of the target stitched image sequence in multiple business dimensions.

[0271] The target feature determination unit 133 is used to obtain the target business dimension features of the target stitched image sequence in multiple business dimensions when each of the multiple business dimensions is determined as the target business dimension.

[0272] The feature fusion unit 134 is used to fuse the target business dimension features of each target input image sequence in multiple business dimensions to obtain the target fusion sequence features of the target stitched image sequence.

[0273] The specific implementation methods of the multi-dimensional feature extraction unit 131, the target dimension determination unit 132, the target feature determination unit 133, and the feature fusion unit 134 can be found in the above description. Figure 5 The description of step S203 in the corresponding embodiment will not be repeated here.

[0274] These include multiple business dimensions such as channel dimension, height dimension, and width dimension;

[0275] The target dimension determination unit 132 includes: a target dimension determination subunit, a first feature determination subunit, a pooling processing subunit, a second feature determination subunit, and a weighted processing subunit;

[0276] The target dimension determination subunit is used to determine the business dimensions other than the channel dimension among multiple business dimensions as business dimensions to be processed if the channel dimension is determined as the target business dimension among multiple business dimensions; the first business dimension among the business dimensions to be processed is the height dimension, and the second business dimension among the business dimensions to be processed is the width dimension.

[0277] The first feature determination subunit is used to obtain the first business dimension feature of the target stitched image sequence in the first business dimension from the business dimension features of the target stitched image sequence in multiple business dimensions, and to obtain the second business dimension feature of the target stitched image sequence in the second business dimension.

[0278] The pooling processing subunit is used to divide the business feature map composed of the first business dimension features on the first business dimension and the second business dimension features on the second business dimension into blocks to obtain multiple block feature maps corresponding to the business feature map. The multiple block feature maps are pooled to obtain pooled features of the multiple block feature maps. The pooled features of the multiple block feature maps are used as feature weights of the target stitched image sequence on the target business dimension.

[0279] The second feature determination subunit is used to obtain the third business dimension feature of the target stitched image sequence in the target business dimension from the business dimension features of the target stitched image sequence in multiple business dimensions;

[0280] The weighted processing subunit is used to perform weighted processing on the third business dimension features in the target business dimension based on the feature weights in the target business dimension, so as to obtain the target business dimension features of the target stitched image sequence in the target business dimension.

[0281] The specific implementation methods of the target dimension determination subunit, the first feature determination subunit, the pooling processing subunit, the second feature determination subunit, and the weighted processing subunit can be found above. Figure 5The description of step S203 in the corresponding embodiment will not be repeated here.

[0282] The feature determination module 14 includes: a first matrix processing unit 141, a second matrix processing unit 142, a third matrix processing unit 143, and a latent sequence feature determination unit 144;

[0283] The first matrix processing unit 141 is used to obtain a first weight matrix and perform a dot product operation on the target input image features of each target input image sequence with the first weight matrix to obtain a first target processing matrix corresponding to each target input image sequence.

[0284] The second matrix processing unit 142 is used to obtain the second weight matrix and perform a dot product operation on the target input image features of each target input image sequence and the second weight matrix to obtain the second target processing matrix corresponding to each target input image sequence.

[0285] The third matrix processing unit 143 is used to obtain the third weight matrix, and to perform a dot product operation on the target fusion sequence features and the third weight matrix to obtain the third target processing matrix.

[0286] The latent sequence feature determination unit 144 is used to determine the target attention matrix based on the second target processing matrix and the third target processing matrix, and to determine the target latent sequence features associated with the N target input image sequences based on the target attention matrix and the target fusion sequence features.

[0287] The specific implementation methods of the first matrix processing unit 141, the second matrix processing unit 142, the third matrix processing unit 143, and the latent sequence feature determination unit 144 can be found in the above description. Figure 10 The descriptions in the corresponding embodiments will not be repeated here.

[0288] The latent sequence feature determination unit 144 includes: a target matrix determination subunit, an associated sequence feature determination subunit, and a feature splicing processing subunit;

[0289] The target matrix determination sub-unit is used to perform dot product operations on the second target processing matrix and the third target processing matrix to obtain the target attention matrix;

[0290] The associated sequence feature determination subunit is used to perform a dot product operation on the target attention matrix and the first target processing matrix corresponding to each target input image sequence to obtain the target associated sequence features associated with N target input image sequences;

[0291] The feature concatenation processing subunit is used to concatenate the target-associated sequence features and the target-fused sequence features to obtain the target latent sequence features associated with N target input image sequences.

[0292] The specific implementation methods of the target matrix determination subunit, the associated sequence feature determination subunit, and the feature concatenation processing subunit can be found above. Figure 10 The descriptions in the corresponding embodiments will not be repeated here.

[0293] Optionally, the device further includes: a model training module 16;

[0294] The model training module 16 includes: a training sequence acquisition unit 161, a sample sequence processing unit 162, a sample feature extraction unit 163, and a target model training unit 164.

[0295] The training sequence acquisition unit 161 is used to acquire a first type of input image sequence and a second type of input image sequence for training the initial image processing model; the first type of input image sequence refers to an image sequence that completely contains N input modalities, and the second type of input image sequence refers to an image sequence in which some input modalities are missing in the N input modalities;

[0296] The sample sequence processing unit 162 is used to take the first type of input image sequence and the second type of input image sequence as sample input image sequences, take the first type of spliced ​​image sequence corresponding to the first type of input image sequence and the second type of spliced ​​image sequence corresponding to the second type of input image sequence as sample spliced ​​image sequences, and determine the sample encoding order sequence of the sample spliced ​​image sequence based on the sample sequence splicing order of the sample spliced ​​image sequence.

[0297] The sample feature extraction unit 163 is used to input the sample input image sequence, the sample spliced ​​image sequence, and the sample encoding sequence into the initial image processing model, and the initial image processing model performs feature extraction processing on the sample input image sequence, the sample spliced ​​image sequence, and the sample encoding sequence to obtain the sample latent sequence features associated with the sample input image sequence.

[0298] The target model training unit 164 is used to determine the predicted label feature information associated with the sample input image sequence based on the sample latent sequence features. Based on the predicted label feature information and the sample label feature information of the sample input image sequence, the initial image processing model is trained to obtain the target image processing model. The target image processing model is used to predict the target latent sequence features of the target missing image sequence with missing input modalities from N target input image sequences.

[0299] The specific implementation methods of the training sequence acquisition unit 161, the sample sequence processing unit 162, the sample feature extraction unit 163, and the target model training unit 164 can be found in the above description. Figure 2 The descriptions of steps S101-S104 in the corresponding embodiments will not be repeated here.

[0300] The training sequence acquisition unit 161 includes: a sequence extraction subunit, a sequence partitioning subunit, a modality missing processing subunit, and a training sequence determination subunit;

[0301] The sequence extraction subunit is used to acquire a sequence of candidate sample images associated with N input modalities. Based on the sampling ratio associated with the N input modalities, it extracts a sequence of sampled images associated with the N input modalities from the sequence of candidate sample images. A sequence of sampled images includes N sampled image sequences corresponding to the N input modalities.

[0302] A sequence partitioning subunit is used to divide a sampled stitched image sequence into a first sampled stitched image sequence and a second sampled stitched image sequence based on the input missing ratio associated with N input modalities.

[0303] The modality missing processing subunit is used to process the N sampled image sequences corresponding to the N input modalities in the second sampled stitched image sequence according to the modality missing information associated with the N input modalities, so as to obtain the sampled stitched missing image sequence corresponding to the second sampled stitched image sequence; the number of sequences in the sampled stitched missing image sequence is consistent with the number of sequences in the second sampled stitched image sequence;

[0304] The training sequence determination subunit is used to take the first sampled and stitched image sequence as the first type of input image sequence for training the initial image processing model, and to take the sampled and stitched missing image sequence as the second type of input image sequence for training the initial image processing model.

[0305] The specific implementation methods of the sequence extraction subunit, sequence partitioning subunit, modality missing handling subunit, and training sequence determination subunit can be found above. Figure 2 The descriptions of steps S101-S104 in the corresponding embodiments will not be repeated here.

[0306] The sample feature extraction unit 163 includes: a first sample extraction subunit, a second sample extraction subunit, and a sample feature determination subunit;

[0307] The first sample extraction subunit is used to input the sample input image sequence, the sample stitched image sequence, and the sample encoding order sequence into the initial image processing model. The initial image processing model performs the first feature extraction processing on each sample input image sequence based on the sample encoding order sequence to obtain the sample input image features of each sample input image sequence.

[0308] The second sample extraction subunit is used to perform second feature extraction processing on the sample stitched image sequence in multiple business dimensions to obtain the sample business dimension features of the sample stitched image sequence in multiple business dimensions, and to fuse the sample business dimension features in multiple business dimensions to obtain the sample fused sequence features of the sample stitched image sequence.

[0309] The sample feature determination subunit is used to determine the sample attention matrix for attention calculation based on the sample input image features of each sample input image sequence, and to determine the sample latent sequence features associated with the sample input image sequence based on the sample attention matrix and the sample fusion sequence features.

[0310] The specific implementation methods of the first sample extraction subunit, the second sample extraction subunit, and the sample feature determination subunit can be found above. Figure 2 The descriptions of steps S101-S104 in the corresponding embodiments will not be repeated here.

[0311] Further, please see Figure 13 , Figure 13 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Figure 13 As shown, the computer device 1000 can be a user terminal, for example, the one described above. Figure 1 The terminal device 10a in the corresponding embodiment can also be a server, for example, as described above. Figure 1The server 2000 in the corresponding embodiment will not be limited here. For ease of understanding, this application takes a computer device 1000 as a user terminal as an example. The computer device 1000 may include: a processor 1001, a network interface 1004, and a memory 1005. In addition, the computer device 1000 may also include: a user interface 1003, and at least one communication bus 1002. The communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk storage device. The memory 1005 may also optionally be at least one storage device located remotely from the aforementioned processor 1001. Figure 13 As shown, the memory 1005, which is a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application.

[0312] The network interface 1004 in the computer device 1000 can also provide network communication functions, and the optional user interface 1003 can further include a display screen and a keyboard. It should be noted that when the computer device 1000 is a server, the user interface 1003 does not include the aforementioned display screen and keyboard. Figure 13 In the computer device 1000 shown, the network interface 1004 provides network communication functionality; the user interface 1003 is mainly used to provide an input interface for the user; and the processor 1001 can be used to call the device control application stored in the memory 1005 to execute the aforementioned... Figure 2 as well as Figure 5 The description of the image data processing method in the corresponding embodiments can also be performed as described above. Figure 11 The description of the image data processing apparatus in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated.

[0313] Furthermore, it should be noted that this application embodiment also provides a computer-readable storage medium, which stores a computer program executed by the aforementioned image data processing device. The computer program includes computer instructions, and when the processor executes the computer instructions, it can execute the aforementioned... Figure 2 as well as Figure 5The description of the image data processing method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments related to this application, please refer to the description of the method embodiments of this application.

[0314] Furthermore, it should be noted that this application also provides a computer program product or computer program, which may include computer instructions, which may be stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor may execute the computer instructions, causing the computer device to perform the aforementioned actions. Figure 2 as well as Figure 5 The description of the image data processing method in the corresponding embodiments is already provided and will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer program products or computer program embodiments related to this application, please refer to the description of the method embodiments of this application.

[0315] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0316] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.

[0317] The modules in the device of this application embodiment can be merged, divided, and deleted according to actual needs.

[0318] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0319] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0320] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. An image data processing method, characterized in that, include: Obtain N target input image sequences and the corresponding target stitched image sequences; The target stitched image sequence refers to the image sequence obtained by stitching together the N target input image sequences according to the target sequence stitching order; N is a positive integer greater than 1; Obtain the target encoding order sequence corresponding to the splicing order of the target sequence, and perform a first feature extraction process on each of the N target input image sequences based on the target encoding order sequence to obtain the target input image features of each target input image sequence; The target stitched image sequence is subjected to a second feature extraction process on multiple business dimensions to obtain the target business dimension features of the target stitched image sequence on multiple business dimensions. The target business dimension features on multiple business dimensions are then fused to obtain the target fused sequence features of the target stitched image sequence. Based on the target input image features of each target input image sequence, a target attention matrix is ​​determined for attention calculation. Based on the target attention matrix and the target fusion sequence features, target potential sequence features associated with the N target input image sequences are determined.

2. The method according to claim 1, characterized in that, The method further includes: If the target encoding sequence indicates that there is a target missing image sequence among the N target input image sequences, then the target missing image sequence is generated based on the target latent sequence features; the target missing image sequence refers to an image sequence in which an input mode is missing among the N input modes corresponding to the N target input image sequences; one target input image sequence corresponds to one input mode.

3. The method according to claim 1, characterized in that, The step of obtaining the encoded sequence corresponding to the sequence concatenation order, and performing a first-type feature extraction process on each input image sequence based on the encoded sequence to obtain the first-type image features of each input image sequence includes: The sequence splicing order is obtained, and the binary encoding component encodes the N target input image sequences based on the sequence splicing order to obtain the target encoding order sequence corresponding to the sequence splicing order; The target encoding sequence and each of the N target input image sequences are input into the target image processing model. The target image processing model performs a first feature extraction process on each target input image sequence based on the target encoding sequence to obtain the target input image features of each target input image sequence.

4. The method according to claim 3, characterized in that, The target image processing model includes a first feature extraction component; The step involves inputting the target encoding sequence and each of the N target input image sequences into a target image processing model. The target image processing model then performs a first feature extraction process on each target input image sequence based on the target encoding sequence to obtain the target input image features for each target input image sequence. This includes: N binary codes are obtained from the target encoding sequence, and based on the concatenation order of the target sequence, binary code matching processing is performed on each of the N target input image sequences to obtain the matching code corresponding to each of the N target input image sequences. Each of the N target input image sequences and its corresponding matching code are input into the first feature extraction component. The first feature extraction component then performs a first feature extraction process on each target input image sequence based on the matching code to obtain the target input image features of each target input image sequence.

5. The method according to claim 4, characterized in that, The N target input image sequences include a first input image sequence and a second input image sequence, wherein the matching code corresponding to the first input image sequence is a first matching code, and the matching code corresponding to the second input image sequence is a second matching code; The step of inputting each of the N target input image sequences and the matching code corresponding to each target input image sequence into the first feature extraction component, and having the first feature extraction component perform a first feature extraction process on each target input image sequence based on the matching code to obtain the target input image features of each target input image sequence, includes: The first input image sequence and the first matching code are input into the first feature extraction component, and the first feature extraction component performs a first feature extraction process on the first input image sequence and the first matching code to obtain the first extracted feature corresponding to the first input image sequence and the first encoding vector corresponding to the first matching code. The first extracted feature and the first encoded vector are fused to obtain the first target extracted feature corresponding to the first input image sequence; The second input image sequence and the second matching code are input into the first feature extraction component, and the first feature extraction component performs a first feature extraction process on the second input image sequence and the second matching code to obtain a second extracted feature corresponding to the second input image sequence and a second encoding vector corresponding to the second matching code. The second extracted features and the second encoded vector are fused to obtain the second target extracted features corresponding to the second input image sequence; The first target extraction feature and the second target extraction feature are used as the target input image features of each target input image sequence.

6. The method according to claim 1, characterized in that, The second feature extraction process is performed on the target stitched image sequence across multiple business dimensions to obtain target business dimension features of the target stitched image sequence across these multiple business dimensions. The target business dimension features across these multiple business dimensions are then fused to obtain target fused sequence features of the target stitched image sequence, including: The target stitched image sequence is input into the target image processing model, and the second feature extraction component in the target image processing model performs dimensional feature extraction processing on the target stitched image sequence in multiple business dimensions to obtain the business dimension features of the target stitched image sequence in the multiple business dimensions. The target business dimension is determined from the plurality of business dimensions, and the target business dimension features of the target stitched image sequence on the target business dimension are obtained from the business dimension features of the target stitched image sequence on the plurality of business dimensions. When each of the multiple business dimensions is determined as a target business dimension, the target business dimension features of the target stitched image sequence on the multiple business dimensions are obtained. The target business dimension features of the target stitched image sequence are fused across multiple business dimensions to obtain the target fused sequence features of the target stitched image sequence.

7. The method according to claim 6, characterized in that, The multiple business dimensions include channel dimension, height dimension, and width dimension; The step of determining the target business dimension from the plurality of business dimensions, and obtaining the target business dimension features of the target stitched image sequence in the target business dimension from the business dimension features of the target stitched image sequence in the plurality of business dimensions, includes: If the channel dimension is determined as the target business dimension among the plurality of business dimensions, then the business dimensions other than the channel dimension among the plurality of business dimensions are determined as business dimensions to be processed; the first business dimension among the business dimensions to be processed is the height dimension, and the second business dimension among the business dimensions to be processed is the width dimension. From the business dimension features of the target stitched image sequence in the multiple business dimensions, obtain the first business dimension feature of the target stitched image sequence in the first business dimension, and obtain the second business dimension feature of the target stitched image sequence in the second business dimension; The business feature map formed by the first business dimension feature on the first business dimension and the second business dimension feature on the second business dimension is divided into blocks to obtain multiple block feature maps corresponding to the business feature map. The multiple block feature maps are pooled to obtain pooled features of the multiple block feature maps. The pooled features of the multiple block feature maps are used as feature weights of the target stitched image sequence on the target business dimension. From the business dimension features of the target stitched image sequence in the multiple business dimensions, obtain the third business dimension feature of the target stitched image sequence in the target business dimension; Based on the feature weights in the target business dimension, the third business dimension features in the target business dimension are weighted to obtain the target business dimension features of the target stitched image sequence in the target business dimension.

8. The method according to claim 1, characterized in that, The step of determining a target attention matrix for attention calculation based on the target input image features of each target input image sequence, and determining target latent sequence features associated with the N target input image sequences based on the target attention matrix and target fusion sequence features, includes: Obtain a first weight matrix, and perform a dot product operation between the target input image features of each target input image sequence and the first weight matrix to obtain a first target processing matrix corresponding to each target input image sequence; Obtain the second weight matrix, and perform a dot product operation between the target input image features of each target input image sequence and the second weight matrix to obtain the second target processing matrix corresponding to each target input image sequence; Obtain the third weight matrix, and perform a dot product operation between the target fusion sequence features and the third weight matrix to obtain the third target processing matrix; Based on the second target processing matrix and the third target processing matrix, the target attention matrix is ​​determined, and based on the target attention matrix and the target fusion sequence features, the target latent sequence features associated with the N target input image sequences are determined.

9. The method according to claim 8, characterized in that, The step of determining the target attention matrix based on the second target processing matrix and the third target processing matrix, and determining the target latent sequence features associated with the N target input image sequences based on the target attention matrix and the target fusion sequence features, includes: The second target processing matrix and the third target processing matrix are subjected to a dot product operation to obtain the target attention matrix; Perform a dot product operation on the target attention matrix and the first target processing matrix corresponding to each target input image sequence to obtain the target association sequence features associated with the N target input image sequences; The target-associated sequence features and the target-fused sequence features are concatenated to obtain the target latent sequence features associated with the N target input image sequences.

10. The method according to claim 1, characterized in that, The method further includes: Obtain a first type of input image sequence and a second type of input image sequence for training the initial image processing model; the first type of input image sequence refers to an image sequence that completely contains the N input modalities, and the second type of input image sequence refers to an image sequence in which some input modalities are missing from the N input modalities; The first type of input image sequence and the second type of input image sequence are used as sample input image sequences. The first type of stitched image sequence corresponding to the first type of input image sequence and the second type of stitched image sequence corresponding to the second type of input image sequence are used as sample stitched image sequences. Based on the sample sequence stitching order of the sample stitched image sequences, the sample encoding order sequence of the sample stitched image sequences is determined. The sample input image sequence, the sample stitched image sequence, and the sample encoding order sequence are input into the initial image processing model. The initial image processing model performs feature extraction processing on the sample input image sequence, the sample stitched image sequence, and the sample encoding order sequence to obtain the sample latent sequence features associated with the sample input image sequence. Based on the latent sequence features of the samples, predictive label feature information associated with the sample input image sequence is determined. Based on the predicted label feature information and the sample label feature information of the sample input image sequence, the initial image processing model is trained to obtain the target image processing model. The target image processing model is used to predict the target latent sequence features of the target missing image sequence with missing input modalities from the N target input image sequences.

11. The method according to claim 10, wherein obtaining the first type of input image sequence and the second type of input image sequence for training the initial image processing model comprises: Obtain a sequence of candidate sample images associated with the N input modalities, and extract a sequence of sampled images associated with the N input modalities from the sequence of candidate sample images based on the sampling ratio associated with the N input modalities. A sampled and stitched image sequence includes N sampled image sequences corresponding to the N input modalities; Based on the input missing ratio associated with the N input modalities, the sampled stitched image sequence is divided into a first sampled stitched image sequence and a second sampled stitched image sequence. Based on the modality missing information associated with the N input modalities, input modality missing processing is performed on the N sampled image sequences corresponding to the N input modalities in the second sampled stitched image sequence to obtain the sampled stitched missing image sequence corresponding to the second sampled stitched image sequence; the number of sequences in the sampled stitched missing image sequence is consistent with the number of sequences in the second sampled stitched image sequence; The first sampled and stitched image sequence is used as the first type of input image sequence for training the initial image processing model, and the sampled and stitched missing image sequence is used as the second type of input image sequence for training the initial image processing model.

12. An image data processing apparatus, characterized in that, include: The sequence acquisition module is used to acquire N target input image sequences and the target stitched image sequences corresponding to the N target input image sequences; The target stitched image sequence refers to the image sequence obtained by stitching together the N target input image sequences according to the target sequence stitching order; N is a positive integer greater than 1; The first extraction module is used to obtain the target encoding order sequence corresponding to the splicing order of the target sequence, and to perform a first feature extraction process on each of the N target input image sequences based on the target encoding order sequence to obtain the target input image features of each target input image sequence. The second extraction module is used to perform second feature extraction processing on the target stitched image sequence in multiple business dimensions to obtain the target business dimension features of the target stitched image sequence in the multiple business dimensions, and to perform feature fusion on the target business dimension features in the multiple business dimensions to obtain the target fused sequence features of the target stitched image sequence. The feature determination module is used to determine a target attention matrix for attention calculation based on the target input image features of each target input image sequence, and to determine the target potential sequence features associated with the N target input image sequences based on the target attention matrix and the target fusion sequence features.

13. A computer device, characterized in that, Including memory and processor; The memory is connected to the processor, the memory is used to store computer programs, and the processor is used to invoke the computer programs so that the computer device performs the method according to any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted to be loaded and executed by a processor to cause a computer device having the processor to perform the method of any one of claims 1-11.

15. A computer program product, characterized in that, Includes a computer program / instruction that, when executed by a processor, implements the method according to any one of claims 1-11.