Multi-time-point image analysis method and device, electronic equipment and storage medium

Through the cascade design of feature embedding-progression representation-prior fusion, the ordinal relationship of disease development is explicitly modeled, which solves the problems of insufficient accuracy and interpretability in longitudinal medical image analysis, achieves accurate capture of subtle features and improves classification results, and enhances the adaptability and interpretability of the model.

CN120783065APending Publication Date: 2025-10-14BEIJING INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510798195.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing longitudinal medical image analysis methods have deficiencies in accuracy and interpretability, making them difficult to generalize across different disease areas and adapt to diverse medical scenarios.

Method used

A multi-time point image analysis method is adopted, through a cascade design of feature embedding-progression representation-prior fusion, to explicitly model the ordinal relationship of disease development, use orthogonal attention to extract features, and design a temporal weighted jump connection structure to improve feature recognition sensitivity and interpretability.

Benefits of technology

It achieves accurate capture of tiny feature information in images, improves the accuracy and interpretability of classification results, can be applied to image analysis of any type of task, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120783065A_ABST
    Figure CN120783065A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, in particular to a multi-time-point image analysis method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the feature extraction of a first image and a second image obtained by carrying out the image collection of a target object at different times, and obtaining a first image feature and a second image feature; and according to the first image feature and the second image feature, determining a progress characterization feature representing a change condition of the target object from the first image acquisition time to the second image acquisition time, taking the progress characterization feature as a prior feature, and performing feature fusion with the second image feature to obtain a target feature so as to determine an object category of the target object. According to the invention, through the cascade design of feature embedding-progress characterization-prior fusion, three-stage hierarchical analysis of the features is realized, tiny feature information in the image is accurately captured, and the accuracy of a classification result is further improved. And meanwhile, the interpretability of a classification result is improved by introducing a progress characterization feature as a prior feature.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of image processing, and particularly relates to a multi-time-point image analysis method and device, electronic equipment and storage medium. BACKGROUND

[0002] Longitudinal medical images are images repeatedly taken at different time points for the same patient, such as CT or X-ray images taken at each follow-up during chronic disease screening, which can continuously record the change process of the body. In terms of judging treatment effect and tracking disease development, it can provide more dynamic information than ordinary images taken only once after the disease, helping doctors make more accurate judgments. For example, in chronic diseases such as osteoarthritis, early joint degeneration can be found through longitudinal images, providing an opportunity for timely intervention; in tumor treatment, it can also be used as a non-invasive "biomarker" to evaluate the effect of chemotherapy or radiotherapy. Accurate modeling of the spatio-temporal dynamic changes in longitudinal medical images is a key prerequisite for realizing personalized medicine and individualized treatment decisions.

[0003] However, the related art has defects such as inaccurate analysis results and poor interpretability in the process of analyzing longitudinal medical images. SUMMARY

[0004] Therefore, the present disclosure provides a multi-time-point image analysis method and device, electronic equipment and storage medium, which has strong computing generalization ability, accurate prediction results and interpretability.

[0005] According to a first aspect of the present disclosure, a multi-time-point image analysis method is provided, the method comprising:

[0006] performing feature extraction on the first image and the second image respectively to obtain first image features of the first image and second image features of the second image, the first image and the second image being images obtained by image acquisition of a target object at different acquisition times, the first image acquisition time of the first image being before the second image acquisition time of the second image;

[0007] determining a progress representation feature according to the first image features and the second image features, the progress representation feature being used to represent the change of the target object from the first image acquisition time to the second image acquisition time;

[0008] using the progress representation feature as a prior feature and performing feature fusion with the second image features to obtain a target feature;

[0009] determining an object category of the target object according to the target feature.

[0010] In a possible implementation, the feature extraction is performed on the first image and the second image respectively to obtain first image features of the first image and second image features of the second image, including:

[0011] The first image and the second image are respectively input into the visual backbone network for spatial feature and texture feature extraction to obtain corresponding first input feature maps and second input feature maps;

[0012] The first input feature maps and the second input feature maps are respectively input into the mapper to obtain first image features and second image features related to the prediction task of the object class.

[0013] In a possible implementation, the first input feature maps and the second input feature maps are respectively input into the mapper to obtain first image features and second image features related to the prediction task of the object class, including:

[0014] The first input feature maps and the second input feature maps are respectively input into the mapper for dimension reduction projection to obtain first feature vectors and second feature vectors;

[0015] A first spatio-temporal matrix corresponding to the first input feature maps and a second spatio-temporal matrix corresponding to the second input feature maps are determined;

[0016] The first image features are determined according to the first feature vectors and the first spatio-temporal matrix, and the second image features are determined according to the second feature vectors and the second spatio-temporal matrix.

[0017] In a possible implementation, the progress representation features are determined according to the first image features and the second image features, including:

[0018] The first reconstruction features and the second reconstruction features are determined by performing orthogonal attention calculation on the first image features and the second image features;

[0019] The progress representation features are determined according to the difference between the first reconstruction features and the second reconstruction features.

[0020] In a possible implementation, the first reconstruction features and the second reconstruction features are determined by performing orthogonal attention calculation on the first image features and the second image features, including:

[0021] The first image features and the second image features are respectively split to obtain a first feature token set and a second feature token set;

[0022] Orthogonal attention calculation is performed based on the first feature token set and the second feature token set to obtain a third feature token set and a fourth feature token set;

[0023] The third feature token set and the fourth feature token set are respectively subjected to feature reconstruction to obtain first reconstructed features and second reconstructed features.

[0024] In a possible implementation, the orthogonal attention calculation based on the first feature token set and the second feature token set obtains the third feature token set and the fourth feature token set, including:

[0025] The orthogonal attention calculation is performed on the first feature token set as a query and the second feature token set as keys and values to obtain the third feature token set.

[0026] The orthogonal attention calculation is performed on the second feature token set as a query and the first feature token set as keys and values to obtain the fourth feature token set.

[0027] In a possible implementation, the progress representation feature is taken as a prior feature, and the second image feature is subjected to feature fusion to obtain the target feature, including:

[0028] The cross-attention calculation is performed on the progress representation feature as a query and the second image feature as keys and values to obtain the feature to be fused.

[0029] The channel splicing is performed on the feature to be fused and the second image feature to obtain the target feature.

[0030] According to a second aspect of the present disclosure, a multi-time-point image analysis device is provided, and the device includes:

[0031] The feature extraction module is configured to extract features from the first image and the second image respectively to obtain first image features of the first image and second image features of the second image, the first image and the second image being images obtained by image acquisition of a target object at different acquisition times, the first image acquisition time of the first image being before the second image acquisition time of the second image.

[0032] The feature determination module is configured to determine a progress representation feature according to the first image features and the second image features, the progress representation feature being used to represent a change of the target object from the first image acquisition time to the second image acquisition time.

[0033] The feature fusion module is configured to take the progress representation feature as a prior feature and perform feature fusion with the second image feature to obtain a target feature.

[0034] The category determination module is configured to determine an object category of the target object according to the target feature.

[0035] In a possible implementation, the feature extraction module is further configured to:

[0036] The first image and the second image are input into a visual backbone network respectively to extract spatial features and texture features, and first input feature maps and second input feature maps are obtained respectively;

[0037] The first input feature maps and the second input feature maps are input into a mapper respectively to obtain first image features and second image features related to a prediction task of the object class.

[0038] In a possible implementation, the feature extraction module is further configured to: input the first input feature maps and the second input feature maps into the mapper respectively for dimension reduction projection to obtain first feature vectors and second feature vectors;

[0039] The first input feature maps correspond to a first spatio-temporal matrix, and the second input feature maps correspond to a second spatio-temporal matrix;

[0040] The first image features are determined according to the first feature vectors and the first spatio-temporal matrix, and the second image features are determined according to the second feature vectors and the second spatio-temporal matrix.

[0041] In a possible implementation, the feature determination module is further configured to:

[0042] The first reconstruction features and the second reconstruction features are determined by performing orthogonal attention calculation on the first image features and the second image features.

[0043] The progress representation features are determined according to the difference between the first reconstruction features and the second reconstruction features.

[0044] In a possible implementation, the feature determination module is further configured to:

[0045] The first image features and the second image features are respectively split to obtain a first feature token set and a second feature token set;

[0046] Orthogonal attention calculation is performed based on the first feature token set and the second feature token set to obtain a third feature token set and a fourth feature token set;

[0047] The first reconstruction features and the second reconstruction features are obtained by performing feature reconstruction on the third feature token set and the fourth feature token set respectively.

[0048] In a possible implementation, the feature determination module is further configured to:

[0049] The first feature token set is taken as a query, and the second feature token set is taken as a key and a value to perform orthogonal attention calculation to obtain a third feature token set;

[0050] The second feature token set is taken as a query, and the first feature token set is taken as a key and a value to perform orthogonal attention calculation to obtain a fourth feature token set.

[0051] In a possible implementation, the feature fusion module is further configured to:

[0052] The progress representation feature is taken as a query, and the second image feature is taken as a key and a value to perform cross attention calculation to obtain the feature to be fused;

[0053] The feature to be fused and the second image feature are channel spliced to obtain the target feature.

[0054] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0055] According to a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, which stores computer program instructions, wherein the computer program instructions are executed by a processor to implement the above method.

[0056] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising computer-readable code or a non-volatile computer-readable storage medium carrying computer-readable code, when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device executes the above method.

[0057] In the embodiment of the present disclosure, the method performs feature extraction on a first image and a second image obtained by image acquisition of a target object at different acquisition times, to obtain a first image feature of the first image and a second image feature of the second image, determines a progress representation feature representing a change of the target object from a first image acquisition time to a second image acquisition time according to the first image feature and the second image feature, takes the progress representation feature as a prior feature, and performs feature fusion on the second image feature to obtain a target feature, so as to determine an object category of the target object. The cascade design of feature embedding-progress representation-prior fusion in the embodiment of the present disclosure realizes three-stage hierarchical analysis of features, accurately captures the micro feature information in the image, and further improves the accuracy of the classification result.

[0058] Other features and aspects of the present disclosure will become apparent from the following detailed description of example embodiments, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0059] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate example embodiments, features, and aspects of the present disclosure and serve to explain the principles of the present disclosure.

[0060] Figure 1 A flowchart of a multi-timepoint image analysis method according to an embodiment of the present disclosure is shown;

[0061] Figure 2 A schematic diagram of a multi-timepoint image analysis process according to an embodiment of the present disclosure is shown;

[0062] Figure 3 A schematic diagram of a progression characterization process according to an embodiment of the present disclosure is shown;

[0063] Figure 4 A schematic diagram of a training process of a multi-timepoint image analysis architecture according to an embodiment of the present disclosure is shown;

[0064] Figure 5 A schematic diagram of a training process of another multi-timepoint image analysis architecture according to an embodiment of the present disclosure is shown;

[0065] Figure 6 A schematic diagram of a multi-timepoint image analysis apparatus according to an embodiment of the present disclosure is shown;

[0066] Figure 7 A schematic diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0067] Various exemplary embodiments, features, and aspects of the present disclosure will be described herein below with reference to the accompanying drawings. The same reference numbers in different drawings indicate the same or similar elements. Although various aspects of embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically noted.

[0068] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0069] In addition, for the purpose of convenience and brevity, detailed descriptions of well-known functions and structures incorporated in the present disclosure can be omitted. It will be appreciated that those skilled in the art will be able to devise various modes of implementing the advantageous aspects of the present disclosure without the exercise of inventive faculty and without the aid of further experimentation.

[0070] The multi-time-point image analysis method of the embodiments of the present disclosure can be executed by an electronic device such as a terminal device or a server. The terminal device can be any fixed or mobile terminal such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, and the like. The server can be a single server or a server cluster composed of multiple servers. Any electronic device can realize the multi-time-point image analysis method of the embodiments of the present disclosure by calling computer-readable instructions stored in a memory through a processor.

[0071] In the embodiments of the present disclosure, the multi-time-point image analysis method can be applied in the medical field, that is, longitudinal medical image analysis is performed on medical images of a same target object collected at different times in any scenario such as an imaging diagnosis scenario, a surgical navigation scenario, and a radiotherapy scenario, to analyze physiological characteristics of an object included in the medical images and timely predict abnormal conditions. The target object can be an organ as a whole of a human or an animal such as skin, a limb, a heart, a lung, and a liver, or can also be a local region in the organ.

[0072] Optionally, the method can also be applied in any field such as a computer vision field and an industrial field, to perform image analysis on images of a same object collected at different times in any scenario of the corresponding field, to monitor task processing progress, and the like.

[0073] In the case where the multi-time-point image analysis method of the embodiments of the present disclosure is applied in the medical field, there are image analysis schemes for longitudinal medical image analysis by generalized estimating equations (GEE), linear mixed effects models (LME), group-based trajectory modeling (GBTM), and the like.

[0074] Among them, the Generalized Estimating Equations (GEE) corrects the correlation between multiple measurements of the same subject by assuming a working correlation structure, but it is essentially a marginal model, focusing on estimating the population average response, and lacks the ability to model the unique changes of each individual, so it cannot be directly used for prediction of individual subjects. The Linear Mixed Effects Model (LME) introduces random effects in addition to fixed effects, which apparently supports individualized prediction, but the estimation of random effects for new subjects relies on the empirical Bayes method, which has great randomness and instability, which often leads to the phenomenon of "posterior contraction" when the follow-up time points are limited, that is, the individual prediction is excessively pulled towards the population average level, making it difficult to reflect the true individual heterogeneity. Group-Based Trajectory Modeling (GBTM) assumes that the population is composed of several potential subgroups, fits a typical trajectory for each group through a polynomial curve, and classifies individuals into the most matching trajectory group, but this classification strategy not only ignores the subtle differences between subjects within the group, but also can only provide the average trajectory prediction of the group to which the new individual belongs, making it difficult to provide continuous and individualized prediction results. Although the three methods have their own advantages in revealing population-level trends, handling different data types and correlation structures, they are all difficult to provide sufficiently fine and reliable single-subject-level predictions due to their reliance on population statistical characteristics and insufficient description of individual differences, which limits their application in precision medicine.

[0075] Compared with the above three longitudinal medical image analysis methods, deep learning technology can automatically extract complex spatio-temporal features and more accurately analyze medical images to achieve accurate modeling of disease development trajectories. Currently, longitudinal medical image analysis methods based on deep learning mainly include longitudinal collaborative learning and longitudinal fusion learning.

[0076] In the related technology of longitudinal medical image analysis using a longitudinal collaborative learning framework, multiple time-point images in the clinical diagnosis process can be regarded as interrelated data sequences, and technologies such as contrast learning or deformable registration can be used to explicitly capture the spatio-temporal patterns of pathological evolution.

[0077] Exemplarily, in the field of neuroscience, contrastive learning is first used to extract deep and high-dimensional brain network representations between different scan time points of the same subject, and these spatio-temporal features are analyzed in association with the trajectory of cognitive function decline in Alzheimer's disease patients over a long period of follow-up to identify potential markers of neurodegenerative progression. In ophthalmic image analysis, the related technology uses deformable registration technology to spatially map longitudinal Optical Coherence Tomography (OCT) images at the voxel level, and by quantifying the intensity changes of gray values and the slight deformation of anatomical structures, it ensures that the extracted features are consistent with the real physiological changes of the macular area of the retina during the progression of age-related macular degeneration. In the field of malignant tumor research, the related technology uses a decoupled representation learning method based on contrastive learning to separate the longitudinal image features of patients before and after treatment and in multiple courses into a general pathological pattern shared by the group and an individual-specific biological response component, so as to extract generalizable biomarkers and take into account individualized phenotype differences in the treatment response prediction of esophageal squamous cell carcinoma and other solid tumors. Although these methods have made significant progress in integrating multi-time point image information and improving feature expression capability, they often focus on short-term spatio-temporal contrast and are difficult to capture deep dependence relationships on a longer time scale, so there is still a lack of clinical interpretability in the prediction model.

[0078] In addition, due to the significant differences in pathological mechanisms, image manifestations, and clinical endpoints between neurodegenerative diseases and tumor treatment evaluation, the related technology's longitudinal collaborative learning framework often needs to be highly customized in terms of network structure, loss function, and regularization strategy, making it difficult to perform image analysis through the same image analysis framework. Therefore, this optimization for specific application scenarios further limits its generalization ability across different disease areas.

[0079] In the related technology of longitudinal medical image analysis using a longitudinal fusion learning framework, by deeply mining the dynamic associations between medical images collected at different time points of the same patient, recurrent neural networks (RNN), temporal convolutional networks (TCN), or more flexible Transformer architectures are used to construct spatio-temporal dependence models across time series.

[0080] Exemplarily, in the field of neuroscience, in order to enhance the model's ability to capture the brain image follow-up data of Alzheimer's disease (AD) patients, a longitudinal pooling strategy is introduced in the longitudinal fusion scenario, which adaptively aggregates the RNN hidden state in the time dimension, so that the model can take into account both short-term subtle changes and long-term evolution trends. At the same time, the method based on the Transformer uses the self-attention mechanism to evaluate the global relationship between each scan time point, so as to describe the overall progress trajectory of the disease more finely. However, these methods still face two major challenges in practical application: first, traditional deep learning time series models (such as standard RNN and TCN) often have difficulty in distinguishing pathological features with only minor differences between adjacent time points, resulting in the high-dimensional representation of multiple time points converging to a "collapsed" state, that is, the model output tends to be homogeneous at each time step (such as f t ≈f (t+Δt) ), which cannot highlight the key clinical progress signals; second, due to the uneven interval of medical imaging follow-up, the existing network architecture is insufficient for feature alignment under non-uniform time intervals, which causes the distribution difference between cross-time points to be ignored or misjudged, which not only weakens the model's adaptability to variable follow-up rhythm, but also reduces the robust prediction ability of long-term evolution patterns.

[0081] Therefore, the longitudinal medical image processing method of the related art is usually designed for a specific disease, and lacks universality and cross-scene adaptability. For image analysis for each task, it often relies on customized model architecture, which is difficult to flexibly adapt to different clinical tasks (such as neurodegenerative disease monitoring and tumor efficacy evaluation), limiting its popularization and application in diversified medical scenarios. At the same time, the existing technology generally ignores the ordinal property of the disease progression process, which is a key clinical feature. Disease evolution is essentially a gradual process with strict temporal dependence, while the current method treats image features at different time points as independent samples for processing, which prevents the model from effectively capturing the microscopic progressive changes in the lesion area.

[0082] The technical problem to be solved by the embodiments of the present application is how to provide a multi-time-point image analysis method with high computational efficiency, strong generalization ability, strong robustness and high interpretability.

[0083] To solve the technical problem above, the embodiment of the present application provides a multi-time-point image analysis method, which enhances the recognition sensitivity of the model to subtle pathological changes by explicitly modeling the ordinal relationship of disease development, thereby improving the detection capability of early lesions and critical states, and can be used for image analysis of any type of task. And through the feature extraction mechanism based on orthogonal attention, the stable features and progress-related features in the image are accurately separated; at the same time, according to the prior knowledge that "the higher the time proximity, the stronger the clinical relevance", a skip connection structure with time sequence weighting is designed. This scheme can not only effectively alleviate the feature collapse problem, but also provide reliable and interpretable analysis results.

[0084] Figure 1 A flowchart of a multi-time-point image analysis method according to an embodiment of the present disclosure is shown. As shown in Figure 1 The multi-time-point image analysis method of the embodiment of the present disclosure can include the following steps S10-S40.

[0085] The following describes an electronic device as the execution subject, which can be understood as not only the electronic device itself, but also a module capable of executing computer tasks such as a processor or processing chip in the electronic device.

[0086] Step S10, the electronic device performs feature extraction on the first image and the second image respectively to obtain the first image feature of the first image and the second image feature of the second image.

[0087] In a possible implementation, the embodiment of the present application can obtain the first image and the second image through the electronic device. The first image and the second image can be images obtained by image acquisition on a target object at different acquisition times. The first image acquisition time of the first image is before the second image acquisition time of the second image, and the second image is the target image in the multi-time-point image analysis process, i.e., the second image is used as the target for the electronic device to perform the multi-time-point image analysis process. The first image is a reference image in the multi-time-point image analysis process, i.e., the first image is used to provide reference information in the process of image analysis on the second image.

[0088] In some embodiments, in the application scenario of longitudinal analysis of medical images, the first image and the second image are medical images successively acquired on the same target object, which can be used to record the physiological state of the target object at different times. The target object can be the head, limbs, heart, lungs and other organs of the human body or part of the animal body, or can also be a specific region in any organ. The first image and the second image are the same type of medical images, for example, X-ray images, CT images, or magnetic resonance images, etc.

[0089] Optionally, the first image and the second image can be acquired directly by an electronic device performing the multi-time-point image analysis method of the present application. Alternatively, the first image and the second image can also be acquired by other image acquisition devices and then sent to the electronic device for image analysis.

[0090] In some other embodiments, the first image and the second image can be original images acquired by image acquisition. Alternatively, the first image and the second image can also be images obtained by performing any image preprocessing such as image cropping, normalization processing, and noise filtering on the original images acquired by image acquisition.

[0091] In a possible implementation, after acquiring the first image and the second image, the electronic device can perform feature extraction on the first image and the second image respectively to obtain features representing the physiological information of the target object in the recorded image, so as to determine the first image feature corresponding to the first image and determine the second image feature corresponding to the second image.

[0092] Optionally, the process of feature extraction on the first image and the second image according to the present application can be implemented based on pixel features, edge and texture features, or neural network models.

[0093] Step S20: The electronic device determines a progress representation feature according to the first image feature and the second image feature.

[0094] In a possible implementation, after performing feature extraction on the first image and the second image to obtain the first image feature and the second image feature, the electronic device can further perform progress analysis based on the first image feature and the second image feature to determine a change of the target object from the first image acquisition time to the second image acquisition time. The progress analysis process is used to analyze the change of the target object between the first image and the second image.

[0095] Optionally, the progress representation feature can be determined by performing orthogonal attention calculation on the first image feature and the second image feature. This method can capture more detailed progress signals of the target object between different time points, and can also reduce the interference of irrelevant noise on the representation by using the orthogonal attention strategy, so as to accurately and explainably represent the change of the target object from the first image acquisition time to the second image acquisition time.

[0096] Step S30: The electronic device fuses the progress representation feature as prior feature with the second image feature to obtain a target feature.

[0097] In a possible implementation, after the electronic device determines the progress representation feature according to the first image feature and the second image feature, the progress representation feature can be further taken as a prior feature, and the feature fusion is performed on the second image feature, so as to refer to the change of the target object, and accurately determine the target feature representing the physiological state of the target object at the current time.

[0098] Optionally, the feature fusion can be performed in the manner of directly splicing the progress representation feature and the second image feature, or the progress representation feature and the second image feature can be subjected to cross-attention calculation, and the target feature can be obtained by extracting the features related to the current classification task in the second image feature.

[0099] In step S40, the electronic device determines the object category of the target object according to the target feature.

[0100] In a possible implementation, after the electronic device determines the target feature, the object category of the target object can be determined according to the target feature. The object category can be preset in different application scenarios. For example, when the organ state of a human or an animal is monitored by using the multi-time-point image analysis method, the object category obtained can be physiological state abnormality or physiological state normality. Or, when the state of a chemical reaction is monitored by using the multi-time-point image analysis method, the object category obtained can be a specific reaction stage of the chemical reaction process.

[0101] Optionally, the classification manner can be that the target feature is input into a trained classifier, and the corresponding object category is determined according to the output classification result. The classifier can include a support vector machine, a decision tree classifier, a random forest classifier, or any classifier capable of classifying based on feature information.

[0102] For example, in the medical field, the object category of the target object can be a disease development stage or a current morphological change state of a physiological tissue. In the industrial pipeline scene, the object category of the target object can be a pipeline node. In the chemical experiment field, the object category of the target object can be a reaction stage of a reactant.

[0103] Based on the above technical features, the embodiment of the present disclosure realizes three-stage hierarchical analysis of features by the cascade design of feature embedding-progress representation-prior fusion, accurately captures the micro feature information in the image, and further improves the accuracy of the classification result. Moreover, the feature extraction process can accurately and subtly extract features for images of any category, so that the multi-time point image analysis method of the embodiment of the present disclosure can be applied to any image analysis application scenario, and has strong generalization ability. At the same time, by introducing the progress representation feature as a prior feature, the explainability of the classification result is increased.

[0104] Figure 2 A schematic diagram of a multi-time point image analysis process according to an embodiment of the present disclosure is shown. As shown in Figure 2 The multi-time point image analysis scheme described above in the embodiment of the present disclosure can be implemented through an image analysis framework including four modules, i.e., a feature embedding module, a progress representation module, a feature fusion module, and a classification module.

[0105] Specifically, after obtaining a second image in which a target object category needs to be analyzed and a first image providing reference information for the image analysis process, the feature embedding module can be used to extract features from the first image and the second image respectively, to obtain first image features representing target object information in the first image and second image features representing target object information in the second image.

[0106] After completing feature embedding, the obtained first image features and second image features can be input into the progress representation module to determine a progress representation feature representing the change of the target object from the first image acquisition time to the second image acquisition time by orthogonal attention calculation of the first image features and the second image features in the progress representation module.

[0107] After obtaining the progress representation feature, the progress representation feature can be input into the feature fusion module as a prior feature together with the second image features to perform feature fusion by cross-attention calculation, to obtain a target feature accurately representing the current state of the target object in combination with the development trend.

[0108] Optionally, the object category of the target object can be determined by inputting the target feature into the classification module for classification. In different application scenarios, the classification task performed by the classification module is different, and the object category that can be output is also different.

[0109] In the above step S10, the electronic device of the embodiment of the present disclosure can extract features from the first image and the second image by any method, i.e., the feature embedding module of the embodiment of the present disclosure can embed any feature extraction algorithm.

[0110] In a possible implementation, the feature embedding module of the embodiment of the present application can include a visual backbone network and a mapper. That is, the process of feature extraction through the feature embedding module is to sequentially extract the features of the input images through the visual backbone network and the mapper. Illustratively, the electronic device inputs the first image and the second image into the visual backbone network for spatial feature and texture feature extraction, respectively, to obtain the corresponding first input feature map and second input feature map. Then, the first input feature map and the second input feature map are input into the mapper, respectively, to obtain the first image feature and the second image feature related to the prediction task of the object class.

[0111] In some embodiments, the visual backbone network may, for example, be a deep residual network ResNet, a visual Transformer, or a lightweight convolutional network, etc., which is used to extract rich spatial and texture features from the input first image and second image through convolution or the like to obtain the corresponding first input feature map and second input feature map.

[0112] In some embodiments, the mapper generally includes several fully connected layers or convolutional layers, and may even be combined with a self-attention mechanism, to project the high-dimensional features output by the visual backbone network into a more compact and task-related embedding space. In this process, the mapper is not only responsible for feature dimension reduction and feature fusion, but also can strengthen the expression ability of information related to the final classification task through activation functions and regularization methods.

[0113] Illustratively, after the spatial feature and texture feature extraction through the visual backbone network, the first input feature map corresponding to the first image and the second input feature map corresponding to the second image are obtained, the first input feature map and the second input feature map are input into the mapper for dimension reduction projection to obtain the first feature vector and the second feature vector. A first spatio-temporal matrix corresponding to the first input feature map is determined, and a second spatio-temporal matrix corresponding to the second input feature map is determined. The first image feature is determined according to the first feature vector and the first spatio-temporal matrix, and the second image feature is determined according to the second feature vector and the second spatio-temporal matrix. Wherein, the first spatio-temporal matrix is used to represent the temporal position information and the spatial position information corresponding to the first image, and the second spatio-temporal matrix is used to represent the temporal position information and the spatial position information corresponding to the second image.

[0114] Optionally, the process of determining the first image feature according to the first feature vector and the first spatio-temporal matrix may be to calculate the sum of the first feature vector and the product of the first spatio-temporal matrix and the unit matrix. Correspondingly, the process of determining the second image feature according to the second feature vector and the second spatio-temporal matrix may be to calculate the sum of the second feature vector and the product of the second spatio-temporal matrix and the unit matrix.

[0115] Illustratively, in the process of determining the first image feature or the second image feature, the first image or the second image I i(1) where i ∈ {prior (first image), curr (second image)} is input to the feature embedding module The original images are image-processed and projected to obtain one-dimensional vectors E i ∈ R D , i ∈ {prior, curr}, R is a set of real numbers, and D is the number of elements in the one-dimensional vector, that is, the projection process can be implemented by formula (1).

[0116]

[0117] Next, in order to explicitly encode the spatiotemporal information at the scanning time point in the embedding, the embodiment of the application can construct a time and space embedding matrix to obtain a first spatiotemporal matrix st prior and a second spatiotemporal matrix st curr by formula (2).

[0118] ST = {st prior ; st curr} ∈ R 2×D (2)

[0119] Each row vector st prior ,st curr contains two parts of features, that is, the time sequence position information (such as the scanning interval) and the spatial position information (such as the center coordinates of a specific object region in the target object) corresponding to the same object region time point in the target object. Finally, feature fusion can be performed by an element-by-element multiplication and element-by-element addition fusion mechanism according to formula (3) to obtain the first image feature and the second image feature.

[0120]

[0121] where 1 L is a unit matrix matching the dimension of the first feature vector or the second feature vector. The first feature vector or the second feature vector E i as the original image embedding is combined with the first spatiotemporal matrix or the second spatiotemporal matrix st i as the spatiotemporal embedding by bit alignment to generate a parameterized context-aware vector T i (that is, the first image feature and the second image feature). During the entire training process, the feature embedding module and the spatiotemporal embedding parameters {st prior ; st curr} can be optimized cooperatively by back propagation to ensure that the final embedding can accurately align the pathological changes at different time points and maintain spatial continuity and anatomical consistency, thereby providing reinforced and context-dependent representations for subsequent temporal analysis or prediction tasks.

[0122] That is, the entire feature embedding module and the first and second spatio-temporal vectors described above can all be trained end-to-end in advance, and the learnable parameters inside them can be automatically optimized through backpropagation during training, so that these parameterized embeddings can be used in the application process to more accurately represent the image information at each time point and its spatio-temporal evolution relationship, so as to extract accurate and effective image features.

[0123] In step S20 described above, the progress representation module of the embodiment of the present application can determine the progress representation features corresponding to the first image and the second image in any manner. Optionally, the electronic device in the embodiment of the present application can reduce the interference of irrelevant noise on the features by performing orthogonal attention calculation, so as to extract features with high relevance to the object classification task, high accuracy and high interpretability.

[0124] Optionally, the electronic device can determine the first reconstruction feature and the second reconstruction feature by performing orthogonal attention calculation on the first image feature and the second image feature after determining the first image feature and the second image feature. Then, the progress representation feature is determined according to the difference between the first reconstruction feature and the second reconstruction feature. The orthogonal self-attention module calculates by projecting the input feature into an orthogonal space, and then maps the result back to the original space to obtain the output feature.

[0125] In some embodiments, the first reconstruction feature can be a feature obtained by directly performing orthogonal attention calculation on the first image feature as a query and the second image feature as a key and a value. The second reconstruction feature can be a feature obtained by directly performing orthogonal attention calculation on the second image feature as a query and the first image feature as a key and a value.

[0126] In other embodiments, the electronic device can split the first image feature and the second image feature respectively to obtain a first feature token set and a second feature token set. Orthogonal attention calculation is performed based on the first feature token set and the second feature token set to obtain a third feature token set and a fourth feature token set. The third feature token set and the fourth feature token set are reconstructed respectively to obtain the first reconstruction feature and the second reconstruction feature.

[0127] In other words, the core process of the progression characterization module can be divided into three consecutive steps: the first step is the "soft split" operation, which performs a differentiable split of the first image feature and the second image feature of the input according to a predefined segmentation strategy to ensure that the subsequent attention calculations of each segment are mutually orthogonal in different subspaces. The second step is the orthogonal attention calculation. The module independently executes the self-attention mechanism within each subspace, while forcing the attention matrices between each subspace to be orthogonal to each other, thereby minimizing the information redundancy between different feature subsegments and highlighting important representations related to disease progression. The third step uses a "reconstruction" operation, which is the inverse transformation of the "soft split", to recombine the orthogonal attention outputs of multiple subspaces into an integrated vector of the same dimension as the original input.

[0128] First, the progress characterization module can be operated through a “soft split” d=D / n Each input one-dimensional feature vector T prior (first image feature), T curr (Second image feature)∈R D It can be split into n context tokens of length d=D / n, so as to establish a granularity-controllable long-range feature interaction on the n×n attention map. Wherein, R is a set of real numbers, D is the number of elements in the first image feature and the second image feature before soft splitting, n is the number of tokens included in the first feature token set and the second feature token set after splitting, and d is the number of elements included in each token. Subsequently, the embodiment of the present application introduces inverse cosine similarity to measure the difference between tokens, and on this basis constructs a segmented orthogonal attention (OrA) module, whose core form can be expressed by formula (4).

[0129]

[0130] Where Q (query), K (key), and V (value) are the first feature token set or the second feature token set, corresponding to the three main components of the attention mechanism. σ is the softmax function, and J is an all-one matrix. By subtracting the attention weighted output, the difference signals between different time points are explicitly strengthened and redundant information is suppressed. In other words, the above orthogonal attention calculation process uses the first feature token set as the query, and the second feature token as the key and value for orthogonal attention calculation to obtain the third feature token set. The second feature token set is used as the query, and the first feature token is used as the key and value for orthogonal attention calculation to obtain the fourth feature token set.

[0131] Furthermore, after completing the interaction between the two sets of tokens, the progress representation module in the electronic device can use the "reconstruction" operation, which is the inverse transformation of the above-mentioned "soft splitting" algorithm, to restore the attention results of the current time point and the previous time point into the high-dimensional first reconstruction feature. and a second reconstructed feature and the final progress representation feature is generated by element-wise subtraction of the first reconstructed feature and the second reconstructed feature through equation (5).

[0132]

[0133] Figure 3 A schematic diagram of a progress representation process according to an embodiment of the present disclosure is shown. As shown, after determining the first image feature and the second image feature through step S10, the above two image features are input into the progress representation module to perform soft classification, obtaining the first feature token set and the second feature token set corresponding respectively. Then, orthogonal attention calculation is performed on the first feature token set and the second feature token set respectively, obtaining the third feature token set and the fourth feature token set. Then, feature reconstruction is performed on the third feature token set and the fourth feature token set respectively, obtaining the first reconstructed feature and the second reconstructed feature. The final progress representation feature is obtained by subtracting the first reconstructed feature and the second reconstructed feature. Figure 3

[0134] In some possible implementations, the progress representation module in the embodiment of the present application can be pre-trained, that is, the parameters thereof can be determined through training, and the accuracy of the output progress representation feature is directly used in the application process. In the training process, the progress constraint loss L is supervised to ensure the reliability of the determined parameters.

[0135] Optionally, in the process of training the progress representation module, the ordinal perception progress constraint loss L PC is used to supervise the model parameters. Wherein, is an ordinal constraint loss function with clinical interpretability, which is designed based on contrast learning, aiming to pull the spatial distance of positive sample pairs and push the spatial distance of negative sample pairs. The positive sample pair is a sample pair with similar representation in the feature space, which is set as the target object with the same progress object classification result here; on the contrary, the negative sample pair is a sample pair with dissimilar representation in the feature space, which is set as the target object with different progress object classification results here. The ordinal property is reflected in the dynamic penalty coefficient w ib , which gives different penalty weights according to the difference between different progress object categories according to exponential decay, which is set as equation (6) here.

[0136]

[0137] where β controls the decay rate, α is the adjustment factor, ΔY i represents the object category of the target progress object i, and ΔY​b object class representing target progress object b, ΔY i = ΔY b indicates that target progress objects i and b have the same progress object class, and both constitute positive sample pairs with each other; ΔY i ≠ ΔY b indicates that target progress objects i and b have different progress object classes, and both constitute negative sample pairs with each other. These parameters constitute dynamic weights w ib under the contrast learning framework, which are subsequently used for similarity representation between target object progress representation features.

[0138] Therefore, the ordinal perception loss l(i,j) depends on supervised contrast learning and can be represented by formula (7).

[0139]

[0140] where τ controls the degree of discreteness of the representation, ensuring that the model not only focuses on the similarity of positive sample pairs of the same level, but also reasonably punishes the difference of negative sample pairs of different levels, so that the learned progress vector accurately captures the subtle and key spatio-temporal changes in the change process of the target object while maintaining spatial continuity and feature consistency. is a similarity calculation operation, and its result represents the distance between the progress representation feature p i of the target object i obtained by formula (5) and the progress representation feature p j of the target object j. w ib is obtained from formula (6), which is a dynamic progress weight, multiplied by S(p i ,p j ) in formula (7) based on different progress levels. l(i,j) represents the calculation of formula (7) for a certain positive sample target object j of the target object i, then for the set J(i) composed of all positive sample target objects of the target object i in the data set, formula (7) is calculated, and finally the loss of each target object is integrated by formula (8).

[0141]

[0142] ensuring that the model not only focuses on the similarity of positive sample pairs of the same level, but also reasonably punishes the difference of negative sample pairs of different levels—so that the learned progress representation feature accurately captures the subtle and key spatio-temporal changes in the change process of the target object while maintaining spatial continuity and feature consistency.

[0143] In step S30, the embodiment of the present application can fuse the progress representation feature and the second image feature through cross-attention calculation, and further complete the whole feature fusion process through feature splicing. Illustratively, the electronic device can perform cross-attention calculation on the progress representation feature as a query, and the second image feature as a key and a value, to obtain a to-be-fused feature. The to-be-fused feature and the second image feature are spliced in a channel to obtain a target feature.

[0144] Therefore, the feature fusion module of the embodiment of the present application can realize the time sequence feature fusion through the skip connection structure with time sequence weighting and the segmented cross-attention mechanism.

[0145] The feature fusion architecture is constructed based on the time proximity prior, based on the prior knowledge that the closer to the current image analysis time, the more important the current object classification result is, and the progress representation feature P is used as a feature prior vector. In this stage, the same tokenization processing procedure as the segmented orthogonal attention can be used to process the progress representation feature P and the second image feature T and fuse P and the current visit feature T curr , but replace the orthogonal attention with cross-attention. In the attention calculation process, the second image feature T curr is used as the key K and the value V in the attention mechanism, the progress representation feature P is used as the query Q, and the recent time sequence information is reserved through the cross-scale connection. Finally, the cross-attention output is spliced in a channel with T curr to form the final fused target feature. Finally, the target feature is input into the classification module to output the result through the classification module.

[0146] Optionally, one or more of the above-mentioned modules for image analysis in the embodiment of the present application can be trained together. In the training process, the parameter adjustment of the progress representation module is supervised by the progress constraint loss , and the final classification module is supervised by the classification loss . In the scenario of training the progress representation module and the classification module together, the final training overall optimization target is the weighted sum of and .

[0147]

[0148] Based on the above technical features, the embodiment of the present application designs a three-stage progressive processing architecture for longitudinal feature extraction of images. Through the cascade design of feature embedding, progress representation and prior fusion, hierarchical analysis of longitudinal image spatio-temporal features is realized. Compared with the traditional single-stage processing method, this architecture can exhibit excellent adaptability in subsequent different classification tasks and has cross-modal generalization ability. At the same time, the progress representation process effectively solves the feature confusion problem caused by the consistency of object structure in the image through the setting of the segmented orthogonal attention mechanism. Through difference-driven feature decoupling, the detection sensitivity of the model to the micro features of the target object is improved, and more accurate image features are extracted.

[0149] Figure 4 A training process schematic diagram of a multi-time point image analysis architecture according to an embodiment of the present disclosure is shown. As shown in Figure 4 The target object in the embodiment of the present application can be a knee joint, i.e., the first image and the second image are images obtained by image acquisition of the knee joint. Further, the knee joint information in the second image can be analyzed by the image analysis architecture of the present application to determine the state or degradation stage of the knee joint, etc. Or, the development stage of knee arthritis can also be directly predicted.

[0150] Optionally, for the image analysis architecture for multi-time point image analysis of knee joint images, the training process includes determination of a sample set, data enhancement, feature embedding, progress representation, prior feature fusion, and internal validation for supervised constraint by a loss function until the model training is completed.

[0151] In some embodiments, the input of the image analysis framework is a knee joint X-ray two-dimensional image, and the object category classified by the classifier is used to represent the severity of knee osteoarthritis. Knee osteoarthritis (KOA) is a global disease and is the main cause of chronic disability in adults over 60 years old. The Kellgren Lawrence grading (KLG) system is used for severity assessment in clinical practice, which can classify a single joint into one of five grades, 0 indicating normal and 4 indicating the most severe. Accurate diagnosis and prediction of KLG are crucial for early intervention, which can significantly improve the quality of life of patients.

[0152] The process of sample set determination can include collecting the left and right knees of the human body to obtain sample images, and extracting the knee joint lesion area in the sample images. Further, the horizontal flip of the right knee joint lesion area is similar to the left knee, which is convenient for subsequent model processing. Optionally, the knee joint lesion area can be pixel normalized, and two images of the same knee at different time points are input as a first sample image and a second sample image.

[0153] Further, the data augmentation process can be to uniformly resample the first sample image and the second sample image to a size of 1024*1024, use data augmentation operations such as contrast change, translation, etc., and finally randomly crop to a size of 896*896 to input into the framework.

[0154] The feature embedding process inputs the first sample image and the second sample image after data augmentation processing as paired double-time-point clinic images, and obtains visual embedding features through a 2-dimensional visual embedding backbone network with the same architecture and shared parameters, and generates feature representations with spatiotemporal perception properties through spatiotemporal coding operations.

[0155] The progress representation process is used to decouple the feature representations of the double-time-point determined by the feature embedding process by segmented orthogonal attention, perform post-soft splitting operations, calculate the inverse cosine similarity, and reconstruct the progress representation features after subtraction. The progress vector receives the optimization of the ordinal progress constraint in this stage, and the penalty coefficient in the constraint controls the constraint degree before the ordinal progress, and the progress vector distribution in the latent representation space.

[0156] The prior feature fusion process takes the progress representation features as the prior condition of the decision output stage, and uses a clinically designed skip connection to retain recent features (i.e., the features of the second sample image), and inputs the fused features into the classifier. The ordinal loss is used to optimize the final output of the decision. According to different application scenarios, the output result can be a knee state classification result or a knee lesion state.

[0157] Further, internal validation can be performed through an internal training set during the training process. After the training is completed, the model trained in the above steps can be loaded to test the new input paired different time point knee X-ray object images.

[0158] Figure 5 A training process schematic diagram of another image analysis architecture according to an embodiment of the present disclosure is shown. As shown in Figure 5 The target object in the embodiment of the present application can be an esophagus, and the first image and the second image are images obtained by image acquisition of the esophagus. Further, the esophageal information in the second image can be analyzed by the image analysis architecture of the present application to determine the state or degradation stage of the esophagus. Alternatively, the development stage of esophagitis can also be directly predicted.

[0159] Optionally, for the image analysis architecture for image analysis of esophagus images, the training process includes determination of a sample set, data augmentation, feature embedding, progress representation, prior feature fusion, and internal validation for supervised constraint through a loss function until the model training is completed.

[0160] In some embodiments, the input of the image analysis framework is a contrast-enhanced CT three-dimensional image of the esophagus, and the object class finally classified by the classifier can be used to represent the physiological state of the esophagus or the development stage of esophageal squamous cell carcinoma. Esophageal squamous cell carcinoma (ESCC) accounts for about 90% of esophageal cancer cases in China, and neoadjuvant chemoradiotherapy (nCRT) combined with esophagectomy is the standard treatment for locally advanced ESCC. There is evidence that for patients who achieve pathological complete response (pCR) after nCRT, a watchful waiting strategy is more appropriate than esophagectomy. Therefore, preoperative prediction of individual ESCC patient treatment response (pCR and non-pCR) is clinically significant for individualized treatment decisions.

[0161] wherein the process of determining the sample set can include extracting the cancerous region of the esophagus before and after nCRT treatment, respectively performing pixel normalization, and inputting the contrast-enhanced CT images before and after treatment as a pair of first sample images and second sample images.

[0162] Further, the data augmentation process can be to uniformly resample the interval of the first sample image and the second sample image to (1, 0.76, 0.76) mm size, use contrast change, translation, etc. Data augmentation operations, and finally crop to 32*64*64 size paired input into the framework.

[0163] The feature embedding process inputs the first sample image and the second sample image after data enhancement processing as a pair of double-time-point clinic images, and respectively inputs the same architecture and parameter-shared 2-dimensional visual embedding backbone network to obtain visual embedding features, and generates feature representations with spatiotemporal perception properties via spatiotemporal coding operations.

[0164] The progress representation process is used to decouple the feature representations of the double-time-point determined by the feature embedding process via segmented orthogonal attention, perform post-soft splitting operations, calculate the inverse cosine similarity, and reconstruct the progress representation features after subtraction. The progress vector receives the optimization of the ordinal progress constraint in this stage, and the penalty coefficient in the constraint controls the constraint degree before the ordinal progress, and the progress vector distribution in the latent representation space.

[0165] The prior feature fusion process takes the progress representation features as the prior condition of the decision output stage, designs a skip connection based on clinical prior to retain recent features (i.e., the features of the second sample image), and inputs the fused progress representation features into the classifier, and uses an unbalanced classification loss to optimize the final output of the decision.

[0166] Further, internal validation can be performed by an internal training set during the training process. After the training is completed, the model trained in the above steps can be loaded to test the object images of the paired pre-and post-treatment esophageal cancer contrast-enhanced CT inputted newly.

[0167] In the above two embodiments, the training process is consistent in different classification task scenarios, but due to the differences in image modalities and analysis objects, it is often necessary to train or fine-tune a special model for each classification task to ensure the best performance in the respective field. Of course, a general base model can also be pre-trained on large-scale, multi-modal data, and then fine-tuned for specific diseases downstream with a small amount of labeled data to balance the generalization ability and special performance of the model.

[0168] Optionally, the above training process can be implemented through a preset end-to-end training strategy, for example, the model training framework set by the embodiments of the present application can be used for training. Wherein, the end-to-end training process through the improved training framework can be parameter updating through the AdamW optimizer (β1=0.9, β2=0.999) with a weight decay coefficient of 0.03, and the training process adopts a cosine annealing learning rate scheduling strategy of 50 epochs. According to the characteristics of different classification tasks, the training process can adopt differentiated training configurations: for example, in the task of evaluating the state of knee osteoarthritis, the initial learning rate can be set to 5×10-6, the batch size is 16, and the backbone network loads the ImageNet pre-training weight provided by TorchVision0.19.1. In the esophageal squamous cell carcinoma treatment response prediction task, in order to alleviate the problem of data scarcity, a learning rate of 3×10-4 and a batch size of 64 can be used, and the backbone network is pre-trained and initialized based on the auxiliary discrimination task. The model adopts a balanced loss combination of λ1=λ2=0.5, wherein the progress constraint loss L PC In the knee osteoarthritis evaluation task, the hyperparameter setting is α=2, β=2, and the temperature parameter τ is fixed at 0.07 to control the feature discretization degree. At the network architecture level, the token number of the segment attention mechanism is uniformly set to 16, and the dropout rate of the fully connected layer is 0.4 to prevent overfitting. The experiment is implemented under the PyTorch2.4.1 framework. The training scheme dynamically adjusts the learning strategy and regularization parameters to ensure stable convergence and generalization performance of the model on different types of image data.

[0169] The multi-time point image analysis framework provided by the embodiments of the present application has significant technical advantages and good clinical applicability in the field of medical image analysis. Specifically, the framework has achieved better analysis performance than the prior art in multiple disease types and multi-modal image data through multi-center clinical trials. At the same time, the image analysis framework of the embodiments of the present application can use a large number of existing single-time models as the visual backbone network of the framework, and through multi-center clinical trials, it has shown significant improvement in clinical indicators such as AUC (p<0.001).

[0170] Taking osteoarthritis as an example, when the image analysis framework described in the present application is used to perform KLG grading on knee osteoarthritis on a clinical data set containing 4,796 longitudinal images, the average accuracy reaches 78.63% (95% confidence interval is 76.23% to 81.00%), which is significantly higher than the existing automatic evaluation method for osteoarthritis based on X-ray images (statistical significance level p<0.001) and traditional single-time point image analysis models (p<0.001). This shows that the embodiments of the present application have stronger discrimination ability and time series modeling ability in processing long-term follow-up (longitudinal) image data.

[0171] In another applicable scenario, the embodiments of the present application are also applicable to the prediction task of solid tumor treatment response. In multi-center clinical image data from four medical centers, a total of 208 patients with locally advanced esophageal squamous cell carcinoma (ESCC), the application of the framework of the present application to analyze the enhanced CT images has an AUC (area under the ROC curve) of 87.90% (95% confidence interval is 85.61% to 90.18%) for predicting treatment response, which is significantly better than the current mainstream artificial or traditional model analysis method based on contrast-enhanced CT images (p<0.001), and is statistically significantly better than the traditional single-time point model (p<0.001). The above results show that the framework of the present application also has strong robustness and discrimination ability in dealing with unbalanced classification problems.

[0172] Further, the image analysis framework provided by the embodiments of the present application has good compatibility and scalability, and can directly nest or integrate a large number of existing single-time point image analysis models as a visual backbone network. Without changing the structure of the backbone network, by introducing the multi-time series modeling mechanism proposed in the framework of the present application, significant improvements in key clinical evaluation indicators such as AUC and accuracy can be achieved in multiple disease scenarios, fully demonstrating the universality and practicality of the framework.

[0173] In summary, the image analysis framework of the embodiments of the present application has significantly better performance than the prior art in multi-modal, multi-time point medical image analysis tasks, and has good transferability and model integration capability, and has very high clinical transformation potential and industrial application value.

[0174] Based on the above technical features, the embodiment of the present application introduces a dynamic ordinal constraint mechanism (L PC ) into the training process to integrate the ordinal characteristics of the dynamic development of the target object into the deep learning framework, which is significantly better than the non-ordinal progress loss. At the same time, the feature fusion strategy based on time proximity and the progress prior fusion strategy conform to the field rules of the classification task, which significantly improves the time modeling capability of the multi-time-point image analysis architecture.

[0175] At the same time, the multi-time-point image analysis architecture of the embodiment of the present application can adopt a flexible dual-path processing architecture of the same architecture. In the case where different input images share the same semantics, the parameters can be shared to save energy consumption. In the case where the input images do not share semantics, independent parameters can be used to improve the model accuracy. And in the progress representation stage, the innovative tokenization-reconstruction mechanism makes the model maintain the effectiveness of the global feature dependency while maintaining the local feature dependency.

[0176] In addition, the multi-time-point image analysis architecture of the embodiment of the present application adopts a modular design, supports plug-and-play of mainstream visual backbone networks (ResNet / ViT, etc.), and users can flexibly configure according to the data characteristics of different classification tasks. And can provide a standardized preprocessing interface to automatically complete the target object region detection-intensity normalization process in the image, which can be directly used for training or testing of the framework, effectively shortening the deployment time.

[0177] Finally, the multi-time-point image analysis architecture of the embodiment of the present application can combine the prior features of the classification task for object classification, which significantly improves the accuracy of the model. And by introducing a dynamic penalty coefficient (wib) into the training process, the model decision-making process conforms to the rules of the development of the target object, and has strong interpretability.

[0178] Figure 6 A schematic diagram of a multi-time-point image analysis apparatus according to an embodiment of the present disclosure is shown. As Figure 6 shown, the multi-time-point image analysis apparatus of the embodiment of the present application can include:

[0179] The feature extraction module 60 is configured to perform feature extraction on the first image and the second image respectively to obtain first image features of the first image and second image features of the second image, the first image and the second image being images obtained by image acquisition of the target object at different acquisition times, the first image acquisition time of the first image being before the second image acquisition time of the second image.

[0180] The feature determination module 61 is configured to determine a progress representation feature according to the first image feature and the second image feature, the progress representation feature being used to represent a change of the target object from a first image capturing time to a second image capturing time.

[0181] The feature fusion module 62 is configured to perform feature fusion on the progress representation feature and the second image feature to obtain a target feature, the progress representation feature being used as a prior feature.

[0182] The category determination module 63 is configured to determine an object category of the target object according to the target feature.

[0183] In a possible implementation, the feature extraction module 60 is further configured to:

[0184] input the first image and the second image into a visual backbone network respectively to extract spatial features and texture features, and obtain a first input feature map and a second input feature map respectively;

[0185] input the first input feature map and the second input feature map into a mapper respectively to obtain a first image feature and a second image feature related to a prediction task of the object category.

[0186] In a possible implementation, the feature extraction module 60 is further configured to:

[0187] input the first input feature map and the second input feature map into the mapper respectively for dimension reduction projection to obtain a first feature vector and a second feature vector;

[0188] determine a first space-time matrix corresponding to the first input feature map and a second space-time matrix corresponding to the second input feature map;

[0189] determine the first image feature according to the first feature vector and the first space-time matrix, and determine the second image feature according to the second feature vector and the second space-time matrix.

[0190] In a possible implementation, the feature determination module 61 is further configured to:

[0191] determine a first reconstructed feature and a second reconstructed feature by performing orthogonal attention calculation on the first image feature and the second image feature;

[0192] determine the progress representation feature according to a difference between the first reconstructed feature and the second reconstructed feature.

[0193] In a possible implementation, the feature determination module 61 is further configured to:

[0194] split the first image feature and the second image feature respectively to obtain a first feature token set and a second feature token set;

[0195] performing orthogonal attention calculation based on the first feature token set and the second feature token set to obtain a third feature token set and a fourth feature token set;

[0196] perform feature reconstruction on the third feature token set and the fourth feature token set respectively to obtain a first reconstructed feature and a second reconstructed feature.

[0197] In a possible implementation, the feature determination module 61 is further configured to:

[0198] perform orthogonal attention calculation on the first feature token set as a query and the second feature token set as a key and a value to obtain the third feature token set;

[0199] perform orthogonal attention calculation on the second feature token set as a query and the first feature token set as a key and a value to obtain the fourth feature token set.

[0200] In a possible implementation, the feature fusion module 62 is further configured to:

[0201] perform cross attention calculation on the progress representation feature as a query and the second image feature as a key and a value to obtain a feature to be fused;

[0202] perform channel splicing on the feature to be fused and the second image feature to obtain a target feature.

[0203] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, they will not be repeated here.

[0204] The embodiments of the present disclosure also propose a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions are executed by a processor to implement the above method. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0205] The embodiments of the present disclosure also propose an electronic device, comprising: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.

[0206] The embodiments of the present disclosure also provide a computer program product, comprising computer readable code or a non-volatile computer readable storage medium carrying computer readable code, when the computer readable code is run in the processor of the electronic device, the processor in the electronic device executes the above method.

[0207] Figure 7A schematic diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 7 The electronic device 1900 includes a processing component 1922, further including one or more processors, and a memory resource represented by a memory 1932, for storing instructions executable by the processing component 1922, such as an application program. The application program stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above method.

[0208] The electronic device 1900 can further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0209] In an exemplary embodiment, a non-transitory computer readable storage medium, such as the memory 1932 including computer program instructions, is also provided, which can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.

[0210] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0211] Computer readable storage media can be tangible storage media which can retain and store instructions for use by an instruction execution device. Computer readable storage media can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer readable storage media include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0212] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0213] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0214] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0215] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0216] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0217] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0218] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative, and not restrictive, of the disclosed embodiments. Many modifications and variations of the described embodiments are possible, and all such modifications and variations are intended to be within the scope of the described embodiments. The description used herein is intended to best explain the principles of the various embodiments, the practical application, and the best mode of using the present disclosure, and to enable others skilled in the art to understand the disclosure, various embodiments, and the application, devices, and apparatuses.

Claims

1. A multi-time point image analysis method, characterized in that: The method comprises: performing feature extraction on the first image and the second image respectively to obtain a first image feature of the first image and a second image feature of the second image, wherein the first image and the second image are images acquired by acquiring images of the target object at different acquisition times, and the first image acquisition time of the first image is before the second image acquisition time of the second image; determining a progress characterization feature based on the first image feature and the second image feature, wherein the progress characterization feature is used to characterize a change in the target object from the time when the first image is acquired to the time when the second image is acquired; The progress representation feature is used as a priori feature, and is fused with the second image feature to obtain a target feature; An object category of the target object is determined according to the target feature.

2. The method according to claim 1, characterized in that The extracting features from the first image and the second image respectively to obtain first image features of the first image and second image features of the second image includes: Inputting the first image and the second image into a visual backbone network to extract spatial features and texture features respectively, to obtain corresponding first input feature maps and second input feature maps; The first input feature map and the second input feature map are respectively input into a mapper to obtain a first image feature and a second image feature related to the prediction task of the object category.

3. The method according to claim 2, characterized in that The step of inputting the first input feature map and the second input feature map into a mapper to obtain first image features and second image features related to the object category prediction task includes: Inputting the first input feature map and the second input feature map into a mapper for dimensionality reduction projection respectively to obtain a first eigenvector and a second eigenvector; Determine a first spatiotemporal matrix corresponding to the first input feature map, and a second spatiotemporal matrix corresponding to the second input feature map; The first image feature is determined according to the first eigenvector and the first spatiotemporal matrix, and the second image feature is determined according to the second eigenvector and the second spatiotemporal matrix.

4. The method according to claim 1, wherein The determining of the progress representation feature according to the first image feature and the second image feature includes: Determining a first reconstructed feature and a second reconstructed feature by performing an orthogonal attention calculation on the first image feature and the second image feature; A progression characterizing feature is determined based on a difference between the first reconstructed feature and the second reconstructed feature.

5. The method according to claim 4, characterized in that The determining of the first reconstruction feature and the second reconstruction feature by performing orthogonal attention calculation on the first image feature and the second image feature includes: Separating the first image feature and the second image feature to obtain a first feature token set and a second feature token set; Performing orthogonal attention calculation based on the first feature token set and the second feature token set to obtain a third feature token set and a fourth feature token set; Feature reconstruction is performed on the third feature token set and the fourth feature token set respectively to obtain a first reconstructed feature and a second reconstructed feature.

6. The method according to claim 4, characterized in that The orthogonal attention calculation is performed based on the first feature token set and the second feature token set to obtain a third feature token set and a fourth feature token set, including: Performing orthogonal attention calculation on the first feature token set as a query and the second feature token set as a key and a value to obtain a third feature token set; The second feature token set is used as a query, and the first feature token set is used as a key and a value to perform orthogonal attention calculation to obtain a fourth feature token set.

7. The method according to claim 1, characterized in that The step of taking the progress representation feature as a priori feature and fusing it with the second image feature to obtain a target feature includes: Using the progress representation feature as a query and the second image feature as a key and a value to perform a cross attention calculation to obtain a feature to be fused; Channel splicing is performed on the feature to be fused and the second image feature to obtain the target feature.

8. A multi-time point image analysis device, characterized in that: The device comprises: a feature extraction module, configured to perform feature extraction on the first image and the second image, respectively, to obtain a first image feature of the first image and a second image feature of the second image, wherein the first image and the second image are images acquired by acquiring images of the target object at different acquisition times, and the first image acquisition time of the first image is before the second image acquisition time of the second image; a feature determination module, configured to determine a progress characterization feature based on the first image feature and the second image feature, wherein the progress characterization feature is used to characterize a change in a target object from a time when the first image is captured to a time when the second image is captured; a feature fusion module, configured to use the progress representation feature as a priori feature and perform feature fusion with the second image feature to obtain a target feature; A category determination module is used to determine the object category of the target object according to the target feature.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method according to any one of claims 1 to 7 when executing the instructions stored in the memory.

10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Nursing parameter adaptation evaluation system and method for alimentary canal state image recognition

    CN121726062A