Image sequence processing method and device, electronic equipment and storage medium

By performing multi-label classification processing on the description text of medical imaging sequences, and using the adjusted pre-trained model to perform target prediction tasks, the problem of low single-label classification efficiency in the existing technology is solved, and more comprehensive and accurate medical imaging sequence classification is achieved, and the utilization efficiency of image data is improved.

CN120198743AInactive Publication Date: 2025-06-24MANTEIA TECH CO LTD

Patent Information

Application Number
CN202510676725.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The classification of traditional Chinese medicine imaging sequences in the prior art is mainly single-label classification and artificial classification, resulting in low classification efficiency and only interpreting medical imaging sequences from the dimension of a single label, inadequate interpretation and inaccurate classification, resulting in the problem of low utilization rate of medical imaging.

Method used

By obtaining the description text of the target image sequence, input it into a pre-trained model adjusted according to the description characteristics of the reference image sequence, the target prediction task is performed to determine multiple tags of the target image sequence, and the target image sequence is classified from multiple dimensions.

Benefits of technology

Multi-dimensional classification of medical imaging sequences is realized, comprehensive interpretation of image sequences and classification accuracy are improved, the utilization efficiency of medical imaging is significantly improved, and the time and cost of manual classification are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198743A_ABST
    Figure CN120198743A_ABST
Patent Text Reader

Abstract

The invention discloses an image sequence processing method and device, electronic equipment and a storage medium, and relates to the field of medical science and technology and the field of artificial intelligence. The method comprises the following steps: acquiring a description text of a target image sequence; inputting the description text into the target model; a target prediction task is executed on the description text through the target model, and the target prediction task is used for determining N labels of the target image sequence according to known words in the description text and sentences in the description text; and taking the N tags as a classification result of the target image sequence. According to the method and the device, the problems of low classification efficiency of the medical image sequence and low medical image utilization rate caused by the fact that the classification of the medical image sequence is mainly single-label classification and manual classification in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical technology. Specifically, it relates to a method, apparatus, electronic device, and storage medium for processing image sequences. Background Art

[0002] With the development of medical imaging technology and the improvement of medical standards, the accumulation of medical imaging data has increased rapidly, providing rich image information for clinical diagnosis and becoming an indispensable part of clinical diagnosis and treatment. Medical imaging data is huge and diverse in type. When doctors receive patients and conduct diagnosis and treatment, they need to spend a lot of time reading and analyzing medical imaging information, which causes certain pressure on doctors' daily work. In clinical applications, being able to quickly and accurately classify image sequences helps to conduct a more in-depth analysis of the image sequences and improve the utilization efficiency of the image sequences.

[0003] Currently, the classification of medical image sequences in the prior art is mainly single-label classification and manual classification, resulting in low classification efficiency of medical image sequences. Especially when the number of medical image sequences is large, not only a large amount of time is required for classification, but also the medical image sequences are interpreted only from the dimension of a single label, and the interpretation of medical image sequences is not sufficient and the classification is not completely accurate, resulting in the problem of low utilization rate of medical images.

[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] This application provides a method, apparatus, electronic device, and storage medium for processing image sequences, so as to at least solve the technical problem that the classification of medical image sequences in the prior art is mainly single-label classification and manual classification, resulting in low classification efficiency of medical image sequences. Especially when the number of medical image sequences is large, not only a large amount of time is required for classification, but also the medical image sequences are interpreted only from the dimension of a single label, and the interpretation of medical images is not sufficient and the classification is inaccurate, resulting in low utilization rate of medical images.

[0006] According to one aspect of the present application, a method for processing an image sequence is provided, including: obtaining a description text of a target image sequence, where the target image sequence includes at least one medical image; inputting the description text into a target model, where the target model is a model obtained by adjusting a pre-trained model according to the description features of a reference image sequence, and the pre-trained model is used to predict unknown words in a sentence based on known words in the sentence and to predict the association relationship between different sentences; performing a target prediction task on the description text through the target model, where the target prediction task is used to determine N labels of the target image sequence based on the known words in the description text and the sentences in the description text, where N is an integer greater than or equal to 1, and the N labels are used to classify the target image sequence from N dimensions; taking the N labels as the classification result of the target image sequence.

[0007] Optionally, inputting the description text into the target model includes: preprocessing the description text, where the preprocessing includes at least one of the following processing operations: a symbol processing operation for deleting the first type of symbols in the description text and retaining the second type of symbols in the description text, where the first type of symbols are symbols irrelevant to the image features of the medical image; the second type of symbols are symbols relevant to the image features of the medical image; a word segmentation processing operation for dividing the sentences in the description text into multiple words; a stop word filtering operation for deleting the stop words in the description text; a format conversion operation for converting the description text into a target text format; inputting the description text after the preprocessing is completed into the target model.

[0008] Optionally, the training process of the pre-trained model includes the following steps: obtaining training text; constructing a first training task and a second training task based on the training text, where the first training task is used to train the neural network's ability to predict the full text content of the training text based on partial sentences of the training text, and the second training task is used to train the neural network's ability to predict whether there is a context association relationship between any two different sentences in the training text; performing the first training task and the second training task on a preset neural network multiple times to obtain the pre-trained model.

[0009] Optionally, constructing the first training task based on the training text includes: randomly selecting a target proportion of words in the training text as target words; replacing the target words in the training text with a preset character to obtain a first text, where the preset character is a character without actual semantics; constructing the first training task based on the target words and the first text, where the first training task is used to train the neural network's ability to predict the target words based on the non-target words in the first text.

[0010] Optionally, construct a second training task based on the training text, including: randomly selecting at least two sentences from the training text to form a sentence combination; setting a relevance label for the sentence combination, where the relevance label is used to characterize whether there is a context correlation relationship between at least two sentences in the sentence combination; constructing a second training task based on the sentence combination and the corresponding relevance label of the sentence combination, where the second training task is used to train the neural network to predict the ability of whether there is a context correlation relationship between different sentences in the sentence combination.

[0011] Optionally, the training steps of the target model include: obtaining a reference image sequence with known corresponding labels and the description information of the reference image sequence; performing length processing on the description information of the reference image sequence, where the length processing is used to adjust the length of the description information of the reference image sequence to a preset length; adjusting the pre-trained model according to the description information after the length processing is completed and the label corresponding to the reference image sequence until the pre-trained model enters the target state, where the pre-trained model in the target state has a label prediction accuracy for the reference image sequence based on the description information after the length processing greater than a preset threshold; using the pre-trained model that enters the target state as the target model.

[0012] Optionally, performing length processing on the description information of the reference image sequence includes: in the case where it is detected that the length of the description information is less than the preset length, expanding the length of the description information to the preset length by filling preset characters at the end of the description information; in the case where it is detected that the length of the description information is greater than the preset length, reducing the length of the description information to the preset length by deleting characters one by one in reverse order starting from the end of the description information.

[0013] Optionally, adjusting the pre-trained model according to the description information after the length processing is completed and the label corresponding to the reference image sequence until the pre-trained model enters the target state includes: performing a first transfer learning operation and a second transfer learning operation on the pre-trained model according to the description information after the length processing is completed, where the first transfer learning operation is used to train the pre-trained model to predict the full text content of the description information based on partial sentences of the description information, and the second transfer learning operation is used to train the pre-trained model to predict the ability of whether there is a context correlation relationship between any two different sentences in the description information; performing gradient training on the pre-trained model that has completed the first transfer learning operation and the second transfer learning operation according to the label corresponding to the reference image sequence until the pre-trained model enters the target state, where the gradient training is used to calculate the label prediction accuracy of the pre-trained model for the reference image sequence and adjust the model parameters of the pre-trained model according to the label prediction accuracy.

[0014] According to another aspect of the present application, there is also provided a processing device for an image sequence, which includes: an acquisition unit for acquiring a description text of a target image sequence, where the target image sequence includes at least one medical image; an input unit for inputting the description text into a target model, where the target model is a model obtained by adjusting a pre-trained model according to the description features of a reference image sequence, and the pre-trained model is used to predict unknown words in a sentence based on known words in the sentence and to predict the association relationship between different sentences; a first processing unit for performing a target prediction task on the description text through the target model, where the target prediction task is used to determine N labels of the target image sequence based on the known words in the description text and the sentences in the description text, where N is an integer greater than or equal to 1, and the N labels are used to classify the target image sequence from N dimensions; and a second processing unit for using the N labels as the classification result of the target image sequence.

[0015] According to another aspect of the present application, there is also provided a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and when the computer program runs, it causes the device where the computer-readable storage medium is located to execute the above-mentioned processing method for the image sequence.

[0016] According to another aspect of the present application, there is also provided an electronic device, where the electronic device includes one or more processors and a memory, and the memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, it causes the one or more processors to execute the above-mentioned processing method for the image sequence.

[0017] As can be seen from the above, different from the traditional classification method which usually can only assign one label to a medical image sequence, the present application can classify the target image sequence from multiple dimensions (N labels) by performing the target prediction task, which includes but is not limited to image type, acquisition direction, whether enhanced, etc. Multi-label classification makes the interpretation of medical images more comprehensive and accurate, helps clinicians and researchers to more accurately locate and analyze medical image data, and improves the utilization value of medical images.

[0018] In addition, by using the target model, the present application realizes the automatic classification of the description text of the target image sequence. Compared with manual classification, automatic classification greatly saves time and reduces labor costs. Especially when the number of medical image sequences is large, it can significantly improve the classification efficiency. At the same time, automatic classification avoids the subjective errors that may be introduced by manual operations and improves the objectivity and consistency of the classification results.

[0019] It should be noted that the target model of this application is a model obtained by fine-tuning a pre-trained model to adapt to the classification task of medical image sequences. The pre-trained model has rich language understanding and context awareness capabilities. Through fine-tuning, it can quickly adapt to specific medical image description features, avoiding the long process of training a model from scratch, thus ensuring the efficiency and practicality of the method. Adopting the fine-tuning strategy of the BERT model makes this technical solution have good adaptability and scalability. This means that with the emergence of new types of image sequences or new classification schemes, the model can continuously learn and adapt through further fine-tuning, continuously optimize the classification performance, and reduce the costs of technology update and maintenance.

[0020] In summary, the technical solution of this application not only solves the problems of low efficiency and single classification in the prior art, but also improves the efficiency and accuracy of medical image data management by introducing multi-label classification and an automated processing mechanism, promotes the efficient utilization of medical images in clinical and research, and thus produces significant beneficial effects in the field of medical imaging. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings described herein are used to provide a further understanding of this application and constitute a part of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0022] Figure 1 is a flowchart of a method for processing an image sequence according to an embodiment of this application;

[0023] Figure 2 is a schematic diagram of a device for processing an image sequence according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.

[0025] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] It should also be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) collected in the present application are information and data authorized by the user or fully authorized by all parties. And the processing of the relevant data, such as collection, storage, use, processing, transmission, provision, disclosure and application, etc., all comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set between the present system and relevant users or institutions. Before obtaining relevant information, a request for obtaining needs to be sent to the aforementioned users or institutions through the interface, and after receiving the consent information fed back by the aforementioned users or institutions, the relevant information is obtained.

[0027] According to an embodiment of the present application, an embodiment of a method for processing an image sequence is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0028] Figure 1 is a flowchart of a method for processing an image sequence according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:

[0029] Step S101, obtaining a description text of a target image sequence.

[0030] In step S101, the target image sequence includes at least one medical image.

[0031] Optionally, an image classification system may serve as the execution subject of the method for processing an image sequence in the embodiments of the present application. Among them, the image classification system may be a software system or an embedded system combining software and hardware. Those skilled in the art should be aware that the execution subject of the method for processing the image sequence in the present application may also be other forms of systems, devices, or apparatuses. The present application does not make a special limitation on the specific manifestation form of the execution subject. For the convenience of description, the image classification system is used as the execution subject for the description of the solution hereinafter.

[0032] Optionally, the target image sequence refers to a data set composed of a series of consecutive medical images in the field of medical imaging. In magnetic resonance imaging (MRI), an image sequence may include the results of multiple scans, which may be performed at different time points, different sectional directions, or different imaging modes (such as T1-weighted, T2-weighted, diffusion-weighted, etc.). Therefore, an image sequence may include multiple medical images, and each image corresponds to specific scan parameters and conditions. The description text refers to the description information about the image sequence stored in a DICOM (Digital Imaging and Communications in Medicine) file, usually found in the DICOM tag (0008,103E). These descriptions contain the key features of the image sequence, such as the sequence type, parameter settings (such as TR, TE), scan direction (such as axial, coronal, sagittal), whether contrast enhancement has been performed, the specific name of the sequence, etc. The description text is an important basis for the model in the present application to perform classification prediction.

[0033] Optionally, a medical image refers to image data obtained through various medical imaging technologies (such as MRI, CT, X-ray, ultrasound, etc.) for diagnosis and research. In the context of the present application, it mainly refers to the images obtained through the MRI technology. These images are generated by a magnetic resonance imaging device under specific parameter settings and can provide detailed information about the human tissue structure, function, or pathological state.

[0034] Step S102: Input the description text into the target model.

[0035] In step S102, the target model is a model obtained by adjusting a pre-trained model according to the description features of the reference image sequence. The pre-trained model is used to predict unknown words in a sentence based on known words in the sentence and to predict the association relationships between different sentences.

[0036] Optionally, the pre-trained model is a language model trained on a large-scale unlabeled text dataset. The pre-trained model has learned rich language representations and context understanding capabilities through training tasks. In the context of this application, the pre-trained model can be a model obtained by pre-training on clinical text data and is capable of understanding and processing terms and descriptions specific to the medical field.

[0037] Optionally, the target model is customized from the pre-trained model through a fine-tuning process for solving specific tasks. In this application, the target model is obtained by fine-tuning the pre-trained model using a dataset annotated with the description features of the image sequence to make it more suitable for the multi-label classification task of medical image sequences. During the fine-tuning process, the parameters learned by the model will be adjusted according to the features of the medical image description text so that it can more accurately predict the multi-dimensional labels of the image sequence. Additionally, the description features refer to the Series Description attribute extracted from the DICOM file, which contains key information such as the type, parameters, and orientation of the image sequence. These description features are used as model inputs to train the model to recognize and predict the multi-label classification of medical image sequences.

[0038] It should be noted that through fine-tuning, the pre-trained model can be "customized" to adapt to specific downstream tasks, such as the medical image sequence classification in this application. During the fine-tuning process, the learning rate of the model is usually low to avoid overwriting the general features learned on the large-scale dataset. The model is trained according to the annotated data of the medical image sequence description features, and by adjusting the weights and biases of the model, it can predict the correct multi-label classification results based on the input description text.

[0039] Step S103, perform a target prediction task on the description text through the target model.

[0040] In step S103, the target prediction task is used to determine N labels of the target image sequence based on the known words and sentences in the description text, where N is an integer greater than or equal to 1, and the N labels are used to classify the target image sequence from N dimensions.

[0041] Optionally, the target prediction task refers to using the trained target model to process the input description text to predict multiple relevant labels of the image sequence. These labels can include the type of the image (such as T1-weighted, T2-weighted, diffusion-weighted, etc.), the orientation (such as transverse plane, coronal plane, sagittal plane, etc.), whether contrast enhancement has been performed, and other possible feature descriptions. The key to the target prediction task is that the model can identify patterns related to the medical image sequence features based on the keywords and sentence structures in the description text, thereby performing accurate multi-dimensional classification.

[0042] In the image sequence description text, known words refer to those words that can directly or indirectly indicate image features, such as "T1", "Coronal", "+C" (contrast enhancement), etc. These words are the basis for the model to make predictions. A sentence, on the other hand, is an information unit in the description text. It may consist of one or more known words and represents a complete description of the image sequence. When performing the target prediction task, the model will consider the positions and context relationships of these keyword vocabularies in the sentence to improve the accuracy of the prediction.

[0043] It should be noted that N tags refer to multiple classification tags that the model can predict for the target image sequence. N is an integer greater than or equal to 1. This means that an image sequence may be assigned multiple tags, and each tag corresponds to a different classification dimension. For example, an image sequence may be marked as "T1 weighted", "coronal plane", and "contrast enhancement" at the same time. This multi-label classification method can provide more comprehensive and detailed image information compared to traditional single-label classification, and has significant value for clinical diagnosis and research.

[0044] Step S104: Use the N tags as the classification result of the target image sequence.

[0045] In step S104, classification is the process of classifying the target image sequence according to the N tags. In this application, classification is achieved through the analysis and prediction of the target model based on the information in the description text. Each image sequence will be assigned one or more tags, and these tags constitute a multi-dimensional feature description of the image sequence, which helps to accurately distinguish and manage different types of medical image data.

[0046] In summary, the technical solution of this application not only solves the problems of low efficiency and single classification in the prior art, but also improves the efficiency and accuracy of medical image data management by introducing multi-label classification and an automated processing mechanism, promotes the efficient utilization of medical images in clinical and research, and thus produces significant beneficial effects in the field of medical imaging.

[0047] In an alternative embodiment, inputting the description text into the target model includes: preprocessing the description text, where the preprocessing includes at least one of the following processing operations:

[0048] Symbol processing operation: used to delete the first type of symbols in the description text and retain the second type of symbols in the description text. The first type of symbols are symbols that are irrelevant to the image features describing the medical image; the second type of symbols are symbols that are relevant to the image features describing the medical image;

[0049] Word segmentation processing operation: used to divide the sentences in the description text into multiple words;

[0050] A stop word filtering operation for deleting stop words in the description text;

[0051] A format conversion operation for converting the description text into a target text format;

[0052] Input the description text after preprocessing into the target model.

[0053] Optionally, the first type of symbols refers to symbols that are irrelevant to the image features describing the medical image, such as certain special characters (e.g., punctuation marks), numbers, or units, which may not carry information about the image characteristics in the text. During preprocessing, such symbols are deleted to reduce noise and avoid interfering with the model's prediction. On the contrary, the second type of symbols refers to symbols that are directly related to the image features describing the medical image, such as "+C" (indicating contrast enhancement), etc. These symbols are crucial for understanding the characteristics of the image sequence. Therefore, during preprocessing, these symbols are retained to ensure that the model can receive all key information.

[0054] Optionally, word segmentation processing refers to decomposing the sentences in the description text into individual words or lexical units. This is crucial for word-based deep learning models because the model needs to process words as units. By performing word segmentation, the independence of each word and the coherence of the context can be ensured, providing clear input for the model's training and prediction.

[0055] In addition, stop words refer to words that frequently appear in the text but do not carry specific meanings or contribute to the model's prediction, such as "de", "shi", "zai", etc. During the preprocessing stage, these words are filtered out to reduce the dimensionality of the input data, avoid the model investing unnecessary computational resources in these words during training, and thus improve the prediction efficiency and accuracy of the model. Format conversion refers to converting the description text into a specific format required by the target model. For example, by adding special tokens (such as [CLS] and [SEP]) to delimit the start and end of the sentence, and standardizing the text, such as unifying the case, to ensure the consistency of the input.

[0056] Through the above preprocessing steps, the description text is processed into a cleaner and more structured form for use by the target model. The purpose of doing this is to improve the model's ability to understand the text, reduce prediction errors caused by non-critical information or format inconsistencies, ensure that the model can extract information directly related to the medical image features from the input text, and then perform accurate multi-label classification. With the preprocessed text as the input, the model can more effectively perform feature extraction and classification prediction, which plays a key role in improving the automation degree and accuracy of medical image sequence classification.

[0057] For example, when extracting the Dicom Series Description (corresponding to the above-described descriptive text), first, the Series Description attribute value of the MRI sequence is extracted from each DICOM file. The descriptive text describes the modal characteristics of the sequence. The selected SD entries are classified and marked as one of the predefined 5 categories, such as T1-weighted imaging sequence, T1-weighted imaging enhanced sequence, T2-weighted imaging sequence, Diffusion-weighted imaging sequence. For those SD sequences that cannot be directly classified, they will be marked as unknown. Among them, T1, T1C, T2, and DWI images refer to different types of magnetic resonance imaging (MRI) scan sequences, and different types of magnetic resonance imaging will provide different information. Transverse, Coronal, Sagittal, and Oax refer to different cross-sectional views in medical images. Transverse (cross-section) refers to the horizontal cross-section, Coronal (coronal plane) refers to the anterior-posterior cross-section, Sagittal (sagittal plane) divides the body into left and right sides, and Oax (Oblique Axial Plane) refers to the oblique plane, as shown in Table 1.

[0058] Table 1

[0059]

[0060] Optionally, Table 2 is an example of the tags of the Dicom sequence.

[0061] Table 2

[0062]

[0063] Optionally, for the Dicom Series Description extracted from the Dicom sequence, data preprocessing is performed as follows:

[0064] Removing special symbols: Taking the descriptive text of the DICOM sequence as input, remove the irrelevant special symbols in the text, such as underscores and colons. Retain specific symbols, such as +C, which represents contrast enhancement in the Dicom sequence description.

[0065] Word segmentation: Segment the cleaned text through a word segmentation algorithm for subsequent processing.

[0066] Stop word filtering: Remove the stop words in the text to improve the quality and relevance of the text data.

[0067] Standardize the text by converting all text to a unified case form, for example, convert all to lowercase, to reduce data inconsistency.

[0068] Among them, Table 3 is an example of the descriptive text after preprocessing.

[0069] Table 3

[0070]

[0071] In an alternative embodiment, the training process of the pre-trained model includes the following steps: Obtain training text; construct a first training task and a second training task based on the training text, where the first training task is used to train the neural network's ability to predict the full text content of the training text based on partial sentences of the training text, and the second training task is used to train the neural network's ability to predict whether there is a contextual association relationship between any two different sentences in the training text; perform the first training task and the second training task on the preset neural network multiple times to obtain the pre-trained model.

[0072] Optionally, in the pre-training stage, the model requires a large amount of unlabeled text data as training materials. These text data can come from various sources, such as books, news articles, clinical text data, and so on. The extensiveness and diversity of the text data are crucial for the model to learn general language representations.

[0073] Optionally, the first training task is to randomly mask some words in the training text and then train the model to predict these masked words. The specific steps are as follows:

[0074] Randomly select a certain proportion (for example, 15%) of words from the training text for masking and replace the words with a special marker [MASK]; retain the remaining words as the input of the model; the goal of the model is to predict the original content of the masked words based on the context. This process trains the model's ability to understand the relationships between words within the text and the context, which is very useful for capturing key features in medical image descriptions.

[0075] Optionally, the second training task aims to train the model's ability to predict whether there is a contextual association relationship between two sentences. The specific steps are as follows:

[0076] Randomly select two sentences from the training text. In one case, the two sentences are consecutive, i.e., they appear immediately one after another in the original text; in another case, the second sentence is drawn from other random positions in the training text and has no direct context relationship with the first sentence. The model needs to determine whether the second sentence is the next sentence of the first sentence. This task trains the model to understand the logical relationships and text coherence between sentences, which is crucial for processing and understanding the composite information in medical image descriptions.

[0077] It should be noted that the training of the pre-trained model is an iterative process, and the first training task and the second training task need to be executed on the neural network multiple times so that the model can learn rich language structures and semantic representations from a large amount of data. In this process, the model will continuously optimize its parameters to improve the prediction performance on the two tasks. Finally, the pre-trained model can deeply understand the words and sentences in the text, providing a powerful language processing ability for subsequent fine-tuning of specific tasks.

[0078] Through the above training process, the pre-trained model can learn the ability to extract features from the text, predict words, and understand the associations between sentences, and these abilities will play an important role in the multi-label classification task of medical image sequence descriptions.

[0079] In an optional embodiment, constructing the first training task based on the training text includes: randomly selecting a target proportion of words in the training text as target words; replacing the target words in the training text with a preset character to obtain a first text, where the preset character is a character without actual semantics; constructing the first training task based on the target words and the first text, where the first training task is used to train the neural network's ability to predict the target words based on the non-target words in the first text.

[0080] Optionally, selecting target words is the first step in constructing the first training task. Specifically, a certain percentage (e.g., 15%) of words are randomly selected from the training text as target words, and these words will be masked, while the goal of the model is to predict these masked words based on the context information. After the target words are selected, the processing system of the image sequence can replace these words with a preset masking symbol [MASK]. The text after replacement is called the first text, which contains some masked words and complete context information. For example, the original text might be "T1-weighted MRI is particularly useful for [noun] diseases of the [noun] system", where "brain" and "nervous" are randomly selected as target words. The first text after replacement then becomes "T1-weighted MRI is particularly useful for [MASK] diseases of the [MASK] system".

[0081] Optionally, the first training task is constructed based on the masked first text and the original target words. The model receives the first text as input, and its goal is to predict the masked words, i.e., the target words. The model can learn the relationship between words and context through its internal self-attention mechanism and feed-forward neural network to improve the prediction accuracy. For example, in the above example, the model will try to predict the word that should be filled in at the [MASK] position based on "T1-weighted MRI is particularly useful for" and "diseases of the system".

[0082] After constructing the first training task, the neural network will be trained to predict the ability to obtain target words based on the non-target words (i.e., unmasked words) in the first text. The model iteratively optimizes its parameters to minimize the difference between the prediction result and the actual target words, thereby improving its performance on the first training task. As the training progresses, the model will gradually learn to predict words given the context, and this ability is crucial for processing medical image description texts, identifying, and classifying image attributes.

[0083] In an alternative embodiment, a second training task is constructed based on the training text, including: randomly selecting at least two sentences from the training text to form a sentence combination; setting a relevance label for the sentence combination, where the relevance label is used to characterize whether there is a context relevance relationship between at least two sentences in the sentence combination; constructing a second training task based on the sentence combination and the corresponding relevance label of the sentence combination, where the second training task is used to train the neural network's ability to predict whether there is a context relevance relationship between different sentences in the sentence combination.

[0084] Optionally, to construct the second training task, first, in the training dataset (such as the training text), two sentences are randomly selected to form a sentence combination. These can be two consecutive sentences from the same paragraph, or one sentence from the current paragraph and another randomly selected sentence that are not consecutive in the dataset.

[0085] Subsequently, a relevance label is set for each sentence combination to characterize whether there is a context relevance relationship between the two sentences in the sentence combination. If the two sentences are indeed consecutive in the original dataset, then the model should be able to understand the logical relationship between them, and at this time, the relevance label is "next sentence" (positive example).

[0086] Conversely, if the two sentences are not consecutive in the dataset, then there may not be a direct context relevance between them, and at this time, the relevance label is "not the next sentence" (negative example). For example, for the sentence combination "T1-weighted MRI is particularly useful for brain diseases" and "T2-weighted images are often used for identifying strokes", if these two sentences appear consecutively in the original data, they will be marked as having a context relevance relationship; if they come from different parts of the dataset, they will be marked as not having a context relevance relationship.

[0087] Optionally, constructing the second training task is a process of combining the above sentence combinations with relevance labels. The model will receive the sentence combinations as input and attempt to predict the relevance labels based on its internal structure and semantic information. That is, the model needs to determine whether the given two sentences appear consecutively in the original text to understand the logical relationship and context coherence between the sentences. After constructing the second training task, the neural network will be trained to predict the ability to determine whether there is a context relevance relationship between different sentences in the sentence combination. This training process gradually improves the model's understanding and prediction ability of sentence relevance by continuously iterating and optimizing the model parameters to minimize the difference between the prediction result and the actual relevance label. This ability is particularly important for processing medical image description texts because understanding the logical relationship between sentences helps capture the implicit medical features and information in the description, providing a more comprehensive semantic understanding for subsequent multi-label classification tasks of image sequences.

[0088] For example, the following example sentences are used to illustrate the relevant content of the first training task and the second training task:

[0089] Sentence 1: The T1-weighted imaging sequence is mainly used to observe the anatomical structure of the brain;

[0090] Sentence 2: The scanning direction of this sequence is axial;

[0091] Sentence 3: The patient needs to fast for 6 hours before the examination to avoid gastrointestinal interference.

[0092] Regarding the construction of the first training task, first, sentences or paragraphs are extracted from a large amount of text data as the input text. Secondly, a certain proportion (e.g., 15%) of the words in the input text are randomly selected for masking. The masking method is to replace the original word with a special marker [MASK]. Example: "The T1-weighted imaging sequence is mainly used to observe the [MASK] structure of the brain;" where "anatomical" is masked. The first training task is to predict the masked word "anatomical" based on the context ("The T1-weighted imaging sequence is mainly used to observe the" and "structure;").

[0093] In addition, for each masked word, its original form is retained as the target for the model to predict. The processed input text is input into the neural network. The neural network will make predictions for each masked position, outputting a probability distribution indicating the probability of all possible words appearing at that position. Calculate the loss between the model prediction result and the actual masked word. After completing the pre-training, the pre-trained model can predict words based on the context, which is very useful for tasks such as text filling and question answering in medical image descriptions.

[0094] Regarding the construction of the second training task, the first step is to construct sentence pairs. For example, select sentence one as Sentence A. Then select sentence two and sentence three respectively as Sentence B to construct two sentence pairs.

[0095] Secondly, generate labels: For sentence one and sentence two, both describe the same image sequence (T1-weighted imaging), and they are developed from two perspectives of function and scanning direction respectively, with a close logical connection semantically. Therefore, it is labeled as "is the next sentence" (label is 1). For sentence one and sentence three, although both belong to medical-related content, the former focuses on the use of the image sequence itself, while the latter discusses the preparations before the examination. The topics and semantic spans of the two are relatively large, lacking context coherence. Therefore, it is labeled as "is not the next sentence" (label is 0).

[0096] Then, for each sentence pair, it is necessary to mark according to the requirements of the neural network. For example, use [CLS] to mark the start and [SEP] to separate Sentence A and Sentence B.

[0097] Input example: [CLS] The T1-weighted imaging sequence is mainly used to observe the anatomical structure of the brain [SEP] The scanning direction of this sequence is axial [SEP] (label is 1);

[0098] [CLS] The T1-weighted imaging sequence is mainly used to observe the anatomical structure of the brain [SEP] The patient needs to fast for 6 hours before the examination to avoid gastrointestinal interference [SEP] (label is 0).

[0099] Input the constructed sentence pairs into the neural network. The neural network will output a probability value indicating whether the two sentences of the sentence pair are consecutive. If the neural network believes that the sentence pair is consecutive, it will output a probability close to 1; if not, it will output a probability close to 0. Subsequently, calculate the loss between the probability predicted by the neural network and the actual label. Adjust the parameters of the neural network through gradient descent so that the neural network can better predict the relationship between sentences.

[0100] In an alternative embodiment, the training steps of the target model include: obtaining a reference image sequence with known corresponding labels and the description information of the reference image sequence; performing length processing on the description information of the reference image sequence, where the length processing is used to adjust the length of the description information of the reference image sequence to a preset length; adjusting the pre-trained model according to the description information after the length processing and the label corresponding to the reference image sequence until the pre-trained model enters the target state, where the pre-trained model in the target state has a label prediction accuracy for the reference image sequence based on the description information after the length processing greater than a preset threshold; using the pre-trained model that enters the target state as the target model.

[0101] Optionally, to train the target model, first, it is necessary to collect a reference image sequence with known multi-labels and its description information. These labels usually include the modality features of the images (such as T1, T2, DWI, etc.), orientation information (such as Transverse, Coronal, Sagittal, etc.), and other relevant information (such as Contrast, etc.). The description information of the reference image sequence is extracted from the Series Description attribute of the DICOM file and contains metadata such as the type, parameters, and orientation of the sequence.

[0102] Secondly, since the pre-trained model has a fixed sequence length limit when processing the input, it is necessary to process the length of the description information of the reference image sequence. Specifically, if the length of the description information exceeds the preset length (for example, 15 tokens), it needs to be truncated, only keeping the first 15 tokens; if the length of the description information is shorter than the preset length, it needs to be padded by adding special characters (such as [PAD]) until it reaches the preset length. This process of standardizing the length is to ensure that all input data has the same processing format for the model, thus avoiding computational complexity and efficiency issues caused by inconsistent sequence lengths.

[0103] Optionally, adjust the pre-trained model to make it more accurate in predicting the multi-labels of the image sequence. The adjustment steps include:

[0104] Step 1, input the processed description information into the model, and the model will output a series of prediction probabilities corresponding to each possible label.

[0105] Step 2, calculate the gap (loss function) between the model prediction result and the actual label, and use an optimization algorithm (such as gradient descent) to update the parameters of the model to minimize this gap.

[0106] Step 3, repeat the process from Step 1 to Step 2, continuously iterating the training of the model until the label prediction accuracy of the model for the description information exceeds the preset threshold. At this time, the model is considered to have entered the "target state".

[0107] Once the pre-trained model reaches the target state through adjustment, that is, the multi-label prediction accuracy for the description information of the reference image sequence exceeds the preset threshold, this indicates that the model has been trained on the current dataset and can efficiently perform multi-label classification of the image sequence. At this time, the model is saved as the "target model" for subsequent inference and classification tasks. The target model has the ability to process medical image description texts and identify multi-labels, and can automatically and accurately classify new image sequences, reducing manual intervention and improving the efficiency and accuracy of medical image management.

[0108] In an alternative embodiment, length processing is performed on the description information of the reference image sequence, including: when it is detected that the length of the description information is less than the preset length, the length of the description information is amplified to the preset length by padding preset characters at the end of the description information; when it is detected that the length of the description information is greater than the preset length, the length of the description information is reduced to the preset length by deleting characters one by one in the reverse order from the end of the description information.

[0109] Optionally, when the length of the description information (usually calculated in tokens, i.e., word or sub-word units) is less than the preset maximum length of the pre-trained model, it is necessary to pad preset characters at the end of the description information to reach the preset length. These preset characters usually refer to special tokens, such as [PAD], which represent padding symbols in the neural network and have no actual semantic content. In this way, all input data will have the same length, facilitating unified processing by the model. For example, if the preset length is 15 and the description information has only 8 tokens, then 7 [PAD] will be padded at the end of the description information to reach a length of 15 tokens.

[0110] Optionally, when the length of the description information exceeds the preset length, it needs to be truncated to ensure that the input length does not exceed the model's limit. Truncation usually starts from the end of the description information and deletes characters one by one in the reverse order until the preset length is reached. Since DICOM sequence descriptions are usually short and the key information often lies in the front part of the description, truncating from the end can retain the main content of the description as much as possible while reducing unnecessary length. For example, if the description information has 20 tokens and the preset length is 15, then the last 5 tokens will be deleted, and the first 15 tokens will be retained as the input.

[0111] It should be noted that standardizing the length ensures that all input data has a unified processing format for the model, avoiding the complexity and efficiency issues of the model when processing data of different lengths. Secondly, fixed-length inputs help improve computational efficiency, especially in a parallel computing environment where the model can utilize computing resources more effectively. Additionally, by restricting the input length, the model is forced to focus on the core part of the description information, which helps reduce the risk of overfitting and makes the model's performance more generalizable.

[0112] In an alternative embodiment, the pre-trained model is adjusted based on the description information after completion length processing and the tags corresponding to the reference image sequence until the pre-trained model enters the target state, including: performing a first transfer learning operation and a second transfer learning operation on the pre-trained model based on the description information after completion length processing, wherein the first transfer learning operation is used to train the pre-trained model to predict the full text content of the description information based on partial sentences of the description information, and the second transfer learning operation is used to train the pre-trained model to predict whether there is a context correlation relationship between any two different sentences in the description information; performing gradient training on the pre-trained model that has completed the first transfer learning operation and the second transfer learning operation based on the tags corresponding to the reference image sequence until the pre-trained model enters the target state, wherein the gradient training is used to calculate the label prediction accuracy of the pre-trained model for the reference image sequence and adjust the model parameters of the pre-trained model according to the label prediction accuracy.

[0113] Optionally, the purpose of the first transfer learning operation is to train the pre-trained model to predict the full text content based on partial sentences of the description information. Specifically, in the description of a medical image sequence, there may be multiple sentences, each sentence describing a different aspect of the sequence, such as modality, orientation, contrast, etc. By masking partial sentences in the description information, the model is trained to predict these masked sentences, thereby understanding the complete content of the description information. This operation helps the model learn the skill of inferring the overall description from partial information, which is particularly important for processing DICOM sequence descriptions that may be missing or incomplete information.

[0114] The second transfer learning operation focuses on training the model to predict whether there is a context correlation relationship between any two different sentences in the description information. In the context of the description of a medical image sequence, the second transfer learning operation pays more attention to the logical connection and coherence between sentences. By constructing sentence pairs and setting positive examples (sentence pairs with context correlation) and negative examples (sentence pairs without context correlation), the model is trained to distinguish which sentences have logical and semantic associations and which do not. This helps the model learn sentence-level semantic representations and how to perform more accurate classification based on the information flow between sentences.

[0115] Once the above two transfer learning operations are completed, the model will perform gradient training based on the tags of the reference image sequence. This involves calculating the difference (i.e., loss) between the multi-labels predicted by the model and the actual labels, and using the backpropagation algorithm and an optimizer (such as Adam) to adjust the model parameters to reduce this difference. This process is repeated until the label prediction accuracy of the model for the reference image sequence reaches a preset threshold, that is, it enters the target state. The target state usually means that the model has achieved the best classification performance on the test set and can predict the multi-labels of the sequence with a high accuracy.

[0116] In summary, the adjustment process of the entire pre-trained model is to first perform two transfer learning operations to enhance the model's ability to predict sentences and understand context associations, and then through label-based gradient training, adjust the model to the "target state" where it can accurately predict multi-labels of the image sequence. This process makes full use of the deep language representations already learned by the pre-trained model and combines specific domain training data, enabling the model to exhibit optimal performance in the multi-label classification task of medical image sequences.

[0117] According to another aspect of the present application, there is also provided a processing device for image sequences, wherein, Figure 2 is a schematic diagram of a processing device for image sequences according to an embodiment of the present application, as Figure 2 shown, the device includes: an acquisition unit 201, an input unit 202, a first processing unit 203, and a second processing unit 204.

[0118] Optionally, the acquisition unit 201 is configured to acquire a description text of a target image sequence, where the target image sequence includes at least one medical image; the input unit 202 is configured to input the description text into a target model, where the target model is a model obtained by adjusting a pre-trained model according to the description features of a reference image sequence, and the pre-trained model is used to predict unknown words in a sentence based on known words in the sentence and to predict the association relationships between different sentences; the first processing unit 203 is configured to perform a target prediction task on the description text through the target model, where the target prediction task is used to determine N labels of the target image sequence based on the known words in the description text and the sentences in the description text, where N is an integer greater than or equal to 1, and the N labels are used to classify the target image sequence from N dimensions; the second processing unit 204 is configured to use the N labels as the classification result of the target image sequence.

[0119] Optionally, the input unit 202 includes: a preprocessing subunit, configured to preprocess the description text, where the preprocessing includes at least one of the following processing operations:

[0120] Symbol processing operation, configured to delete the first type of symbols in the description text and retain the second type of symbols in the description text, where the first type of symbols are symbols irrelevant to the image features describing the medical image; the second type of symbols are symbols relevant to the image features describing the medical image;

[0121] Word segmentation processing operation, configured to divide the sentences in the description text into multiple words;

[0122] Stop word filtering operation, configured to delete the stop words in the description text;

[0123] Format conversion operation for converting the description text into the target text format;

[0124] Input the description text after preprocessing into the target model.

[0125] Optionally, the processing device for the image sequence further includes: a first acquisition unit for acquiring training text; a task construction unit for constructing a first training task and a second training task based on the training text, where the first training task is used to train the neural network's ability to predict the full text content of the training text based on partial sentences of the training text, and the second training task is used to train the neural network's ability to predict whether there is a context correlation relationship between any two different sentences in the training text; a task execution unit for repeatedly executing the first training task and the second training task on a preset neural network to obtain a pre-trained model.

[0126] Optionally, the task construction unit includes: a first processing subunit, a second processing subunit, and a first construction subunit. Among them, the first processing subunit is used to randomly select a target proportion of words in the training text as target words; the second processing subunit is used to replace the target words in the training text with preset characters to obtain a first text, where the preset characters are characters without actual semantics; the first construction subunit is used to construct a first training task based on the target words and the first text, where the first training task is used to train the neural network's ability to predict the target words based on the non-target words in the first text.

[0127] Optionally, the task construction unit includes: a third processing subunit for randomly selecting at least two sentences from the training text to form a sentence combination; a fourth processing subunit for setting a relevance label for the sentence combination, where the relevance label is used to characterize whether there is a context correlation relationship between at least two sentences in the sentence combination; a second construction subunit for constructing a second training task based on the sentence combination and the relevance label corresponding to the sentence combination, where the second training task is used to train the neural network's ability to predict whether there is a context correlation relationship between different sentences in the sentence combination.

[0128] Optionally, the processing device for the image sequence further includes: a second acquisition unit, configured to acquire a reference image sequence with known corresponding tags and description information of the reference image sequence; a length processing unit, configured to perform length processing on the description information of the reference image sequence, where the length processing is used to adjust the length of the description information of the reference image sequence to a preset length; a model adjustment unit, configured to adjust the pre-trained model according to the description information after the length processing and the tags corresponding to the reference image sequence until the pre-trained model enters a target state, where the pre-trained model in the target state has a tag prediction accuracy for the reference image sequence based on the description information after the length processing greater than a preset threshold; and a third processing unit, configured to use the pre-trained model that enters the target state as the target model.

[0129] Optionally, the length processing unit includes: a padding subunit, configured to, when detecting that the length of the description information is less than the preset length, amplify the length of the description information to the preset length by padding preset characters at the end of the description information; and a deletion subunit, configured to, when detecting that the length of the description information is greater than the preset length, reduce the length of the description information to the preset length by deleting characters one by one in reverse order starting from the end of the description information.

[0130] Optionally, the model adjustment unit includes: an execution subunit for transfer learning operations, configured to perform a first transfer learning operation and a second transfer learning operation on the pre-trained model according to the description information after the length processing, where the first transfer learning operation is used to train the pre-trained model's ability to predict the full text content of the description information based on partial sentences of the description information, and the second transfer learning operation is used to train the pre-trained model's ability to predict whether there is a context association relationship between any two different sentences in the description information; and a training subunit, configured to perform gradient training on the pre-trained model that has completed the first transfer learning operation and the second transfer learning operation according to the tags corresponding to the reference image sequence until the pre-trained model enters the target state, where the gradient training is used to calculate the tag prediction accuracy of the pre-trained model for the reference image sequence and adjust the model parameters of the pre-trained model according to the tag prediction accuracy.

[0131] According to another aspect of the present application, there is also provided a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and when the computer program runs, it causes the device where the computer-readable storage medium is located to execute the above-mentioned processing method for the image sequence.

[0132] According to another aspect of the present application, an electronic device is further provided. The electronic device includes one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-described method for processing an image sequence.

[0133] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.

[0134] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0135] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the units or modules can be in an electrical or other form.

[0136] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0137] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0138] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0139] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A method for processing an image sequence, characterized in that Including: Obtain a description text of a target image sequence, where the target image sequence includes at least one medical image; Input the description text into a target model, where the target model is a model obtained by adjusting a pre-trained model according to the description features of a reference image sequence, and the pre-trained model is used to predict unknown words in a sentence based on known words in the sentence and to predict the association relationship between different sentences; Execute a target prediction task on the description text through the target model, where the target prediction task is used to determine N labels of the target image sequence based on known words in the description text and sentences in the description text, where N is an integer greater than or equal to 1, and the N labels are used to classify the target image sequence from N dimensions; Use the N labels as the classification result of the target image sequence.

2. The method according to claim 1, wherein Inputting the description text into the target model includes: Preprocess the description text, where the preprocessing includes at least one of the following processing operations: Symbol processing operation, which is used to delete the first type of symbols in the description text and retain the second type of symbols in the description text, where the first type of symbols are symbols irrelevant to the image features describing the medical image; the second type of symbols are symbols relevant to the image features describing the medical image; Word segmentation processing operation, which is used to divide the sentences in the description text into multiple words; Stop word filtering operation, which is used to delete the stop words in the description text; Format conversion operation, which is used to convert the description text into a target text format; Input the description text after the preprocessing into the target model.

3. The method according to claim 1, wherein The training process of the pre-trained model includes the following steps: Obtain training text; Construct a first training task and a second training task based on the training text, where the first training task is used to train a neural network's ability to predict the full text content of the training text based on partial sentences of the training text, and the second training task is used to train a neural network's ability to predict whether there is a context association relationship between any two different sentences in the training text; Execute the first training task and the second training task on a preset neural network multiple times to obtain the pre-trained model.

4. The method according to claim 3, wherein Constructing the first training task based on the training text includes: Randomly select a target proportion of words in the training text as target words; Replace the target words in the training text with a preset character to obtain a first text, where the preset character is a character without actual semantics; Construct the first training task based on the target words and the first text, where the first training task is used to train a neural network's ability to predict the target words based on non-target words in the first text.

5. The method according to claim 3, wherein Constructing the second training task based on the training text includes: Randomly select at least two sentences from the training text to form a sentence combination; Set relevance tags for the sentence combination, where the relevance tags are used to characterize whether there is a context relevance relationship between at least two sentences in the sentence combination; Construct a second training task based on the sentence combination and the relevance tags corresponding to the sentence combination, where the second training task is used to train a neural network to predict the ability of whether there is a context relevance relationship between different sentences in the sentence combination.

6. The method according to claim 1, wherein The training steps of the target model include: Obtain a reference image sequence with known corresponding tags and description information of the reference image sequence; Perform length processing on the description information of the reference image sequence, where the length processing is used to adjust the length of the description information of the reference image sequence to a preset length; Adjust the pre-trained model according to the description information after the length processing and the tags corresponding to the reference image sequence until the pre-trained model enters a target state, where the pre-trained model in the target state has a tag prediction accuracy for the reference image sequence based on the description information after the length processing greater than a preset threshold; Use the pre-trained model that enters the target state as the target model.

7. The method according to claim 6, wherein Performing length processing on the description information of the reference image sequence includes: When it is detected that the length of the description information is less than the preset length, expand the length of the description information to the preset length by filling preset characters at the end of the description information; When it is detected that the length of the description information is greater than the preset length, reduce the length of the description information to the preset length by deleting characters one by one in reverse order starting from the end of the description information.

8. The method according to claim 6, wherein Adjusting the pre-trained model according to the description information after the length processing and the tags corresponding to the reference image sequence until the pre-trained model enters a target state includes: Perform a first transfer learning operation and a second transfer learning operation on the pre-trained model according to the description information after the length processing, where the first transfer learning operation is used to train the pre-trained model to predict the full text content of the description information based on partial sentences of the description information, and the second transfer learning operation is used to train the pre-trained model to predict whether there is a context relevance relationship between any two different sentences in the description information; Perform gradient training on the pre-trained model that has completed the first transfer learning operation and the second transfer learning operation according to the tags corresponding to the reference image sequence until the pre-trained model enters the target state, where the gradient training is used to calculate the tag prediction accuracy of the pre-trained model for the reference image sequence and adjust the model parameters of the pre-trained model according to the tag prediction accuracy.

9. An apparatus for processing an image sequence, characterized in that, Include: An acquisition unit for acquiring a description text of a target image sequence, where the target image sequence includes at least one medical image; An input unit for inputting the description text into a target model, where the target model is a model obtained by adjusting a pre-trained model according to the description features of a reference image sequence, and the pre-trained model is used to predict unknown words in a sentence based on known words in the sentence and to predict the association relationship between different sentences; A first processing unit for performing a target prediction task on the description text through the target model, where the target prediction task is used to determine N labels of the target image sequence based on known words in the description text and sentences in the description text, where N is an integer greater than or equal to 1, and the N labels are used to classify the target image sequence from N dimensions; A second processing unit for using the N labels as the classification result of the target image sequence.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where when the computer program runs, it causes the device where the computer-readable storage medium is located to execute the image sequence processing method according to any one of claims 1 to 8.

11. An electronic device, characterized in that, Comprising one or more processors and a memory, the memory is used to store one or more programs, where when the one or more programs are executed by the one or more processors, it causes the one or more processors to execute the image sequence processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Natural language processing method and device and computer equipment

    CN114528919A

  • Method and device for training text auditing model

    CN114970540A

  • Hotspot information marking method and system based on pre-trained intelligent label model

    CN117807225A

  • Image data and AI auxiliary film reading model matching system and method

    CN118098516A

  • Systems and methods for semi-supervised extraction of text classification information

    US20220229984A1

Cited By

  • Nuclear magnetic resonance imaging sequence standard quality control method and device and storage medium

    CN121562602A