Model training method, terminal equipment and computer readable storage medium
The training model of knowledge labels is obtained through unsupervised clustering, and the problem of time-consuming interpretation of medical images is solved, and a high-matching text report is generated, which reduces the cost of manual labeling and improves model accuracy.
Patent Information
- Application Number
- CN202510386850.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-29
AI Technical Summary
Medical imaging interpretation and reporting are time-consuming and highly dependent on expertise, which puts a heavy burden on clinicians.
The knowledge label of the sample report is obtained through unsupervised clustering, and the cluster label of the sample image and the report are trained to narrow the gap between text features and visual features and generate text reports with high matching degree with medical images.
It reduces the cost of manual labeling, improves the accuracy and adaptability of model training, and generates text reports with high matching degree with the input image.
Smart Images

Figure CN120388197A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of image processing, and particularly relates to a model training method, a terminal device, and a computer-readable storage medium. Background Art
[0002] Medical imaging can visually present internal organs, tissues, and structures, which plays a crucial role in disease diagnosis and potential disease screening in modern medicine. However, the process of interpreting medical images and writing reports is both time-consuming and highly dependent on professional knowledge, imposing a heavy burden on clinicians. Therefore, automatically generating text reports based on medical images is an urgent problem to be solved. Summary of the Invention
[0003] Embodiments of this application provide a model training method, a terminal device, and a computer-readable storage medium, which can generate a text report with a high degree of matching with a medical image.
[0004] In a first aspect, embodiments of this application provide a model training method, including:
[0005] Obtain a sample set, where the sample set includes multiple sample data; wherein, each sample data includes a sample image and a sample report corresponding to the sample image;
[0006] Obtain a clustering label of the sample report in the sample data; wherein, the clustering label is obtained by performing clustering processing on multiple sample reports in the sample set;
[0007] Train a first model according to the sample image in the sample data and the clustering label of the sample report in the sample data; wherein, the first model is used to generate a text report corresponding to an input image.
[0008] In embodiments of this application, a model is trained according to the knowledge label (clustering label) of the sample image and the sample report, so as to align the text feature with the visual feature and narrow the gap between the text and the image. The model trained by the model training method of the embodiments of this application can generate a text report with a high degree of matching with the input image. In addition, the process of obtaining the knowledge label is equivalent to unsupervised learning, and potential knowledge can be extracted from the text report without additional labels, reducing the cost of manual annotation; the process of training the model according to the knowledge label of the sample image and the sample report is equivalent to supervised learning, using the knowledge label to guide the model training and narrowing the difference between the text feature and the visual feature, which helps to improve the training accuracy of the model; the above-mentioned way of combining unsupervised learning with supervised learning can not only reduce the sample annotation cost and improve the adaptability of model training, but also improve the training accuracy of the model.
[0009] In a possible implementation of the first aspect, the method further includes:
[0010] Generating a vector representation corresponding to each sample report in the sample set;
[0011] Performing clustering processing on multiple sample reports in the sample set according to the vector representation corresponding to each sample report to obtain a clustering label corresponding to each sample report.
[0012] In a possible implementation of the first aspect, the performing clustering processing on multiple sample reports in the sample set according to the vector representation corresponding to each sample report to obtain a clustering label corresponding to each sample report includes:
[0013] Performing dimensionality reduction processing on the vector representation corresponding to the sample report to obtain a reduced-dimensional vector;
[0014] Performing clustering processing on multiple sample reports in the sample set according to the reduced-dimensional vector to obtain a clustering label corresponding to each sample report.
[0015] Compared with the implementation without dimensionality reduction processing, in the above implementation, through dimensionality reduction processing, the computational complexity brought by high-dimensional embedding vectors can be reduced, which helps to improve the efficiency of clustering processing.
[0016] In a possible implementation of the first aspect, the first model includes a first network and a second network;
[0017] The training of the first model according to the sample images in the sample data and the clustering labels of the sample reports in the sample data includes:
[0018] Inputting the sample images in the sample data into the first network to output the visual features of the sample images;
[0019] Inputting the visual features of the sample images into the second network to obtain the first text report;
[0020] Calculating a first loss value according to the clustering label corresponding to the sample report in the sample data and the visual features of the sample images;
[0021] Calculating a second loss value according to the true probability value of each word in the sample report and the predicted probability value of each word in the first text report;
[0022] Adjusting the parameters of the first model according to the first loss value and the second loss value.
[0023] In the embodiments of the present application, the process of training a model based on the knowledge tags of the sample images and sample reports is equivalent to supervised learning. Using the knowledge tags to guide the model training can reduce the difference between the text features and visual features, which helps to improve the training accuracy of the model.
[0024] In a possible implementation manner of the first aspect, adjusting the parameters of the first model according to the first loss value and the second loss value includes:
[0025] Inputting the sample report in the sample data and the first text report into the pre-trained second model, and outputting a first score representing the similarity between the sample report and the first text report;
[0026] Calculating a third loss value according to the first score;
[0027] Adjusting the model parameters of the first model according to the first loss value, the second loss value, and the third loss value.
[0028] In the embodiments of the present application, the first loss value can characterize the difference between the text features and visual features, the second loss value can characterize the difference between single words in the predicted text and the true text, and the third loss value can characterize the global semantic difference between the predicted text and the true text. Adjusting the model parameters of the first model according to the first loss value, the second loss value, and the third loss value enables the first model to learn more comprehensive information, which helps to improve the training accuracy of the first model.
[0029] In a possible implementation manner of the first aspect, adjusting the model parameters of the first model according to the first loss value, the second loss value, and the third loss value includes:
[0030] Adjusting the model parameters of the first network according to the first loss value;
[0031] Adjusting the model parameters of the second network according to the first loss value, the second loss value, and the third loss value.
[0032] In a possible implementation manner of the first aspect, the second network includes an encoder and a decoder, and the encoder includes a plurality of first detection heads, and different detection heads correspond to feature spaces of different scales;
[0033] Inputting the visual features of the sample image into the second network to obtain the first text report includes:
[0034] Inputting the visual features into each of the first detection heads of the encoder respectively, and obtaining a first feature vector output by each of the first detection heads;
[0035] Generate a second feature vector based on the first feature vectors output by each of the first detection heads;
[0036] Input the second feature vector into the decoder to obtain the first text report.
[0037] In the above implementation, the multi-head attention mechanism is adopted. Different detection heads capture the feature details in different-dimensional feature spaces. Different detection heads can focus on different parts and semantic information of the input text, learn more complex feature representations, and thus understand the text content more comprehensively, enabling the model to handle more complex natural language processing tasks. In addition, multiple detection heads process in parallel, which helps to improve the training and inference speed of the model.
[0038] In a possible implementation of the first aspect, the method further includes:
[0039] Obtain an image to be processed;
[0040] Input the image to be processed into the trained first model, and output the text report corresponding to the image to be processed.
[0041] In a second aspect, an embodiment of the present application provides a model training device, including:
[0042] An acquisition unit, configured to acquire a sample set, where the sample set includes multiple pieces of sample data; wherein, each piece of sample data includes a sample image and a sample report corresponding to the sample image;
[0043] The acquisition unit is further configured to acquire the clustering label of the sample report in the sample data; wherein, the clustering label is obtained by performing clustering processing on multiple sample reports in the sample set;
[0044] A training unit, configured to train a first model according to the sample image in the sample data and the clustering label of the sample report in the sample data; wherein, the first model is used to generate a text report corresponding to the input image.
[0045] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the model training method described in any one of the first aspects above is implemented.
[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the model training method described in any one of the first aspects above is implemented.
[0047] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device is caused to execute the model training method described in any one of the above first aspects.
[0048] It can be understood that for the beneficial effects of the above second aspect to fifth aspect, reference can be made to the relevant descriptions in the above first aspect, which will not be elaborated herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 is a schematic diagram of a training architecture provided by an embodiment of the present application;
[0051] Figure 2 is a schematic diagram of a training architecture provided by another embodiment of the present application;
[0052] Figure 3 is a schematic diagram of a training architecture provided by another embodiment of the present application;
[0053] Figure 4 is a schematic flowchart of the model training method provided by an embodiment of the present application;
[0054] Figure 5 is a structural block diagram of the model training device provided by an embodiment of the present application;
[0055] Figure 6 is a schematic diagram of the structure of the terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0057] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0058] It should also be understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0059] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.
[0060] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0061] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.
[0062] Medical imaging can provide a visual representation of internal organs, tissues, and structures, which plays a crucial role in disease diagnosis and potential disorder screening in modern medicine. However, the process of interpreting medical images and writing reports is both time-consuming and highly dependent on professional knowledge, imposing a heavy burden on clinicians. Therefore, automatically generating text reports based on medical images is an urgent problem to be solved.
[0063] Based on this, the embodiments of this application provide a model training method. In the embodiments of this application, knowledge labels of sample reports can be obtained through unsupervised clustering, which is equivalent to obtaining the text features of sample reports; then, a model is trained according to the sample images and the knowledge labels of the sample reports so that the text features are aligned with the visual features, narrowing the gap between the text and the image. The model trained by the model training method of the embodiments of this application can generate a text report with a relatively high matching degree with the input image.
[0064] First, introduce the model involved in the embodiments of this application. Refer to Figure 1 , which is a schematic diagram of the training architecture provided by the embodiments of this application. AsFigure 1 As shown, the training architecture includes a first model, a second model, and a third model. Among them, the first model in the embodiments of the present application is used to generate a text report corresponding to the input image. The second model is used to calculate the similarity score between the input data. The third model is used to obtain the clustering label of the input data.
[0065] It can be understood that in the model application stage, only the first model is used to generate a text report corresponding to the input image. In the model training stage, the second model and the third model are used to assist in training the first model.
[0066] For example, in the model training stage, the input of the first model is the sample image in the sample data, and the output is the predicted text report of the sample image. The input data of the second model is the sample report in the sample data, and the output is the clustering label of the sample report. The input of the third model is the sample report in the sample data and the text report output by the first model, and the output is the similarity score between the real sample report and the predicted text report.
[0067] In the embodiments of the present application, the clustering label output by the second model is equivalent to prior knowledge, which helps to promote the first model to learn visual features related to prior knowledge, thereby bridging the difference between visual features and text features. The third model can compare the global semantics between the real text report and the predicted text report, thereby helping to promote the first model to learn long sentences and complex sentences.
[0068] In one embodiment, the first model may include a first network and a second network. Among them, the first network is used to extract the features of the input image. The second network is used to generate a text report corresponding to the input image according to the features output by the first network.
[0069] Exemplarily, refer to Figure 2 , which is a schematic diagram of the training architecture provided by another embodiment of the present application. As Figure 2 shown, in the training process of the first model, the output of the third model and the output of the first network are used to calculate the first loss value; the output of the second network is used to calculate the second loss value; the output of the second model is used to calculate the third loss value; the first loss value, the second loss value, and the third loss value jointly adjust the parameters of the first model.
[0070] In one implementation, the first network may adopt a convolutional neural network with shared weights.
[0071] In one implementation, the second network may adopt a transformer model. Among them, the transformer model may include an encoder and a decoder.
[0072] Optionally, the encoder and decoder of the transformer model may include a multi-head attention module for extracting feature information of different scales from the input data, so as to capture detailed features of different dimensions.
[0073] In one implementation, the second model may include a vector embedding module and a feature comparison module. Among them, the vector embedding module is used to represent the input data as a vector, and the feature comparison module is used to calculate a similarity score based on the vector output by the vector embedding module.
[0074] Optionally, the vector embedding module in the second model may adopt the S-Bert model. This model, namely the Sentence-Bert model, is a sentence embedding model based on the Bidirectional Encoder Representations from Transformers (Bert). The S-Bert model can well understand the semantic and syntactic structures of sentences, capture subtle semantic differences in the text, and thus perform well in various semantic-related tasks.
[0075] Optionally, the feature comparison module may use cosine similarity to calculate the similarity score. Of course, it can be understood that other similarity algorithms, such as Euclidean distance, Mahalanobis distance, Pearson correlation coefficient, etc., may also be used, and the embodiments of the present application do not make specific limitations on this.
[0076] In one implementation, the third model may adopt a knowledge distillation model. Specifically, the knowledge distillation model may include an embedding module, a dimensionality reduction module, and a clustering module. Among them, the embedding module is used to represent the input data as a vector. The dimensionality reduction module is used to perform dimensionality reduction processing on the vector output by the embedding module. The clustering module is used to perform clustering processing on the input data according to the vector output by the dimensionality reduction module to obtain the clustering label of the input data.
[0077] Exemplarily, refer to Figure 3 , which is a schematic diagram of a training architecture provided by another embodiment of the present application. As Figure 3 shown, the training architecture may include a first model, a second model, and a third model. Among them, the first model includes a first network and a second network. The first network includes a convolutional neural network (CNN) and a first dimensionality reduction module. The second network includes an encoder and a decoder. The second model includes a vector embedding module and a feature comparison module.
[0078] The third model includes a report embedding module, a second dimensionality reduction module, and a knowledge clustering module.
[0079] Based on Figure 3 the training architecture shown, a training process of the first model may include:
[0080] Input the sample reports in the sample data into the third model. The report embedding module in the third model represents the sample reports as vectors. The second dimensionality reduction module in the third model performs dimensionality reduction on the vectors output by the report embedding module. The knowledge clustering module of the third model performs clustering processing based on the vectors after dimensionality reduction output by the second dimensionality reduction module, and outputs the clustering labels corresponding to the sample reports.
[0081] Input the sample images in the sample data into the first network of the first model. The CNN in the first network extracts the visual features in the sample images. The first dimensionality reduction module in the first network performs dimensionality reduction on the visual features. The encoder in the second network further extracts features from the visual features output by the CNN of the first network. The decoder in the second network predicts the text report corresponding to the sample image according to the features output by the encoder.
[0082] Input the sample reports and the text reports output by the first model into the second model. The vector embedding module in the second model represents the sample reports and the text reports as vectors respectively. The feature comparison module in the second model performs feature comparison on the vectors output by the vector embedding module, and outputs the similarity score.
[0083] Calculate the first loss value according to the visual features after dimensionality reduction output by the first dimensionality reduction module of the first network of the first model and the clustering labels output by the third model. Calculate the second loss value according to the text reports output by the first model. Calculate the third loss value according to the similarity scores output by the third model. Adjust the parameters of the first model according to the first loss value, the second loss value, and the third loss value.
[0084] Iteratively train in this way until the trained first model is obtained.
[0085] Based on the above training architecture, the training method of the first model is introduced below.
[0086] See Figure 4 , which is a schematic flow chart of the model training method provided by the embodiments of the present application. As an example but not a limitation, the method may include the following steps:
[0087] S101, Obtain a sample set, where the sample set includes multiple pieces of sample data.
[0088] Wherein, each piece of sample data includes a sample image and the sample report corresponding to the sample image.
[0089] It can be understood that the first model in the embodiments of the present application can be a single-type model, that is, it is only used to generate a text report for a certain type of medical image. The first model can also be a general model, that is, it can be used to generate text reports for multiple types of medical images. Whether the first model is a single-type model or a general model is related to the sample data used in the model training stage. In other words, if only a certain type of medical image is included in the sample set during the model training stage, the trained first model is a single-type model; if multiple types of medical images are included in the sample set during the model training stage, the trained first model is a general model.
[0090] In some application scenarios, a sample data can include multiple sample images and a sample report. For example, in the application scenario of B-ultrasound detection, a sample data can include B-ultrasound images (sample images) corresponding to multiple angles respectively and a sample report generated based on the multiple B-ultrasound images.
[0091] S102, obtain the clustering label of the sample report in the sample data.
[0092] As Figures 1 - 3 shown, the clustering label of the sample report can be generated by a third model.
[0093] In one implementation, the process of the third model generating the clustering label can be processed offline. During the model training stage, directly obtain the clustering label obtained by offline processing. This method can reduce the time of model training and is beneficial to improving the efficiency of model training.
[0094] Among them, the clustering label is obtained by clustering multiple sample reports in the sample set. For example, the sample set includes 100 sample data, and the third model pre-generates the clustering labels corresponding to the sample reports in the 100 sample data respectively.
[0095] In one embodiment, the method further includes:
[0096] Generate the vector representation corresponding to each sample report in the sample set;
[0097] Cluster the multiple sample reports in the sample set according to the vector representation corresponding to each sample report to obtain the clustering label corresponding to each sample report.
[0098] In one implementation, the vector representation corresponding to the sample report can be dimension-reduced first to obtain a reduced-dimensional vector; then cluster the multiple sample reports in the sample set according to the reduced-dimensional vector to obtain the clustering label corresponding to each sample report.
[0099] Compared with the implementation method without dimensionality reduction, in the above implementation method, through dimensionality reduction, the computational complexity brought by high-dimensional embedding vectors can be reduced, which helps to improve the efficiency of clustering processing.
[0100] In one example, step S102 can be implemented based on Figure 3 the third model in the shown training architecture. Specifically, the sample reports in the sample data are input into the third model. The report embedding module in the third model represents the sample reports as vectors. The second dimensionality reduction module in the third model performs dimensionality reduction on the vectors output by the report embedding module. The knowledge clustering module of the third model performs clustering processing based on the dimensionality-reduced vectors output by the second dimensionality reduction module, and outputs the clustering labels corresponding to the sample reports.
[0101] Optionally, the report embedding module can adopt the bag-of-words model (BOW), the term frequency-inverse document frequency (TF-IDF) module, or the S-Bert model. Among them, the BOW model represents the sample report as a set of constituent words. The TF-IDF model calculates the importance of a word in the sample report based on its frequency in the document and its inverse frequency in the corpus. The S-Bert model uses a pre-trained language model to embed the sample report into a vector representation.
[0102] The S-Bert model can encode sentences into fixed-length vector representations, effectively capturing the semantic information of sentences. This enables it to quickly and accurately perform feature extraction and representation learning on sentences when dealing with various natural language processing tasks. Compared with traditional models such as the bag-of-words model and the TF-IDF model, the S-Bert model can better process text data with complex semantics and has reached the current optimal level in tasks such as text matching and text entailment. Therefore, preferably, in the embodiments of the present application, the report embedding module of the third model can adopt the S-Bert model.
[0103] Optionally, the second dimensionality reduction module in the third model can adopt the Uniform Manifold Approximation and Projection (UMAP) method. UMAP is a non-linear dimensionality reduction algorithm based on manifold learning. It can reduce high-dimensional data to a low-dimensional space while preserving the internal structure of the data, which helps to reduce the computational complexity brought by high-dimensional embedding vectors.
[0104] It can be understood that the knowledge clustering module in the third model aims to group similar sample reports together to extract potential prior knowledge from the sample reports.
[0105] Optionally, the knowledge clustering module may adopt the K-means algorithm. The goal of the K-means algorithm is to divide a given data set into K clusters, such that the data points within each cluster have a relatively high similarity, while the data points between different clusters have a relatively low similarity. It iteratively calculates the distances between data points and cluster centers, and continuously adjusts the positions of the cluster centers to achieve the optimal clustering effect. The steps of the K-means algorithm may include:
[0106] 1) Initialize the cluster centers: Randomly select K data points from the data set as the initial cluster centers.
[0107] 2) Assign data points to clusters: For each data point in the data set, calculate its distances from the K cluster centers (usually using the Euclidean distance), and assign the data point to the cluster where the nearest cluster center is located.
[0108] 3) Update the cluster centers: For each cluster, calculate the mean of all data points in the cluster, and use the mean as the new cluster center.
[0109] 1) Repeat steps 2) and 3): Continuously repeat the steps of assigning data points to clusters and updating the cluster centers until the cluster centers no longer change or a preset number of iterations is reached.
[0110] In the embodiments of the present application, the data set in the K-means algorithm includes the vectors corresponding to the sample reports of each sample data in the sample set (the vectors after dimensionality reduction output by the second dimensionality reduction module), and the data points refer to the vectors corresponding to a sample report in the data set.
[0111] The K-means algorithm is a commonly used unsupervised clustering algorithm. Through the above process of obtaining cluster labels, multiple sample reports in the sample set can be divided into K groups, which is equivalent to learning the latent prior knowledge in each group of sample reports in an unsupervised clustering manner. The prior knowledge may include the writing style of the report, the associations between sentences in the report, etc. Since no labels of samples need to be pre-annotated during the unsupervised learning process, that is, latent prior knowledge can be extracted from sample reports without additional sample labels, the workload brought by manual annotation is reduced, which provides convenience for obtaining a large number of rich samples and helps improve the adaptability of model training.
[0112] S103. Train a first model according to the sample images in the sample data and the cluster labels of the sample reports in the sample data.
[0113] It can be understood that the purpose of training the first model is to enable the first model to learn the association relationship between the text features of the text report and the visual features of the image, so as to minimize the difference between the predicted text report and the input image during the model application stage.
[0114] In one embodiment, S103 may include:
[0115] Input the sample image in the sample data into the first network, and output the visual features of the sample image;
[0116] Input the visual features of the sample image into the second network to obtain the first text report;
[0117] Calculate the first loss value according to the clustering label corresponding to the sample report in the sample data and the visual features of the sample image;
[0118] Calculate the second loss value according to the true probability value of each word in the sample report and the predicted probability value of each word in the first text report;
[0119] Adjust the parameters of the first model according to the first loss value and the second loss value.
[0120] In one example, the step of obtaining the first text report may be implemented based on Figure 3 the first model in the training architecture shown. Specifically, input the sample image in the sample data into the first network of the first model. The CNN in the first network extracts the visual features in the sample image, and the first dimensionality reduction module of the first network performs dimensionality reduction processing on the visual features; the encoder in the second network further extracts features from the visual features output by the CNN of the first network, and the decoder in the second network predicts the text report (the first text report) corresponding to the sample image according to the features output by the encoder.
[0121] Optionally, the CNN may adopt the ResNet-101 model for melodies.
[0122] As described in the embodiment of S101, a sample data may include multiple sample images. In this case, each sample image needs to be input into the first network respectively to output the visual features of each sample image, and then the visual features of multiple sample images in this sample data are concatenated into global features.
[0123] For example, when a sample data S1 includes two sample images, input the two sample images into the first network respectively to output the visual feature V1 of the first sample image and the visual feature V2 of the second sample image; use a convolutional kernel to perform average pooling processing on the visual feature V1 and the visual feature V2 respectively to obtain the visual feature V1' and the visual feature V2'; concatenate the visual feature V1' and the visual feature V2' together to obtain the global feature V corresponding to the sample data S1 avg .
[0124] Since the dimension of the global feature is greater than that of a single sample image, in order to calculate the loss value, optionally, the visual feature of the sample image is first dimensionally reduced to obtain the dimensionally reduced visual feature V. a ′ vg ; Calculate the first loss value according to the clustering label corresponding to the sample report and the dimensionally reduced visual feature.
[0125] Based on Figure 3 the training architecture shown, the dimensionality reduction processing of the visual feature of the sample image can be implemented through the first dimensionality reduction module in the first network.
[0126] Through the dimensionality reduction processing in the above manner, the visual feature output by the first network can be matched with the dimension of a single sample image, which is conducive to the subsequent calculation of the loss value. It can be understood that if a sample data only includes a single sample image, in this case, the dimensionality reduction processing can be not performed.
[0127] Optionally, the first loss value can be calculated by the following formula:
[0128]
[0129] where, t i represents the i-th clustering label, and K is the number of clustering labels. S f is the SoftMax function. L kmve represents the first loss value.
[0130] Optionally, in the step of inputting the visual feature of the sample image into the second network, the dimensionally reduced visual feature can be input into the second network, or the visual feature without dimensionality reduction processing can be input into the second network.
[0131] Since V avg has a higher dimension and contains more comprehensive details in the visual feature compared with V a ′ vg . Therefore, selecting the visual feature V avg without dimensionality reduction processing as the input of the second network helps the second network obtain more comprehensive features, so as to generate a more matching text report.
[0132] Optionally, the second loss value can be calculated by the following formula:
[0133]
[0134] where, y i is the true probability value of the i-th word, and y i is 0 or 1. p i is the predicted probability of the i-th word. n is the total number of words in the text report. LTF Represents the second loss value.
[0135] In one implementation, the encoder of the second network of the first model may include multiple first detection heads, and different detection heads correspond to feature spaces of different scales. Correspondingly, the steps of obtaining the first text report may include:
[0136] Input the visual features into each first detection head of the encoder respectively to obtain the first feature vectors output by each first detection head;
[0137] Generate second feature vectors according to the first feature vectors output by each first detection head;
[0138] Input the second feature vectors into the decoder to obtain the first text report.
[0139] Exemplarily, the second network may adopt a Transformer model, where the second network includes an encoder and a decoder. In the encoder, the global features output by the first network are converted into queries (Q), keys (K), and values (V), and are respectively input into each first detection head; each first detection head is used to calculate the scaled dot-product attention between Q, K, and V. Multiple first detection heads process in parallel to respectively capture the details of the feature spaces of different dimensions; finally, the first feature vectors output by each first detection head are concatenated into second feature vectors. Correspondingly, the decoder may include multiple second detection heads for capturing the details of the feature spaces of different dimensions.
[0140] In the above implementation, the multi-head attention mechanism is adopted, and the feature details of the feature spaces of different dimensions are captured through multiple detection heads. Different detection heads can pay attention to different parts and semantic information of the input text, learn more complex feature representations, thereby more comprehensively understanding the text content, enabling the model to handle more complex natural language processing tasks. In addition, multiple detection heads process in parallel, which helps to improve the training and inference speed of the model.
[0141] In one implementation, the steps of adjusting the parameters of the first model may include:
[0142] Input the sample reports and the first text report in the sample data into the pre-trained second model, and output the first score representing the similarity between the sample report and the first text report;
[0143] [ Calculate the third loss value according to the first score;
[0144] Adjust the model parameters of the first model according to the first loss value, the second loss value, and the third loss value.
[0145] In one example, step S102 may be based on Figure 3Implementation of the third model in the shown training architecture. Specifically, the sample report and the text report output by the first model (the first text report) are input into the second model. The vector embedding module in the second model represents the sample report and the text report as vectors respectively, and the feature comparison module in the second model performs feature comparison on the vectors output by the vector embedding module and outputs a similarity score. Among them, the vector embedding module can represent each sentence in the sample report and the text report as a sentence vector. The feature comparison module can calculate the similarity between sentence vectors.
[0146] Optionally, the third loss value can be calculated by the following formula:
[0147]
[0148] Where S j represents the similarity score between the j-th sentence in the sample report and the j-th sentence in the text report output by the first model. L SC represents the third loss value.
[0149] In the above example, the similarity score characterizes the degree of difference between sentences, so the third loss value can characterize the global semantic difference between the real report and the predicted report.
[0150] In one implementation, the steps of adjusting the parameters of the first model may include:
[0151] Adjusting the model parameters of the first network according to the first loss value;
[0152] Adjusting the model parameters of the second network according to the first loss value, the second loss value, and the third loss value.
[0153] For example, the first loss value, the second loss value, and the third loss value can be weighted, the gradient of the weighted value can be calculated using the gradient descent method, and the model parameters can be adjusted according to the gradient direction. Since the second network does not participate in the calculation of the first loss value, the gradient of the first loss value for the second network is 0 during backpropagation adjustment, that is, the first loss value is only used to adjust the model parameters of the first network, while the first loss value, the second loss value, and the third loss value are used to adjust the model parameters of the second network.
[0154] In the embodiments of the present application, the first loss value can characterize the difference between text features and visual features, the second loss value can characterize the difference between individual words in the predicted text and the real text, and the third loss value can characterize the global semantic difference between the predicted text and the real text. Adjusting the model parameters of the first model according to the first loss value, the second loss value, and the third loss value enables the first model to learn more comprehensive information, thereby helping to improve the training accuracy of the first model.
[0155] In addition, the second model is pre-trained, which can effectively improve the training efficiency while ensuring the training accuracy.
[0156] In the embodiments of the present application, the model is trained according to the knowledge labels (clustering labels) of the sample images and sample reports to align the text features with the visual features and narrow the gap between the text and the image. The model obtained by training through the model training method of the embodiments of the present application can generate a text report with a relatively high matching degree with the input image. In addition, the process of obtaining the knowledge labels is equivalent to unsupervised learning, and potential knowledge can be extracted from the text reports without additional labels, reducing the cost of manual annotation; the process of training the model according to the knowledge labels of the sample images and sample reports is equivalent to supervised learning, using the knowledge labels to guide the model training and narrowing the difference between the text features and the visual features, which helps to improve the training accuracy of the model; the combination of the above unsupervised learning and supervised learning can not only reduce the sample annotation cost and improve the adaptability of the model training, but also improve the training accuracy of the model.
[0157] In one embodiment, after obtaining the trained first model, the model application stage may include:
[0158] Obtain the image to be processed; input the image to be processed into the trained first model, and output the text report corresponding to the image to be processed.
[0159] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution is prior or subsequent, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0160] Corresponding to the model training method described in the above embodiments, Figure 5 is the structural block diagram of the model training device provided by the embodiments of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.
[0161] Refer to Figure 5 and the device 5 includes:
[0162] An acquisition unit, configured to acquire a sample set, where the sample set includes multiple sample data; wherein, each sample data includes a sample image and a sample report corresponding to the sample image.
[0163] The acquisition unit is further configured to acquire the clustering label of the sample report in the sample data; wherein, the clustering label is obtained by clustering multiple sample reports in the sample set.
[0164] A training unit for training a first model according to the sample images in the sample data and the clustering labels of the sample reports in the sample data; wherein, the first model is used to generate a text report corresponding to the input image.
[0165] Optionally, the obtaining unit is further configured to:
[0166] Generate a vector representation corresponding to each sample report in the sample set;
[0167] Perform clustering processing on multiple sample reports in the sample set according to the vector representation corresponding to each sample report, and obtain a clustering label corresponding to each sample report.
[0168] Optionally, the obtaining unit is further configured to:
[0169] Perform dimensionality reduction processing on the vector representation corresponding to the sample report to obtain a dimensionality-reduced vector;
[0170] Perform clustering processing on multiple sample reports in the sample set according to the dimensionality-reduced vector, and obtain a clustering label corresponding to each sample report.
[0171] Optionally, the training unit is further configured to:
[0172] Input the sample images in the sample data into the first network, and output the visual features of the sample images;
[0173] Input the visual features of the sample images into the second network to obtain the first text report;
[0174] Calculate a first loss value according to the clustering label corresponding to the sample report in the sample data and the visual features of the sample images;
[0175] Calculate a second loss value according to the true probability value of each word in the sample report and the predicted probability value of each word in the first text report;
[0176] Adjust the parameters of the first model according to the first loss value and the second loss value.
[0177] Optionally, the training unit is further configured to:
[0178] Input the sample report in the sample data and the first text report into a pre-trained second model, and output a first score representing the similarity between the sample report and the first text report;
[0179] Calculate a third loss value according to the first score;
[0180] Adjust the model parameters of the first model according to the first loss value, the second loss value and the third loss value.
[0181] Optionally, the training unit is further configured to:
[0182] Adjust the model parameters of the first network according to the first loss value;
[0183] Adjust the model parameters of the second network according to the first loss value, the second loss value, and the third loss value.
[0184] Optionally, the training unit is further configured to:
[0185] Input the visual features into each of the first detection heads of the encoder respectively to obtain first feature vectors output by each of the first detection heads;
[0186] Generate a second feature vector according to the first feature vectors output by each of the first detection heads;
[0187] Input the second feature vector into the decoder to obtain the first text report.
[0188] It should be noted that for the information interaction, execution process, etc. between the above-mentioned device / units, since they are based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part specifically, and details will not be elaborated here.
[0189] In addition, Figure 5 The shown model training device can be a software unit, a hardware unit, or a unit combining software and hardware built into an existing terminal device, can also be integrated into the terminal device as an independent pendant, or can exist as an independent terminal device.
[0190] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, can also exist as individual physical units, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments, and details will not be elaborated here.
[0191] Figure 6 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. AsFigure 6 As shown, the terminal device 6 of this embodiment includes: at least one processor 60 ( Figure 6 only one is shown in the figure), a processor, a memory 61, and a computer program 62 stored in the memory 61 and executable on the at least one processor 60. When the processor 60 executes the computer program 62, the steps in any of the above-mentioned model training method embodiments are implemented.
[0192] The terminal device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art can understand that Figure 6 merely examples of the terminal device 6, and do not constitute a limitation on the terminal device 6. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0193] The processor 60 may be a central processing unit (CPU), and the processor 60 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0194] In some embodiments, the memory 61 may be an internal storage unit of the terminal device 6, such as the hard disk or memory of the terminal device 6. In some other embodiments, the memory 61 may also be an external storage device of the terminal device 6, such as a plug-in hard disk equipped on the terminal device 6, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 61 may also include both the internal storage unit and the external storage device of the terminal device 6. The memory 61 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 61 may also be used to temporarily store data that has been output or will be output.
[0195] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0196] An embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executed.
[0197] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device capable of carrying the computer program code to the device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.
[0198] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0199] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0200] In the embodiments provided in the present application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.
[0201] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0202] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A model training method, characterized in that, Including: Obtain a sample set, where the sample set includes multiple pieces of sample data; among them, each piece of sample data includes a sample image and a sample report corresponding to the sample image; Obtain the clustering labels of the sample reports in the sample data; among them, the clustering labels are obtained by clustering multiple sample reports in the sample set; Train a first model according to the sample images in the sample data and the clustering labels of the sample reports in the sample data; among them, the first model is used to generate a text report corresponding to the input image.
2. The model training method according to claim 1, wherein, The method further includes: Generate a vector representation corresponding to each sample report in the sample set; Cluster the multiple sample reports in the sample set according to the vector representation corresponding to each sample report to obtain the clustering label corresponding to each sample report.
3. The model training method according to claim 2, wherein The clustering the multiple sample reports in the sample set according to the vector representation corresponding to each sample report to obtain the clustering label corresponding to each sample report includes: Perform dimensionality reduction processing on the vector representation corresponding to the sample report to obtain a dimensionality-reduced vector; Cluster the multiple sample reports in the sample set according to the dimensionality-reduced vector to obtain the clustering label corresponding to each sample report.
4. The model training method according to claim 1, wherein The first model includes a first network and a second network; The training the first model according to the sample images in the sample data and the clustering labels of the sample reports in the sample data includes: Input the sample images in the sample data into the first network to output the visual features of the sample images; Input the visual features of the sample images into the second network to obtain a first text report; Calculate a first loss value according to the clustering label corresponding to the sample report in the sample data and the visual features of the sample images; Calculate a second loss value according to the true probability value of each word in the sample report and the predicted probability value of each word in the first text report; Adjust the parameters of the first model according to the first loss value and the second loss value.
5. The model training method according to claim 4, wherein The adjusting the parameters of the first model according to the first loss value and the second loss value includes: Input the sample reports in the sample data and the first text report into a pre-trained second model to output a first score representing the similarity between the sample report and the first text report; Calculate a third loss value according to the first score; Adjust the model parameters of the first model according to the first loss value, the second loss value, and the third loss value.
6. The model training method according to claim 5, wherein The adjusting the model parameters of the first model according to the first loss value, the second loss value, and the third loss value includes: Adjust the model parameters of the first network according to the first loss value; Adjust the model parameters of the second network according to the first loss value, the second loss value, and the third loss value.
7. The model training method according to claim 4, wherein The second network includes an encoder and a decoder, and the encoder includes multiple first detection heads, and different detection heads correspond to feature spaces of different scales; The inputting the visual features of the sample images into the second network to obtain the first text report includes: Input the visual features into each of the first detection heads of the encoder respectively, and obtain the first feature vectors output by each of the first detection heads; Generate second feature vectors according to the first feature vectors output by each of the first detection heads; Input the second feature vectors into the decoder to obtain the first text report.
8. The model training method according to any one of claims 1 to 7, characterized in that The method further includes: Obtain an image to be processed; Input the image to be processed into the trained first model, and output the text report corresponding to the image to be processed.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 8 is implemented.