Training Method of Classification Model, Image Classification Method, Device, Equipment and Medium

By performing feature extraction and feature fusion of sample medical images in the local classification model of the client, and combining federated learning and attention mechanisms, the problem of insufficient data of high-quality samples is solved, and the training effect and parameter accuracy of the image classification model are improved.

CN115239675BActive Publication Date: 2025-07-08PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210908820.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-07-08
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

The existing image classification methods rely on neural network models, but it is difficult to obtain high-quality sample data, resulting in poor model training results.

Method used

By using local classification models on the client, it is used to perform feature extraction, feature fusion and self-attention calculation of sample medical images, and combined with federated learning and attention mechanisms, the model parameters are optimized to avoid overfitting problems.

Benefits of technology

The training effect of the model is improved, the multimodal fusion ability of image features is enhanced, the accuracy of model parameters is improved, and the overfitting problem of local models is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115239675B_ABST
    Figure CN115239675B_ABST
Patent Text Reader

Abstract

The present application provides a training method for a classification model, an image classification method and device, equipment, and medium, belonging to the field of artificial intelligence technology. The method is applied to a client and includes: obtaining a sample medical image; extracting features from the sample medical image through the convolutional layer of a local classification model to obtain sample image features; performing attention calculation on the sample image features through the cross-attention mechanism and self-attention mechanism of the local classification model to obtain target fusion image features; calculating a model loss value by the prediction layer of the local classification model for the target fusion image features, and updating the original model parameters received by the local classification model to local model parameters according to the model loss value; sending the local model parameters to the server side; downloading target model parameters from the server side; and updating the local model parameters according to the downloaded target model parameters to train the local classification model. The present application can improve the training effect of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a training method for a classification model, an image classification method, device, equipment and medium. Background Art

[0002] Most current image classification methods rely on neural network models to implement. The training of neural network models often requires good sample data, and it is difficult to obtain a large amount of high-quality sample data, which affects the training effect of the model. Therefore, how to improve the training effect of the model has become an urgent technical problem to be solved. Summary of the Invention

[0003] The main purpose of the embodiments of this application is to propose a training method for a classification model, an image classification method, device, equipment and medium, aiming to improve the training effect of the model.

[0004] To achieve the above object, in the first aspect of the embodiments of this application, a training method for a classification model is proposed, which is applied to a client. The client stores a pre-trained local classification model. The method includes:

[0005] Obtain a sample medical image;

[0006] Extract features from the sample medical image through the convolutional layer of the local classification model to obtain sample image features;

[0007] Fuse the features of the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features;

[0008] Perform self-attention calculation on the initial fused image features through the self-attention mechanism of the local classification model to obtain target fused image features;

[0009] Calculate the loss value of the model through the prediction layer of the local classification model for the target fused image features, and update the original model parameters received by the local classification model to local model parameters according to the model loss value;

[0010] Send the local model parameters to the server side;

[0011] Download target model parameters from the server side;

[0012] Update the local model parameters according to the downloaded target model parameters to train the local classification model.

[0013] In some embodiments, the step of obtaining a sample medical image includes:

[0014] Obtain an original medical image;

[0015] Perform dimensionality transformation on the original medical image to obtain the sample medical image.

[0016] In some embodiments, the sample image features include first image features and second image features. The step of performing feature fusion on the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features includes:

[0017] Perform positional encoding on the first image features to obtain first encoded feature vectors;

[0018] Obtain second image features corresponding to the first image features according to a preset image cross-relationship;

[0019] Perform embedding processing on the first encoded feature vectors through the cross-attention mechanism to obtain first embedding vectors, and perform embedding processing on the second image features through the cross-attention mechanism to obtain second embedding vectors;

[0020] Perform similarity calculation on the first embedding vectors and the second embedding vectors through the cross-attention mechanism to obtain feature similarity values;

[0021] Perform fusion processing on the first embedding vectors and the second embedding vectors according to the feature similarity values to obtain the initial fused image features.

[0022] In some embodiments, the step of calculating a model loss value for the target fused image features through the prediction layer of the local classification model and updating the original model parameters received by the local classification model to local model parameters according to the model loss value includes:

[0023] Perform splicing processing on the target fused image features to obtain target classification image features;

[0024] Perform classification probability calculation on the target classification image features through the prediction function of the prediction layer and a reference classification label to obtain a predicted classification value;

[0025] Perform screening processing on the reference classification label according to the predicted classification value to obtain a predicted label;

[0026] Perform loss calculation on the predicted label and the original label of the sample medical image to obtain a model loss value;

[0027] Update the original model parameters received by the local classification model to local model parameters according to the model loss value.

[0028] To achieve the above object, a second aspect of the embodiments of the present application proposes a training method for a classification model, which is applied to the server side. The method includes:

[0029] Sending preset original model parameters to the client;

[0030] Obtaining local model parameters sent by multiple clients; wherein, the local model parameters are obtained according to the training method described in the first aspect;

[0031] Training the global classification model on the server side according to the local model parameters to obtain target model parameters; wherein, the target model parameters are used for the client to download, so that the client updates the local model parameters according to the downloaded target model parameters.

[0032] To achieve the above object, a third aspect of the embodiments of the present application proposes an image classification method, which is applied to the client. The image classification method includes:

[0033] Obtaining a target medical image to be classified;

[0034] Performing image preprocessing on the target medical image to obtain an initial medical image;

[0035] Inputting the initial medical image into the local classification model for prediction processing to obtain the target category of the target medical image, wherein the local classification model is trained according to the training method described in the first aspect.

[0036] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a training device for a classification model, which is applied to the client. The client stores a pre-trained local classification model. The device includes:

[0037] A sample image acquisition module, configured to acquire a sample medical image;

[0038] A feature extraction module, configured to extract features from the sample medical image through the convolutional layer of the local classification model to obtain sample image features;

[0039] A feature fusion module, configured to fuse the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features;

[0040] A self-attention calculation module, configured to perform self-attention calculation on the initial fused image features through the self-attention mechanism of the local classification model to obtain target fused image features;

[0041] A loss calculation module, configured to calculate a loss of the target fusion image feature through a prediction layer of the local classification model, obtain a model loss value, and update the original model parameters received by the local classification model to local model parameters according to the model loss value;

[0042] A parameter sending module, configured to send the local model parameters to the server side;

[0043] A parameter downloading module, configured to download target model parameters from the server side;

[0044] A parameter updating module, configured to update the local model parameters according to the downloaded target model parameters to train the local classification model.

[0045] To achieve the above object, a fifth aspect of the embodiments of the present application provides an image classification device, which is applied to a client. The image classification device includes:

[0046] A target image acquisition module, configured to acquire a target medical image to be classified;

[0047] An image preprocessing module, configured to perform image preprocessing on the target medical image to obtain an initial medical image;

[0048] A classification module, configured to input the initial medical image into a local classification model for prediction processing to obtain a target category of the target medical image, where the local classification model is trained according to the training device described in the fourth aspect.

[0049] To achieve the above object, a sixth aspect of the embodiments of the present application provides an electronic device, which includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the method described in the first aspect, or the method described in the second aspect, or the method described in the third aspect.

[0050] To achieve the above object, a seventh aspect of the embodiments of the present application provides a storage medium, which is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in the first aspect, or the method described in the second aspect, or the method described in the third aspect.

[0051] The training method of the classification model proposed in this application, the image classification method, device, equipment and medium. By obtaining sample medical images and extracting features of the sample medical images through the convolutional layer of the local classification model to obtain sample image features, it can extract relatively complex image information through a deep learning model, so as to incorporate image data of different modalities into the model training process. Further, feature fusion is performed on the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features; and self-attention calculation is performed on the initial fused image features through the self-attention mechanism of the local classification model to obtain target fused image features, which can better perform multimodal fusion of different image features and improve the training effect of the model. Further, loss calculation is performed on the target fused image features through the prediction layer of the local classification model to obtain a model loss value, and the original model parameters received by the local classification model are updated to local model parameters according to the model loss value. By introducing the attention mechanism into the local classification model, the attention parameters of the local classification model can be effectively optimized, and the accuracy of the obtained local model parameters can be improved. Finally, the local model parameters are sent to the server side, the target model parameters are downloaded from the server side, and the local model parameters are updated according to the downloaded target model parameters to train the local classification model. By means of federated modeling, the overfitting problem of the local classification model on the client side can be effectively avoided. At the same time, combining multimodal fusion, the attention mechanism and federated learning can improve the training effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a flowchart of the training method of the classification model provided by an embodiment of this application;

[0053] Figure 2 is Figure 1 a flowchart of step S101 in

[0054] Figure 3 is Figure 1 a flowchart of step S103 in

[0055] Figure 4 is Figure 1 a flowchart of step S105 in

[0056] Figure 5 is another flowchart of the training method of the classification model provided by an embodiment of this application;

[0057] Figure 6 is a flowchart of the image classification method provided by an embodiment of this application;

[0058] Figure 7 is a schematic structural diagram of the training device of the classification model provided by an embodiment of this application;

[0059] Figure 8 It is a schematic structural diagram of an image classification device provided by an embodiment of the present application;

[0060] Figure 9 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0061] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0062] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing embodiments of the present application and are not intended to limit the present application.

[0064] First, several nouns involved in the present application are analyzed:

[0065] Artificial intelligence (AI): It is a new technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, sense the environment, acquire knowledge and use knowledge to obtain the best results of theories, methods, technologies and application systems.

[0066] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information intent recognition, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and linguistic research related to language computing, etc.

[0067] Information Extraction: A text processing technology that extracts factual information such as specified types of entities, relationships, events, etc. from natural language texts and forms structured data output. Information extraction is a technology for extracting specific information from text data. Text data is composed of some specific units, such as sentences, paragraphs, and passages. Text information is precisely composed of some small specific units, such as characters, words, phrases, sentences, paragraphs, or combinations of these specific units. Extracting noun phrases, personal names, place names, etc. from text data are all text information extraction. Of course, the information extracted by text information extraction technology can be various types of information.

[0068] Federated Learning: Also known as collaborative learning and consortium learning. Federated Learning is a machine learning framework that can effectively help multiple institutions conduct data usage and machine learning modeling while meeting the requirements of user privacy protection, data security, and government regulations. As a distributed machine learning paradigm, federated learning can effectively solve the data silo problem, enabling participating parties to jointly model without sharing data, and can technically break data silos and achieve AI collaboration. Federated learning has three major components: data sources, federated learning systems, and users. Under the federated learning system, each data source party performs data preprocessing, jointly builds a machine learning model, and feeds back the output results to the users.

[0069] Attention Mechanism: The attention mechanism enables neural networks to have the ability to focus on subsets of their inputs (or features), select specific inputs, and can be applied to any type of input regardless of its shape. In the case of limited computing power, the attention mechanism is a resource allocation scheme that is one of the main means to solve the problem of information overload, allocating computing resources to more important tasks.

[0070] Multimodal fusion: It refers to integrating or fusing two or more biometric technologies, leveraging the unique advantages of multiple biometric technologies, and combining data fusion technology to make the authentication and recognition process more accurate and secure. The main difference from traditional single biometric methods is that multimodal biometric technology can collect different biometric features (such as fingerprints, finger veins, faces, iris images, etc.) through independent or combined collectors of multiple acquisition methods, and identify and authenticate by analyzing and judging the feature values of multiple biometric methods.

[0071] Most current image classification methods rely on neural network models to achieve. The training of neural network models often requires good sample data, and it is difficult to obtain a large amount of high-quality sample data, which affects the training effect of the model. Therefore, how to improve the training effect of the model has become a technical problem to be solved urgently.

[0072] Based on this, the embodiments of the present application provide a training method for a classification model, an image classification method and device, equipment, and medium, aiming to improve the training effect of the model.

[0073] The training method for the classification model, the image classification method and device, equipment, and medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the recommendation method in the embodiments of the present application is described.

[0074] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in theory, method, technology, and application systems.

[0075] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0076] The training method of the classification model, the image classification method, device, equipment and medium provided by the embodiments of the present application relate to the field of artificial intelligence technology. The training method of the classification model provided by the embodiments of the present application can be applied to a terminal, or to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the recommendation method, etc., but is not limited to the above forms.

[0077] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0078] The classification model and the image classification method of the embodiments of the present application are applicable to a federated system. The entire federated system framework consists of two components: a server and multiple clients (such as laptop computers, smart phones, tablet computers, etc.). Each client is provided with a local classification model, and the local classification model is used to classify the multi-modal medical images received by the client.

[0079] Figure 1 is an optional flowchart of the training method of the classification model provided by the embodiments of the present application, applied to a client, Figure 1 The method in may include but is not limited to steps S101 to S108.

[0080] Step S101, obtain sample medical images;

[0081] Step S102: Extract features from the sample medical image through the convolutional layer of the local classification model to obtain sample image features;

[0082] Step S103: Perform feature fusion on the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features;

[0083] Step S104: Perform self-attention calculation on the initial fused image features through the self-attention mechanism of the local classification model to obtain target fused image features;

[0084] Step S105: Calculate the loss of the target fused image features through the prediction layer of the local classification model to obtain the model loss value, and update the original model parameters received by the local classification model to local model parameters according to the model loss value;

[0085] Step S106: Send the local model parameters to the server side;

[0086] Step S107: Download the target model parameters from the server side;

[0087] Step S108: Update the local model parameters according to the downloaded target model parameters to train the local classification model.

[0088] Steps S101 to S108 illustrated in the embodiments of the present application, by obtaining the sample medical image and extracting features from the sample medical image through the convolutional layer of the local classification model to obtain sample image features, can extract relatively complex image information through the deep learning model, so as to incorporate different modalities of image data in the model training process. Further, perform feature fusion on the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features; and perform self-attention calculation on the initial fused image features through the self-attention mechanism of the local classification model to obtain target fused image features, which can better perform multi-modal fusion of different image features and improve the training effect of the model. Further, calculate the loss of the target fused image features through the prediction layer of the local classification model to obtain the model loss value, and update the original model parameters received by the local classification model to local model parameters according to the model loss value. By introducing the attention mechanism into the local classification model, the attention parameters of the local classification model can be effectively optimized, and the accuracy of the obtained local model parameters can be improved. Finally, send the local model parameters to the server side, download the target model parameters from the server side, and update the local model parameters according to the downloaded target model parameters to train the local classification model. Through the federated modeling method, the overfitting problem of the local classification model of the client can be effectively avoided. At the same time, combining multi-modal fusion, the attention mechanism and federated learning can improve the training effect of the model.

[0089] Please refer to Figure 2 , in some embodiments, step S101 may include but is not limited to steps S201 to S202:

[0090] Step S201, obtain the original medical image;

[0091] Step S202, perform dimensionality transformation on the original medical image to obtain the sample medical image.

[0092] In step S201 of some embodiments, when obtaining the original medical image, it can be obtained from an existing medical database, or can be obtained by camera shooting or other means, without limitation.

[0093] In some medical application scenarios, the original medical image is a medical image, and the type of the object included in the original medical image is a lesion, that is, the part of the body where a disease occurs. It should be noted that a medical image refers to internal tissues obtained in a non-invasive manner for medical treatment or medical research. For example, images of the stomach, abdomen, heart, knee, brain, such as images generated by medical instruments such as Computed Tomography (CT), Magnetic Resonance Imaging (MRI), ultrasonic (US), X-ray images, electroencephalograms, and optical photography.

[0094] In step S202 of some embodiments, since the original medical image includes grayscale images, three-dimensional images, etc., the original medical image often presents as multi-modal. Therefore, it is necessary to perform dimensionality transformation on the original medical image so that the original medical images of different modalities are in the same dimension to obtain the sample medical image. Specifically, when performing dimensionality transformation on the original medical image, various processing methods such as adjusting the range of image grayscale values, resampling or oversampling the original medical image, or data augmentation of the original medical image can be used to make the original medical images of different modalities in the same dimension.

[0095] Furthermore, in order to improve the image processing efficiency, one or a combination of the above processing methods can be used for image processing without limitation. For example, in a specific image processing process, first perform oversampling on the original medical image, and then adjust the grayscale value of the original medical image after oversampling to obtain the sample medical image.

[0096] In step S102 of some embodiments, each client presets a local classification model, which includes a convolutional layer, multiple attention modules, and a prediction layer. First, the convolutional layer of the local classification model extracts features from the sample medical images, capturing the image texture information of each sample medical image, including the extraction of shallow edge structure information to deep texture semantic structure information, so as to obtain sample image features. Among them, the convolutional layer can be a shallow convolutional layer. In some image classification scenarios, the specific convolutional operation process can include Gaussian blur, bilateral filtering, etc., without limitation.

[0097] Please refer to Figure 3 , in some embodiments, the sample image features include first image features and second image features, and step S103 may include but is not limited to steps S301 to S305:

[0098] Step S301, perform position encoding on the first image features to obtain a first encoded feature vector;

[0099] Step S302, obtain the second image features corresponding to the first image features according to the preset image cross-relationship;

[0100] Step S303, perform embedding processing on the first encoded feature vector through the cross-attention mechanism to obtain a first embedded vector, and perform embedding processing on the second image features through the cross-attention mechanism to obtain a second embedded vector;

[0101] Step S304, calculate the similarity between the first embedded vector and the second embedded vector through the cross-attention mechanism to obtain a feature similarity value;

[0102] Step S305, perform fusion processing on the first embedded vector and the second embedded vector according to the feature similarity value to obtain an initial fused image feature.

[0103] In step S301 of some embodiments, the local classification model includes multiple attention modules, each attention module includes a cross-attention layer and a self-attention layer, and each attention module is used to perform attention calculation on a first image feature. Specifically, for a certain first image feature, first perform position encoding on the first image feature through the attention module, so that the first image feature is mapped to a preset vector space to obtain a first encoded feature vector.

[0104] In step S302 of some embodiments, in order to better capture the image information of different sample medical images, enable the use of the similarity between the image information of different sample medical images for model training, and improve the training effect of the model, it is necessary to obtain a preset image cross-relationship. This image cross-relationship can be a preset image mapping relationship, which is determined by computer programming or manual presetting. For example, randomly pair sample medical images in pairs to obtain image pairs, and construct an image cross-relationship based on the image pairs. There is a cross-relationship between the two sample medical images defined as the image pair. When calculating the attention for the sample medical features corresponding to a certain sample medical image, it is necessary to extract and incorporate the sample image features of another sample medical image corresponding to this sample medical image. Therefore, the second image feature corresponding to the first image feature can be obtained according to the preset image cross-relationship.

[0105] In step S303 of some embodiments, the first encoded feature vector is embedded through a cross-attention mechanism, mapping the first encoded feature vector from a high-dimensional vector space to a low-dimensional vector space to obtain a first embedded vector. Similarly, the second encoded feature vector is embedded through a cross-attention mechanism, mapping the second encoded feature vector from a high-dimensional vector space to a low-dimensional vector space to obtain a second embedded vector.

[0106] In step S304 of some embodiments, when calculating the similarity between the first embedded vector and the second embedded vector through a cross-attention mechanism, vector extraction is performed on the first feature map corresponding to the first embedded vector, extracting the feature vectors at each channel position on the first feature map. Similarly, vector extraction is performed on the second feature map corresponding to the second embedded vector, extracting the feature vectors at each channel position on the second feature map. A cross-operation is performed on the feature vectors on the first feature map and the feature vectors on the second feature map to obtain the first correlation degree between the first embedded vector and the second embedded vector and the second correlation degree of the feature vectors on the cross-path. The first correlation degree and the second correlation degree are weighted and calculated to obtain a feature similarity value.

[0107] In step S305 of some embodiments, according to the magnitude of the feature similarity value, the first embedded vector and the second embedded vector with feature similarity values greater than a preset threshold are vector-concatenated to fuse the image information of the first embedded vector and the second embedded vector, so that the obtained initial fused image features contain rich context information. By adding context information to the initial fused image features, it is possible to better enhance the image local feature representation and the image pixel-level representation, enabling multi-modal fusion between different image features and improving the training effect of the model.

[0108] In a specific application scenario, the sample image features include a first image feature M and a second image feature N. There is an image intersection relationship between the first image feature M and the second image feature N. The first image feature M is obtained by extracting features from a certain sample medical image through a convolutional layer, and the second image feature N is obtained by extracting features from another sample medical image through a convolutional layer. The attention module includes an attention module P and an attention module Q. Among them, the attention module P is used to perform attention calculation on the first image feature M, and the attention module Q is used to perform attention calculation on the second image feature N. According to the above steps S301 to S305, the attention module P performs position encoding on the first image feature M to obtain a first encoded feature vector m, and the attention module Q performs position encoding on the second image feature N to obtain a second encoded feature vector n. Then, the cross-attention layer of the attention module P performs cross-attention calculation on the first encoded feature vector m and the second image feature N to obtain a first fused image feature, and the cross-attention layer of the attention module Q performs cross-attention calculation on the second encoded feature vector n and the first image feature M to obtain a second fused image feature.

[0109] In the above steps S301 to S305, the feature fusion of the sample image features is performed through the cross-attention mechanism of the local classification model, which can better perform multi-modal fusion of different image features, enabling the embodiments of the present application to jointly model multiple image features of multiple clients and improving the training effect of the model.

[0110] In step S104 of some embodiments, self-attention calculation is performed on the initial fused image feature through the self-attention mechanism of the self-attention layer of the local classification model, which specifically captures the local image features of the initial fused image feature, performs importance ranking on the local image features, focuses on the local image information with higher importance, and obtains the target fused image feature. Among them, the importance of the local image features can be determined based on parameters such as image gray values and image pixel values, or can be determined by other means, without limitation.

[0111] Please refer to Figure 4 , in some embodiments, step S105 may include but is not limited to steps S401 to S405:

[0112] Step S401, perform splicing processing on the target fused image feature to obtain a target classification image feature;

[0113] Step S402, calculate the classification probability of the target classification image feature through the prediction function of the prediction layer and the reference classification label to obtain a predicted classification value;

[0114] Step S403, perform screening processing on the reference classification label according to the predicted classification value to obtain a predicted label;

[0115] Step S404: Calculate the loss between the predicted label and the original label of the sample medical image to obtain the model loss value;

[0116] Step S405: Update the original model parameters received by the local classification model to local model parameters according to the model loss value.

[0117] In step S401 of some embodiments, when splicing the target fusion image features generated from different sample medical images, methods such as vector addition or vector splicing can be used. For example, adding multiple target fusion image features vectorially to obtain target classification image features, which fuse various image features from different sample medical images and can better meet the requirements of multi-modal fusion.

[0118] In step S402 of some embodiments, the prediction function can be a softmax function, etc. The reference classification labels can be medical labels commonly used in the medical field or others, without limitation. Specifically, a probability distribution is created on each preset reference classification label through the softmax function, and according to the probability distribution, a predicted classification value is obtained, which can reflect the possibility that the target classification image feature belongs to each reference classification label.

[0119] In step S403 of some embodiments, since the predicted classification value can reflect the possibility that the target classification image feature belongs to each reference classification label, when the predicted classification value of a certain reference classification label is higher, it indicates that the target classification image feature is more likely to belong to that reference classification label. Therefore, when screening the reference classification labels according to the predicted classification value, the reference classification label with the largest predicted classification value is selected as the predicted label.

[0120] In step S404 of some embodiments, calculate the loss between the predicted label and the original label of the sample medical image according to a preset loss function to obtain the model loss value. Specifically, the loss function can be a cross-entropy loss function, etc. The calculation process can be expressed as

[0121]

[0122] where L is the model loss value, N is the number of sample medical images, y ic is the sign function. If the original label of the sample medical image is the same as the predicted label, c takes 1; if the original label of the sample medical image is different from the predicted label, c takes 0. P ic is the predicted classification value that the sample medical image belongs to the predicted label.

[0123] In step S405 of some embodiments, backpropagation is performed on the model loss value, and the model parameters of the local classification model are adjusted according to the model loss value so that the model loss value meets the preset iteration condition. For example, the model parameters are adjusted so that the model loss value is less than the preset loss threshold, the current model parameters are extracted, the current model parameters are used as the local model parameters, and the original model parameters received by the local classification model are updated to the local model parameters.

[0124] In step S106 of some embodiments, the local model parameters are sent to the server side through the federated system, so that the server side can aggregate the local model parameters of all clients, perform weighted calculation on all the local model parameters according to the preset weight parameters to obtain the comprehensive model parameters, and calculate the average value of the comprehensive model parameters according to the total number of clients to obtain the current model parameters. The server side uses the obtained current model parameters to train the global classification model and generates model performance data. The current model performance data is compared with the previous model performance data (for example, the model performance data obtained by training the global classification model using the original model parameters). If the current model performance data is better, the previous model parameters (such as the original model parameters) are updated to the current model parameters to obtain the target model parameters. If the current model performance data is inferior to the previous model performance data, the previous model parameters are used as the target model parameters.

[0125] In step S107 of some embodiments, the target model parameters are downloaded from the server side through the federated system; wherein, the target model parameters are obtained by the server side updating the preset original model parameters according to the local model parameters sent by multiple clients.

[0126] It should be noted that according to the application scenario and the actual data situation, both the global classification model and the local recommendation model can be trained into various deep learning classification models based on the attention mechanism, such as Deep Neural Network (DNN), Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), etc., without limitation.

[0127] In step S108 of some embodiments, the local model parameters are updated according to the downloaded target model parameters, thereby training the local classification model. Specifically, when updating the local model parameters according to the downloaded target model parameters, the attention mechanism can also be introduced to train the local classification model and optimize each attention parameter of the local classification model so that the model parameters of the local classification model are optimal.

[0128] Specifically, when updating the local model parameters according to the downloaded target model parameters and training the local classification model, the original medical image is obtained again, and the model is trained according to the obtained original medical image. This training process is basically the same as the processing process of the above steps S102 to S105 and will not be elaborated here.

[0129] The training method of the classification model according to the embodiments of the present application obtains sample medical images, extracts sample image features from the sample medical images through the convolutional layer of the local classification model, and can extract relatively complex image information through a deep learning model so as to incorporate image data of different modalities into the model training process. Further, feature fusion is performed on the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features; and self-attention calculation is performed on the initial fused image features through the self-attention mechanism of the local classification model to obtain target fused image features, which can preferably perform multimodal fusion on different image features and jointly model the image features of multiple clients through the method of federated learning, improving the training effect of the model. Further, loss calculation is performed on the target fused image features through the prediction layer of the local classification model to obtain a model loss value, and the original model parameters received by the local classification model are updated to local model parameters according to the model loss value. By introducing the attention mechanism into the local classification model, the attention parameters of the local classification model can be effectively optimized, and the accuracy of the obtained local model parameters can be improved. Finally, the local model parameters are sent to the server side, the target model parameters are downloaded from the server side, and the local model parameters are updated according to the downloaded target model parameters to train the local classification model. By means of federated modeling, the problem of overfitting of the local classification model of the client can be effectively avoided. At the same time, combining multimodal fusion, the attention mechanism and federated learning can improve the training effect of the model.

[0130] Figure 5 is another optional flowchart of the training method of the classification model provided by the embodiments of the present application, which is applied to the server side. Figure 5 The method in may include but is not limited to steps S501 to S503.

[0131] Step S501: Send the preset original model parameters to the client;

[0132] Step S502: Obtain the local model parameters sent by multiple clients; wherein, the local model parameters are obtained according to the training method of the first aspect embodiment.

[0133] Step S503, training the global classification model of the server according to the local model parameters to obtain target model parameters; wherein the target model parameters are used for downloading by the client, so that the client updates the local model parameters according to the downloaded target model parameters.

[0134] In step S501 of some embodiments, the server sends preset original model parameters to the client through network communication, so that the client can initialize the local classification model through the original model parameters.

[0135] In step S502 of some embodiments, after the local classification model processes the acquired original medical image and generates local model parameters, the server obtains the local model parameters sent by multiple clients through network communication, wherein the local model parameters of the client need to be obtained from the client through communication. Acquiring data in this way can reduce communication costs.

[0136] In step S503 of some embodiments, the server can aggregate the local model parameters of all clients, and perform weighted calculation on all local model parameters according to preset weight parameters to obtain comprehensive model parameters, and average the comprehensive model parameters according to the total number of clients to obtain current model parameters. The server uses the current model parameters to train the global classification model and generate model performance data, and compares the current model performance data with the previous model performance data (for example, the model performance data obtained by training the global classification model using the original model parameters). If the current model performance data is better, the previous model parameters (such as the original model parameters) are updated to the current model parameters to obtain the target model parameters. If the current model performance data is not as good as the previous model performance data, the previous model parameters are used as the target model parameters. The target model parameters are used for downloading by the client so that the client updates the local model parameters according to the downloaded target model parameters.

[0137] The training method of the classification model in the embodiment of the present application can effectively solve the problem of overfitting in the training process of the local classification model due to too little sample data by obtaining the local model parameters output by the client. The server side can more conveniently process and analyze the local model parameters of all clients, and introduce an attention mechanism to adjust the target model parameters to determine the optimal target model parameters. Combining the attention mechanism with federated learning can improve the training effect of the model.

[0138] Figure 6 is another optional flow chart of the image classification method provided in the embodiment of the present application, which is applied to the client. Figure 6The method in [it] may include but is not limited to steps S601 to S603.

[0139] Step S601: Obtain a target medical image to be classified.

[0140] Step S602: Perform image preprocessing on the target medical image to obtain an initial medical image.

[0141] Step S603: Input the initial medical image into a local classification model for prediction processing to obtain the target category of the target medical image, where the local classification model is trained according to the training method of the first aspect embodiment.

[0142] In step S601 of some embodiments, the target medical image to be classified can be obtained in various ways such as by camera shooting or magnetic resonance imaging. The target medical image can be a three-dimensional image or a two-dimensional image, without limitation.

[0143] In step S602 of some embodiments, since the target medical image includes grayscale images, three-dimensional images, etc., the target medical image often presents as multi-modal. Therefore, it is necessary to perform image preprocessing on the target medical image to make different modal target medical images in the same dimension that can meet the requirements of classification prediction, and obtain the initial medical image. Specifically, the image preprocessing process includes one or more of adjusting the grayscale value range of the target medical image, resampling or oversampling the target medical image, and data augmentation of the target medical image, without limitation.

[0144] In step S603 of some embodiments, the initial medical image is input into the local classification model. The convolutional layer of the local classification model extracts features from the initial medical image to obtain initial image features. Then, the attention module of the local classification model calculates the attention of the initial medical image, including cross-attention calculation of the initial medical image to fuse the image features of different initial medical images to obtain the first fused image features. Then, self-attention calculation is performed on the first fused image features to extract the local information of the first fused image features to obtain local image features. Finally, the label probability of the local image features is calculated through the prediction function (such as the softmax function, etc.) of the prediction layer and the reference classification label to obtain the target probability value, and the reference classification label with the highest target probability value is selected as the target classification label. The label information of the target classification label includes the target category of the target medical image.

[0145] The image classification method according to the embodiments of the present application obtains a target medical image to be classified, and performs dimensionality transformation on the target medical image to obtain an initial medical image, which enables the image dimension of the target medical image to meet the requirements of inputting into the local classification model. Further, inputting the initial medical image into the local classification model, and performing multi-modal feature fusion and attention calculation on the initial medical image through the local classification model can better obtain the image information of the initial medical image, so as to more accurately predict the image category of the target medical image and improve the accuracy of image classification.

[0146] Please refer to Figure 7 , in some embodiments, the embodiments of the present application further provide a training device for a classification model, which is applied to a client. The client stores a pre-trained local classification model, and can implement the above classification model training method. The device includes:

[0147] A sample image acquisition module 701, configured to acquire a sample medical image;

[0148] A feature extraction module 702, configured to extract features from the sample medical image through the convolutional layer of the local classification model to obtain sample image features;

[0149] A feature fusion module 703, configured to perform feature fusion on the sample image features through the cross-attention mechanism of the local classification model to obtain initial fusion image features;

[0150] A self-attention calculation module 704, configured to perform self-attention calculation on the initial fusion image features through the self-attention mechanism of the local classification model to obtain target fusion image features;

[0151] A loss calculation module 705, configured to calculate the loss of the target fusion image features through the prediction layer of the local classification model to obtain a model loss value, and update the original model parameters received by the local classification model to local model parameters according to the model loss value;

[0152] A parameter sending module 706, configured to send the local model parameters to the server side;

[0153] A parameter downloading module 707, configured to download target model parameters from the server side;

[0154] A parameter updating module 708, configured to update the local model parameters according to the downloaded target model parameters to train the local classification model.

[0155] In some embodiments, the sample image acquisition module 701 includes:

[0156] An image acquisition unit, configured to acquire an original medical image;

[0157] A dimensionality transformation unit, configured to perform dimensionality transformation on an original medical image to obtain a sample medical image.

[0158] In some embodiments, the sample image features include first image features and second image features, and the feature fusion module 703 includes:

[0159] An encoding unit, configured to perform position encoding on the first image features to obtain a first encoded feature vector;

[0160] A feature acquisition unit, configured to obtain second image features corresponding to the first image features according to a preset image cross-relationship;

[0161] An embedding unit, configured to perform embedding processing on the first encoded feature vector through a cross-attention mechanism to obtain a first embedding vector, and perform embedding processing on the second image features through the cross-attention mechanism to obtain a second embedding vector;

[0162] A similarity calculation unit, configured to calculate the similarity between the first embedding vector and the second embedding vector through a cross-attention mechanism to obtain a feature similarity value;

[0163] A fusion unit, configured to fuse the first embedding vector and the second embedding vector according to the feature similarity value to obtain an initial fused image feature.

[0164] In some embodiments, the loss calculation module 705 includes:

[0165] A splicing unit, configured to splice the target fused image features to obtain target classification image features;

[0166] A probability calculation unit, configured to calculate the classification probability of the target classification image features through the prediction function of the prediction layer and the reference classification label to obtain a predicted classification value;

[0167] A screening unit, configured to screen the reference classification label according to the predicted classification value to obtain a predicted label;

[0168] A loss calculation unit, configured to calculate the loss between the predicted label and the original label of the sample medical image to obtain a model loss value;

[0169] An update unit, configured to update the original model parameters received by the local classification model to local model parameters according to the model loss value.

[0170] The specific implementation manner of the training device of this classification model is basically the same as the specific embodiments of the above classification model training method, and will not be elaborated here.

[0171] Please refer to Figure 8, an embodiment of the present application further provides an image classification device, which is applied to a client and can implement the above image classification method. The device includes:

[0172] A target image acquisition module 801, configured to acquire a target medical image to be classified;

[0173] An image preprocessing module 802, configured to perform image preprocessing on the target medical image to obtain an initial medical image;

[0174] A classification module 803, configured to input the initial medical image into a local classification model for prediction processing to obtain the target category of the target medical image, where the local classification model is trained by the training device according to the above embodiment.

[0175] The specific implementation manner of this image classification device is basically the same as that of the specific embodiment of the above image classification method, and will not be elaborated here.

[0176] An embodiment of the present application further provides an electronic device. The electronic device includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the above classification model training method or image classification method. The electronic device can be any intelligent terminal including a tablet computer, a vehicle-mounted computer, etc.

[0177] Please refer to Figure 9 , Figure 9 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0178] A processor 901, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0179] A memory 902, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the classification model training method or image classification method of the embodiments of the present application;

[0180] The input / output interface 903 is used to implement information input and output;

[0181] The communication interface 904 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0182] The bus 905 transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0183] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 achieve communication connections with each other inside the device through the bus 905.

[0184] The embodiment of this application also provides a storage medium. The storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned training method of the classification model or the image classification method.

[0185] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.

[0186] The training method of the classification model, the image classification method and device, equipment, and medium provided by the embodiments of the present application obtain sample medical images and extract features of the sample medical images through the convolutional layer of the local classification model to obtain sample image features, enabling the extraction of relatively complex image information through a deep learning model so as to incorporate different modalities of image data during the model training process. Further, feature fusion is performed on the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features; and self-attention calculation is performed on the initial fused image features through the self-attention mechanism of the local classification model to obtain target fused image features, which can better perform multi-modal fusion of different image features and jointly model the image features of multiple clients through the way of federated learning, improving the training effect of the model. Further, loss calculation is performed on the target fused image features through the prediction layer of the local classification model to obtain a model loss value, and the original model parameters received by the local classification model are updated to local model parameters according to the model loss value. By introducing the attention mechanism into the local classification model, the attention parameters of the local classification model can be effectively optimized, improving the accuracy of the obtained local model parameters. Finally, the local model parameters are sent to the server side, the target model parameters are downloaded from the server side, and the local model parameters are updated according to the downloaded target model parameters to train the local classification model. Through the way of federated modeling, the overfitting problem of the local classification model of the client can be effectively avoided. At the same time, combining multi-modal fusion, the attention mechanism, and federated learning can improve the training effect of the model.

[0187] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0188] Those skilled in the art can understand that Figure 1-6 the technical solutions shown in do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figure, or combine some steps, or different steps.

[0189] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0190] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0191] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0192] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (individual) of the following" or a similar expression means any combination of these items, including any combination of single items (individuals) or plural items (individuals). For example, at least one (individual) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0193] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0194] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0195] In addition, the functional units in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0196] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs and other various media that can store programs.

[0197] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.

Claims

1. A training method for a classification model, characterized in that, Applied to a client, where the client stores a pre-trained local classification model, the method includes: Obtain a sample medical image; Extract features from the sample medical image through the convolutional layer of the local classification model to obtain sample image features; Perform feature fusion on the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features; Perform self-attention calculation on the initial fused image features through the self-attention mechanism of the local classification model to obtain target fused image features; Perform loss calculation on the target fused image features through the prediction layer of the local classification model to obtain a model loss value, and update the original model parameters received by the local classification model to local model parameters according to the model loss value; Send the local model parameters to the server side; Download target model parameters from the server side; Update the local model parameters according to the downloaded target model parameters to train the local classification model; The sample image features include a first image feature and a second image feature. The step of performing feature fusion on the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features includes: Perform position encoding on the first image feature to obtain a first encoded feature vector; Obtain a second image feature corresponding to the first image feature according to a preset image cross-relationship; Perform embedding processing on the first encoded feature vector through the cross-attention mechanism to obtain a first embedding vector, and perform embedding processing on the second image feature through the cross-attention mechanism to obtain a second embedding vector; Perform similarity calculation on the first embedding vector and the second embedding vector through the cross-attention mechanism to obtain a feature similarity value; Perform fusion processing on the first embedding vector and the second embedding vector according to the feature similarity value to obtain the initial fused image features.

2. The training method according to claim 1, characterized in that The step of obtaining the sample medical image includes: Obtain an original medical image; Perform dimensionality transformation on the original medical image to obtain the sample medical image.

3. The training method according to any one of claims 1 to 2, characterized in that The step of performing loss calculation on the target fused image features through the prediction layer of the local classification model to obtain a model loss value, and updating the original model parameters received by the local classification model to local model parameters according to the model loss value includes: Perform splicing processing on the target fused image features to obtain target classification image features; Perform classification probability calculation on the target classification image features through the prediction function of the prediction layer and a reference classification label to obtain a predicted classification value; Perform screening processing on the reference classification label according to the predicted classification value to obtain a predicted label; Perform loss calculation on the predicted label and the original label of the sample medical image to obtain a model loss value; Update the original model parameters received by the local classification model to local model parameters according to the model loss value.

4. A training method for a classification model, characterized in that, Applied to a server side, the method includes: Send preset original model parameters to the client; Obtain the local model parameters sent by multiple said clients; wherein, the local model parameters are obtained according to the training method described in any one of claims 1 to 3; Train the global classification model of the server side according to the local model parameters to obtain target model parameters; wherein, the target model parameters are used for the client to download, so that the client updates the local model parameters according to the downloaded target model parameters.

5. An image classification method, characterized in that, Applied to a client, the image classification method includes: Obtain a target medical image to be classified; Perform image preprocessing on the target medical image to obtain an initial medical image; Input the initial medical image into the local classification model for prediction processing to obtain the target category of the target medical image, wherein the local classification model is trained according to the training method described in any one of claims 1 to 3.

6. A training device for a classification model, which is applied to a client. The client stores a pre-trained local classification model, and is characterized in that The device includes: A sample image acquisition module for acquiring sample medical images; A feature extraction module for extracting features from the sample medical image through the convolutional layer of the local classification model to obtain sample image features; A feature fusion module for fusing the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features; A self-attention calculation module for performing self-attention calculation on the initial fused image features through the self-attention mechanism of the local classification model to obtain target fused image features; A loss calculation module for calculating the loss of the target fused image features through the prediction layer of the local classification model to obtain a model loss value, and updating the original model parameters received by the local classification model to local model parameters according to the model loss value; A parameter sending module for sending the local model parameters to the server side; A parameter downloading module for downloading target model parameters from the server side; A parameter updating module for updating the local model parameters according to the downloaded target model parameters to train the local classification model; The sample image features include a first image feature and a second image feature, and the step of fusing the sample image features through the cross-attention mechanism of the local classification model to obtain initial fused image features includes: Perform position encoding on the first image feature to obtain a first encoded feature vector; Obtain a second image feature corresponding to the first image feature according to a preset image cross relationship; Perform embedding processing on the first encoded feature vector through the cross-attention mechanism to obtain a first embedding vector, and perform embedding processing on the second image feature through the cross-attention mechanism to obtain a second embedding vector; Perform similarity calculation on the first embedding vector and the second embedding vector through the cross-attention mechanism to obtain a feature similarity value; Fuse the first embedding vector and the second embedding vector according to the feature similarity value to obtain the initial fused image features.

7. An image classification device, applied to a client, characterized in that The image classification device includes: A target image acquisition module for acquiring a target medical image to be classified; An image preprocessing module for preprocessing the target medical image to obtain an initial medical image; A classification module for inputting the initial medical image into a local classification model for prediction processing to obtain the target category of the target medical image, wherein the local classification model is trained by the training device according to Claim 6.

8. An electronic device, characterized in that, The electronic device includes a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for realizing connection communication between the processor and the memory. When the program is executed by the processor, it implements the training method according to any one of Claims 1 to 3, or the training method according to Claim 4, or the steps of the image classification method according to Claim 5.

9. A storage medium, which is a computer-readable storage medium for computer-readable storage, and is characterized in that, The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the training method according to any one of Claims 1 to 3, or the training method according to Claim 4, or the steps of the image classification method according to Claim 5.

Citation Information

Patent Citations

  • Traffic signal lamp sensing method and device, equipment and storage medium

    CN114694123A

  • Federal learning method and device based on self-supervision, equipment and storage medium

    CN114792139A