Image classification method and apparatus, and electronic device, computer-readable storage medium and computer program product

By modal conversion and modal fusion of classified images, multimodal vectors are generated, which solves the problem of low classification accuracy caused by low information in single-modal images, and achieves higher image classification accuracy.

WO2025118783A1PCT designated stage expired Publication Date: 2025-06-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/120338
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-05
Filing Date
2024-09-23
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

In the information classification, the classification accuracy is low due to the low amount of single-modal image information.

Method used

By performing modal conversion of different modalities on the classified images, multimodal information is obtained, and modal fusion is performed to generate multimodal vectors for image classification.

Benefits of technology

The accuracy of image classification is improved, and the amount of information of vectors to be classified is enhanced through the fusion of multimodal information, and the accuracy of category prediction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024120338_12062025_PF_FP_ABST
    Figure CN2024120338_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are an image classification method and apparatus, and an electronic device, a computer-readable storage medium and a computer program product. The method comprises: acquiring an image to be classified which comprises a plurality of sub-images to be classified, and performing modal conversion of different modalities on said image, so as to obtain modal information of said image which corresponds to each modality; performing modal fusion on the modal information, so as to obtain a multi-modal vector corresponding to said image; encoding said image on the basis of the multi-modal vector, so as to obtain sub-vectors to be classified which correspond to said sub-images on a one-to-one basis; and classifying said image on the basis of said sub-vectors which correspond to said sub-images, so as to obtain the category of said image.
Need to check novelty before this filing date? Find Prior Art

Description

Image classification method, device, electronic device, computer-readable storage medium, and computer program product

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on the Chinese patent application with application number 202311661392.1 and application date December 5, 2023, and claims the priority of the Chinese patent application. The entire content of the Chinese patent application is hereby introduced into this application as a reference. Technical Field

[0003] The present application relates to the field of computer technology, and in particular to an image classification method, device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0004] Artificial Intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0005] In related technologies, information classification is usually performed directly based on the to-be-classified vector of the to-be-classified image to obtain the category of the to-be-classified image. Since the multiple to-be-classified sub-images included in the to-be-classified image will affect the category of the to-be-classified image, and the amount of information corresponding to the single-modal to-be-classified image is low, the classification accuracy of the information is low.

[0006] Summary of the Invention

[0007] The embodiments of the present application provide an image classification method, an image classification model training method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which can effectively improve the classification accuracy of information.

[0008] The technical solution of the embodiment of the present application is implemented as follows:

[0009] The present invention provides an image classification method, including:

[0010] Acquiring an image to be classified comprising a plurality of sub-images to be classified, and performing modality conversion of different modalities on the image to be classified to obtain modality information corresponding to each modality of the image to be classified;

[0011] Performing modal fusion on each of the modal information to obtain a multimodal vector corresponding to the image to be classified;

[0012] Encoding the image to be classified based on the multimodal vector to obtain sub-vectors to be classified corresponding to each sub-image to be classified;

[0013] The image to be classified is classified based on the sub-vector to be classified corresponding to each sub-image to be classified to obtain the category of the image to be classified.

[0014] The present invention provides an image classification device, comprising:

[0015] an acquisition module configured to acquire an image to be classified comprising a plurality of sub-images to be classified, and perform modality conversion of the image to be classified into different modalities to obtain modality information corresponding to each modality of the image to be classified;

[0016] a modality fusion module configured to perform modality fusion on each of the modal information to obtain a multimodal vector corresponding to the image to be classified;

[0017] an encoding module configured to encode the image to be classified based on the multimodal vector to obtain a sub-vector to be classified corresponding to each sub-image to be classified;

[0018] The category prediction module is configured to classify the image to be classified based on the sub-vector to be classified corresponding to each sub-image to be classified, so as to obtain the category of the image to be classified.

[0019] The present invention provides a method for training an image classification model, including:

[0020] Acquire an image sample to be classified including a plurality of sub-samples to be classified, and a sample label carried by the image sample to be classified;

[0021] Encoding the image sample to be classified to obtain a sample vector corresponding to the image sample to be classified, wherein the sample vector includes a subsample vector corresponding to each subsample to be classified;

[0022] For each of the sub-samples to be classified, determining the sample label as the initial label of the sub-sample to be classified, and calling the classification model to perform category prediction on the sub-sample to be classified based on the sub-sample vector corresponding to the sub-sample to be classified, to obtain a predicted probability of the sub-sample to be classified;

[0023] Determining target labels for the subsamples to be classified based on the predicted probabilities of the subsamples to be classified;

[0024] Training the classification model based on the predicted probability, the initial label, and the target label to obtain a target classification model;

[0025] The target classification model is used to perform category prediction on each sub-image to be classified in the image to be classified, and obtain the prediction probability of each sub-image to be classified.

[0026] The present invention provides a training device for an image classification model, comprising:

[0027] A sample acquisition module is configured to acquire an image sample to be classified, comprising a plurality of sub-samples to be classified, and a sample label carried by the image sample to be classified;

[0028] a sample encoding module configured to encode the image sample to be classified to obtain a sample vector corresponding to the image sample to be classified, wherein the sample vector includes a subsample vector corresponding to each subsample to be classified;

[0029] a sample prediction module configured to, for each of the sub-samples to be classified, determine the sample label as the initial label of the sub-sample to be classified, and call the classification model to perform category prediction on the sub-sample to be classified based on the sub-sample vector corresponding to the sub-sample to be classified, to obtain a predicted probability of the sub-sample to be classified;

[0030] a label determination module, configured to determine a target label for each of the sub-samples to be classified based on the predicted probability of each of the sub-samples to be classified;

[0031] The target training module trains the classification model in combination with the predicted probability, the initial label and the target label to obtain a target classification model; wherein the target classification model is used to perform category prediction on each sub-image to be classified in the image to be classified to obtain the predicted probability of each sub-image to be classified.

[0032] An embodiment of the present application provides an electronic device, including:

[0033] a memory configured to store computer-executable instructions or computer programs;

[0034] The processor is configured to implement the image classification method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.

[0035] An embodiment of the present application provides an electronic device, including:

[0036] a memory configured to store computer-executable instructions or computer programs;

[0037] The processor is configured to execute the computer-executable instructions or computer programs stored in the memory to implement the training method of the image classification model provided in the embodiment of the present application.

[0038] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which is configured to cause a processor to execute and implement the image classification method provided in the embodiment of the present application.

[0039] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which is configured to cause a processor to execute and implement the training method of the image classification model provided in the embodiment of the present application.

[0040] The present invention provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer program or computer-executable instructions, causing the electronic device to perform the image classification method described in the present invention.

[0041] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the image classification model training method described in the present invention.

[0042] The embodiments of the present application have the following beneficial effects:

[0043] By performing modal conversion of different modalities on an image to be classified including a plurality of sub-images to be classified, modal information of each modality corresponding to the image to be classified is obtained, and each modal information is modally fused to obtain a multimodal vector corresponding to the image to be classified, and based on the multimodal vector, the image to be classified is encoded to obtain a vector to be classified corresponding to the image to be classified, and for each sub-image to be classified, based on the sub-vector to be classified corresponding to the sub-image to be classified, a category prediction is performed on the sub-image to be classified to obtain a category prediction probability corresponding to the sub-image to be classified, and based on the category prediction probability corresponding to each sub-image to be classified, the image to be classified is classified to obtain a category of the image to be classified. In this way, by performing modal fusion on the modal information obtained by performing modal conversion of the image to be classified into different modalities, the obtained multimodal vector can accurately reflect the modal characteristics of the image to be classified from different modalities, and based on the multimodal vector, the vector to be classified is obtained by encoding the image to be classified, thereby effectively improving the information content of the vector to be classified, and based on the vector to be classified corresponding to each sub-image to be classified, the category of the sub-image to be classified is predicted. Since the vector to be classified on which the category prediction is based has a more comprehensive amount of information, the category prediction probability obtained by the prediction is more accurate. Through the accurate category prediction probability, the image to be classified is classified, thereby effectively improving the classification accuracy of the information. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] FIG1 is a schematic diagram of the architecture of an information classification system provided in an embodiment of the present application;

[0045] FIG2 is a schematic diagram of the structure of an electronic device for information classification provided in an embodiment of the present application;

[0046] FIG3 is a schematic diagram of the structure of an electronic device for training an image classification model provided in an embodiment of the present application;

[0047] FIG4 is a flowchart of an image classification method according to an embodiment of the present application;

[0048] FIG5 is a second flow chart of the image classification method provided in an embodiment of the present application;

[0049] FIG6 is a schematic diagram of the principle of modality fusion provided in an embodiment of the present application;

[0050] FIG7 is a third flow chart of the image classification method provided in an embodiment of the present application;

[0051] FIG8 is a fourth flow chart of the image classification method provided in an embodiment of the present application;

[0052] FIG9 is a flowchart of a training method for an image classification model according to an embodiment of the present application;

[0053] FIG10 is a schematic diagram of the structure of a classification model provided in an embodiment of the present application;

[0054] FIG11 is a second flow chart of a method for training an image classification model according to an embodiment of the present application;

[0055] FIG12 is a third flow chart of the training method of the image classification model provided in an embodiment of the present application;

[0056] FIG13 is a fourth flow chart of a method for training an image classification model according to an embodiment of the present application;

[0057] FIG14 is a schematic diagram showing the principle of the image classification method provided in an embodiment of the present application;

[0058] FIG15 is a first schematic diagram of the classification effect of the image classification method provided in an embodiment of the present application;

[0059] FIG16 is a second schematic diagram of the classification effect of the image classification method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0060] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0061] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0062] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0064] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0065] 1) Artificial Intelligence (AI): This is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI disciplines. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0066] 2) Computer Vision (CV): Computer vision is the science of enabling machines to "see." Specifically, it refers to machine vision, where cameras and computers replace the human eye in identifying and measuring objects, and further image processing is performed to transform the computer's image into one more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can extract information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the field of vision, such as the Swin Transformer, ViT, V-MOE, and MAE, can be fine-tuned to quickly and widely apply to specific downstream tasks. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / action recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0067] 3) Machine Learning (ML): This is a multidisciplinary interdisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning. Pretrained models are the latest development in deep learning, integrating these techniques.

[0068] 4) Convolutional Neural Networks (CNN): These are a type of feedforward neural network (FNN) that incorporates convolutional computations and possesses a deep structure. They are a representative algorithm for deep learning. CNNs possess representation learning capabilities and can perform shift-invariant classification of input images based on their hierarchical structure.

[0069] 5) In response to: used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0070] 6) Convolutional Layer: Each convolutional layer in a convolutional neural network consists of several convolutional units, whose parameters are optimized through a backpropagation algorithm. The purpose of the convolution operation is to extract different features from the input. The first convolutional layer may only extract low-level features such as edges, lines, and corners. More layers of the network can iteratively extract more complex features from these low-level features.

[0071] 7) Pooling Layer: After encoding in the convolutional layer, the output feature map is passed to the pooling layer for feature selection and information filtering. The pooling layer contains a pre-defined pooling function, which replaces the result of a single point in the feature map with the feature map statistics of its adjacent area. The pooling layer selects the pooling area in the same way as the convolution kernel scans the feature map, controlled by the pooling size, stride, and padding.

[0072] 8) Fully-Connected Layer: A fully-connected layer in a convolutional neural network is equivalent to a hidden layer in a traditional feedforward neural network. It is located at the end of the hidden layer of a convolutional neural network and only transmits signals to other fully-connected layers. Feature maps lose their spatial topology in fully-connected layers, becoming flattened into vectors and passing through activation functions.

[0073] 9) Immune receptors: The immune system is one of the most important life-sustaining systems in an organism, capable of detecting and eliminating pathogens and abnormal cells in the body. The main components of the immune system are white blood cells and molecules closely related to them. Immune receptors are an important component of the immune system and are a class of proteins specifically used to recognize and bind to antigen molecules. The main function of immune receptors is to detect and bind to antigens, which is a key step for the immune system to identify and eliminate pathogens and abnormal cells in the body. Immune receptors also have the function of regulating immune responses. Immune responses must be carried out at the appropriate time and extent to avoid damage to host tissues. Immune receptors are a class of proteins that can recognize and bind to antigen molecules, including T cell receptors and B cell receptors. Their main function is to detect and eliminate pathogens and abnormal cells in the body. They can also regulate the duration and intensity of immune responses to maintain immune balance.

[0074] 10) Immune Repertoire (IR): Also known as the immune receptor repertoire, it is the sum of all functionally diverse T cells and B cells in an individual's immune system at a given time. The diversity of the immune repertoire is specifically reflected in the diversity of B cell receptors (BCRs) or T cell receptors (TCRs) formed by the permutations and combinations of V, D, J, and C gene segments. BCRs consist of two heavy chains and a light chain, while TCR chains consist of α and β chains. The human immune system is composed of innate immunity and adaptive immunity. Adaptive immunity is the immune response generated by T cells and B cells after contact with pathogens (antigens), which can recognize and eliminate specific pathogens. The adaptive immune receptor repertoire (abbreviated as immune repertoire), as a collection of T cell receptors (TCRs) or B cell receptors (BCRs), records past and ongoing adaptive immune responses. The state of the immune repertoire can be used to reflect the state of immune activity and an individual's response to infectious diseases, autoimmune diseases, and tumor-related pathogens. A large adaptive immune receptor repertoire (AIRR) consists of a large number of adaptive immune receptors (AIR, TCR or BCR). Assume that, given N AIRR, its mathematical representation is {IR1, IR2, ..., IR N Each AIRR contains M AIRs, expressed as It should be noted that the M in the immune repertoire of different individuals is very different. At the same time, the labels of the N immune repertoires are defined as {Y1, Y2, ..., Y N}, where Y∈{0,…,C}, C represents the number of categories, for example, if it is a two-class problem, then C=1. In addition, each AIR has a corresponding frequency, which is expressed as Frequency represents the strength of the immune response to a specific antigen. In the embodiment of the present application, a mapping function Y=F(IR) can be established to convert the immune repertoire IR into the immune state Y.

[0075] 11) Gene: The complete nucleotide sequence required to produce a polypeptide chain or functional RNA. Genes support the basic structure and function of life. They store all the information of life. The interdependence of environment and genetics determines important physiological processes such as reproduction, cell division, and protein synthesis. All life phenomena, including the birth, growth, decline, disease, aging, and death of organisms, are related to genes. They are also the internal factors that determine life and health. Therefore, genes have dual properties: materiality (form of existence) and information (fundamental nature).

[0076] 12) Modal: Modality generally refers to a single mode of perception or expression. It can refer to any single method of information transmission or interaction. In the fields of information technology, human-computer interaction, and communications, modality can refer to different forms of information expression, such as text, images, sound, and touch. For example, in a communication process, language (oral or written) can be considered a modality. Modality emphasizes a single channel of information or interaction, such as communication through text alone.

[0077] 13) Multimodal: This refers to information processing or interaction involving two or more different modalities. In multimodal interaction, information can be transmitted simultaneously or sequentially through multiple sensory channels, such as vision (images, video), hearing (speech, music), and touch (vibration, pressure). Multimodal interaction aims to leverage the strengths of different modalities to improve the efficiency, accuracy, and user experience of information transmission.

[0078] 14) Medical images: Medical images refer to images captured using specific imaging technologies and containing medical information. Medical images are generally used to record and analyze the physiological characteristics, structural properties, or behavioral patterns of medical subjects. The medical information contained in medical images is diverse and can be macroscopic structures of medical subjects, such as body shape, color, texture, etc., or microscopic structures, such as cell structure, tissue arrangement, chemical molecular structure, and drug molecular structure. In the field of medical recognition, common medical information includes but is not limited to palm prints, fingerprints, irises, facial features, amino acid molecular structures, cell molecular structures, chemical molecular structures, and drug molecular structures. The type of image to be classified can be used to indicate the category of medical information contained in the image to be classified.

[0079] During the implementation of the embodiments of this application, the applicant discovered that the related technology has the following problems:

[0080] See Table 1 below, which is a schematic table comparing the performance of the embodiments of the present application and related technologies on datasets 1 and 2, respectively.

[0081] Table 1 Schematic comparison of performance of the embodiment of the present application and related technologies

[0082] As can be seen from the table above, the present application is significantly superior to Related Technology 1 and Related Technology 2 in all indicators. In related technologies, for information classification, the image to be classified is usually classified directly based on the to-be-classified vector of the image to be classified to obtain the category of the image to be classified. Since the multiple to-be-classified sub-images included in the image to be classified will affect the category of the image to be classified, and the amount of information corresponding to the unimodal image to be classified is relatively low, the classification accuracy of the information is relatively low.

[0083] Embodiments of the present application provide an image classification method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can effectively improve the classification accuracy of information. The following describes an exemplary application of the information classification system provided by embodiments of the present application.

[0084] See Figure 1, which is a schematic diagram of the architecture of the information classification system 100 provided in an embodiment of the present application. The terminal (terminal 400 is shown as an example) is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0085] The terminal 400 is configured to allow the user to use a client 410 to display the category of the image to be classified on a graphical interface 410-1 (graphic interface 410-1 is shown as an example). The terminal 400 and the server 200 are connected to each other via a wired or wireless network.

[0086] In some embodiments, the server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, a car terminal, etc., but is not limited to this. The electronic device provided in the embodiment of the present application can be implemented as a terminal or as a server. The terminal and the server can be directly or indirectly connected by wired or wireless communication, which is not limited in the embodiment of the present application.

[0087] In some embodiments, the server 200 obtains the image to be classified, and performs modal conversion of different modalities on the image to be classified to obtain modal information, and encodes the image to be classified based on the modal information to obtain the vector to be classified, and performs category prediction on the sub-image to be classified based on the sub-vector to be classified to obtain the category prediction probability, and classifies the image to be classified based on the prediction probability of each category to obtain the category of the image to be classified, and sends the category to the terminal 400.

[0088] In other embodiments, the terminal 400 obtains the image to be classified, and performs modal conversion of different modalities on the image to be classified to obtain modal information, and encodes the image to be classified based on the modal information to obtain the vector to be classified, and performs category prediction on the sub-image to be classified based on the sub-vector to be classified to obtain the category prediction probability, and classifies the image to be classified based on the prediction probability of each category to obtain the category of the image to be classified, and sends the category to the server 200.

[0089] Referring to Figure 2, Figure 2 is a schematic diagram of the structure of an electronic device 500 for information classification provided in an embodiment of the present application, wherein the electronic device 500 shown in Figure 2 can be the server 200 or the terminal 400 in Figure 1, and the electronic device 500 shown in Figure 2 includes: at least one processor 430, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together through a bus system 440. It can be understood that the bus system 440 is configured to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus system 440 in Figure 2.

[0090] The processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0091] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 430.

[0092] The memory 450 includes a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0093] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0094] Operating system 451, including system programs configured to handle various basic system services and perform hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., configured to implement various basic services and handle hardware-based tasks;

[0095] The network communication module 452 is configured to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).

[0096] In some embodiments, the image classification device provided in the embodiments of the present application can be implemented in software. FIG2 shows an image classification device 455 stored in memory 450. The image classification device 455 can be software in the form of a program or plug-in, and includes the following software modules: an acquisition module 4551, a modality fusion module 4552, an encoding module 4553, and a category prediction module 4554. These modules are logical and can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.

[0097] Referring to FIG3 , FIG3 is a schematic diagram of the structure of an electronic device 600 for training an image classification model provided in an embodiment of the present application, wherein the electronic device 600 shown in FIG3 may be the server 200 or the terminal 400 in FIG1 , and the electronic device 600 shown in FIG3 includes: at least one processor 530, a memory 550, and at least one network interface 520. The various components in the electronic device 600 are coupled together via a bus system 540. It is understood that the bus system 540 is configured to enable connection and communication between these components. In addition to including a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all the various buses are labeled as the bus system 540 in FIG3 .

[0098] The processor 530 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0099] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 530.

[0100] The memory 550 includes a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.

[0101] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0102] Operating system 551, including system programs configured to handle various basic system services and perform hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., configured to implement various basic services and handle hardware-based tasks;

[0103] The network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).

[0104] In some embodiments, the training device for the image classification model provided in the embodiments of the present application can be implemented in software. FIG3 shows a training device 555 for the image classification model stored in the memory 550. The training device 555 can be software in the form of a program or plug-in, and includes the following software modules: a sample acquisition module 5551, a sample encoding module 5552, a sample prediction module 5553, a label determination module 5554, and a target training module 5555. These modules are logical and can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.

[0105] In other embodiments, the image classification device provided in the embodiments of the present application can be implemented in hardware. As an example, the image classification device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the image classification method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs) or other electronic components.

[0106] In some embodiments, the terminal or server can implement the image classification method provided in the embodiments of the present application by running a computer program or computer executable instructions. For example, the computer program can be a native program (for example, a dedicated information classification program) or a software module in the operating system, for example, an information classification module that can be embedded in any program (such as an instant messaging client, a photo album program, an electronic map client, a navigation client); for example, it can be a native application (APP, Application), that is, a program that needs to be installed in the operating system to run. In short, the above-mentioned computer program can be any form of application, module or plug-in.

[0107] In some embodiments, the embodiments of the present application can be implemented in the following manner: obtaining information to be classified including multiple sub-information to be classified, and performing modal conversion of the information to be classified into different modalities to obtain modal information corresponding to each modality of the information to be classified; performing modal fusion on each modal information to obtain a multimodal vector corresponding to the information to be classified; encoding the information to be classified based on the multimodal vector to obtain a sub-vector to be classified corresponding to each sub-information to be classified; for each sub-information to be classified, classifying the information to be classified based on the sub-vector to be classified corresponding to the sub-information to be classified to obtain a category of the information to be classified.

[0108] In some embodiments, the information to be classified may be an image to be classified. The image classification method provided in the embodiments of the present application will be described below using the information to be classified being an image to be classified as an example.

[0109] The image classification method provided in the embodiment of the present application will be described in conjunction with the exemplary application and implementation of the server or terminal provided in the embodiment of the present application.

[0110] Refer to Figure 4, which is a flow chart of the image classification method provided in an embodiment of the present application. It will be explained in conjunction with steps 101 to 105 shown in Figure 4. The image classification method provided in an embodiment of the present application can be implemented by a server or a terminal alone, or by a server and a terminal in collaboration. The following will be explained using the server alone as an example.

[0111] In step 101 , an image to be classified including a plurality of sub-images to be classified is obtained, and modality conversion of different modalities is performed on the image to be classified to obtain modality information corresponding to each modality of the image to be classified.

[0112] In some embodiments, the image to be classified includes multiple sub-images to be classified, and the sub-images to be classified are constituent elements of the image to be classified. For example, the image to be classified A includes sub-images A1, A2, and A3.

[0113] In some embodiments, modality indicates the perceptual channel of information. Different modal information corresponds to different perceptual channels. For example, the video modality information corresponds to the video perceptual channel, while the audio modality information corresponds to the audio perceptual channel. The above-mentioned modality conversion is the process of converting the image to be classified in the original modality into the modal information of the target modality.

[0114] In some embodiments, the above-mentioned modal conversion of different modalities of the image to be classified to obtain the modal information of each modality corresponding to the image to be classified can be achieved as follows: when the image to be classified is multimodal information, each target modality in the image to be classified is obtained, and for each target modality, the modal information of the target modality is extracted from the image to be classified to obtain the modal information corresponding to the target modality of the image to be classified; when the image to be classified is single-modal information, multiple target modalities are obtained, and for each target modality, the modal conversion of the target modality is performed on the image to be classified to obtain the modal information corresponding to the target modality of the image to be classified.

[0115] As an example, in a text classification application scenario, an image to be classified can be obtained by performing image conversion on the text to be classified (single modal information). The text to be classified includes multiple sub-texts to be classified. For example, if the text to be classified is "A fingerprint is a whorl, and the coordinates of the starting point, end point, bifurcation, and intersection of the skin ridge are respectively:...", the multiple sub-texts to be classified included in the text to be classified may be "whorl," "coordinates of the starting point of the skin ridge," "coordinates of the end point," "coordinates of the bifurcation," etc. A modal conversion of the audio modality of the medical image to be classified is performed to obtain the audio to be classified in the audio modality corresponding to the image to be classified. A modal conversion of the text modality of the image to be classified is performed to obtain the text to be classified in the text modality corresponding to the image to be classified.

[0116] As an example, in an application scenario of an immune receptor library to be classified, the image to be classified may be an image containing the immune receptor library to be classified, wherein the immune receptor library to be classified includes multiple immune receptors, and different immune receptors correspond to different image regions in the image to be classified. The image regions corresponding to the immune receptors are sub-images to be classified in the image to be classified. The immune receptor library to be classified in the image to be classified is modally converted from an amino acid sequence modality to obtain amino acid sequence modality information of the immune receptor library to be classified including the amino acid sequence modality. The immune receptor library to be classified in the image to be classified is modally converted from a gene segment modality to obtain gene segment modality information of the immune receptor library to be classified including the gene segment modality.

[0117] As an example, in the application scenario of video classification, the image to be classified can be obtained by splicing multiple video frames to be classified in the video to be classified (multimodal information), and the image to be classified includes multiple video frames to be classified. For example, the image to be classified is modally converted to an audio mode to obtain the audio to be classified of the audio mode corresponding to the image to be classified, and the image to be classified is modally converted to a video mode to obtain the video to be classified of the video mode corresponding to the image to be classified.

[0118] As an example, in an application scenario of video classification, a medical image to be classified including multiple video frames to be classified is obtained, and modal conversion of different modalities is performed on the image to be classified to obtain modal information corresponding to each modality of the image to be classified; each modal information is modally fused to obtain a multimodal vector corresponding to the image to be classified; based on the multimodal vector, the image to be classified is encoded to obtain a vector to be classified corresponding to the image to be classified, and the vector to be classified includes sub-vectors to be classified corresponding to each video frame to be classified; for each video frame to be classified, based on the sub-vectors to be classified corresponding to the video frame to be classified, the category of the video frame to be classified is predicted to obtain the category prediction probability corresponding to the video frame to be classified; based on the category prediction probability corresponding to each video frame to be classified, the image to be classified is classified to obtain the category of the image to be classified.

[0119] As an example, in an application scenario of an immune receptor library, an image to be classified of an immune receptor library to be classified including multiple immune receptors to be classified is obtained, and modal conversion of different modalities is performed on the image to be classified including the immune receptor library to be classified to obtain modal information of each modality corresponding to the immune receptor library to be classified; each modal information is modally fused to obtain a multimodal vector corresponding to the immune receptor library to be classified; based on the multimodal vector, the immune receptor library to be classified is encoded to obtain a vector to be classified corresponding to the immune receptor library to be classified, and the vector to be classified includes a sub-vector to be classified corresponding to each immune receptor to be classified; for each immune receptor to be classified, based on the sub-vector to be classified corresponding to the immune receptor to be classified, a category prediction is performed on the immune receptor to be classified to obtain a category prediction probability corresponding to the immune receptor to be classified; based on the category prediction probability corresponding to each immune receptor to be classified, the image to be classified including the immune receptor library to be classified is classified to obtain a category of the immune receptor library to be classified.

[0120] In some embodiments, the image to be classified of the immune receptor library to be classified, which includes multiple immune receptors to be classified, is a medical image to be classified. A medical image refers to an image captured by a specific imaging technology and containing medical information. Medical images are generally used to record and analyze the physiological characteristics, structural properties or behavioral patterns of medical subjects. The medical information contained in medical images is diverse and can be the macroscopic structure of the medical subject, such as body shape, color, texture, etc., or the microscopic structure, such as cell structure, tissue arrangement, chemical molecular structure, drug molecular structure, etc. In the field of medical recognition, common medical information includes but is not limited to palm prints, fingerprints, irises, facial features, amino acid molecular structures, cell molecular structures, chemical molecular structures, drug molecular structures, etc. The type of the image to be classified can be used to indicate the category of the medical information contained in the image to be classified.

[0121] As an example, in an application scenario of text classification, a medical text to be classified including multiple sub-medical texts to be classified is obtained, and the medical text to be classified is converted into an image to be classified, and modal conversion of different modalities is performed on the image to be classified to obtain modal information corresponding to each modality of the image to be classified; each modal information is modally fused to obtain a multimodal vector corresponding to the image to be classified; based on the multimodal vector, the image to be classified is encoded to obtain a vector to be classified corresponding to the image to be classified, and the vector to be classified includes a sub-vector to be classified corresponding to each sub-image to be classified; for each sub-image to be classified, based on the sub-vector to be classified corresponding to the sub-image to be classified, a category prediction is performed on the sub-image to be classified to obtain a category prediction probability corresponding to the sub-image to be classified; based on the category prediction probability corresponding to each sub-image to be classified, the image to be classified is classified to obtain a category of the image to be classified.

[0122] In this way, by obtaining an image to be classified including multiple sub-images to be classified, and performing modal conversion of different modalities on the image to be classified, the modal information of the image to be classified corresponding to each modality is obtained, thereby realizing the modal conversion of the image to be classified, so that the image to be classified can be multi-dimensionally characterized from multiple different modal dimensions, effectively enhancing the amount of information in the image to be classified.

[0123] In step 102, each modal information is modally fused to obtain a multimodal vector corresponding to the image to be classified.

[0124] In some embodiments, the above-mentioned modal fusion refers to the process of converting multiple modal information into a multi-modal vector.

[0125] In some embodiments, modal fusion refers to the process of integrating information from different modalities (such as text, images, and audio) within a multimodal data processing system. Each modality typically contains different aspects or characteristics of the data, and the goal of modal fusion is to combine these different information sources to produce a more comprehensive and richer data representation.

[0126] In some embodiments, the multimodal vector includes multimodal sub-vectors corresponding to each modal information, see Figure 5, which is a second flow chart of the image classification method provided in an embodiment of the present application. Step 102 shown in Figure 4 can be implemented through steps 1021 to 1025 shown in Figure 5.

[0127] In step 1021, each modal information is encoded respectively to obtain a modal vector corresponding to each modal information.

[0128] In some embodiments, the above step 1021 can be implemented as follows: for each modal information, calling the encoding layer, encoding the modal information, and obtaining the modal vector corresponding to the modal information.

[0129] In some embodiments, the encoding layer includes a convolutional layer, which is composed of several convolutional units. The parameters of each convolutional unit are optimized via a backpropagation algorithm. The purpose of the convolution operation is to extract different features of the input. The first convolutional layer may only extract low-level features such as edges, lines, and corners. A network with more layers can iteratively extract more complex features from these low-level features.

[0130] As an example, see Figure 6, which is a schematic diagram of the principle of modal fusion provided by an embodiment of the present application. The modal information 61 shown in Figure 6 is gene modal information, and the modal information 62 shown in Figure 6 is sequence modal information. The modal information 61 is encoded to obtain the modal vector h corresponding to the modal information 61. g, encode the modal information 62 and obtain the modal vector h corresponding to the modal information 62 s .

[0131] In some embodiments, after the above step 1021, the following steps 1022 to 1025 can be performed respectively for each modal information.

[0132] In step 1022 , the modal information is determined as target modal information, and each modal information except the target modal information is determined as reference modal information.

[0133] As an example, the modal information corresponding to each modality of the image to be classified includes modal information A, modal information B and modal information C. When processing modal information A, modal information A is determined as the target modal information, and each modal information (modal information B and modal information C) except the target modal information A is determined as the reference modal information.

[0134] Continuing with the example, when processing modal information B, modal information B is determined as target modal information, and each modal information (modal information A and modal information C) except the target modal information B is determined as reference modal information.

[0135] Continuing with the above example, when processing modal information C, modal information C is determined as target modal information, and each modal information (modal information A and modal information B) except the target modal information C is determined as reference modal information.

[0136] As an example, referring to FIG. 6 , when processing is performed on modal information 61 , modal information 61 is determined as target modal information, and modal information 62 is determined as reference modal information.

[0137] As an example, referring to FIG. 6 , when processing is performed on modal information 62 , the modal information 62 is determined as target modal information, and the modal information 61 is determined as reference modal information.

[0138] In step 1023, a vector conversion is performed on the modal vector corresponding to the target modal information to obtain a target modal vector corresponding to the target modal information.

[0139] In some embodiments, the above-mentioned vector conversion is also called feature linear transformation, which refers to the process of performing linear transformation on the feature to be processed (modal vector) to obtain the target vector (target modal vector).

[0140] In some embodiments, vector transformation refers to the process of performing a series of mathematical operations on the processed features (i.e., modal vectors) in the fields of data processing and machine learning to change their form or spatial position, thereby obtaining a new vector representation (i.e., target modal vector). This process usually involves linear transformations in linear algebra, which can help improve the separability of data, reduce dimensions, or adapt to specific model requirements.

[0141] As an example, referring to FIG6 , when the target modal information is modal information 61, the modal vector h corresponding to the modal information 61 is g Perform vector conversion to obtain the target modal vector h′ corresponding to the target modal information 61 g .

[0142] As an example, referring to FIG6 , when the target modal information is modal information 62, the modal vector h corresponding to the modal information 62 is s Perform vector conversion to obtain the target modal vector h′ corresponding to the target modal information 62 s .

[0143] In some embodiments, the above-mentioned step 1023 can be implemented as follows: obtain a target conversion vector corresponding to the target modal information, and a target correction vector corresponding to the target modal information, the modal vector corresponding to the target modal information having the same vector dimensions as the corresponding target conversion vector and the target correction vector; multiplying the modal vector corresponding to the target modal information and the target conversion vector to obtain the target vector corresponding to the target modal information; adding the target conversion vector and the target correction vector to obtain the correction vector corresponding to the target modal information; and performing vector activation on the target correction vector to obtain the target modal vector corresponding to the target modal information.

[0144] As an example, referring to FIG6 , when the target modal information is modal information 61 , the expression of the target modal vector corresponding to the target modal information can be: h′ g =ReLu(W g *h g +b g ) (1)

[0145] Where h′ g It is used to indicate the target modal vector corresponding to the modal information 61, h g It is used to indicate the modal vector corresponding to the modal information 61, W g Used to indicate the target transformation vector, b g Used to indicate the target correction vector.

[0146] As an example, referring to FIG6 , when the target modal information is modal information 62 , the expression of the target modal vector corresponding to the target modal information may be: h′g =ReLu(W s *h g +b s ) (2)

[0147] Where h′ g It is used to indicate the target modal vector corresponding to the modal information 62, h g It is used to indicate the modal vector corresponding to the modal information 62, W g Used to indicate the target transformation vector, b g Used to indicate the target correction vector.

[0148] In some embodiments, vector activation refers to the process of applying a nonlinear function to a vector in machine learning and deep learning. The purpose is to increase the nonlinearity of the model, thereby improving the model's ability to approximate complex functions. In the context described above, vector activation refers to obtaining the target conversion vector and target correction vector corresponding to the target modal information, and then performing a nonlinear transformation on these vectors to obtain the final target modal vector.

[0149] For example, in a speech recognition system, audio signals need to be converted into text. It is desirable to optimize the representation of the audio features through vector transformation and correction. A transformation matrix is ​​designed to map the original 128-dimensional MFCC vectors into a new feature space. This transformation matrix is ​​also 128-dimensional because we want to maintain the feature dimensionality for ease of subsequent processing. A 128-dimensional correction vector is also defined to fine-tune the transformed features to compensate for information lost during the transformation or enhance specific features. The 128-dimensional modality vector is multiplied by the 128-dimensional target transformation matrix. This multiplication transforms the original audio features into a new feature space, producing a 128-dimensional target vector. This target vector may better represent the speech signal, making it more suitable for subsequent recognition tasks. The target transformation vector is added to the target correction vector to produce a 128-dimensional correction vector. This correction vector combines the mapping effect of the transformation matrix and the fine-tuning effect of the correction vector. Vector activation is applied to the correction vector. This process may include applying a nonlinear function, such as a ReLU or Sigmoid function, to further optimize the feature representation. The activated vector, also known as the target modality vector, not only contains the original audio information but also incorporates additional information obtained through transformation and correction. This final target modality vector will be used as input to the speech recognition model in the hope of improving recognition accuracy.

[0150] Thus, by obtaining the target transformation vector and target correction vector corresponding to the target modal information and ensuring that these vectors have the same dimensions as the modal vector corresponding to the target modal information, we can perform a series of data processing steps to optimize and improve the representation of the modal information. First, the modal vector corresponding to the target modal information is multiplied by the target transformation vector. This step can effectively map the original features to a new feature space, generating a target vector corresponding to the target modal information, which may better reflect the inherent structure and relationship of the data. Next, by adding the target transformation vector and the target correction vector, we obtain a correction vector. This correction vector takes into account additional adjustment factors, which can be a subtle adjustment or compensation for the original features. Finally, the correction vector is vector activated, a nonlinear transformation that can further optimize the representation of the features. The final target modal vector contains not only the original modal information, but also the transformed and corrected information. This processing flow significantly improves the modal vector's ability to represent the target task, helping to improve model performance and prediction accuracy.

[0151] In step 1024 , the modal vector corresponding to the target modal information and the modal vectors corresponding to each reference modal information are fused to obtain a fused modal vector corresponding to the target modal information.

[0152] In some embodiments, the above-mentioned vector fusion includes linear fusion and nonlinear fusion. The above-mentioned vector fusion refers to a processing process of fusing multiple features to be fused to obtain at least one fused feature.

[0153] In some embodiments, vector fusion is a feature fusion technology that involves the process of combining multiple feature vectors from different sources or different dimensions into one or more new feature vectors. It is used to merge multiple independent feature vectors into a single feature vector, or to fuse multiple feature vectors into a set of new feature vectors. There are two main types: linear fusion and nonlinear fusion. Linear fusion refers to combining multiple feature vectors together by linear combination. Common linear fusion methods include weighted averaging, principal component analysis (PCA), linear discriminant analysis (LDA), etc. In linear fusion, each original feature vector is transformed by a linear transformation (such as a weight matrix), and then the fused feature vector is obtained by linear combination. Nonlinear fusion involves using nonlinear functions to merge multiple feature vectors. This fusion method can better capture the complex relationship between features. Common nonlinear fusion methods include neural networks, deep learning models, kernel function methods, etc.

[0154] In some embodiments, the above-mentioned step 1024 can be implemented as follows: obtaining a reference conversion vector corresponding to the target modal information and a reference correction vector corresponding to the target modal information, the modal vector corresponding to the target modal information having the same vector dimensions as the corresponding reference conversion vector and the reference correction vector, respectively; performing a vector transposition on the modal vector corresponding to the target modal information to obtain a transposed vector corresponding to the target modal information; multiplying the transposed vector by the corresponding reference conversion vector to obtain a first reference vector, and multiplying the first reference vector by the modal vector to obtain a second reference vector; adding the reference correction vector corresponding to the target modal information to the second reference vector to obtain a fused modal vector corresponding to the target modal information.

[0155] As an example, referring to FIG6 , when the target modal information is modal information 61 , the expression of the fused modal vector corresponding to the target modal information may be:

[0156] Among them, z g It is used to indicate the fused modal vector corresponding to the modal information 61, It is used to indicate the transposed vector corresponding to the modal information 61, h s It is used to indicate the modal vector corresponding to the modal information 62, W gs Used to indicate the reference conversion vector corresponding to the modal information 61, b gs Used to indicate the reference correction vector corresponding to the modal information 61.

[0157] As an example, referring to FIG6 , when the target modal information is modal information 62 , the expression of the fused modal vector corresponding to the target modal information may be:

[0158] Among them, z s It is used to indicate the fused modal vector corresponding to the modal information 62, It is used to indicate the transposed vector corresponding to the modal information 62, h g It is used to indicate the modal vector corresponding to the modal information 61, W sg Used to indicate the reference conversion vector corresponding to the modal information 62, b sg Used to indicate the reference correction vector corresponding to the modal information 62.

[0159] In this way, by obtaining the reference transformation vector and reference correction vector corresponding to the target modal information and ensuring that the dimensions of these vectors match the modal vector corresponding to the target modal information, multi-source heterogeneous data can be effectively integrated. Specifically, after transposing the modal vector corresponding to the target modal information, it is multiplied by the reference transformation vector to obtain a first reference vector. The modal vector is then multiplied by the first reference vector to obtain a second reference vector. Finally, the reference correction vector is added to the second reference vector. The resulting fused modal vector not only retains the core features of the original modal information but also enhances its accuracy and robustness through correction and transformation, thereby improving the fusion quality of the modal information and providing more accurate and comprehensive input for subsequent data analysis and processing. This enhances the model's adaptability and generalization capabilities for complex scenarios, helping to improve the performance of prediction and classification tasks. It optimizes data resource utilization and reduces computational resource waste through effective vector operations and fusion strategies. This provides a more reliable data foundation for subsequent advanced applications such as feature extraction and sentiment analysis.

[0160] In step 1025, the fused modal vector and the target modal vector are fused to obtain a multimodal sub-vector corresponding to the target modal information.

[0161] As an example, referring to FIG6 , when the target modal information is modal information 61, the expression of the multimodal sub-vector corresponding to the target modal information can be: g =σ(z g )*h′ g (5)

[0162] Among them, g Used to indicate the multimodal subvector corresponding to the modal information 61, h′ g , used to indicate the fusion modal vector corresponding to the modal information 61, h′ g Used to indicate the target modal vector corresponding to the modal information 61.

[0163] As an example, referring to FIG6 , when the target modal information is modal information 62 , the expression of the multimodal subvector corresponding to the target modal information may be: s =σ(z s )*h′ s (6)

[0164] Among them, s It is used to indicate the multi-modal sub-vector corresponding to the modal information 62, z s Used to indicate the fused modal vector corresponding to the modal information 62, h′ s Used to indicate the target modal vector corresponding to the modal information 62.

[0165] In this way, by fusing the modal information of each modality, the multimodal vector corresponding to the image to be classified is obtained, which significantly enhances the information content of the multimodal vector, and lays a strong data support for the subsequent encoding of the image to be classified based on the multimodal vector.

[0166] In step 103, the image to be classified is encoded based on the multimodal vector to obtain a sub-vector to be classified corresponding to each sub-image to be classified.

[0167] In some embodiments, the vector to be classified includes sub-vectors to be classified corresponding one-to-one to each sub-image to be classified.

[0168] As an example, suppose there is a dataset of products. Each product has the following categorical attributes: product category, brand, and color. Our goal is to create a vector to be classified, which contains a subvector to be classified for each attribute. Product Category: Original Information: Products might be electronics, furniture, clothing, etc. Encoding Method: One-Hot Encoding can be used to convert each category into a vector. Electronics: [1, 0, 0], furniture: [0, 1, 0], clothing: [0, 0, 1]. The corresponding subvector to be classified is a three-dimensional vector. For example, the subvector for electronics is [1, 0, 0]. Brand: Original Information: Product brands might be Huawei, Apple, Xiaomi, etc. Encoding Method: One-Hot Encoding is also used to convert each brand into a vector. For example: Huawei: [1, 0, 0], Apple: [0, 1, 0], Xiaomi: [0, 0, 1]. The corresponding subvector to be classified is a three-dimensional vector. For example, the subvector for Huawei is [1, 0, 0]. Color: Original Information: Product colors might be red, blue, green, etc. Encoding method: Continue using one-hot encoding to convert each color into a vector. For example, the subvectors to be classified corresponding to red: [1, 0, 0], blue: [0, 1, 0], and green: [0, 0, 1] are three-dimensional vectors. For example, the subvector corresponding to red is [1, 0, 0]. Finally, all subvectors to be classified are combined to form a single vector to be classified. For example, for a red Huawei electronic product, its vector to be classified is: [1, 0, 0, 1, 0, 0, 1, 0, 0]. This vector to be classified contains three subvectors to be classified: the product category subvector [1, 0, 0], the brand subvector [1, 0, 0], and the color subvector [1, 0, 0]. In this way, we can convert non-numeric categorical information into numerical vectors that can be processed by machine learning models.

[0169] In some embodiments, the multimodal vector includes multimodal sub-vectors corresponding to each modal information, see Figure 7, which is a flow chart of the image classification method provided in an embodiment of the present application. Step 103 shown in Figure 4 can be implemented through steps 1031 to 1033 shown in Figure 7.

[0170] In step 1031, the vector dimension of each multimodal sub-vector is obtained, and the average value of each vector dimension is determined as the target dimension.

[0171] As an example, the expression for the target dimension above can be:

[0172] Among them, W is used to indicate the target dimension, W1…W N It is used to indicate the vector dimension of each multimodal sub-vector, and N is used to indicate the number of multimodal sub-vectors.

[0173] In some embodiments, multimodal data refers to data from different sources or different types, such as text, images, sounds, etc. Each source or type of data can be represented as a vector, called a multimodal subvector. For example, a database describing a product may contain a vector of image features (image subvector), a vector of text descriptions (text subvector), and a vector of sound descriptions (sound subvector). For each multimodal subvector, its vector dimension is the number of elements in the vector. For example, if an image feature vector contains 100 features, then its vector dimension is 100.

[0174] In step 1032 , the vector dimensions of each multimodal sub-vector are adjusted to the target dimension to obtain the target sub-vector corresponding to each multimodal sub-vector.

[0175] In some embodiments, the target sub-vector and the multimodal sub-vector have a one-to-one correspondence, and the vector dimension of the target sub-vector corresponding to the multimodal sub-vector is equal to the target dimension.

[0176] In some embodiments, by adjusting the vector dimensions of each multimodal sub-vector to the target dimension, the target sub-vector corresponding to each multimodal sub-vector is obtained, thereby achieving dimensional averaging of each multimodal sub-vector.

[0177] As an example, consider image subvector resizing: the original image subvector is [feature 1, feature 2, ..., feature 512]. After dimensionality reduction, the target subvector for the image is [reduced feature 1, reduced feature 2, ..., reduced feature 312]. Similarly, consider text subvector resizing: the original text subvector is [feature 1, feature 2, ..., feature 100]. After dimensionality increase, the target subvector for the text is [original feature 1, original feature 2, ..., original feature 100, additional feature 1, additional feature 2, ..., additional feature 212]. In this process, each multimodal subvector is resized to a vector of target dimension 312. The resized vector is called the target subvector. For example, the target subvector for an image might look like this (just for illustration, not real data): [0.12, -0.03, 0.45, ..., 0.67]; the target subvector for a text might look like this (just for illustration, not real data): [0.22, 0.55, 0.13, 0, 0, ..., 0]. In this way, the target subvectors for both the image and the text are ensured to have the same dimensions, allowing them to be compared and fused in the same space. This allows features from different modalities to be effectively fused in a unified dimension, enabling better multimodal tasks such as multimodal classification, retrieval, or generation.

[0178] In step 1033, each target sub-vector is multiplied to obtain a reference multimodal vector, and the reference multimodal vector is corrected based on the correction parameter to obtain a vector to be classified corresponding to the image to be classified.

[0179] As an example, referring to FIG6 , the expression of the reference multimodal vector can be:

[0180] Among them, O is used to indicate the reference multimodal vector, o g It is used to indicate the multi-modal sub-vector corresponding to the modal information 61, o s It is used to indicate the multi-modal sub-vector corresponding to the modal information 62, Used to indicate the Kronecker product.

[0181] In some embodiments, the above-mentioned correction of the reference multimodal vector based on the correction parameter to obtain the vector to be classified corresponding to the image to be classified can be achieved as follows: obtaining a first correction parameter and a second correction parameter, multiplying the reference multimodal vector by the first correction parameter to obtain the reference vector to be classified corresponding to the image to be classified, and adding the reference vector to be classified to the second correction parameter to obtain the vector to be classified corresponding to the image to be classified.

[0182] As an example, the expression of the vector to be classified corresponding to the above image to be classified can be:

[0183] Among them, h is used to indicate the vector to be classified corresponding to the image to be classified. Used to indicate the reference multimodal vector, o g It is used to indicate the multi-modal sub-vector corresponding to the modal information 61, o s It is used to indicate the multi-modal sub-vector corresponding to the modal information 62, Used to indicate the Kronecker product, W fusion Used to indicate the first correction parameter, b fusion Used to indicate the second calibration parameter.

[0184] As an example, the goal is to fuse features from two different modalities (e.g., image and text) into a unified feature representation, and then use correction parameters to adjust this representation to optimize classification performance. The first correction parameter, W_fusion, is a weight matrix whose dimensions typically match the dimensions of the fused feature vector. This matrix is ​​learned and is used to adjust the size and direction of the fused feature vector to optimize classification performance. The second correction parameter, b_fusion, is a bias vector whose dimensions match the dimensions of the fused feature vector. This vector is also learned and is used to provide additional correction on top of the weighted sum. Consider two feature vectors from different modalities: an image feature vector, o_g, and a text feature vector, o_s. The image feature vector, o_g, and the text feature vector, o_s, are outer-producted. This typically involves element-wise multiplication of the two vectors to produce a new matrix whose dimensions are the dimensions of o_g multiplied by the dimensions of o_s. The resulting outer-product matrix is ​​then multiplied by the weight matrix, W_fusion, to produce an adjusted feature vector. The bias vector, b_fusion, is added to the vector obtained in the previous step to further adjust the feature vector. The adjusted feature vector is nonlinearly transformed through the ReLU (Rectified Linear Unit) activation function to obtain the final feature classification vector h to be classified.

[0185] In this way, features of different modalities, such as images and text, are combined to generate a unified feature classification vector to be classified. In this process, the reference multimodal feature modal vectors are first combined through the outer product operation and multiplied by the first correction parameter, which not only adjusts the size and direction of the features, but also strengthens the important correlation between the features. Subsequently, by adding the second correction parameter, the feature representation can be further optimized to make it more in line with the requirements of the classification task. Using the ReLU activation function to process the result vector helps to introduce nonlinearity, making the feature classification vector richer and more discriminative in representation. The beneficial effect of this correction and fusion strategy is reflected in the improvement of the accuracy and robustness of the classification task, enabling the model to better extract and utilize useful information from multimodal data.

[0186] In this way, by multiplying each target sub-vector, a reference multimodal vector is obtained, and based on the correction parameter, the reference multimodal vector is corrected to obtain the vector to be classified corresponding to the image to be classified, so that the vector to be classified corresponding to the image to be classified is more accurate.

[0187] In step 104, the image to be classified is classified based on the sub-vector to be classified corresponding to each sub-image to be classified to obtain a category of the image to be classified.

[0188] In some embodiments, the above step 104 can be implemented as follows: for each sub-image to be classified, based on the sub-vector to be classified corresponding to the sub-image to be classified, the category of the sub-image to be classified is predicted to obtain the category prediction probability corresponding to the sub-image to be classified; based on the category prediction probability corresponding to each sub-image to be classified, the image to be classified is classified to obtain the category of the image to be classified.

[0189] In some embodiments, the above-mentioned category prediction can be achieved by a target classification model, which can be a machine learning model used to predict the category of a data sample. During the training process, the classification model learns how to map the input vector to the correct category label. In the above-mentioned step 104, for each sub-image to be classified, based on the sub-vector to be classified corresponding to the sub-image to be classified, the category prediction of the sub-image to be classified is performed on the sub-image to be classified, and the category prediction probability corresponding to the sub-image to be classified is obtained. This can be achieved as follows: for each sub-image to be classified, the target classification model is called, and based on the sub-vector to be classified corresponding to the sub-image to be classified, the category prediction of the sub-image to be classified is performed on the sub-image to be classified, and the category prediction probability corresponding to the sub-image to be classified is obtained.

[0190] In some embodiments, category prediction is implemented through a target classification model. Before executing the above step 104, the target classification model can also be obtained in the following manner: obtain an image sample to be classified including multiple sub-samples to be classified, and a sample label carried by the image sample to be classified; encode the image sample to be classified to obtain a sample vector corresponding to the image sample to be classified, and the sample vector includes a sub-sample vector corresponding to each sub-sample to be classified; for each sub-sample to be classified, determine the sample label as the initial label of the sub-sample to be classified, and call the classification model to perform category prediction on the sub-sample to be classified based on the sub-sample vector corresponding to the sub-sample to be classified to obtain the predicted probability of the sub-sample to be classified; based on the predicted probability of each sub-sample to be classified, determine the target label of each sub-sample to be classified; train the classification model in combination with the predicted probability, the initial label and the target label to obtain the target classification model.

[0191] In some embodiments, the sample label carried by the image sample to be classified is used to indicate the category to which the image sample to be classified belongs. The image sample to be classified includes multiple subsamples to be classified. The subsample labels carried by the subsamples to be classified can be the same as or different from the sample labels carried by the corresponding image sample to be classified. The subsample labels carried by each subsample to be classified included in the image sample to be classified are associated with the sample label carried by the image sample to be classified, that is, the subsample labels carried by each subsample to be classified included in the image sample to be classified affect the sample label carried by the image sample to be classified.

[0192] In some embodiments, the sample vector includes a subsample vector corresponding to each subsample to be classified, and the sample vector corresponding to the image sample to be classified is used to indicate the attributes or characteristics of the image sample to be classified.

[0193] In some embodiments, the above encoding of the image samples to be classified to obtain the sample vectors corresponding to the image samples to be classified can be achieved in the following ways: performing modal conversion of different modalities on the image samples to be classified to obtain modal samples of each modality corresponding to the image samples to be classified; performing modal fusion on the samples of each modality to obtain a multimodal sample vector corresponding to the image samples to be classified; based on the multimodal sample vector, encoding the image samples to be classified to obtain the sample vector corresponding to the image samples to be classified.

[0194] In some embodiments, the multimodal sample vector includes a multimodal sample sub-vector corresponding to each modal sample. The modal fusion of each modal sample to obtain a multimodal sample vector corresponding to the image sample to be classified can be achieved as follows: encoding each modal sample to obtain a modal sample vector corresponding to each modal sample; performing the following processing on each modal sample: determining the modal sample as a target modal sample, and determining each modal sample other than the target modal sample as a reference modal sample; performing vector conversion on the modal sample vector corresponding to the target modal sample to obtain a target modal sample vector corresponding to the target modal sample; performing vector fusion on the modal sample vector corresponding to the target modal sample and the modal sample vector corresponding to each reference modal sample to obtain a fused modal sample vector corresponding to the target modal sample; performing vector fusion on the fused modal sample vector and the target modal sample vector to obtain a multimodal sample sub-vector corresponding to the target modal sample.

[0195] In some embodiments, the above-mentioned vector conversion of the modal sample vector corresponding to the target modal sample to obtain the target modal sample vector corresponding to the target modal sample can be achieved in the following manner: obtaining the target conversion vector corresponding to the target modal sample and the target correction vector corresponding to the target modal sample, the modal sample vector corresponding to the target modal sample having the same vector dimensions as the corresponding target conversion vector and the target correction vector, respectively; multiplying the modal sample vector corresponding to the target modal sample and the target conversion vector to obtain the target vector corresponding to the target modal sample; adding the target conversion vector and the target correction vector to obtain the correction vector corresponding to the target modal sample; and performing vector activation on the target correction vector to obtain the target modal sample vector corresponding to the target modal sample.

[0196] In some embodiments, the above-mentioned vector fusion of the modal sample vector corresponding to the target modal sample and the modal sample vector corresponding to each reference modal sample to obtain the fused modal sample vector corresponding to the target modal sample can be achieved as follows: obtaining a reference conversion vector corresponding to the target modal sample and a reference correction vector corresponding to the target modal sample, the modal sample vector corresponding to the target modal sample having the same vector dimensions as the corresponding reference conversion vector and the reference correction vector, respectively; performing vector transposition on the modal sample vector corresponding to the target modal sample to obtain the transposed vector corresponding to the target modal sample; multiplying the transposed vector by the corresponding reference conversion vector to obtain a first reference vector, and multiplying the first reference vector by the modal sample vector to obtain a second reference vector; adding the reference correction vector corresponding to the target modal sample to the second reference vector to obtain the fused modal sample vector corresponding to the target modal sample.

[0197] In some embodiments, the multimodal sample vector includes a multimodal sub-sample vector corresponding to each modal sample. The encoding of the image sample to be classified based on the multimodal sample vector to obtain the sample vector corresponding to the image sample to be classified can be achieved in the following manner: obtaining the vector dimension of each multimodal sub-sample vector, and determining the average value of each vector dimension as the target dimension; adjusting the vector dimension of each multimodal sub-sample vector to the target dimension to obtain the target sub-vector corresponding to each multimodal sub-sample vector; multiplying each target sub-vector to obtain a reference multimodal vector, and correcting the reference multimodal vector based on the correction parameter to obtain the sample vector corresponding to the image sample to be classified.

[0198] As an example, the image sample A to be classified includes sub-sample A1 to be classified, sub-sample A2 to be classified, and sub-sample A3 to be classified. The sample label carried by the image sample A to be classified is category T. The sample label carried by the image sample A to be classified: category T is determined as the initial label of sub-sample A1 to be classified, sub-sample A2 to be classified, and sub-sample A3 to be classified.

[0199] In some embodiments, the classification model may be a convolutional neural network (CNN). CNNs are a type of feedforward neural network (FNN) with a deep structure that incorporates convolutional computations and are a representative algorithm for deep learning. CNNs possess representation learning capabilities and can perform shift-invariant classification on input images based on their hierarchical structure.

[0200] As an example, refer to Figure 10, which is a structural diagram of the classification model provided in an embodiment of the present application. The classification model includes an encoding layer 1 and a decoding layer 2. The encoding layer 1 is called to perform sample encoding on each sub-sample to be classified based on the sub-sample vector corresponding to the sub-sample to be classified to obtain a target sub-sample vector; the decoding layer 2 is called to perform category prediction on the sub-sample to be classified based on the target sub-sample vector to obtain a predicted probability of the sub-sample to be classified.

[0201] In some embodiments, the predicted probability includes sub-prediction probabilities corresponding one-to-one to each candidate category.

[0202] As an example, the expression of the predicted probability can be: P = {P1, ..., P n} (10)

[0203] Among them, P is used to indicate the predicted probability, P1,…,P n It is used to indicate the sub-prediction probability corresponding to each candidate category, and n is used to indicate the number of candidate categories.

[0204] In some embodiments, the target label of the sub-sample to be classified is used to indicate the candidate category to which the sub-sample to be classified belongs.

[0205] In some embodiments, the above-mentioned category prediction probability includes a category probability corresponding to each candidate category, and the category probability is used to indicate the possibility that the sub-image to be classified belongs to the corresponding candidate category.

[0206] In some embodiments, referring to FIG8 , FIG8 is a flowchart diagram four of the image classification method provided in an embodiment of the present application. In step 104 shown in FIG4 , the image to be classified is classified based on the category prediction probability corresponding to each sub-image to be classified, and obtaining the category of the image to be classified can be achieved through steps 1041 to 1043 shown in FIG8 .

[0207] In step 1041 , the weight coefficients corresponding to the sub-images to be classified are obtained.

[0208] In some embodiments, the sub-images to be classified correspond to the weight coefficients in a one-to-one correspondence, and the weight coefficients are used to indicate the importance of the sub-images to be classified in the image to be classified.

[0209] As an example, the sub-images to be classified include sub-image A to be classified and sub-image B to be classified. The weight coefficient corresponding to sub-image A to be classified is A1, and the weight coefficient corresponding to sub-image B to be classified is A2.

[0210] In some embodiments, the weight coefficient is a parameter commonly used in machine learning models, which is used to quantify the contribution of input features (in this case, the sub-image to be classified) to the model output. Each sub-image to be classified has a unique weight coefficient corresponding to it, which is used to indicate the importance of the sub-information in the overall image to be classified. The size of the weight coefficient reflects the degree of influence of the sub-image to be classified on the classification result. A high weight coefficient indicates that the sub-information is very important in the classification decision, while a low weight coefficient means that the sub-information has little influence on the classification result.

[0211] For example, sub-image A to be classified: This is a piece of sub-information in a classification task, such as representing a specific attribute of a product. Sub-image B to be classified: This is sub-information in another classification task, different from sub-image A to be classified, and may represent another attribute of the product. Weight coefficient A1: This is a weight coefficient that corresponds one-to-one with sub-image A to be classified, and is used to indicate the importance of sub-image A to be classified in the image to be classified. Weight coefficient A2: This is a weight coefficient that corresponds one-to-one with sub-image B to be classified, and is used to indicate the importance of sub-image B to be classified in the image to be classified.

[0212] In step 1042 , based on the weight coefficients, the category prediction probabilities corresponding to the sub-images to be classified are weighted and summed to obtain the target prediction probability corresponding to the image to be classified.

[0213] In some embodiments, the target prediction probability includes a target probability corresponding one-to-one to each candidate category.

[0214] As an example, the expression of the target prediction probability corresponding to the above image to be classified can be:

[0215] Among them, p i Used to indicate the target prediction probability corresponding to the image to be classified, Used to indicate the category prediction probability corresponding to each sub-image to be classified. Used to indicate the weight coefficient corresponding to each sub-image to be classified.

[0216] In step 1043 , the candidate category with the highest target probability is determined as the category of the image to be classified.

[0217] As an example, the target probability corresponding to candidate category A is 0.5, the target probability corresponding to candidate category B is 0.6, and the target probability corresponding to candidate category C is 0.1. Then, candidate category B with the largest target probability is determined as the category of the image to be classified.

[0218] In this way, the introduction of weight coefficients can highlight the importance of different sub-information in classification, thereby giving key information a greater influence in the comprehensive prediction and reducing the interference of noise information. Secondly, the target prediction probability after weighted summation is more realistic because it comprehensively considers the importance and probability distribution of each sub-information. Finally, by comparing the target prediction probabilities, the candidate category with the highest probability is determined as the category of the image to be classified. This can ensure the reliability of the classification results and improve the accuracy and robustness of the overall classification. It has significant advantages in multi-attribute decision-making and complex information processing, and helps to optimize the performance and efficiency of information classification.

[0219] In this way, by performing modal conversion of different modalities on the image to be classified including multiple sub-images to be classified, the modal information of each modality corresponding to the image to be classified is obtained, and the modal information is modally fused to obtain a multimodal vector corresponding to the image to be classified. Based on the multimodal vector, the image to be classified is encoded to obtain the vector to be classified corresponding to the image to be classified. For each sub-image to be classified, based on the sub-vector to be classified corresponding to the sub-image to be classified, the category of the sub-image to be classified is predicted to obtain the category prediction probability corresponding to the sub-image to be classified. Based on the category prediction probability corresponding to each sub-image to be classified, the image to be classified is classified to obtain the category of the image to be classified. In this way, by performing modal fusion on the modal information obtained by performing modal conversion of the image to be classified into different modalities, the obtained multimodal vector can accurately reflect the modal characteristics of the image to be classified from different modalities, and based on the multimodal vector, the vector to be classified is obtained by encoding the image to be classified, thereby effectively improving the information content of the vector to be classified, and based on the vector to be classified corresponding to each sub-image to be classified, the category of the sub-image to be classified is predicted. Since the vector to be classified on which the category prediction is based has a more comprehensive amount of information, the category prediction probability obtained by the prediction is more accurate. Through the accurate category prediction probability, the image to be classified is classified, thereby effectively improving the classification accuracy of the information.

[0220] Refer to Figure 9, which is a flow chart of the training method of the image classification model provided in an embodiment of the present application. It will be explained in conjunction with steps 201 to 206 shown in Figure 9. The training method of the image classification model provided in an embodiment of the present application can be implemented by the server or the terminal alone, or by the server and the terminal in collaboration. The following will be explained using the server alone as an example.

[0221] In step 201, an image sample to be classified including a plurality of sub-samples to be classified and a sample label carried by the image sample to be classified are obtained.

[0222] In some embodiments, the sample label carried by the image sample to be classified is used to indicate the category to which the image sample to be classified belongs. The image sample to be classified includes multiple subsamples to be classified. The subsample labels carried by the subsamples to be classified can be the same as or different from the sample labels carried by the corresponding image sample to be classified. The subsample labels carried by each subsample to be classified included in the image sample to be classified are associated with the sample label carried by the image sample to be classified, that is, the subsample labels carried by each subsample to be classified included in the image sample to be classified affect the sample label carried by the image sample to be classified.

[0223] In step 202, the image samples to be classified are encoded to obtain sample vectors corresponding to the image samples to be classified.

[0224] In some embodiments, the sample vector includes a subsample vector corresponding to each subsample to be classified, and the sample vector corresponding to the image sample to be classified is used to indicate the attributes or characteristics of the image sample to be classified.

[0225] In some embodiments, the above encoding of the image samples to be classified to obtain the sample vectors corresponding to the image samples to be classified can be achieved in the following ways: performing modal conversion of different modalities on the image samples to be classified to obtain modal samples of each modality corresponding to the image samples to be classified; performing modal fusion on the samples of each modality to obtain a multimodal sample vector corresponding to the image samples to be classified; based on the multimodal sample vector, encoding the image samples to be classified to obtain the sample vector corresponding to the image samples to be classified.

[0226] In some embodiments, the image sample to be classified includes multiple sub-images to be classified, and the sub-images to be classified are constituent elements of the image sample to be classified. For example, the image sample to be classified A includes sub-image A1 to be classified, sub-image A2 to be classified, and sub-image A3 to be classified.

[0227] In some embodiments, modality indicates the perceptual channel of information. Different modal information corresponds to different perceptual channels. For example, the video modality information corresponds to the video perceptual channel, while the audio modality information corresponds to the audio perceptual channel. The above-mentioned modality conversion is the process of converting the image samples to be classified in the original modality into the modal information of the target modality.

[0228] In some embodiments, the above-mentioned modal conversion of different modalities of the image samples to be classified to obtain the modal information of each modality corresponding to the image samples to be classified can be achieved as follows: when the image samples to be classified are multimodal information, each target modality in the image samples to be classified is obtained, and for each target modality, the modal information of the target modality is extracted from the image samples to be classified to obtain the modal information corresponding to the target modality of the image samples to be classified; when the image samples to be classified are single-modal information, multiple target modalities are obtained, and for each target modality, the modal conversion of the target modality is performed on the image samples to be classified to obtain the modal information corresponding to the target modality of the image samples to be classified.

[0229] As an example, in a text classification application scenario, the image sample to be classified can be converted from the text to be classified (single-modal information), which includes multiple sub-texts to be classified. For example, if the text to be classified is "tumor exists in the liver area", the multiple sub-texts to be classified included in the text to be classified can be "liver area", "existence", and "tumor". The image sample to be classified is modally converted from the audio modality to obtain the audio to be classified in the audio modality corresponding to the image sample to be classified. The image sample to be classified is modally converted from the text modality to obtain the image to be classified in the text modality corresponding to the image sample to be classified.

[0230] As an example, in an application scenario of medical information classification, the image sample to be classified may be an image containing an immune receptor library to be classified, the immune receptor library to be classified including multiple immune receptors, and the region of the immune receptor in the image sample to be classified is a sub-image to be classified in the image sample to be classified. The immune receptor library to be classified is modally converted from an amino acid sequence modality to obtain an amino acid sequence of the immune receptor library to be classified in an amino acid sequence modality. The immune receptor library to be classified is modally converted from a gene segment modality to obtain a gene segment of the immune receptor library to be classified in a gene segment modality.

[0231] In some embodiments, the multimodal sample vector includes a multimodal sample sub-vector corresponding to each modal sample. The modal fusion of each modal sample to obtain a multimodal sample vector corresponding to the image sample to be classified can be achieved as follows: encoding each modal sample to obtain a modal sample vector corresponding to each modal sample; performing the following processing on each modal sample: determining the modal sample as a target modal sample, and determining each modal sample other than the target modal sample as a reference modal sample; performing vector conversion on the modal sample vector corresponding to the target modal sample to obtain a target modal sample vector corresponding to the target modal sample; performing vector fusion on the modal sample vector corresponding to the target modal sample and the modal sample vector corresponding to each reference modal sample to obtain a fused modal sample vector corresponding to the target modal sample; performing vector fusion on the fused modal sample vector and the target modal sample vector to obtain a multimodal sample sub-vector corresponding to the target modal sample.

[0232] In some embodiments, modal fusion refers to the process of integrating information from different modalities (such as text, images, and audio) within a multimodal data processing system. Each modality typically contains different aspects or characteristics of the data, and the goal of modal fusion is to combine these different information sources to produce a more comprehensive and richer data representation.

[0233] In this way, each modal sample is processed as a target modal sample, while the other modal samples are processed as reference modal samples. This method ensures that each modal sample receives attention and allows for personalized vector transformation tailored to the characteristics of each modal sample, thereby enhancing intermodal correlation and complementarity. Transforming the target modal sample vector not only highlights the characteristics of the target modal sample but also enhances its compatibility with other modal samples. By vector fusion of the target modal sample vector with each reference modal sample vector, a fused modal sample vector is obtained that incorporates information from multiple modalities. This step helps integrate information from different modalities, improving the comprehensiveness and accuracy of feature representation. The fused modal sample vector is further vector-fused with the target modal sample vector to obtain a multimodal sample subvector that further incorporates the original information of the target modal sample and the additional information obtained through fusion. Such a multimodal sample subvector not only contains rich feature information but also provides a more accurate and robust data foundation for subsequent classification, recognition, or other advanced tasks, thereby improving the performance of the entire multimodal processing system.

[0234] In some embodiments, the above-mentioned vector conversion of the modal sample vector corresponding to the target modal sample to obtain the target modal sample vector corresponding to the target modal sample can be achieved in the following manner: obtaining the target conversion vector corresponding to the target modal sample and the target correction vector corresponding to the target modal sample, the modal sample vector corresponding to the target modal sample having the same vector dimensions as the corresponding target conversion vector and the target correction vector, respectively; multiplying the modal sample vector corresponding to the target modal sample and the target conversion vector to obtain the target vector corresponding to the target modal sample; adding the target conversion vector and the target correction vector to obtain the correction vector corresponding to the target modal sample; and performing vector activation on the target correction vector to obtain the target modal sample vector corresponding to the target modal sample.

[0235] In some embodiments, vector activation refers to the process of applying a nonlinear function to a vector in machine learning and deep learning. The purpose is to increase the nonlinearity of the model, thereby improving the model's ability to approximate complex functions. In the context described above, vector activation refers to obtaining the target conversion vector and target correction vector corresponding to the target modal information, and then performing a nonlinear transformation on these vectors to obtain the final target modal vector.

[0236] In this way, in the multimodal data processing and conversion task, the target conversion vector and target correction vector corresponding to the target modal sample are obtained, and the vector dimension is kept consistent with the modal sample vector, which has the following beneficial effects: First, by multiplying the modal sample vector with the target conversion vector to obtain the target vector, the original modality can be effectively mapped to the target modal space, which helps to achieve effective conversion and feature preservation between modalities. Secondly, the target conversion vector and the target correction vector are added to form a correction vector, and the mapping result is further fine-tuned, which helps to improve the accuracy and fidelity of the converted target modal sample. Finally, vector activation of the correction vector can activate the dimensions with significant characteristics in the target modal sample vector, thereby generating a sample vector that is more consistent with the target modal feature distribution. This series of processing not only improves the efficiency and accuracy of modal conversion, but also enhances the generalization ability of the model, enabling the model to better adapt to the complex conversion relationship between different modalities, and provides strong support for the fusion, analysis and application of multimodal data.

[0237] In some embodiments, the above-mentioned vector fusion of the modal sample vector corresponding to the target modal sample and the modal sample vector corresponding to each reference modal sample to obtain the fused modal sample vector corresponding to the target modal sample can be achieved as follows: obtaining a reference conversion vector corresponding to the target modal sample and a reference correction vector corresponding to the target modal sample, the modal sample vector corresponding to the target modal sample having the same vector dimensions as the corresponding reference conversion vector and the reference correction vector, respectively; performing vector transposition on the modal sample vector corresponding to the target modal sample to obtain the transposed vector corresponding to the target modal sample; multiplying the transposed vector by the corresponding reference conversion vector to obtain a first reference vector, and multiplying the first reference vector by the modal sample vector to obtain a second reference vector; adding the reference correction vector corresponding to the target modal sample to the second reference vector to obtain the fused modal sample vector corresponding to the target modal sample.

[0238] As an example, suppose you are processing data in both image and text modalities and wish to convert the image modality to text for cross-modal retrieval or analysis. This process begins by obtaining a set of reference transformation vectors and a reference correction vector corresponding to samples in the target modality. For example, if the target modality is text, the reference transformation vectors may contain text features such as word frequency, TF-IDF scores, or embedding representations. These vectors are used to convert image features into text features. The reference correction vectors contain fine-tuned parameters to further optimize the text features after conversion. Now, suppose you have an image modality sample vector that contains features such as color, shape, and texture, and the dimensions of this vector are the same as those of the reference transformation vector and the reference correction vector. The next steps are as follows: Transpose the image modality sample vector to obtain a transposed vector. This step is equivalent to converting a row of image features into a column, which may facilitate subsequent matrix multiplication operations. Then, multiply the transposed vector by the reference transformation vector to obtain a first reference vector. This operation can be viewed as a way to map image features into the space of text features. Next, the first reference vector is multiplied by the original modal sample vector (untransposed image feature vector) to obtain the second reference vector. This is the process of refining the mapping so that the converted text features are more consistent with the semantic content of the original image. Finally, the reference correction vector of the target modal sample is added to the second reference vector to obtain the fused modal sample vector. This fused vector is the final output. It contains both the feature information of the original image and the converted text features, and through the adjustment of the correction vector, it is closer to the desired target text modal features. Through this process, the image modality samples can be converted into text modality samples. Such a conversion is very useful for realizing information fusion between images and texts and cross-modal retrieval tasks.

[0239] In this way, through vector transposition and multiplication operations, the original information of the target modal sample vector and the conversion information provided by the reference transformation are effectively combined, which not only enhances the expressiveness of the sample features, but also helps to reveal the correlation between different modalities. By multiplying the transposed vector with the reference transformation vector to obtain a first reference vector, and multiplying the first reference vector with the modal sample vector to obtain a second reference vector, a deeper feature representation can be extracted, corresponding to the transformation and re-expression of the features, respectively, which helps to separate useful signals and suppress noise. The reference correction vector is added to the second reference vector to obtain a fused modal sample vector, which integrates the correction information and the converted feature information, so that the fused modal sample vector retains the original features while increasing the understanding and representation of the target modal sample features, thereby improving the accuracy and efficiency of modal fusion.

[0240] In some embodiments, the multimodal sample vector includes a multimodal sub-sample vector corresponding to each modal sample. The encoding of the image sample to be classified based on the multimodal sample vector to obtain the sample vector corresponding to the image sample to be classified can be achieved in the following manner: obtaining the vector dimension of each multimodal sub-sample vector, and determining the average value of each vector dimension as the target dimension; adjusting the vector dimension of each multimodal sub-sample vector to the target dimension to obtain the target sub-vector corresponding to each multimodal sub-sample vector; multiplying each target sub-vector to obtain a reference multimodal vector, and correcting the reference multimodal vector based on the correction parameter to obtain the sample vector corresponding to the image sample to be classified.

[0241] As an example, consider a set of multimodal subsample vectors from different modalities, such as an image feature vector and an audio feature vector. Each vector may have different dimensions; an image feature vector may have tens of thousands of dimensions, while an audio feature vector may have hundreds of dimensions. Obtain the vector dimensions of each multimodal subsample vector: For example, there may be an image feature vector with a dimension of 1024 and an audio feature vector with a dimension of 128. Calculate the average of these vector dimensions to determine the target dimension: If we have five multimodal subsample vectors with dimensions of 1024, 128, 256, 512, and 768, respectively, then their average dimension (the target dimension) is (1024 + 128 + 256 + 512 + 768) / 5 = 512. Adjust the vector dimensions of each multimodal subsample vector to the target dimension: Now, adjust each vector to 512 dimensions. This can be achieved through interpolation, dimensionality reduction, or learning a mapping function. The adjusted vector is called the target subvector. Multiply the target subvectors: For example, multiply the adjusted image target subvector with the audio target subvector to produce a single reference multimodal vector. "Multiplication" here can refer to element-wise multiplication or matrix multiplication, depending on how the data is processed. Correct the reference multimodal vector based on correction parameters: The reference multimodal vector may also require some correction to better suit the classification task. These correction parameters, which may be learned from training data, adjust the values ​​of the reference multimodal vector to enhance certain features or suppress noise. The resulting vector is called the sample vector corresponding to the image sample to be classified. This process fuses features from different modalities into a single vector that incorporates information from multiple modalities and is adjusted and corrected to suit the classification task. This multimodal fusion vector more comprehensively represents the original sample, achieving better performance in the classification task.

[0242] This ensures dimensional consistency among multimodal subsample vectors, providing a standardized foundation for subsequent vector operations. Specifically, by adjusting the dimensions of each subsample vector to the target dimension, the corresponding target subvectors are obtained. This not only simplifies the processing of multimodal data but also helps improve data processing accuracy. Multiplying these target subvectors synthesizes a reference multimodal vector that incorporates information from different modalities, thereby enhancing the representation of sample features. This operation helps to explore the inherent connections between different modalities, providing more comprehensive and in-depth information for subsequent classification tasks. Correcting the reference multimodal vector based on correction parameters further optimizes the sample vector, making it more suitable for the classification task. The corrected sample vector improves classification accuracy and robustness while preserving the original information. The entire process, from dimensionality analysis to vector adjustment and then correction, not only optimizes the multimodal data processing process but also effectively improves the performance of the sample vector in the classification task, providing strong support for the final classification results.

[0243] In step 203, for each sub-sample to be classified, the sample label is determined as the initial label of the sub-sample to be classified.

[0244] As an example, the image sample A to be classified includes sub-sample A1 to be classified, sub-sample A2 to be classified, and sub-sample A3 to be classified. The sample label carried by the image sample A to be classified is category T. The sample label carried by the image sample A to be classified: category T is determined as the initial label of sub-sample A1 to be classified, sub-sample A2 to be classified, and sub-sample A3 to be classified.

[0245] In step 204, the classification model is called to perform category prediction on the sub-sample to be classified based on the sub-sample vector corresponding to the sub-sample to be classified, and obtain the predicted probability of the sub-sample to be classified.

[0246] In some embodiments, the classification model may be a convolutional neural network (CNN). CNNs are a type of feedforward neural network (FNN) with a deep structure that incorporates convolutional computations and are a representative algorithm for deep learning. CNNs possess representation learning capabilities and can perform shift-invariant classification on input images based on their hierarchical structure.

[0247] As an example, refer to Figure 10, which is a structural diagram of the classification model provided in an embodiment of the present application. The classification model includes an encoding layer 1 and a decoding layer 2. The above step 204 can be implemented in the following manner: calling the encoding layer 1, based on the subsample vector corresponding to the subsample to be classified, performing sample encoding on each subsample to be classified to obtain a target subsample vector; calling the decoding layer 2, based on the target subsample vector, performing category prediction on the subsample to be classified to obtain a predicted probability of the subsample to be classified.

[0248] In step 205 , based on the predicted probability of each sub-sample to be classified, the target label of each sub-sample to be classified is determined respectively.

[0249] In some embodiments, the predicted probability includes sub-prediction probabilities corresponding one-to-one to each candidate category.

[0250] As an example, the expression of the predicted probability can be: P = {P1, ..., P n} (12)

[0251] Among them, P is used to indicate the predicted probability, P1,…,P n It is used to indicate the sub-prediction probability corresponding to each candidate category, and n is used to indicate the number of candidate categories.

[0252] In some embodiments, the target label of the sub-sample to be classified is used to indicate the candidate category to which the sub-sample to be classified belongs.

[0253] In some embodiments, referring to FIG11 , FIG11 is a second flow chart of the training method of the image classification model provided in an embodiment of the present application. Step 205 shown in FIG9 can be implemented by executing steps 2051A to 2052A shown in FIG11 for each sub-sample to be classified.

[0254] In step 2051A, the maximum sub-prediction probability among the prediction probabilities of the sub-samples to be classified is determined as the reference prediction probability of the sub-samples to be classified.

[0255] Continuing with the above example, when n is equal to 3, the sub-prediction probability corresponding to candidate category A1 is 0.2, the sub-prediction probability corresponding to candidate category A2 is 0.3, and the sub-prediction probability corresponding to candidate category A3 is 0.5. The largest sub-prediction probability of 0.5 among the prediction probabilities of the sub-samples to be classified is determined as the reference prediction probability of the sub-samples to be classified.

[0256] In step 2052A, the candidate category corresponding to the reference prediction probability is determined as the target label of the sub-sample to be classified.

[0257] Continuing with the above example, the candidate category A3 corresponding to the reference predicted probability is determined as the target label of the sub-sample to be classified.

[0258] In this way, by determining the largest sub-prediction probability among the prediction probabilities of the sub-samples to be classified as the reference prediction probability of the sub-samples to be classified, and determining the candidate category corresponding to the reference prediction probability as the target label of the sub-samples to be classified, the candidate category corresponding to the largest sub-prediction probability is selected from multiple existing candidate categories as the target label of the sub-samples to be classified, thereby effectively improving the accuracy of the target label.

[0259] In some embodiments, referring to FIG12 , FIG12 is a flowchart diagram three of the training method of the image classification model provided in an embodiment of the present application. Step 205 shown in FIG9 can traverse i and execute steps 2051B to 2054B shown in FIG12 for each sub-sample to be classified until the target label of the sub-sample to be classified is obtained.

[0260] In step 2051B, based on the predicted probability of each sub-sample to be classified, an i-th update vector is determined, and the i-th update vector is used to perform an i-th update on the initial label of the sub-sample to be classified.

[0261] In some embodiments, 1≤i≤N, where N indicates the maximum number of times the initial labels of the subsamples to be classified are updated.

[0262] In some embodiments, the above step 2051B can be implemented as follows: when i is equal to 1, the first update vector is determined based on the predicted probability of each sub-sample to be classified; when i is greater than or equal to 2, the first update vector is updated for the i-1th time based on the sub-sample vector corresponding to the sub-sample to be classified to obtain the i-th update vector.

[0263] As an example, when i is equal to 1, the first update vector is determined based on the predicted probability of each sub-sample to be classified; when i is equal to 2, the first update vector is updated for the first time based on the sub-sample vector corresponding to the sub-sample to be classified, and the second update vector is obtained; when i is equal to 3, the first update vector is updated for the second time based on the sub-sample vector corresponding to the sub-sample to be classified, and the third update vector is obtained; when i is equal to 4, the first update vector is updated for the third time based on the sub-sample vector corresponding to the sub-sample to be classified, and the fourth update vector is obtained.

[0264] In some embodiments, the above-mentioned prediction probability includes sub-prediction probabilities corresponding one-to-one to each candidate category. The above-mentioned determination of the first update vector based on the prediction probability of each sub-sample to be classified can be achieved in the following manner: for each candidate category, based on the prediction probability of each sub-sample to be classified, a target sub-sample corresponding to the candidate category is selected from multiple sub-samples to be classified, and based on the sub-sample vector corresponding to the target sub-sample, the category update vector corresponding to the candidate category is determined; vector fusion is performed on the update vectors of each category to obtain the first update vector.

[0265] As an example, the expression of the first update vector can be: T1={T 11 ,…,T 1m} (13)

[0266] Among them, T1 is used to indicate the first update vector, T 11 ,…,T1m is used to indicate the category update vector corresponding to the candidate category, and m is used to indicate the number of candidate categories.

[0267] In some embodiments, the sub-prediction probability corresponding to the candidate category in the predicted probability of the target sub-sample is greater than a probability threshold.

[0268] In some embodiments, the above-mentioned selection of target subsamples corresponding to the candidate category from multiple subsamples to be classified based on the predicted probability of each subsample to be classified can be achieved as follows: the following processing is performed on each subsample to be classified: the sub-prediction probability corresponding to the candidate category in the predicted probability of the subsample to be classified is determined as the target sub-probability of the candidate category corresponding to the subsample to be classified; when the target sub-probability is greater than the probability threshold, the subsample to be classified is determined as the target subsample corresponding to the candidate category.

[0269] As an example, for candidate category B1, the predicted probability of the sub-sample to be classified includes the predicted probabilities corresponding to each candidate category. The sub-prediction probability corresponding to candidate category B1 in the predicted probability of the sub-sample to be classified is determined as the target sub-probability of the sub-sample to be classified corresponding to candidate category B1. When the target sub-probability is greater than the probability threshold, the sub-sample to be classified is determined as the target sub-sample corresponding to the candidate category.

[0270] As an example, for candidate category B2, the predicted probability of the sub-sample to be classified includes the predicted probabilities corresponding to each candidate category. The sub-prediction probability corresponding to candidate category B2 in the predicted probability of the sub-sample to be classified is determined as the target sub-probability of the sub-sample to be classified corresponding to candidate category B2. When the target sub-probability is greater than the probability threshold, the sub-sample to be classified is determined as the target sub-sample corresponding to the candidate category.

[0271] In some embodiments, the above-mentioned determination of the category update vector corresponding to the candidate category based on the subsample vector corresponding to the target subsample can be achieved as follows: when the number of target subsamples is one, the subsample vector corresponding to the target subsample is determined as the category update vector corresponding to the candidate category; when the number of target subsamples is multiple, the subsample vectors corresponding to each target subsample are vector fused to obtain the category update vector corresponding to the candidate category.

[0272] In some embodiments, for each candidate category, when the number of target subsamples corresponding to the candidate category is one, the subsample vector corresponding to the target subsample corresponding to the candidate category is determined as the category update vector corresponding to the candidate category; when the number of target subsamples corresponding to the candidate category is multiple, the subsample vectors corresponding to each target subsample corresponding to the candidate category are vector fused to obtain the category update vector corresponding to the candidate category.

[0273] As an example, when the number of target subsamples corresponding to the candidate category is multiple, the category update vector corresponding to the candidate category can be expressed as: Q1 = {Q 11 ,…,Q 1T} (14)

[0274] Among them, Q1 is used to indicate the category update vector corresponding to the candidate category, Q 11 ,…,Q 1T It is used to indicate the subsample vectors corresponding to each target subsample corresponding to the candidate category, and T is used to indicate the number of target subsamples corresponding to the candidate category.

[0275] In some embodiments, the above-mentioned i-1th vector update of the first update vector based on the subsample vector corresponding to the subsample to be classified to obtain the i-th update vector can be achieved as follows: when i is equal to 2, the first update vector and the subsample vector corresponding to the subsample to be classified are weightedly fused to obtain the second update vector; when i is greater than 2, the i-1th update vector and the subsample vector corresponding to the subsample to be classified are weightedly fused to obtain the i-th update vector.

[0276] As an example, the expression of the above i-th update vector can be:

[0277] in, Used to indicate the i-th update vector, Used to indicate the i-1th update vector, Used to indicate the subsample vector corresponding to the subsample to be classified.

[0278] In step 2052B, the initial label of the subsample to be classified is updated for the i-th time by combining the i-th update vector and the subsample vector corresponding to the subsample to be classified, thereby obtaining the i-th updated label of the subsample to be classified.

[0279] In some embodiments, the above-mentioned combination of the i-th update vector and the subsample vector corresponding to the subsample to be classified, performing the i-th label update on the initial label of the subsample to be classified, and obtaining the i-th updated label of the subsample to be classified can be achieved as follows: obtaining the similarity between the i-th update vector and the subsample vector corresponding to the subsample to be classified; when i is equal to 1, based on the similarity, performing a label update on the initial label of the subsample to be classified, and obtaining the 1st updated label of the subsample to be classified; when i is greater than 1, based on the similarity, updating the i-1th updated label of the subsample to be classified, and obtaining the i-th updated label of the subsample to be classified.

[0280] As an example, the similarity between the i-th update vector and the subsample vector corresponding to the subsample to be classified may be expressed as:

[0281] in, Used to indicate the similarity between the i-th update vector and the subsample vector corresponding to the subsample to be classified, mathopargmax1 c Used to indicate the similarity function, Used to indicate the i-th update vector, Used to indicate the subsample vector corresponding to the subsample to be classified.

[0282] In some embodiments, the above-mentioned similarity-based label update of the initial label of the sub-sample to be classified to obtain the first updated label of the sub-sample to be classified can be achieved as follows: obtain the label value of the initial label of the sub-sample to be classified, and standardize the similarity to obtain the standard value corresponding to the similarity, perform weighted summation on the standard value and the label value of the initial label to obtain the target label value of the first updated label; obtain the target label corresponding to the target label value, and determine the target label as the first updated label.

[0283] In some embodiments, the above-mentioned acquisition of the label value of the initial label of the sub-sample to be classified can be achieved in the following manner: obtaining a label value-label mapping relationship, querying the target index entry including the initial label from the label value-label mapping relationship, and determining the label value in the target index entry as the label value of the initial label of the sub-sample to be classified.

[0284] In some embodiments, the above-mentioned acquisition of the target tag corresponding to the target tag value can be achieved in the following manner: obtaining the tag value-tag mapping relationship, querying the reference index entry including the target tag value from the tag value-tag mapping relationship, and determining the reference tag in the reference index entry as the target tag corresponding to the target tag value.

[0285] In step 2053B, when i is greater than 2, and the i-th updated label and the (i-1)-th updated label are the same, the i-th updated label is determined as the target label of the sub-sample to be classified.

[0286] In some embodiments, when i is greater than 2, and the i-th updated label is the same as the i-1-th updated label, that is, the i-th updated label obtained by the i-th label update is the same as the i-1-th updated label obtained by the i-1-th label update, then it means that the label update has stabilized, and at this time the i-th updated label can be determined as the target label of the sub-sample to be classified.

[0287] In step 2054B, when i is equal to 1 and the first updated label is the same as the initial label, the first updated label is determined as the target label of the sub-sample to be classified.

[0288] In some embodiments, 1≤i≤N, where N indicates the maximum number of times the initial labels of the subsamples to be classified are updated.

[0289] In some embodiments, when i is equal to 1 and the first updated label is the same as the initial label, it means that the initial label and the first updated label are consistent, which means that the first label update has confirmed the accuracy of the initial label, then the first updated label can be directly determined as the target label of the sub-sample to be classified.

[0290] In this way, by performing N rounds of updates on the initial labels of the sub-samples to be classified, when i is greater than 2 and the i-th updated label is the same as the i-1-th updated label, the i-th updated label is determined as the target label of the sub-sample to be classified; when i is equal to 1 and the 1st updated label is the same as the initial label, the 1st updated label is determined as the target label of the sub-sample to be classified, thereby accurately improving the accuracy of the target labels of the sub-samples to be classified, so that the target labels can accurately reflect the categories of the sub-samples to be classified.

[0291] In step 206, the classification model is trained based on the predicted probability, the initial label, and the target label to obtain a target classification model.

[0292] In some embodiments, the target classification model is used to perform category prediction on each sub-image to be classified in the image to be classified to obtain a prediction probability of each sub-image to be classified.

[0293] In some embodiments, referring to FIG13 , FIG13 is a flowchart diagram four of the training method of the image classification model provided in an embodiment of the present application. Step 206 shown in FIG9 can be implemented by executing steps 2061 to 2063 shown in FIG13 .

[0294] In step 2061, the predicted probability and the initial label are combined to determine the first loss value of the classification model.

[0295] In some embodiments, the predicted probability includes sub-prediction probabilities corresponding one-to-one to each candidate category. The above-mentioned step 2061 can be implemented as follows: multiply each sub-prediction probability by the category value corresponding to the initial label to obtain a first reference loss value corresponding to each sub-prediction probability, and sum up the first reference loss values ​​to obtain a second reference loss value; obtain the reference probability corresponding to each sub-prediction probability and the reference category value corresponding to the initial label; multiply each reference probability by the reference category value to obtain a third reference loss value corresponding to each reference probability, and sum up the third reference loss values ​​to obtain a fourth reference loss value; sum the second reference loss value and the fourth reference loss value to obtain the first loss value.

[0296] In some embodiments, the sum of the reference probability and the corresponding sub-prediction probability is equal to 1, and the sum of the reference category value and the category value is equal to 1.

[0297] As an example, the expression of the first reference loss value may be: L1=Y*log(p i ) (17)

[0298] Among them, L1 is used to indicate the first reference loss value, Y is used to indicate the category value corresponding to the initial label, and p i Used to indicate the probability of each sub-prediction.

[0299] As an example, the expression of the second reference loss value may be:

[0300] Among them, L2 is used to indicate the second reference loss value, Y*log(p i ) is used to indicate each first reference loss value.

[0301] As an example, the expression of the third reference loss value corresponding to the above reference probability can be: L3=(1-Y i )*log(1-p i ) (19)

[0302] Among them, L3 is used to indicate the third reference loss value corresponding to the reference probability, log(1-p i ) is used to indicate the reference probability, (1-Y i ) is used to indicate the reference category value.

[0303] As an example, the expression of the fourth reference loss value may be:

[0304] Wherein, L4 is used to indicate the fourth reference loss value, (1-Y i )*log(1-p i) is used to indicate the third reference loss value corresponding to each reference probability.

[0305] As an example, the expression of the first loss value can be: L = L2 + L4 (21)

[0306] Among them, L is used to indicate the first loss value, L2 is used to indicate the second reference loss value, and L4 is used to indicate the fourth reference loss value.

[0307] In this way, by combining the predicted probability and the initial label, the first loss value of the classification model is determined, so that the loss of the classification model is considered from the perspective of the initial label, and the corresponding first loss value is obtained, which facilitates the subsequent training of the classification model based on the first loss value to obtain the target classification model.

[0308] In step 2062, the predicted probability and the target label are combined to determine a second loss value of the classification model.

[0309] In some embodiments, the predicted probability includes sub-prediction probabilities corresponding one-to-one to each candidate category. The above-mentioned step 2062 can be implemented as follows: multiply each sub-prediction probability by the category value corresponding to the target label to obtain the first target loss value corresponding to each sub-prediction probability, and sum up the first target loss values ​​to obtain the second target loss value; obtain the reference probability corresponding to each sub-prediction probability and the target category value corresponding to the target label; multiply each reference probability by the target category value to obtain the third target loss value corresponding to each reference probability, and sum up the third target loss values ​​to obtain the fourth target loss value; sum the second target loss value and the fourth target loss value to obtain the second loss value.

[0310] In some embodiments, the sum of the reference probability and the corresponding sub-prediction probability is equal to 1, and the sum of the reference category value and the category value is equal to 1.

[0311] As an example, the expression of the first target loss value may be: F1=Y e *log(p i ) (twenty two)

[0312] Among them, F1 is used to indicate the first target loss value, Y e Used to indicate the category value corresponding to the target label, p i Used to indicate the probability of each sub-prediction.

[0313] As an example, the expression of the second target loss value can be:

[0314] Among them, F2 is used to indicate the second target loss value, Y e *log(p i) is used to indicate the first target loss value.

[0315] As an example, the expression of the third target loss value may be: F3 = (1-Y e )*log(1-p i ) (twenty four)

[0316] Among them, F3 is used to indicate the third target loss value, (1-Y e ) is used to indicate the reference probability corresponding to each sub-prediction probability, log(1-p i ) is used to indicate the target category value corresponding to the target label.

[0317] As an example, the expression of the fourth target loss value may be:

[0318] Among them, F4 is used to indicate the fourth target loss value, (1-Y e )*log(1-p i ) is used to indicate the third target loss value.

[0319] As an example, the expression of the second loss value can be: F=F2+F4 (26)

[0320] Among them, F is used to indicate the second loss value, F2 is used to indicate the second target loss value, and F4 is used to indicate the fourth target loss value.

[0321] In this way, by combining the predicted probability and the target label, the second loss value of the classification model is determined, so that the loss of the classification model is considered from the perspective of the target label, and the corresponding second loss value is obtained, which facilitates the subsequent training of the classification model based on the second loss value to obtain the target classification model.

[0322] In step 2063, the classification model is trained in combination with the first loss value and the second loss value to obtain a target classification model.

[0323] In some embodiments, the above-mentioned combination of the first loss value and the second loss value to train the classification model to obtain the target classification model can be achieved as follows: weighted sum of the first loss value and the second loss value to obtain the target loss value, and based on the target loss value, the classification model is trained to obtain the target classification model.

[0324] As an example, the expression of the target loss value can be: W=L+F (27)

[0325] Among them, W is used to indicate the target loss value, F is used to indicate the second loss value, and L is used to indicate the first loss value.

[0326] In this way, by combining the predicted probability and the initial label, the first loss value of the classification model is determined, so that the loss of the classification model is considered from the perspective of the initial label, and the corresponding first loss value is obtained. By combining the predicted probability and the target label, the second loss value of the classification model is determined, so that the loss of the classification model is considered from the perspective of the target label, and the corresponding second loss value is obtained. The classification model is trained by combining the first loss value and the second loss value to obtain a target classification model, so that the obtained target classification model considers the loss of the classification model from the perspective of the initial label and the target label, thereby effectively improving the classification performance of the target classification model obtained by training.

[0327] In this way, the following process is performed on each subsample to be classified, i, until the target label of the subsample to be classified is obtained: based on the predicted probability of each subsample to be classified, the i-th update vector is determined, and the i-th update vector is used to perform the i-th update on the initial label of the subsample to be classified; combining the i-th update vector and the subsample vector corresponding to the subsample to be classified, the i-th label update is performed on the initial label of the subsample to be classified, and the i-th updated label of the subsample to be classified is obtained; when i is greater than 2 and the i-th updated label and the i-1-th updated label are the same, the i-th updated label is determined as the target label of the subsample to be classified; when i is equal to 1 and the first updated label and the initial label are the same, the first updated label is determined as the target label of the subsample to be classified. This achieves N rounds of label disambiguation of the initial label, making the obtained target label more accurate.

[0328] Below, an exemplary application of the embodiments of the present application in an actual application scenario of an immune receptor repertoire will be described.

[0329] The human immune system is composed of innate immunity and adaptive immunity. Adaptive immunity is an immune response that can recognize and eliminate specific pathogens after T cells and B cells come into contact with pathogens (antigens). The adaptive immune receptor repertoire (abbreviated as immune repertoire), as a collection of human T cell receptors (TCR) or B cell receptors (BCR), records past and ongoing adaptive immune responses. The state of the immune repertoire can be used to reflect the state of immune activity and the individual's response to infectious diseases, autoimmune diseases and tumor-related pathogens. Therefore, the study of the TCR immune repertoire or the BCR immune repertoire can be applied to predicting the state of human immunity after vaccination or infection, predicting the state of human diseases, especially cancer and immune diseases, and designing immunotherapy.

[0330] The image classification model training method and image classification method provided in the embodiments of this application can train accurate sequence-level classifiers without accurate labels. This not only integrates predictions for each TCR / BCR sequence for repertoire-level classification, but also discovers disease-associated TCR / BCR sequences in the repertoire.

[0331] A large adaptive immune receptor repertoire (AIRR) consists of a large number of adaptive immune receptors (AIR, TCR or BCR). Assume that, given N AIRR, its mathematical representation is {IR1, IR2, ..., IR N Each AIRR contains M AIRs, expressed as It should be noted that the M in the immune repertoire of different individuals is very different. At the same time, the labels of the N immune repertoires are defined as {Y1, Y2, ..., Y N}, where Y∈{0,…,C}, C represents the number of categories, for example, if it is a two-class problem, then C=1. In addition, each AIR has a corresponding frequency, which is expressed as Frequency represents the strength of immune response to a specific antigen. In this embodiment, a mapping function Y=F(IR) can be established to convert the immune repertoire IR into the immune state Y. In this embodiment, the repertoire label Y i (Package-level label) assigned to the recipient These pseudo labels are then updated to appropriate values.

[0332] In some embodiments, refer to Figure 14, which is a schematic diagram of the principle of the image classification method provided in the embodiment of the present application. In the actual application stage, the immune receptor group library group 51 (AIRRs) includes an immune receptor group library 511, an immune receptor group library 512 and an immune receptor group library 513, and the immune receptor group library 511, the immune receptor group library 512 and the immune receptor group library 513 each include multiple immune receptors. By calling the encoding layer 52, each immune receptor group library in the immune receptor group library group 51 is encoded to obtain the embedding features of each immune receptor group library, and by calling the classification layer 53, each immune receptor is classified based on the receptor embedding features of each immune receptor in each immune hand group library to obtain the category of each immune receptor; for each immune receptor group library, the aggregation layer 54 is called to aggregate the category of each immune receptor in the immune receptor group library with the frequency corresponding to each immune receptor to obtain the category of the immune receptor group library. In the training stage, the initial classification layer is trained by the label disambiguation layer 55 shown in Figure 14 to obtain the classification layer 53.

[0333] In some embodiments, referring to FIG14 , the encoding layer 52 shown in FIG14 is described below using the immune receptor repertoire 511 as an example. In order to obtain a comprehensive representation of each immune receptor in the immune receptor repertoire 511, for each immune receptor in the immune receptor repertoire 511, a multimodal fusion method can be used to integrate the immune receptor gene (e.g., V(D)J gene segment) and sequence (e.g., amino acid (AA) sequence). The encoding layer 52 can adopt a gated attention mechanism and then perform tensor fusion. Specifically, the gene encoder (Gene Encoder) composed of a trainable embedding layer performs gene encoding on the digitized V(D)J gene segment to obtain a converted digital vector, which is recorded as h g The embedding dimension of the V gene segment (TCRBV11) is 16, and the embedding dimension of the J gene segment (TCRBJ02) is 8. At the same time, the sequence is encoded using a pre-trained sequence encoder (immune receptor pre-training model) to generate a sequence representation of the immune receptor, which is defined as h s , whose embedding dimension can be 512. The immune receptor pre-training model is a BERT-like model, which includes 6 standard transformer layers, each of which contains 4 attention heads.

[0334] In some embodiments, referring to FIG14 , through a gated attention mechanism, the expression of the gene feature Og shown in FIG14 can be: g =σ(z g )*h′ g (28)

[0335] Among them, z g is the converted digital vector h g and immune receptor sequences expressed as h s The bilinear transformation result of z g The expression can be:

[0336] Among them, b gs Weight matrix parameters learned for the gated-based attention mechanism.

[0337] Where h′ g is the converted digital vector h g The linear transformation result, h′ g The expression of h′ can be: g =ReLu(W g *h g +b g ) (30)

[0338] Among them, Wg and b g is the weight matrix parameter learned by the gate-based attention mechanism, ReLu is used to indicate the activation function, and σ is used to indicate the activation function.

[0339] In some embodiments, referring to FIG14 , through a gated attention mechanism, the expression of the sequence feature Os shown in FIG14 can be: s =σ(z s )*h′ s (31)

[0340] Among them, z s is the converted digital vector h g and immune receptor sequences expressed as h s The bilinear transformation result of z s The expression can be:

[0341] Among them, b sg Weight matrix parameters learned for the gated-based attention mechanism.

[0342] Where h′ s h is the sequence of immune receptor s The linear transformation result, h′ s The expression of h′ can be: s =ReLu(W s *h s +b s ) (33)

[0343] Among them, b s and W s is the weight matrix parameter learned by the gate-based attention mechanism, and ReLu is used to indicate the activation function.

[0344] In some embodiments, referring to FIG14 , tensor fusion is used to g and o s The final expression h is obtained by fusion. The final expression h can be expressed as:

[0345] Among them, W fusion and b fusion is a learnable parameter of tensor fusion, is the Kronecker product.

[0346] In some embodiments, referring to FIG14 , the label disambiguation layer 55 shown in FIG14 is described below. After the representation of the receptor is obtained using the encoder, the predicted value of each receptor is calculated by the following method:

[0347] FC receptor is a classifier whose learnable parameters are and is the jth immune receptor based on the i-th immune repertoire The predicted probability of the multimodal representation of .

[0348] Then, the embodiment of the present application defines a set kec-receptor, and selects K immune receptors for each category c∈{0,…,C}. In the e-th round of training, their The values ​​all exceed the threshold θ, and the expression for the set kec-receptor can be:

[0349] Next, the embodiment of the present application uses a momentum-based update method to update the prototype point. Specifically, the representation of the selected K receptors in the eth round is used to update the representation of the c class in the prototype point during the e+1th round of training. The calculation method is as follows:

[0350] where λ∈[0, 1] is the momentum update coefficient.

[0351] Finally, the labeling of each immune receptor (Initially set to the corresponding immune repertoire label Y i ) can be adjusted based on the similarity between the immune receptor and the prototype point in the $e$th training round.

[0352] Among them, γ is also the momentum update coefficient, is the similarity between the immune receptor and the prototype point in the e-th training round, mathopargmax is the similarity function, is defined as:

[0353] In some embodiments, referring to FIG. 14 , the polymerization layer 54 shown in FIG. 14 is described below. i The predicted value of this application embodiment is the corresponding immune receptor The predicted value and corresponding frequency Aggregate according to the following formula:

[0354] Then, the final immune repertoire prediction results were output by using min-max normalization.

[0355] In some embodiments, referring to FIG14, the first loss value and the second loss value shown in FIG14 are described below. The training phase includes a warm-up phase and a label disambiguation phase. In the warm-up phase, the prototype point update and label disambiguation are temporarily suspended. The loss function at this time is calculated by predicting the result. and the initial label The cross entropy loss between them is used to obtain the first loss value. The expression of the first loss value can be:

[0356] In the label disambiguation phase, the prediction results are calculated With adjusted labels The label disambiguation loss L between disambiguation , which is the second loss value, is used to supervise the model parameters. The expression of the second loss value can be:

[0357] In some embodiments, referring to Figure 15, Figure 15 is a schematic diagram of the classification effect of the image classification method provided in an embodiment of the present application. Figure 15 shows the predicted distribution of the first receptor (CMV-related sequence) and the second receptor (other sequence), wherein the predicted probability of the first receptor is mainly distributed in the probability interval of 0 to 0.6, and the predicted probability of the second receptor is mainly distributed in the probability interval of 0.2 to 1.0. Referring to Figure 16, Figure 16 is a schematic diagram of the classification effect of the image classification method provided in an embodiment of the present application. Figure 16 shows the ROC curve for CMV-related receptor recognition. When the true positive rate is 0.25, the corresponding accuracy rate is 0.6.

[0358] Thus, the embodiments of this application rethink the modeling of the immune system classification problem. A new modeling approach is proposed: a simple sequence-level classifier coupled with robust training. The implementation of the embodiments of this application can be used to infer repertoire status based on the special noisy labels converted from repertoire-level labels, which will promote many cutting-edge research fields, such as the treatment of infectious diseases and autoimmune diseases, and the design of cancer immune vaccines.

[0359] In this way, by performing modal conversion of different modalities on the image to be classified including multiple sub-images to be classified, the modal information of each modality corresponding to the image to be classified is obtained, and the modal information is modally fused to obtain a multimodal vector corresponding to the image to be classified. Based on the multimodal vector, the image to be classified is encoded to obtain the vector to be classified corresponding to the image to be classified. For each sub-image to be classified, based on the sub-vector to be classified corresponding to the sub-image to be classified, the category of the sub-image to be classified is predicted to obtain the category prediction probability corresponding to the sub-image to be classified. Based on the category prediction probability corresponding to each sub-image to be classified, the image to be classified is classified to obtain the category of the image to be classified. In this way, by performing modal fusion on the modal information obtained by performing modal conversion of the image to be classified into different modalities, the obtained multimodal vector can accurately reflect the modal characteristics of the image to be classified from different modalities, and based on the multimodal vector, the vector to be classified is obtained by encoding the image to be classified, thereby effectively improving the information content of the vector to be classified, and based on the vector to be classified corresponding to each sub-image to be classified, the category of the sub-image to be classified is predicted. Since the vector to be classified on which the category prediction is based has a more comprehensive amount of information, the category prediction probability obtained by the prediction is more accurate. Through the accurate category prediction probability, the image to be classified is classified, thereby effectively improving the classification accuracy of the information.

[0360] It is understandable that in the embodiments of the present application, when data related to images to be classified is involved, when the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0361] The following continues to describe an exemplary structure of the image classification device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, as shown in Figure 2, the software modules stored in the image classification device 455 in the memory 450 may include: an acquisition module 4551, configured to acquire an image to be classified including multiple sub-images to be classified, and perform modal conversion of different modalities on the image to be classified to obtain modal information corresponding to each modality of the image to be classified; a modal fusion module 4552, configured to perform modal fusion on each modal information to obtain a multimodal vector corresponding to the image to be classified; an encoding module 4553, configured to encode the image to be classified based on the multimodal vector to obtain a vector to be classified corresponding to the image to be classified, wherein the vector to be classified includes a sub-vector to be classified corresponding to each sub-image to be classified; and a category prediction module 4554, used to classify the image to be classified based on the sub-vector to be classified corresponding to each sub-image to be classified to obtain a category of the image to be classified.

[0362] In some embodiments, the multimodal vector includes multimodal sub-vectors corresponding to each of the modal information. The modal fusion module is further configured to encode each of the modal information to obtain a modal vector corresponding to each of the modal information; and perform the following processing on each of the modal information: determining the modal information as the target modal information, and determining each of the modal information other than the target modal information as the reference modal information; performing vector conversion on the modal vector corresponding to the target modal information to obtain a target modal vector corresponding to the target modal information; performing vector fusion on the modal vector corresponding to the target modal information and the modal vector corresponding to each of the reference modal information to obtain a fused modal vector corresponding to the target modal information; and performing vector fusion on the fused modal vector and the target modal vector to obtain a multimodal sub-vector corresponding to the target modal information.

[0363] In some embodiments, the above-mentioned modal fusion module is further configured to obtain a target conversion vector corresponding to the target modal information, and a target correction vector corresponding to the target modal information, the modal vector corresponding to the target modal information having the same vector dimension as the corresponding target conversion vector and the target correction vector; multiplying the modal vector corresponding to the target modal information and the target conversion vector to obtain the target vector corresponding to the target modal information; adding the target conversion vector and the target correction vector to obtain the correction vector corresponding to the target modal information; and performing vector activation on the target correction vector to obtain the target modal vector corresponding to the target modal information.

[0364] In some embodiments, the above-mentioned modal fusion module is further configured to obtain a reference conversion vector corresponding to the target modal information, and a reference correction vector corresponding to the target modal information, the modal vector corresponding to the target modal information having the same vector dimension as the corresponding reference conversion vector and the reference correction vector; perform vector transposition on the modal vector corresponding to the target modal information to obtain a transposed vector corresponding to the target modal information; multiply the transposed vector by the corresponding reference conversion vector to obtain a first reference vector, and multiply the first reference vector by the modal vector to obtain a second reference vector; add the reference correction vector corresponding to the target modal information to the second reference vector to obtain a fused modal vector corresponding to the target modal information.

[0365] In some embodiments, the multimodal vector includes multimodal sub-vectors corresponding to each of the modal information. The encoding module is further configured to obtain the vector dimension of each of the multimodal sub-vectors and determine the average value of each of the vector dimensions as the target dimension; adjust the vector dimension of each of the multimodal sub-vectors to the target dimension to obtain the target sub-vector corresponding to each of the multimodal sub-vectors; multiply each of the target sub-vectors to obtain a reference multimodal vector, and correct the reference multimodal vector based on the correction parameter to obtain the vector to be classified corresponding to the image to be classified.

[0366] In some embodiments, the above-mentioned encoding module is further configured to obtain a first correction parameter and a second correction parameter, multiply the reference multimodal vector by the first correction parameter to obtain a reference vector to be classified corresponding to the image to be classified; and add the reference vector to be classified to the second correction parameter to obtain a vector to be classified corresponding to the image to be classified.

[0367] In some embodiments, the category prediction is achieved through a target classification model, and the above-mentioned image classification device further includes: a training module, configured to obtain an image sample to be classified including multiple sub-samples to be classified, and a sample label carried by the image sample to be classified; encode the image sample to be classified to obtain a sample vector corresponding to the image sample to be classified, and the sample vector includes a sub-sample vector corresponding to each sub-sample to be classified; for each sub-sample to be classified, the sample label is determined as the initial label of the sub-sample to be classified, and the classification model is called to perform category prediction on the sub-sample to be classified based on the sub-sample vector corresponding to the sub-sample to be classified to obtain the predicted probability of the sub-sample to be classified; based on the predicted probability of each sub-sample to be classified, the target label of each sub-sample to be classified is determined respectively; in combination with the predicted probability, the initial label and the target label, the classification model is trained to obtain the target classification model.

[0368] In some embodiments, the category prediction probability includes a category probability corresponding one-to-one to each candidate category, and the category probability is used to indicate the possibility that the sub-image to be classified is the corresponding candidate category. The above-mentioned classification module is also configured to obtain a weight coefficient corresponding to each sub-image to be classified, and the weight coefficient is used to indicate the importance of the sub-image to be classified in the image to be classified; based on each weight coefficient, the category prediction probability corresponding to each sub-image to be classified is weighted and summed to obtain the target prediction probability corresponding to the image to be classified; wherein the target prediction probability includes a target probability corresponding one-to-one to each candidate category; the candidate category with the largest target probability is determined as the category of the image to be classified.

[0369] In some embodiments, the above-mentioned acquisition module is further configured to, when the image to be classified is multimodal information, acquire each target modality in the image to be classified; for each target modality, extract the modal information of the target modality from the image to be classified, and obtain the modal information of the image to be classified corresponding to the target modality; when the image to be classified is single-modal information, acquire multiple target modalities, and for each target modality, perform modal conversion of the target modality on the image to be classified, and obtain the modal information of the image to be classified corresponding to the target modality.

[0370] The following continues to describe an exemplary structure of the training device 555 of the image classification model provided by the embodiment of the present application implemented as a software module. In some embodiments, as shown in Figure 3, the software modules stored in the image classification device 555 of the memory 550 may include: a sample acquisition module 5551, configured to acquire an image sample to be classified including a plurality of sub-samples to be classified, and a sample label carried by the image sample to be classified; a sample encoding module 5552, configured to encode the image sample to be classified to obtain a sample vector corresponding to the image sample to be classified, wherein the sample vector includes a sub-sample vector corresponding to each of the sub-samples to be classified; a sample prediction module 5553, configured to, for each of the sub-samples to be classified, The sample label is determined as the initial label of the sub-sample to be classified, and the classification model is called to perform category prediction on the sub-sample to be classified based on the sub-sample vector corresponding to the sub-sample to be classified, so as to obtain the predicted probability of the sub-sample to be classified; the label determination module 5554 is configured to determine the target label of each sub-sample to be classified based on the predicted probability of each sub-sample to be classified; the target training module 5555 is configured to train the classification model in combination with the predicted probability, the initial label and the target label, so as to obtain a target classification model; wherein, the target classification model is used to perform category prediction on each sub-image to be classified in the image to be classified, so as to obtain the predicted probability of each sub-image to be classified.

[0371] In some embodiments, the predicted probability includes sub-prediction probabilities corresponding one-to-one to each candidate category. The above-mentioned label determination module is also configured to perform the following processing for each sub-sample to be classified: determining the largest sub-prediction probability among the predicted probabilities of the sub-sample to be classified as the reference prediction probability of the sub-sample to be classified; and determining the candidate category corresponding to the reference prediction probability as the target label of the sub-sample to be classified.

[0372] In some embodiments, the label determination module is further configured to traverse i for each of the subsamples to be classified and perform the following processing until the target label of the subsample to be classified is obtained: based on the predicted probability of each of the subsamples to be classified, determine the i-th update vector, and the i-th update vector is used to update the initial label of the subsample to be classified for the i-th time; combine the i-th update vector and the subsample vector corresponding to the subsample to be classified, and perform the i-th label update on the initial label of the subsample to be classified to obtain the i-th updated label of the subsample to be classified; when i is greater than 2, and the i-th updated label and the i-1-th updated label are the same, determine the i-th updated label as the target label of the subsample to be classified; when i is equal to 1, and the 1st updated label and the initial label are the same, determine the 1st updated label as the target label of the subsample to be classified; wherein 1≤i≤N, N is used to indicate the maximum number of updates of the initial label of the subsample to be classified.

[0373] In some embodiments, the label determination module is further configured to, when i is equal to 1, determine the first update vector based on the predicted probability of each of the sub-samples to be classified; and when i is greater than or equal to 2, perform an i-1th vector update on the first update vector based on the sub-sample vector corresponding to the sub-sample to be classified to obtain the i-th update vector.

[0374] In some embodiments, the predicted probability includes a sub-prediction probability corresponding to each candidate category one by one. The above-mentioned label determination module is further configured to select a target sub-sample corresponding to the candidate category from the multiple sub-samples to be classified for each candidate category based on the predicted probability of each sub-sample to be classified, and determine the category update vector corresponding to the candidate category based on the sub-sample vector corresponding to the target sub-sample; wherein the sub-prediction probability corresponding to the candidate category in the predicted probability of the target sub-sample is greater than a probability threshold; and perform vector fusion on each of the category update vectors to obtain the first update vector.

[0375] In some embodiments, the above-mentioned label determination module is further configured to perform the following processing for each of the sub-samples to be classified: determining the sub-prediction probability corresponding to the candidate category in the prediction probability of the sub-sample to be classified as the target sub-probability of the sub-sample to be classified corresponding to the candidate category; when the target sub-probability is greater than the probability threshold, determining the sub-sample to be classified as the target sub-sample corresponding to the candidate category.

[0376] In some embodiments, the label determination module is further configured to, when the number of the target sub-samples is one, determine the sub-sample vector corresponding to the target sub-sample as the category update vector corresponding to the candidate category; and when the number of the target sub-samples is multiple, perform vector fusion on the sub-sample vectors corresponding to each of the target sub-samples to obtain the category update vector corresponding to the candidate category.

[0377] In some embodiments, the above-mentioned label determination module is further configured to perform the following processing for each of the sub-samples to be classified: determining the sub-prediction probability corresponding to the candidate category in the prediction probability of the sub-sample to be classified as the target sub-probability of the sub-sample to be classified corresponding to the candidate category; when the target sub-probability is greater than the probability threshold, determining the sub-sample to be classified as the target sub-sample corresponding to the candidate category.

[0378] In some embodiments, the label determination module is further configured to, when the number of the target sub-samples is one, determine the sub-sample vector corresponding to the target sub-sample as the category update vector corresponding to the candidate category; and when the number of the target sub-samples is multiple, perform vector fusion on the sub-sample vectors corresponding to each of the target sub-samples to obtain the category update vector corresponding to the candidate category.

[0379] In some embodiments, the label determination module is further configured to, when i is equal to 2, perform weighted fusion on the first update vector and the subsample vector corresponding to the subsample to be classified to obtain the second update vector; and when i is greater than 2, perform weighted fusion on the i-1th update vector and the subsample vector corresponding to the subsample to be classified to obtain the i-th update vector.

[0380] In some embodiments, the label determination module is further configured to obtain the similarity between the i-th update vector and the subsample vector corresponding to the subsample to be classified; when i is equal to 1, based on the similarity, the initial label of the subsample to be classified is updated to obtain the first updated label of the subsample to be classified; when i is greater than 1, based on the similarity, the i-1-th updated label of the subsample to be classified is updated to obtain the i-th updated label of the subsample to be classified.

[0381] In some embodiments, the above-mentioned label determination module is further configured to perform the following processing for each of the sub-samples to be classified: determining the largest sub-prediction probability among the prediction probabilities of the sub-samples to be classified as the reference prediction probability of the sub-sample to be classified; and determining the candidate category corresponding to the reference prediction probability as the target label of the sub-sample to be classified.

[0382] In some embodiments, the above-mentioned target training module is further configured to determine the first loss value of the classification model in combination with the predicted probability and the initial label, and to determine the second loss value of the classification model in combination with the predicted probability and the target label; and to train the classification model in combination with the first loss value and the second loss value to obtain the target classification model.

[0383] In some embodiments, the predicted probability includes sub-prediction probabilities corresponding one-to-one to each candidate category. The above-mentioned target training module is further configured to multiply each of the sub-prediction probabilities by the category value corresponding to the initial label to obtain a first reference loss value corresponding to each of the sub-prediction probabilities, and sum up the first reference loss values ​​to obtain a second reference loss value; obtain the reference probability corresponding to each of the sub-prediction probabilities, and the reference category value corresponding to the initial label, the sum of the reference probability and the corresponding sub-prediction probability is equal to 1, and the sum of the reference category value and the category value is equal to 1; multiply each of the reference probabilities by the reference category value to obtain a third reference loss value corresponding to each of the reference probabilities, and sum up the third reference loss values ​​to obtain a fourth reference loss value; sum the second reference loss value and the fourth reference loss value to obtain the first loss value.

[0384] In some embodiments, the predicted probability includes sub-prediction probabilities corresponding one-to-one to each candidate category. The above-mentioned target training module is further configured to multiply each of the sub-prediction probabilities by the category value corresponding to the target label to obtain a first target loss value corresponding to each of the sub-prediction probabilities, and sum up the first target loss values ​​to obtain a second target loss value; obtain the reference probability corresponding to each of the sub-prediction probabilities and the target category value corresponding to the target label, the sum of the reference probability and the corresponding sub-prediction probability is equal to 1, and the sum of the reference category value and the category value is equal to 1; multiply each of the reference probabilities by the target category value to obtain a third target loss value corresponding to each of the reference probabilities, and sum up the third target loss values ​​to obtain a fourth target loss value; sum the second target loss value and the fourth target loss value to obtain the second loss value.

[0385] The present invention provides a computer program product including a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the image classification method described in the present invention.

[0386] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the image classification method provided by an embodiment of the present application, for example, the image classification method shown in Figure 4.

[0387] The present invention provides a computer program product comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the image classification model training method described in the present invention.

[0388] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will execute the image classification method provided by an embodiment of the present application, for example, the training method of the image classification model shown in Figure 9.

[0389] In some embodiments, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), a flash memory, a magnetic surface memory, an optical disk, or a CD-ROM; or it may be various electronic devices including one or any combination of the above memories.

[0390] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0391] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0392] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0393] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0394] In summary, the embodiments of the present application have the following beneficial effects:

[0395] (1) By performing modal conversion of different modalities on an image to be classified including a plurality of sub-images to be classified, modal information of each modality corresponding to the image to be classified is obtained, and each modal information is modally fused to obtain a multimodal vector corresponding to the image to be classified, and based on the multimodal vector, the image to be classified is encoded to obtain a vector to be classified corresponding to the image to be classified, and for each sub-image to be classified, based on the sub-vector to be classified corresponding to the sub-image to be classified, the category of the sub-image to be classified is predicted to obtain a category prediction probability corresponding to the sub-image to be classified, and based on the category prediction probability corresponding to each sub-image to be classified, the image to be classified is classified to obtain a category of the image to be classified. In this way, by performing modal fusion on the modal information obtained by performing modal conversion of the image to be classified into different modalities, the obtained multimodal vector can accurately reflect the modal characteristics of the image to be classified from different modalities, and based on the multimodal vector, the vector to be classified is obtained by encoding the image to be classified, thereby effectively improving the information content of the vector to be classified, and based on the vector to be classified corresponding to each sub-image to be classified, the category of the sub-image to be classified is predicted. Since the vector to be classified on which the category prediction is based has a more comprehensive amount of information, the category prediction probability obtained by the prediction is more accurate. Through the accurate category prediction probability, the image to be classified is classified, thereby effectively improving the classification accuracy of the information.

[0396] (2) By obtaining an image to be classified including multiple sub-images to be classified and performing modal conversion of different modalities on the image to be classified, the modal information of the image to be classified corresponding to each modality is obtained, thereby realizing the modal conversion of the image to be classified. As a result, the image to be classified can be multi-dimensionally characterized from multiple different modal dimensions, effectively enhancing the amount of information of the image to be classified.

[0397] (3) By fusing the modal information of each modality, the multimodal vector corresponding to the image to be classified is obtained, which significantly enhances the information content of the multimodal vector, and lays a strong data support for the subsequent encoding of the image to be classified based on the multimodal vector.

[0398] (4) By multiplying each target sub-vector, a reference multimodal vector is obtained, and based on the correction parameter, the reference multimodal vector is corrected to obtain the vector to be classified corresponding to the image to be classified, so that the vector to be classified corresponding to the image to be classified is more accurate.

[0399] (5) By determining the largest sub-prediction probability among the prediction probabilities of the sub-samples to be classified as the reference prediction probability of the sub-samples to be classified, and determining the candidate category corresponding to the reference prediction probability as the target label of the sub-samples to be classified, the candidate category corresponding to the largest sub-prediction probability is selected from multiple existing candidate categories as the target label of the sub-samples to be classified, thereby effectively improving the accuracy of the target label.

[0400] (6) By determining the largest sub-prediction probability among the prediction probabilities of the sub-samples to be classified as the reference prediction probability of the sub-samples to be classified, and determining the candidate category corresponding to the reference prediction probability as the target label of the sub-samples to be classified, the candidate category corresponding to the largest sub-prediction probability is selected from multiple existing candidate categories as the target label of the sub-samples to be classified, thereby effectively improving the accuracy of the target label.

[0401] (7) By performing N rounds of updates on the initial labels of the sub-samples to be classified, when i is greater than 2 and the i-th updated label is the same as the i-1-th updated label, the i-th updated label is determined as the target label of the sub-sample to be classified; when i is equal to 1 and the first updated label is the same as the initial label, the first updated label is determined as the target label of the sub-sample to be classified, thereby accurately improving the accuracy of the target labels of the sub-samples to be classified, so that the target labels can accurately reflect the categories of the sub-samples to be classified.

[0402] (8) By performing N rounds of updates on the initial labels of the sub-samples to be classified, when i is greater than 2 and the i-th updated label is the same as the i-1-th updated label, the i-th updated label is determined as the target label of the sub-sample to be classified; when i is equal to 1 and the first updated label is the same as the initial label, the first updated label is determined as the target label of the sub-sample to be classified, thereby accurately improving the accuracy of the target labels of the sub-samples to be classified, so that the target labels can accurately reflect the categories of the sub-samples to be classified.

[0403] (9) By combining the predicted probability and the initial label, the first loss value of the classification model is determined, thereby considering the loss of the classification model from the perspective of the initial label and obtaining the corresponding first loss value. By combining the predicted probability and the target label, the second loss value of the classification model is determined, thereby considering the loss of the classification model from the perspective of the target label and obtaining the corresponding second loss value. The classification model is trained by combining the first loss value and the second loss value to obtain the target classification model, so that the obtained target classification model considers the loss of the classification model from the perspective of the initial label and the perspective of the target label, thereby effectively improving the classification performance of the target classification model obtained by training.

[0404] (10) The following processing is performed on each sub-sample to be classified by traversing i respectively until the target label of the sub-sample to be classified is obtained: based on the predicted probability of each sub-sample to be classified, the i-th update vector is determined, and the i-th update vector is used to perform the i-th update on the initial label of the sub-sample to be classified; combining the i-th update vector and the sub-sample vector corresponding to the sub-sample to be classified, the i-th label update is performed on the initial label of the sub-sample to be classified, and the i-th update label of the sub-sample to be classified is obtained; when i is greater than 2, and the i-th update label and the i-1-th update label are the same, the i-th update label is determined as the target label of the sub-sample to be classified; when i is equal to 1, and the 1st update label and the initial label are the same, the 1st update label is determined as the target label of the sub-sample to be classified. In this way, N rounds of label disambiguation are performed on the initial label, so that the accuracy of the obtained target label is higher.

[0405] (11) The present embodiment rethinks the modeling of the immune system classification problem and proposes a new modeling approach: a simple sequence-level classifier coupled with robust training. The implementation of the present embodiment can be used to infer the repertoire state from the special noisy labels converted from the repertoire-level labels, which will promote many cutting-edge research fields such as the treatment of infectious diseases and autoimmune diseases, and the design of cancer immune vaccines.

[0406] (12) By obtaining the target transformation vector and target correction vector corresponding to the target modal information and ensuring that these vectors are consistent with the modal vector dimensions corresponding to the target modal information, we can perform a series of data processing steps to optimize and improve the representation of modal information. First, the modal vector corresponding to the target modal information is multiplied by the target transformation vector. This step can effectively map the original features to a new feature space and generate a target vector corresponding to the target modal information. This vector may better reflect the intrinsic structure and relationship of the data. Then, by adding the target transformation vector and the target correction vector, we obtain a correction vector. This correction vector takes into account additional adjustment factors, which can be a subtle adjustment or compensation of the original features. Finally, the correction vector is vector activated, which is a nonlinear transformation that can further optimize the representation of the features. The final target modal vector not only contains the original modal information, but also integrates the transformed and corrected information. This processing flow significantly improves the modal vector's representation ability for the target task, which helps to improve the performance of the model and the accuracy of prediction.

[0407] (13) By obtaining the reference transformation vector and reference correction vector corresponding to the target modal information and ensuring that the dimensions of these vectors match the modal vector corresponding to the target modal information, multi-source heterogeneous data can be effectively integrated. Specifically, after the modal vector corresponding to the target modal information is transposed, it is multiplied by the reference transformation vector to obtain the first reference vector, and then the modal vector is multiplied by the first reference vector to obtain the second reference vector. Finally, the reference correction vector is added to the second reference vector. The resulting fused modal vector not only retains the core features of the original modal information, but also enhances its expression accuracy and robustness through correction and transformation, thereby improving the fusion quality of modal information and providing more accurate and comprehensive input for subsequent data analysis and processing. It enhances the adaptability and generalization ability of the model to complex scenarios, which helps to improve the performance of prediction and classification tasks. It optimizes the utilization of data resources and reduces the waste of computing resources through effective vector operations and fusion strategies. It provides a more reliable data foundation for subsequent advanced applications such as feature extraction and sentiment analysis.

[0408] (14) Features of different modalities, such as images and text, are combined to generate a unified feature classification vector to be classified. In this process, the reference multimodal feature modal vectors are first combined through the outer product operation and multiplied by the first correction parameter, which not only adjusts the size and direction of the features, but also strengthens the important correlation between the features. Subsequently, by adding the second correction parameter, the feature representation can be further optimized to make it more in line with the requirements of the classification task. Using the ReLU activation function to process the result vector helps to introduce nonlinearity, making the feature classification vector richer and more discriminative in representation. The beneficial effect of this correction and fusion strategy is reflected in improving the accuracy and robustness of the classification task, enabling the model to better extract and utilize useful information from multimodal data.

[0409] (15) The introduction of weight coefficients can highlight the importance of different sub-information in classification, thereby giving key information a higher influence in comprehensive prediction and reducing the interference of noise information. Secondly, the target prediction probability after weighted summation is more in line with the actual situation because it comprehensively considers the importance and probability distribution of each sub-information. Finally, by comparing the target prediction probabilities, the candidate category with the highest probability is determined as the category of the image to be classified, which can ensure the reliability of the classification results and improve the accuracy and robustness of the overall classification. It has significant advantages in multi-attribute decision-making and complex information processing, and helps to optimize the performance and efficiency of information classification.

[0410] (16) Each modal sample is processed as a target modal sample, and other modal samples are processed as reference modal samples. This method ensures that each modal sample can be paid attention to, and can perform personalized vector transformation based on the characteristics of each modal sample, thereby improving the correlation and complementarity between modalities. Vector transformation of the target modal sample vector can not only highlight the characteristics of the target modal sample, but also enhance its fusion with other modal samples through this transformation. By vector fusion of the target modal sample vector with each reference modal sample vector, a fused modal sample vector that fuses multiple modal information can be obtained. This step helps to integrate information between different modalities and improves the comprehensiveness and accuracy of feature representation. The fused modal sample vector is vector-fused again with the target modal sample vector, and the resulting multimodal sample sub-vector further fuses the original information of the target modal sample and the additional information obtained through fusion. Such a multimodal sample sub-vector not only contains rich feature information, but also provides a more accurate and robust data basis for subsequent classification, recognition or other advanced tasks, thereby helping to improve the performance of the entire multimodal processing system.

[0411] (17) In the multimodal data processing and conversion task, the target conversion vector and target correction vector corresponding to the target modal sample are obtained, and the vector dimension is kept consistent with the modal sample vector, which has the following beneficial effects: First, by multiplying the modal sample vector with the target conversion vector to obtain the target vector, the original modality can be effectively mapped to the target modal space, which helps to achieve effective conversion and feature preservation between modalities. Secondly, the target conversion vector and the target correction vector are added to form a correction vector, and the mapping result is further fine-tuned, which helps to improve the accuracy and fidelity of the converted target modal sample. Finally, vector activation of the correction vector can activate the dimensions with significant characteristics in the target modal sample vector, thereby generating a sample vector that is more consistent with the target modal feature distribution. This series of processing not only improves the efficiency and accuracy of modal conversion, but also enhances the generalization ability of the model, enabling the model to better adapt to the complex conversion relationship between different modalities, and provides strong support for the fusion, analysis and application of multimodal data.

[0412] (18) Through vector transposition and multiplication operations, the original information of the target modal sample vector and the conversion information provided by the reference transformation are effectively combined, which not only enhances the expressiveness of the sample features, but also helps to reveal the correlation between different modalities. By multiplying the transposed vector with the reference transformation vector to obtain the first reference vector, and multiplying the first reference vector with the modal sample vector to obtain the second reference vector, a deeper feature representation can be extracted, corresponding to the transformation and re-expression of the features, respectively, which helps to separate useful signals and suppress noise. The reference correction vector is added to the second reference vector to obtain a fused modal sample vector, which integrates the correction information and the converted feature information, so that the fused modal sample vector retains the original features while increasing the understanding and representation of the target modal sample features, thereby improving the accuracy and efficiency of modal fusion.

[0413] (19) It can ensure the dimensional consistency of multimodal sub-sample vectors and provide a standardized basis for subsequent vector operations. Specifically, by adjusting the dimension of each sub-sample vector to the target dimension, the corresponding target sub-vectors are obtained, which not only simplifies the processing process of multimodal data, but also helps to improve the accuracy of data processing. By multiplying these target sub-vectors, a reference multimodal vector can be synthesized, which integrates the information of different modalities, thereby enhancing the representation ability of sample features. Such an operation helps to explore the intrinsic connection between different modalities and provides more comprehensive and in-depth information for subsequent classification tasks. Correcting the reference multimodal vector based on the correction parameters can further optimize the sample vector to make it more suitable for the requirements of the classification task. The corrected sample vector improves the accuracy and robustness of classification while retaining the original information. The process from dimensional analysis to vector adjustment to correction not only optimizes the multimodal data processing process, but also effectively improves the performance of the sample vector in the classification task, providing strong support for the final classification results.

[0414] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. An image classification method, performed by an electronic device, comprising: Acquire an image to be classified including a plurality of sub-images to be classified, and perform modality conversion of different modalities on the image to be classified to obtain modality information of each modality corresponding to the image to be classified; Perform modal fusion on each of the modal information to obtain a multimodal vector corresponding to the image to be classified; Based on the multimodal vector, the image to be classified is encoded to obtain a sub-vector to be classified corresponding to each sub-image to be classified; Based on the sub-vectors to be classified corresponding to the sub-images to be classified, the images to be classified are classified to obtain the categories of the images to be classified.

2. The method according to claim 1, wherein: The multimodal vector includes multimodal sub-vectors corresponding to the respective modal information, and the modal fusion of the respective modal information to obtain the multimodal vector corresponding to the image to be classified includes: Encode each of the modal information respectively to obtain a modal vector corresponding to each of the modal information; The following processing is performed for each of the modal information: Determining the modal information as target modal information, and determining each modal information except the target modal information as reference modal information; Performing vector conversion on the modal vector corresponding to the target modal information to obtain a target modal vector corresponding to the target modal information; Performing vector fusion on the modal vector corresponding to the target modal information and the modal vectors corresponding to each of the reference modal information to obtain a fused modal vector corresponding to the target modal information; The fused modal vector and the target modal vector are vector-fused to obtain a multimodal sub-vector corresponding to the target modal information.

3. The method according to claim 1 or 2, wherein: The performing vector conversion on the modal vector corresponding to the target modal information to obtain the target modal vector corresponding to the target modal information includes: Obtaining a target conversion vector corresponding to the target modal information and a target correction vector corresponding to the target modal information, wherein the modal vector corresponding to the target modal information has the same vector dimension as the corresponding target conversion vector and the target correction vector; Multiplying the modal vector corresponding to the target modal information and the target conversion vector to obtain a target vector corresponding to the target modal information; Adding the target conversion vector and the target correction vector to obtain a correction vector corresponding to the target modal information; Vector activation is performed on the target correction vector to obtain a target modal vector corresponding to the target modal information.

4. The method according to claim 1 or 2, wherein: The step of fusing the modal vector corresponding to the target modal information and the modal vectors corresponding to each of the reference modal information to obtain a fused modal vector corresponding to the target modal information includes: Acquire a reference conversion vector corresponding to the target modal information and a reference correction vector corresponding to the target modal information, wherein the modal vector corresponding to the target modal information has the same vector dimension as the corresponding reference conversion vector and the reference correction vector; Performing vector transposition on the modal vector corresponding to the target modal information to obtain a transposed vector corresponding to the target modal information; Multiplying the transposed vector by the corresponding reference transformation vector to obtain a first reference vector, and multiplying the first reference vector by the modal vector to obtain a second reference vector; A reference correction vector corresponding to the target modal information is added to the second reference vector to obtain a fused modal vector corresponding to the target modal information.

5. The method according to claim 1, wherein: The multimodal vector includes multimodal sub-vectors corresponding to each of the modal information, and encoding the image to be classified based on the multimodal vector to obtain sub-vectors to be classified corresponding to each of the sub-images to be classified, including: Obtaining a vector dimension of each of the multimodal sub-vectors, and determining an average value of each of the vector dimensions as a target dimension; Adjusting the vector dimensions of each of the multimodal sub-vectors to the target dimensions, respectively, to obtain target sub-vectors corresponding to each of the multimodal sub-vectors; Each of the target sub-vectors is multiplied to obtain a reference multimodal vector, and based on a correction parameter, the reference multimodal vector is corrected to obtain a vector to be classified corresponding to the image to be classified, wherein the vector to be classified includes sub-vectors to be classified corresponding one-to-one to each of the sub-images to be classified.

6. The method according to any one of claims 1 to 5, wherein: The step of correcting the reference multimodal vector based on the correction parameter to obtain the vector to be classified corresponding to the image to be classified includes: Obtaining a first correction parameter and a second correction parameter, and multiplying the reference multimodal vector by the first correction parameter to obtain a reference vector to be classified corresponding to the image to be classified; The reference vector to be classified is added to the second correction parameter to obtain the vector to be classified corresponding to the image to be classified.

7. The method according to claim 1, wherein: The category prediction is implemented by a target classification model. Before the image to be classified is classified based on the sub-vector to be classified corresponding to the sub-image to be classified and the category of the image to be classified is obtained, the method further includes: Acquire an image sample to be classified including a plurality of sub-samples to be classified, and a sample label carried by the image sample to be classified; Encoding the image sample to be classified to obtain a sample vector corresponding to the image sample to be classified, wherein the sample vector includes a subsample vector corresponding to each subsample to be classified; For each of the sub-samples to be classified, the sample label is determined as the initial label of the sub-sample to be classified, and the classification model is called to perform category prediction on the sub-sample to be classified based on the sub-sample vector corresponding to the sub-sample to be classified, so as to obtain the prediction probability of the sub-sample to be classified. Rate; Determining the target label of each of the sub-samples to be classified based on the predicted probability of each of the sub-samples to be classified; The classification model is trained in combination with the predicted probability, the initial label and the target label to obtain the target classification model.

8. The method according to claim 1, wherein: The category prediction probability includes a category probability corresponding to each candidate category one by one, and the category probability is used to indicate the possibility that the sub-image to be classified is the corresponding candidate category. For each sub-image to be classified, based on the sub-vector to be classified corresponding to each sub-image to be classified, the image to be classified is classified to obtain the category of the image to be classified, including: Based on the sub-vector to be classified corresponding to the sub-image to be classified, performing category prediction on the sub-image to be classified to obtain a category prediction probability corresponding to the sub-image to be classified; Obtaining weight coefficients corresponding to the sub-images to be classified, respectively, where the weight coefficients are used to indicate the importance of the sub-images to be classified in the image to be classified; Based on the weight coefficients, weighted summation is performed on the category prediction probabilities corresponding to the sub-images to be classified, so as to obtain the target prediction probability corresponding to the image to be classified; The target prediction probability includes the target probability corresponding to each candidate category one by one; The candidate category with the highest target probability is determined as the category of the image to be classified.

9. The method according to claim 1, wherein: The performing modality conversion of different modalities on the images to be classified to obtain modality information of the images to be classified corresponding to the modalities respectively includes: When the image to be classified is multimodal information, obtaining each target modality in the image to be classified; For each of the target modalities, extracting modality information of the target modality from the image to be classified, to obtain modality information of the image to be classified corresponding to the target modality; When the image to be classified is single-modal information, a plurality of the target modalities are obtained, and for each of the target modalities, the image to be classified is subjected to modality conversion of the target modality to obtain modality information of the image to be classified corresponding to the target modality.

10. A method for training an image classification model, the method comprising: Acquire an image sample to be classified including a plurality of sub-samples to be classified, and a sample label carried by the image sample to be classified; Encoding the image sample to be classified to obtain a sample vector corresponding to the image sample to be classified, wherein the sample vector includes a subsample vector corresponding to each subsample to be classified; For each of the sub-samples to be classified, the sample label is determined as the initial label of the sub-sample to be classified, and the classification model is called to perform category prediction on the sub-sample to be classified based on the sub-sample vector corresponding to the sub-sample to be classified, so as to obtain the prediction probability of the sub-sample to be classified; Determining the target label of each of the sub-samples to be classified based on the predicted probability of each of the sub-samples to be classified; Training the classification model by combining the predicted probability, the initial label and the target label to obtain a target classification model; The target classification model is used to perform category prediction on each sub-image to be classified in the image to be classified, and obtain the prediction probability of each sub-image to be classified.

11. The method according to claim 10, wherein: The step of determining the target label of each of the sub-samples to be classified based on the predicted probability of each of the sub-samples to be classified comprises: For each of the sub-samples to be classified, traverse i and perform the following processing until the target label of the sub-sample to be classified is obtained: Based on the predicted probability of each of the sub-samples to be classified, determining an i-th update vector, wherein the i-th update vector is used to perform an i-th update on the initial label of the sub-sample to be classified; Combining the i-th update vector and the subsample vector corresponding to the subsample to be classified, performing an i-th label update on the initial label of the subsample to be classified to obtain an i-th updated label of the subsample to be classified; When i is greater than 2, and the i-th updated label is the same as the i-1-th updated label, the i-th updated label is determined as the target label of the sub-sample to be classified; When i is equal to 1, and the first updated label is the same as the initial label, the first updated label is determined as the target label of the sub-sample to be classified; Wherein, 1≤i≤N, and N is used to indicate the maximum number of updates of the initial label of the sub-sample to be classified.

12. The method according to claim 10 or 11, wherein: The step of determining the i-th update vector based on the predicted probability of each of the sub-samples to be classified comprises: When i is equal to 1, determining a first update vector based on the predicted probability of each of the sub-samples to be classified; When i is greater than or equal to 2, the first update vector is updated for an i-1th time based on the subsample vector corresponding to the subsample to be classified to obtain the i-th update vector.

13. The method according to any one of claims 10 to 12, wherein: The predicted probability includes sub-prediction probabilities corresponding to each candidate category one by one, and determining the first update vector based on the predicted probability of each sub-sample to be classified includes: For each of the candidate categories, based on the predicted probability of each of the sub-samples to be classified, a target sub-sample corresponding to the candidate category is selected from the multiple sub-samples to be classified, and based on the sub-sample vector corresponding to the target sub-sample, a category update vector corresponding to the candidate category is determined; The sub-prediction probability corresponding to the candidate category in the prediction probability of the target sub-sample is greater than the probability threshold; Vector fusion is performed on each of the category update vectors to obtain the first update vector.

14. The method according to any one of claims 10 to 13, wherein: The selecting, from the plurality of sub-samples to be classified based on the predicted probability of each of the sub-samples to be classified, a target sub-sample corresponding to the candidate category, comprises: The following processing is performed for each of the sub-samples to be classified: Determine the sub-prediction probability corresponding to the candidate category in the prediction probability of the sub-sample to be classified as the target sub-probability of the sub-sample to be classified corresponding to the candidate category; When the target sub-probability is greater than the probability threshold, the sub-sample to be classified is determined as the target sub-sample corresponding to the candidate category.

15. The method according to any one of claims 10 to 13, wherein: The determining, based on the subsample vector corresponding to the target subsample, a category update vector corresponding to the candidate category includes: When the number of the target sub-sample is one, determining the sub-sample vector corresponding to the target sub-sample as the category update vector corresponding to the candidate category; When there are multiple target sub-samples, vector fusion is performed on the sub-sample vectors corresponding to the target sub-samples to obtain a category update vector corresponding to the candidate category.

16. The method according to any one of claims 10 to 13, wherein: The step of performing an i-1th vector update on the first update vector based on the subsample vector corresponding to the subsample to be classified to obtain the i-th update vector includes: When i is equal to 2, weighted fusion is performed on the first update vector and the subsample vector corresponding to the subsample to be classified to obtain a second update vector; When i is greater than 2, the i-1th update vector and the subsample vector corresponding to the subsample to be classified are weightedly fused to obtain the i-th update vector.

17. The method according to any one of claims 10 to 13, wherein: The step of combining the i-th update vector and the subsample vector corresponding to the subsample to be classified, performing an i-th label update on the initial label of the subsample to be classified, and obtaining the i-th updated label of the subsample to be classified includes: Obtaining the similarity between the i-th update vector and the sub-sample vector corresponding to the sub-sample to be classified; When i is equal to 1, based on the similarity, the initial label of the sub-sample to be classified is updated to obtain the first updated label of the sub-sample to be classified; When i is greater than 1, based on the similarity, the i-1th updated label of the sub-sample to be classified is updated to obtain the i-th updated label of the sub-sample to be classified.

18. The method according to claim 10, wherein: The predicted probability includes sub-prediction probabilities corresponding to each candidate category one by one, and the target labels of each sub-sample to be classified are determined based on the predicted probability of each sub-sample to be classified, including: The following processing is performed for each of the sub-samples to be classified: Determine the largest sub-prediction probability among the prediction probabilities of the sub-samples to be classified as the reference prediction probability of the sub-samples to be classified; The candidate category corresponding to the reference prediction probability is determined as the target label of the sub-sample to be classified.

19. The method according to claim 10, wherein: The step of training the classification model by combining the predicted probability, the initial label and the target label to obtain a target classification model includes: Determine a first loss value of the classification model by combining the predicted probability and the initial label, and determine a second loss value of the classification model by combining the predicted probability and the target label; The classification model is trained in combination with the first loss value and the second loss value to obtain the target classification model.

20. The method according to any one of claims 10 to 19, wherein: The predicted probability includes sub-prediction probabilities corresponding to each candidate category one by one, and combining the predicted probability and the initial label to determine the first loss value of the classification model includes: Multiplying each of the sub-prediction probabilities by the category value corresponding to the initial label to obtain a first reference loss value corresponding to each of the sub-prediction probabilities, and summing the first reference loss values ​​to obtain a second reference loss value; Obtain reference probabilities corresponding to the sub-prediction probabilities and reference category values ​​corresponding to the initial labels, wherein the sum of the reference probabilities and the corresponding sub-prediction probabilities is equal to 1, and the sum of the reference category value and the category value is equal to 1; Multiplying each of the reference probabilities by the reference category value to obtain a third reference loss value corresponding to each of the reference probabilities, and summing the third reference loss values ​​to obtain a fourth reference loss value; The second reference loss value and the fourth reference loss value are summed to obtain the first loss value.

21. The method according to any one of claims 10 to 19, wherein: The predicted probability includes sub-prediction probabilities corresponding to each candidate category one by one, and combining the predicted probability and the target label to determine the second loss value of the classification model includes: Multiplying each of the sub-prediction probabilities by the category value corresponding to the target label to obtain a first target loss value corresponding to each of the sub-prediction probabilities, and summing the first target loss values ​​to obtain a second target loss value; Obtain reference probabilities corresponding to the sub-prediction probabilities and target category values ​​corresponding to the target labels, wherein the sum of the reference probabilities and the corresponding sub-prediction probabilities is equal to 1, and the sum of the reference category value and the category value is equal to 1; Multiplying each of the reference probabilities by the target category value respectively to obtain a third target loss value corresponding to each of the reference probabilities, and summing the third target loss values ​​to obtain a fourth target loss value; The second target loss value and the fourth target loss value are summed to obtain the second loss value.

22. An image classification device, the device comprising: an acquisition module, configured to acquire an image to be classified including a plurality of sub-images to be classified, and perform modality conversion of different modalities on the image to be classified to obtain modality information of each modality corresponding to the image to be classified; A modality fusion module, configured to perform modality fusion on each of the modal information to obtain a multi-modal vector corresponding to the image to be classified; an encoding module configured to encode the image to be classified based on the multimodal vector to obtain a sub-vector to be classified corresponding to each sub-image to be classified; The category prediction module is configured to classify the image to be classified based on the sub-vector to be classified corresponding to each sub-image to be classified, so as to obtain the category of the image to be classified.

23. A training device for an image classification model, the device comprising: A sample acquisition module, configured to acquire an image sample to be classified including a plurality of sub-samples to be classified, and a sample label carried by the image sample to be classified; A sample encoding module, configured to encode the image sample to be classified to obtain a sample vector corresponding to the image sample to be classified, wherein the sample vector includes a subsample vector corresponding to each subsample to be classified; The sample prediction module is configured to determine the sample label as the initial label of the sub-sample to be classified for each sub-sample to be classified, and call the classification model to perform category prediction on the sub-sample to be classified based on the sub-sample vector corresponding to the sub-sample to be classified, so as to obtain the prediction probability of the sub-sample to be classified; A label determination module, configured to determine a target label for each of the sub-samples to be classified based on the predicted probability of each of the sub-samples to be classified; The target training module is configured to train the classification model in combination with the predicted probability, the initial label and the target label to obtain a target classification model; wherein the target classification model is used to perform category prediction on each sub-image to be classified in the image to be classified to obtain the predicted probability of each sub-image to be classified.

24. An electronic device, comprising: A memory for storing computer executable instructions or computer programs; A processor, configured to implement the method according to any one of claims 1 to 21 when executing computer executable instructions or computer programs stored in the memory.

25. A computer-readable storage medium storing computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program, when executed by a processor, implement the method according to any one of claims 1 to 21.

26. A computer program product, comprising a computer program or computer executable instructions, wherein the computer program or the computer executable instructions, when executed by a processor, implements the method according to any one of claims 1 to 21.

Citation Information

Patent Citations

  • Vehicle detection model training method and device, vehicle detection method and electronic equipment

    CN114399657A

  • Image classification method based on spatial transformation network and convolutional neural network

    CN115631377A

  • Table information extraction method and device, electronic equipment and readable storage medium

    CN116343248A

  • Text information extraction method and device, storage medium and computer equipment

    CN116503877A

  • Image object classification method, system and computer readable medium

    US20230067231A1