Classification Method, Device, Equipment and Storage Medium Based on Multimodal Learning Model
By extracting image and text features and fusion of features, the problem of low accuracy in traditional image classification methods is solved, and a higher image classification accuracy is achieved.
Patent Information
- Application Number
- CN202210422050.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-04-21
AI Technical Summary
Traditional image processing methods fail to effectively utilize text features in the image when classifying images, resulting in low classification accuracy.
The classification method based on multimodal learning model is adopted, and image features and text features are extracted and mapped into the dimension space of the pre-constructed feature fusion model, the fusion feature vector is calculated, and finally the pre-trained classification model is input for classification.
Improve the accuracy of image classification, ensure the integrity of each information in the image, and enhance the classification ability of using pre-trained models.
Smart Images

Figure CN114708461B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a classification method, device, equipment and storage medium based on a multi-modal learning model. Background Art
[0002] In traditional image processing, only the more prominent image features in the image are often used for a series of business operations, without considering the text features carried in the image, which results in a large classification deviation in the performance of image classification and poor accuracy of image processing. Summary of the Invention
[0003] The present invention provides a classification method, device, equipment and storage medium based on a multi-modal learning model, and its main purpose is to solve the problem of low accuracy in image classification.
[0004] To achieve the above object, a classification method based on a multi-modal learning model provided by the present invention includes:
[0005] Obtain an image to be processed, and extract the image features of the image to be processed;
[0006] Extract the text features of the image to be processed to obtain a text feature set;
[0007] Map the image features and the text feature set in the dimension space of a pre-constructed feature fusion model to obtain a feature fusion space, and calculate the fusion feature vectors of the image features and each text feature in the feature fusion space respectively;
[0008] Input the fusion feature vectors into a pre-trained classification model to perform a classification operation, and classify the target category of the image to be processed.
[0009] Optionally, the extracting the image features of the image to be processed includes:
[0010] Perform image normalization processing on the image to be processed to obtain a normalized image;
[0011] Perform convolution pooling operation on the normalized image by using a pre-constructed convolutional neural network model based on the DenseNet algorithm to obtain the image features.
[0012] Optionally, the performing image normalization processing on the image to be processed to obtain a normalized image includes:
[0013] According to the pixel matrix size of the image to be processed, cut the image to be processed into multiple pixel blocks;
[0014] Extract the grayscale value of each pixel block, and calculate the grayscale mean and grayscale variance of each pixel block;
[0015] Use the grayscale mean, grayscale variance of each pixel block, and the preset initial grayscale value and grayscale standard deviation to readjust the grayscale value of each pixel block, and integrate the grayscale values of each pixel block to obtain a standardized image.
[0016] Optionally, the extraction of text features from the image to be processed includes:
[0017] Extract the text in the image to be processed;
[0018] Use a pre-constructed tokenizer to perform tokenization on the text to obtain a tokenized text;
[0019] Use a preset word vector conversion model to convert the tokenized text into word vectors to obtain the text features.
[0020] Optionally, the mapping of the image features and the text feature set in the dimension space of a pre-constructed feature fusion model to obtain a feature fusion space includes:
[0021] The dimension space of the feature fusion model determines the starting point and ending point of the image features;
[0022] Successively map each text feature of the text feature set in the dimension space;
[0023] Summarize the image features and each text feature in the dimension space to obtain the feature fusion space.
[0024] Optionally, the calculation of the fusion feature vectors of the image features and each text feature in the feature fusion space includes:
[0025] The fusion feature vectors of the image features and each text feature in the feature fusion space can be calculated by the following formula:
[0026] F = D 3 (D 1 (F Image ) + D 2 (F Text ));
[0027] Where D 1 , D 2 , D 3 are the fully connected layers of the feature fusion model, F Image is the image feature, F Text is the text feature, and F is the fusion feature vector.
[0028] Optionally, inputting the fused feature vector into a pre-trained classification model to perform classification operations includes:
[0029] Calculating the probability values of the fused feature vector hitting a plurality of preset classification labels;
[0030] Sorting the probability values and extracting the classification labels corresponding to a preset number of probability values with the top rankings;
[0031] Obtaining the target category of the image to be processed according to the classification labels.
[0032] To solve the above problems, the present invention also provides a classification device based on a multi-modal learning model, and the device includes:
[0033] A feature acquisition module, configured to acquire an image to be processed and extract the image features of the image to be processed; extract the text features of the image to be processed to obtain a text feature set;
[0034] A feature fusion module, configured to map the image features and the text feature set in the dimension space of a pre-constructed feature fusion model to obtain a feature fusion space, and calculate the fused feature vectors of the image features and each text feature in the feature fusion space respectively;
[0035] A classification module, configured to input the fused feature vector into a pre-trained classification model to perform classification operations and classify the target category of the image to be processed.
[0036] To solve the above problems, the present invention also provides an electronic device, and the electronic device includes:
[0037] At least one processor; and,
[0038] A memory communicatively connected to the at least one processor; wherein,
[0039] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned classification method based on a multi-modal learning model.
[0040] To solve the above problems, the present invention also provides a computer-readable storage medium, and at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned classification method based on a multi-modal learning model.
[0041] In the embodiments of the present invention, by extracting the image features and text features of an image, the integrity of each piece of information in the image can be ensured. Additionally, by fusing the extracted image features and text features, the feature information contained in the fused features after fusion can be made more complete, and at the same time, the accuracy of classification using a pre-trained classification model is improved. Therefore, the classification method, device, equipment, and storage medium based on a multi-modal learning model proposed by the present invention can solve the problem of low accuracy in image classification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 FIG. is a schematic flowchart of a classification method based on a multi-modal learning model provided by an embodiment of the present invention;
[0043] Figure 2 FIG. is a functional module diagram of a classification device based on a multi-modal learning model provided by an embodiment of the present invention;
[0044] Figure 3 FIG. is a schematic structural diagram of an electronic device for implementing the classification method based on a multi-modal learning model provided by an embodiment of the present invention.
[0045] The implementation, functional features, and advantages of the objectives of the present invention will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] An embodiment of the present application provides a classification method based on a multi-modal learning model. The execution subject of the classification method based on a multi-modal learning model includes, but is not limited to, at least one of an electronic device such as a server, a terminal, etc. that can be configured to execute the method provided by the embodiment of the present application. In other words, the classification method based on a multi-modal learning model can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0048] Refer to Figure 1As shown in the figure, it is a schematic flowchart of a classification method based on a multi-modal learning model provided by an embodiment of the present invention. In this embodiment, the classification method based on the multi-modal learning model includes:
[0049] Step S1: Obtain an image to be processed and extract the image features of the image to be processed.
[0050] In the embodiment of the present invention, the image to be processed refers to an image that needs to be recognized and classified. For example, in an e-commerce scenario, the image to be processed may be a product picture, and the product picture may include multiple items. The present invention can identify a target item from multiple items in the product picture.
[0051] Specifically, in the embodiment of the present invention, the image features of the image to be processed can be extracted through a convolutional neural network model based on the DenseNet algorithm.
[0052] As an embodiment of the present invention, the extraction of the image features of the image to be processed includes: performing image normalization processing on the image to be processed to obtain a normalized image; performing a convolutional pooling operation on the normalized image by using a pre-constructed convolutional neural network model based on the DenseNet algorithm to obtain the image features.
[0053] In the embodiment of the present invention, the DenseNet algorithm refers to a neural network that can maximize the use of the image features of an image for convolutional processing. Using the DenseNet algorithm can effectively reduce the vanishing gradient during image convolutional processing, can strengthen the transmission of image features, can more effectively utilize image features, and can reduce the number of calculation parameters.
[0054] Specifically, the performing image normalization processing on the image to be processed to obtain a normalized image includes:
[0055] According to the pixel matrix size of the image to be processed, the image to be processed is sliced into multiple pixel blocks;
[0056] Extract the gray value of each pixel block, and calculate the gray mean value and gray variance of each pixel block;
[0057] Use the gray mean value, the gray variance of each pixel block, a preset gray initial value, and a gray standard deviation to re-adjust the gray value of each pixel block, and integrate the gray values of each pixel block to obtain a normalized image.
[0058] In the embodiment of the present invention, the following formula can be used to calculate the gray mean value mean of each pixel block:
[0059]
[0060] Wherein, W is the width of the size of the pixel matrix, H is the height of the size of the pixel matrix, and I(x, y) is the gray value of the image to be processed at (x, y);
[0061] In the embodiments of the present invention, the following formula can be used to calculate the gray variance var of each pixel block:
[0062]
[0063] In the embodiments of the present invention, the following formula can be used to readjust the gray value of each pixel block
[0064]
[0065] Wherein, m 0 is the initial gray value, and v 0 is the standard deviation of the gray value.
[0066] Step S2: Extract the text features of the image to be processed to obtain a text feature set.
[0067] In the embodiments of the present invention, the text features refer to the word vector features obtained by converting the text extracted from the image to be processed using a pre-constructed word vector conversion model.
[0068] As an embodiment of the present invention, the extraction of the text features in the image to be processed includes: extracting the text in the image to be processed; performing word segmentation on the text using a pre-constructed word segmenter to obtain a segmented text; and converting the segmented text into word vectors using a preset word vector conversion model to obtain the text features. For example, a word vector conversion model based on the bert network can be used to convert the text extracted from the image to be processed into word vectors.
[0069] In the embodiments of the present invention, before converting the segmented text into word vectors using a preset word vector conversion model, it may further include: removing the stop words in the segmented text using a preset stop word list.
[0070] Step S3: Map the image features and the text feature set in the dimension space of a pre-constructed feature fusion model to obtain a feature fusion space, and calculate the fusion feature vectors of the image features and each text feature in the feature fusion space respectively.
[0071] In the embodiments of the present invention, the dimension space refers to a multi-dimensional space for implementing the mapping operation of the image features and the text features.
[0072] In the embodiments of the present invention, the text features and the image features are mapped to the same feature dimension space to realize the feature fusion of the image features and the text features.
[0073] In an embodiment of the present invention, mapping the image feature and the text feature set in the dimension space of a pre-constructed feature fusion model to obtain a feature fusion space includes:
[0074] The dimension space of the feature fusion model determines the starting point and the ending point of the image feature;
[0075] Sequentially perform mapping of each text feature in the text feature set in the dimension space;
[0076] Summarize the image feature and each text feature in the dimension space to obtain the feature fusion space.
[0077] In an embodiment of the present invention, the fusion feature vector of the image feature and each text feature in the feature fusion space can be calculated by the following formula:
[0078] F = D 3 (D 1 (F Image ) + D 2 (F Text ));
[0079] Where D 1 , D 2 , D 3 are the fully connected layers of the feature fusion model, F Image is the image feature, F Text is the text feature, and F is the fusion feature vector.
[0080] Step S4: Input the fusion feature vector into a pre-trained classification model to perform a classification operation, and classify the target category of the image to be processed.
[0081] In an embodiment of the present invention, the pre-trained classification model refers to a classification algorithm model obtained after training a pre-constructed classification model using a large amount of training corpus related to this solution.
[0082] Specifically, inputting the fusion feature vector into a pre-trained classification model to perform a classification operation includes: calculating the probability values of the fusion feature vector hitting a preset plurality of classification labels; sorting the probability values, and extracting the classification labels corresponding to a preset number of probability values with the top rankings; obtaining the target category of the image to be processed according to the classification labels.
[0083] In an embodiment of the present invention, the target category refers to the specific category to which the image to be processed belongs. For example, in the field of product classification, if the image is a top and the text descriptions are pure cotton, black, and loose, then the target categories in the above example are "top & pure cotton", "top & black", and "top & loose".
[0084] For example, in the field of product classification, after fusing the image features and text features in the product image, the specific target category of the product image can be classified by using a large number of training corpora related to this solution to a pre-constructed classification model.
[0085] In an embodiment of the present invention, by extracting the image features and text features of the image, the integrity of each piece of information in the image can be ensured. In addition, by fusing the extracted image features and text features, the feature information contained in the fused features can be made more complete, and at the same time, the accuracy of classification using a pre-trained classification model is improved.
[0086] As Figure 2 shown, it is a functional module diagram of a classification device based on a multi-modal learning model provided by an embodiment of the present invention.
[0087] The classification device 100 based on the multi-modal learning model of the present invention can be installed in an electronic device. According to the functions achieved, the classification device 100 based on the multi-modal learning model can include a feature acquisition module 101, a feature fusion module 102, and a classification module 103. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.
[0088] In this embodiment, the functions of each module / unit are as follows:
[0089] The feature acquisition module 101 is used to acquire the image to be processed and extract the image features of the image to be processed; extract the text features of the image to be processed to obtain a text feature set.
[0090] In an embodiment of the present invention, the image to be processed refers to an image that needs to be recognized and classified. For example, in an e-commerce scenario, the image to be processed can be a product picture, and the product picture may include multiple items. The present invention can identify the target item from the multiple items in the product picture.
[0091] Specifically, in an embodiment of the present invention, the image features of the image to be processed can be extracted by a convolutional neural network model based on the DenseNet algorithm.
[0092] As an embodiment of the present invention, extracting the image features of the image to be processed includes: performing image normalization processing on the image to be processed to obtain a normalized image; and performing convolution pooling operations on the normalized image by using a pre-constructed convolutional neural network model based on the DenseNet algorithm to obtain the image features.
[0093] In the embodiment of the present invention, the DenseNet algorithm refers to a neural network that can maximize the use of image features for convolution processing. Using the DenseNet algorithm can effectively reduce the vanishing gradient during image convolution processing, strengthen the transmission of image features, more effectively utilize image features, and reduce the number of calculation parameters.
[0094] Specifically, performing image normalization processing on the image to be processed to obtain a normalized image includes:
[0095] According to the size of the pixel matrix of the image to be processed, splitting the image to be processed into multiple pixel blocks;
[0096] Extracting the gray value of each pixel block, and calculating the gray mean value and gray variance of each pixel block;
[0097] Using the gray mean value, the gray variance, the preset gray initial value, and the gray standard deviation of each pixel block to readjust the gray value of each pixel block, and integrating the gray values of each pixel block to obtain a normalized image.
[0098] In the embodiment of the present invention, the following formula can be used to calculate the gray mean value mean of each pixel block:
[0099]
[0100] where W is the width of the pixel matrix size, H is the height of the pixel matrix size, and I(x, y) is the gray value of the image to be processed at (x, y);
[0101] In the embodiment of the present invention, the following formula can be used to calculate the gray variance var of each pixel block:
[0102]
[0103] In the embodiment of the present invention, the following formula can be used to readjust the gray value of each pixel block
[0104]
[0105] where m 0 is the gray initial value, v 0is the standard deviation of the grayscale.
[0106] In the embodiment of the present invention, the text feature refers to the word vector feature obtained by converting the text extracted from the image to be processed using a pre-constructed word vector conversion model.
[0107] As an embodiment of the present invention, extracting the text feature from the image to be processed includes: extracting the text in the image to be processed; performing word segmentation on the text using a pre-constructed word segmenter to obtain segmented text; and converting the segmented text into word vectors using a preset word vector conversion model to obtain the text feature. For example, a word vector conversion model based on the bert network can be used to convert the text extracted from the image to be processed into word vectors.
[0108] In the embodiment of the present invention, before converting the segmented text into word vectors using a preset word vector conversion model, it may further include: removing stop words in the segmented text using a preset stop word list.
[0109] The feature fusion module 102 is configured to map the image feature and the text feature set in the dimension space of a pre-constructed feature fusion model to obtain a feature fusion space, and calculate the fusion feature vectors of the image feature and each text feature in the feature fusion space respectively;
[0110] In the embodiment of the present invention, the dimension space refers to a multi-dimensional space for implementing the mapping operation of the image feature and the text feature.
[0111] In the embodiment of the present invention, the text feature and the image feature are mapped to the same feature dimension space to realize the feature fusion of the image feature and the text feature.
[0112] In the embodiment of the present invention, mapping the image feature and the text feature set in the dimension space of a pre-constructed feature fusion model to obtain a feature fusion space includes:
[0113] The dimension space of the feature fusion model determines the starting point and the ending point of the image feature;
[0114] Sequentially perform mapping of each text feature in the text feature set in the dimension space;
[0115] Summarize the image feature and each text feature in the dimension space to obtain the feature fusion space.
[0116] In the embodiment of the present invention, the fusion feature vectors of the image feature and each text feature in the feature fusion space can be calculated by the following formula:
[0117] F = D 3 (D 1 (FImage ) + D 2 (F Text ));
[0118] Where D 1 、D 2 、D 3 are the fully connected layers of the feature fusion model, F Image is the image feature, F Text is the text feature, and F is the fused feature vector.
[0119] The classification module 103 is configured to input the fused feature vector into a pre-trained classification model to perform a classification operation, and classify the target category of the image to be processed.
[0120] In the embodiments of the present invention, the pre-trained classification model refers to a classification algorithm model obtained after training a pre-constructed classification model using a large number of training corpora related to this solution.
[0121] Specifically, inputting the fused feature vector into the pre-trained classification model to perform a classification operation includes: calculating the probability values of the fused feature vector hitting a preset plurality of classification labels; sorting the probability values, and extracting the classification labels corresponding to a preset number of probability values with the top rankings; obtaining the target category of the image to be processed according to the classification labels.
[0122] In the embodiments of the present invention, the target category refers to the specific category to which the image to be processed belongs. For example, in the field of commodity classification, if the image is a top and the text description is pure cotton, black, and loose, then the target categories in the above example are "top & pure cotton", "top & black", and "top & loose".
[0123] For example, in the field of commodity classification, after fusing the image features and text features in the commodity image, the specific target category of the commodity image can be classified by using a large number of training corpora related to this solution to a pre-constructed classification model.
[0124] In the embodiments of the present invention, by extracting the image features and text features of the image, the integrity of each piece of information in the image can be ensured. In addition, by fusing the extracted image features and text features, the feature information contained in the fused feature can be made more complete, and at the same time, the accuracy of classification using the pre-trained classification model is improved.
[0125] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing a classification method based on a multi-modal learning model provided by an embodiment of the present invention.
[0126] The electronic device 1 may include a processor 10, a memory 11, a communication bus 12, and a communication interface 13. It may also include a computer program stored in the memory 11 and executable on the processor 10, such as a classification program based on a multimodal learning model.
[0127] Among them, in some embodiments, the processor 10 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. By running or executing programs or modules stored in the memory 11 (such as executing a classification program based on a multimodal learning model, etc.), and calling data stored in the memory 11, it performs various functions of the electronic device and processes data.
[0128] The memory 11 includes at least one type of readable storage medium, which includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical discs, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device, such as the mobile hard disk of the electronic device. In some other embodiments, the memory 11 may also be an external storage device of the electronic device, such as a plug-in mobile hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can be used not only to store application software installed on the electronic device and various types of data, such as the code of a classification program based on a multimodal learning model, etc., but also to temporarily store data that has been output or will be output.
[0129] The communication bus 12 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection communication between the memory 11 and at least one processor 10, etc.
[0130] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), and is generally used to establish a communication connection between this electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device and to display a visual user interface.
[0131] Figure 3 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0132] For example, although not shown, the electronic device may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or an inverter, and a power status indicator. The electronic device may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0133] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0134] The classification program based on the multi-modal learning model stored in the memory 11 in the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:
[0135] Obtain an image to be processed, and extract the image features of the image to be processed;
[0136] Extract the text features of the image to be processed to obtain a text feature set;
[0137] Map the image features and the text feature set in the dimension space of a pre-constructed feature fusion model to obtain a feature fusion space, and calculate the fusion feature vectors of the image features and each text feature in the feature fusion space respectively;
[0138] Input the fusion feature vectors into a pre-trained classification model to perform a classification operation, and classify the target category of the image to be processed.
[0139] Specifically, for the specific implementation method of the above instructions by the processor 10, reference can be made to the description of the relevant steps in the corresponding embodiments of the attached drawings, which will not be elaborated here.
[0140] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).
[0141] The present invention also provides a computer-readable storage medium, and the readable storage medium stores a computer program, which when executed by a processor of an electronic device, can implement:
[0142] Obtain an image to be processed, and extract the image features of the image to be processed;
[0143] Extract the text features of the image to be processed to obtain a text feature set;
[0144] Map the image features and the text feature set in the dimension space of a pre-constructed feature fusion model to obtain a feature fusion space, and calculate the fusion feature vectors of the image features and each text feature in the feature fusion space respectively;
[0145] Input the fusion feature vectors into a pre-trained classification model to perform a classification operation, and classify the target category of the image to be processed.
[0146] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.
[0147] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] In addition, in each embodiment of the present invention, each functional module may be integrated in a processing unit, may be physically present separately for each unit, or two or more units may be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.
[0149] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0150] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.
[0151] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, an application service layer, etc.
[0152] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results in theory, methods, technologies, and application systems.
[0153] In addition, obviously, the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as first and second are used to represent names and do not represent any specific order.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A classification method based on a multi-modal learning model, characterized in that, the method includes: obtaining an image to be processed, and extracting the image features of the image to be processed; extracting the text features of the image to be processed to obtain a text feature set; determining a starting point and an ending point of the image features in the dimension space of a pre-constructed feature fusion model, sequentially performing mapping of each text feature in the text feature set in the dimension space, summarizing the image features and each text feature in the dimension space to obtain a feature fusion space, calculating the image features through a first fully connected layer of the feature fusion model, calculating the text features through a second fully connected layer, and calculating the fusion feature vectors of the image features and each text feature in the feature fusion space through a third fully connected layer; inputting the fusion feature vectors into a pre-trained classification model to perform a classification operation, and classifying the target category of the image to be processed.
2. The classification method based on a multi-modal learning model according to claim 1, characterized in that, the extracting the image features of the image to be processed includes: performing image normalization processing on the image to be processed to obtain a normalized image; performing convolution pooling operations on the normalized image by using a pre-constructed convolutional neural network model based on the DenseNet algorithm to obtain the image features.
3. The classification method based on a multi-modal learning model according to claim 2, characterized in that, the performing image normalization processing on the image to be processed to obtain a normalized image includes: according to the pixel matrix size of the image to be processed, splitting the image to be processed into a plurality of pixel blocks; extracting the gray values of each pixel block, and calculating the gray mean value and gray variance of each pixel block; using the gray mean value, the gray variance of each pixel block, a preset gray initial value and a gray standard deviation to re-adjust the gray values of each pixel block, and integrating the gray values of each pixel block to obtain a normalized image.
4. The classification method based on a multi-modal learning model according to claim 1, characterized in that, the extracting the text features in the image to be processed includes: extracting the text in the image to be processed; performing word segmentation processing on the text by using a pre-constructed word segmenter to obtain a segmented text; converting the segmented text into word vectors by using a preset word vector conversion model to obtain the text features.
5. The classification method based on a multi-modal learning model according to claim 1, characterized in that, the calculating the image features through a first fully connected layer of the feature fusion model, calculating the text features through a second fully connected layer, and calculating the fusion feature vectors of the image features and each text feature in the feature fusion space through a third fully connected layer includes: the fusion feature vectors of the image features and each text feature in the feature fusion space can be calculated by the following formula: F = D 3 (D 1 (F Image ) + D 2 (F Text )); Among them, D 1 , D 2 , D 3 are the fully connected layers of the feature fusion model, F Image is the image feature, F Text is the text feature, and F is the fused feature vector.
6. The classification method based on a multi-modal learning model according to claim 1, characterized in that, Performing a classification operation in the pre-trained classification model by inputting the fusion feature vector includes: Calculating the probability values of the fusion feature vector hitting a plurality of preset classification labels; Sorting the probability values, and extracting the classification labels corresponding to a preset number of probability values with the top rankings; Obtaining the target category of the image to be processed according to the classification labels.
7. A classification device based on a multi-modal learning model, characterized in that the device includes: A feature acquisition module, configured to acquire an image to be processed, and extract the image features of the image to be processed; extract the text features of the image to be processed to obtain a text feature set; A feature fusion module, configured to determine a starting point and an ending point of the image features in the dimension space of a pre-constructed feature fusion model, sequentially perform mapping on each text feature of the text feature set in the dimension space, summarize the image features and each text feature in the dimension space to obtain a feature fusion space, calculate the image features through a first fully-connected layer of the feature fusion model, calculate the text features through a second fully-connected layer, and calculate the fusion feature vectors of the image features and each text feature in the feature fusion space through a third fully-connected layer; A classification module, configured to input the fusion feature vector into a pre-trained classification model to perform a classification operation, and classify the target category of the image to be processed.
8. An electronic device, characterized in that the electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the classification method based on a multi-modal learning model according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that when the computer program is executed by a processor, it implements the classification method based on a multi-modal learning model according to any one of claims 1 to 6.
Citation Information
Patent Citations
Skin disease image classification system based on multi-modal data input
CN111444960A
Illegal image recognition method, system and equipment
CN114140673A