Image identification and classification method, system, equipment and medium

By integrating Vision Transformer and MoE mechanisms to dynamically select expert networks, the method addresses inefficiencies in existing models, enhancing adaptability and efficiency in image recognition and classification tasks.

CN120318591APending Publication Date: 2025-07-15INSPUR SMART TECH (NANJING) CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510540423.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing deep neural networks have problems with fixed computational complexity and cannot be dynamically adjusted in image recognition and classification tasks, resulting in insufficient computational redundancy and performance in simple and complex scenarios.

Method used

Combining the Vision Transformer and Mixture of Experts mechanisms, by dynamically selecting the expert network that is most suitable for the current task, using the gating mechanism and sparse activation mechanism, only some expert networks are activated for reasoning, and feature aggregation is performed by combining the multi-head self-attention mechanism and the feedforward network.

Benefits of technology

It significantly improves the flexibility and adaptability of the model, reduces the consumption of computing resources, and improves the accuracy and efficiency of image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318591A_ABST
    Figure CN120318591A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition and classification method, system and device and a medium, and belongs to the field of image model optimization. The method comprises the following steps: acquiring a to-be-processed image, and carrying out normalization and zooming processing on the image to generate a target image; performing block processing and vectorization processing on the target image to generate a plurality of image blocks containing one-dimensional vectors as initial feature vectors; inputting the initial feature vector into a hybrid expert module, and calculating a routing weight of an expert network through a gating mechanism; activating the expert network for processing the data by adopting a sparse activation mechanism based on the calculated routing weight; outputting a feature vector according to the initial feature vector by using the MoE layer based on the activated expert network; and performing aggregation processing on the feature vectors through a multi-head self-attention mechanism and a feedforward network, and generating a classification result of the to-be-processed image by using a classifier. The invention provides an efficient, flexible and excellent-performance solution for the field of image recognition and classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image model optimization, and more specifically relates to an image recognition and classification method, system, device and medium. Background Art

[0002] In the field of computer vision, deep neural networks have become the core solution for image recognition and classification tasks. Mainstream models such as convolutional neural networks (CNNs) and vision transformers (ViTs) have achieved remarkable progress in benchmark tasks such as ImageNet classification and object detection by stacking dense amounts of parameters and computing units. For example, the ResNet series enables feature reuse through deep residual connections, while Vision Transformer and its variants utilize self-attention mechanisms to model global dependencies. However, such dense networks have inherent drawbacks: their computational complexity and model capacity are fixed before training and cannot be dynamically adjusted according to the complexity of the input data.

[0003] Specifically, for images containing simple scenes or common objects (such as images of a single object taken centered), the network still needs to activate all parameters for inference, resulting in a large amount of redundant computation; while for images containing complex backgrounds, multi-scale targets or occluded scenes (such as dense urban street scenes), a fixed-capacity model may reduce the classification accuracy due to insufficient representational ability.

[0004] Existing improvement schemes mainly focus on static model compression or cascaded inference, but have significant limitations. For example, knowledge distillation trains a small student network to mimic a large teacher network, but sacrifices the adaptability to complex samples; dynamic routing networks (such as dynamic convolutions, conditional computation) attempt to select computation paths on demand, but most methods still rely on heuristics designed manually and do not break through the paradigm of dense computation. At the same time, the sparse mixture-of-experts (MoE) models developed in the field of natural language processing provide a new paradigm: GShard, Switch Transformer, etc. dynamically route the input to at least a small number of expert subnetworks through a gating mechanism, significantly reducing the computational amount while maintaining performance. However, directly migrating such methods to the field of computer vision faces challenges: the high-dimensional continuous nature of image data makes it difficult to design routing strategies, and the architecture of the expert networks needs to be adapted to the spatial hierarchy of visual features. Summary of the Invention

[0005] To address the above problems, the purpose of the present invention is to provide an image recognition and classification method, system, device, and medium, which combines the Vision Transformer and the Mixture of Experts mechanism. By dynamically selecting the expert network most suitable for the current task, the flexibility and adaptability of the model are significantly improved, thereby enhancing the accuracy and efficiency of image recognition, while greatly reducing the consumption of computing resources.

[0006] To achieve the above object, the present invention is realized through the following technical solutions: In a first aspect, an embodiment of the present application provides an image recognition and classification method, including: Obtain the image to be processed, perform normalization and scaling processing on the image to generate a target image; Perform block processing and vectorization processing on the target image to generate multiple image blocks containing one-dimensional vectors as initial feature vectors; Input the initial feature vectors into the mixture of experts module, and calculate the routing weights of the expert networks through a gating mechanism; Based on the calculated routing weights, activate the expert networks for processing data using a sparse activation mechanism; Based on the activated expert networks, use the MoE layer of the mixture of experts module to output feature vectors according to the initial feature vectors; Aggregate the feature vectors through a multi-head self-attention mechanism and a feed-forward network, and use a classifier to generate the classification result of the image to be processed.

[0007] In an optional embodiment, the obtaining the image to be processed, performing normalization and scaling processing on the image to generate a target image includes: Perform normalization processing on the image to be processed to map the image pixel value range to a specific interval; Scale the image to be processed to a unified image size to generate a target image.

[0008] In an optional embodiment, the performing block processing and vectorization processing on the target image to generate multiple image blocks containing one-dimensional vectors as initial feature vectors includes: Divide the target image into multiple image blocks of a fixed size, and regard each image block as a three-dimensional matrix; Through matrix transformation, flatten each three-dimensional matrix into a one-dimensional vector containing position information, and append position encoding to obtain the initial feature vectors.

[0009] In an optional embodiment, the inputting the initial feature vectors into the mixture of experts module and calculating the routing weights of the expert networks through a gating mechanism includes: Input the initial feature vector into the mixture of experts module, and use the Softmax function to calculate the probability distribution of the expert network adaptation scores to determine the routing weight vectors of the expert networks; The specific formula used is as follows: g ( x ) = softmax( Wgx ) Wherein, x is the initial feature vector, Wg is the gating network weight matrix, g ( x ) is the probability distribution of the expert network adaptation scores; in g ( x ), g ( x ) i represents the routing weight vector of the i-th expert network.

[0010] In an alternative embodiment, the expert network for processing data is activated by using a sparse activation mechanism based on the calculated routing weights, including: Sort the routing weight vectors of each expert network in descending order, select the k expert networks with the highest weights, and set the routing weight vectors of the remaining patent networks to zero.

[0011] In an alternative embodiment, the activated expert network outputs a feature vector based on the initial feature vector by using the mixture of experts module MoE layer, including: Calculate the feature vector MoE( x ) of the initial feature vector x ) through the following formula:

[0012] Wherein, is the output of the i-th expert network, E is the total number of expert networks, and only the k expert networks with the highest activation weights are activated, where k < E. In an alternative embodiment, aggregating the feature vectors through the multi-head self-attention mechanism and the feed-forward network, and using a classifier to generate the classification result of the image to be processed, including: Aggregate the feature vectors through the multi-head self-attention mechanism and the feed-forward network into global features; Input the global feature data into a Softmax classifier to calculate the probability distribution of each image category of the image to be processed, and output the final classification label.

[0013] In a second aspect, an image recognition and classification system provided by an embodiment of the present application further includes: An image preprocessing module, configured to obtain an image to be processed, perform normalization and scaling processing on the image, and generate a target image; An image vectorization module, configured to perform block processing and vectorization processing on the target image, generate multiple image blocks containing one-dimensional vectors, and use them as initial feature vectors; A routing weight calculation module, configured to input the initial feature vectors into a mixture-of-experts module, and calculate the routing weights of the expert networks through a gating mechanism; A network activation module, configured to output feature vectors based on the initial feature vectors by using the MoE layer of the mixture-of-experts module based on the activated expert networks; A feature recognition module, configured to output feature vectors based on one-dimensional vectors by using the MoE layer of the mixture-of-experts module based on the activated expert networks; A classification module, configured to perform aggregation processing on the feature vectors through a multi-head self-attention mechanism and a feed-forward network, and generate a classification result of the image to be processed by using a classifier.

[0014] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the image recognition and classification method described in any one of the above are implemented.

[0015] In a fourth aspect, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the image recognition and classification method described in any one of the above are implemented.

[0016] From the above technical solutions, the following advantages of the present invention can be seen: In the image recognition and classification method provided by the present application, first, a gating mechanism is used to evaluate the adaptability of input features to expert networks, and only a small number of the most relevant experts are activated for inference, significantly reducing the full-parameter operation redundancy of dense networks; secondly, local spatial information is retained through block vectorization and positional encoding, and global feature aggregation is achieved in combination with a multi-head self-attention mechanism, maintaining the representation ability of complex scenarios while compressing the computational complexity; finally, the modular-designed expert networks support flexible expansion, enabling the system to dynamically match computing resources according to the input complexity, ensuring the efficient processing of simple images and improving the classification accuracy of complex scenarios through multi-expert collaboration, providing a new solution with both adaptability and scalability for computer vision tasks.

[0017] By only activating some expert networks (Experts) to process input data, the present application significantly reduces unnecessary calculations. Compared with traditional densely activated models (such as the standard Transformer architecture), the computational complexity is greatly reduced while maintaining the model performance.

[0018] The gating mechanism of this application can dynamically select the most suitable expert network according to the characteristics of the input image, enabling the model to flexibly adapt to different types of image tasks (such as image classification with different resolutions and different scenarios). This dynamic adjustment ability significantly improves the adaptability of the model to diverse tasks.

[0019] This application performs normalization and scaling processing on the image, limits the pixel values to a specific interval, and adjusts the image to a unified size. This helps to reduce the differences in image data, provides a more stable and standard input for subsequent processing, and enhances the generalization ability and robustness of the model.

[0020] This application performs block division and vectorization processing on the target image, divides the image into blocks of a fixed size and converts them into one-dimensional vectors containing position information, which can effectively extract the local features of the image. At the same time, position encoding helps the model capture the spatial structure information of the image and improves the accuracy of feature expression.

[0021] The hybrid expert module of this application combines the gating mechanism and the sparse activation mechanism, selects the top k expert networks with the highest weights for activation by calculating the routing weights, and zeros out the weights of the remaining networks. This approach avoids all expert networks processing data simultaneously, reduces the computational amount and memory occupancy, and improves the utilization efficiency of computing resources.

[0022] The hybrid expert module of this application outputs feature vectors according to the activated expert networks. Different expert networks can learn the features of different aspects of the image, and through weighted combination, a richer and more comprehensive feature expression can be generated, enhancing the model's learning and representation ability of image features.

[0023] This application aggregates the feature vectors through the multi-head self-attention mechanism and the feed-forward network, converts them into global features, and then uses the Softmax classifier to calculate the probability distribution of each image category and output the final classification label. This approach can fully exploit the global information of the image and improve the accuracy and reliability of classification. Brief Description of the Drawings

[0024] To more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 It is a schematic flowchart of the image recognition and classification method provided by this application.

[0026] Figure 2Schematic diagram of the image recognition and classification system provided by this application.

[0027] Figure 3 Schematic diagram of the electronic device provided by this application. Detailed implementation manners

[0028] In the following, the specific steps of the image recognition and classification method will be described in detail, and various embodiments of the present disclosure will be described more comprehensively. The present disclosure can have various embodiments, and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents, and / or alternative solutions that fall within the spirit and scope of the various embodiments of the present disclosure.

[0029] To facilitate a clear description of the technical solutions of the embodiments of this application, hereinafter, some terms and technologies involved in the embodiments of this application will be briefly introduced: Expert Network: In the MoE architecture, each expert network is an independent sub-network responsible for processing specific types of input data or features.

[0030] Gating Mechanism: In the MoE architecture, a mechanism used to dynamically select the most suitable expert network for processing the current input. The gating mechanism distributes the input to different expert networks according to the weights of the input features.

[0031] Sparse Activation: In the MoE architecture, only part of the expert networks are activated to process the input data, thereby reducing the consumption of computing resources.

[0032] Softmax: A commonly used classifier activation function for converting the model output into a probability distribution.

[0033] Hereinafter, the term "comprising" or "may comprise" that can be used in various embodiments of the present disclosure indicates the presence of the disclosed functions, operations, or elements, and does not limit the addition of one or more functions, operations, or elements. In addition, as used in various embodiments of the present disclosure, the terms "comprising", "having" and their cognates are only intended to indicate specific features, numbers, steps, operations, elements, components, or combinations of the foregoing items, and should not be construed as first excluding the existence or addition of the possibility of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items.

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0035] Please refer to Figure 1 The following is a flowchart of a method for image recognition and classification in a specific embodiment. The method includes: S1: Obtain the image to be processed, perform normalization and scaling processing on the image, and generate a target image.

[0036] In the specific implementation, the image to be processed is first preprocessed, including normalization and scaling processing. Specifically, the pixel value range of the image to be processed is adjusted to a specific interval, and the image size is unified to generate a target image to ensure the consistency of subsequent processing.

[0037] S2: Perform block processing and vectorization processing on the target image to generate multiple image blocks containing one-dimensional vectors as initial feature vectors.

[0038] In the specific implementation, the target image is divided into multiple image blocks, and the size of each image block is 16×16 pixels. From the perspective of a matrix, these image blocks can be regarded as a three-dimensional matrix of 16 × 16×3. Through matrix transformation, these image blocks are flattened into one-dimensional vectors and position encoding is added to retain the spatial information of the image. After a picture is cut into several three-dimensional matrices, through matrix transformation, it becomes several one-dimensional vectors containing position information.

[0039] Exemplarily, the target image is first divided into small blocks (Patches) of 16×16 pixels, and each small block is regarded as a three-dimensional matrix of 16×16×3.

[0040] Then, each small block is flattened into a one-dimensional vector through matrix transformation, and position encoding (retaining spatial information) is added to form a sequence of feature vectors containing position information (i.e., Token sequence) as the initial feature vector.

[0041] S3: Input the initial feature vector into the mixture-of-experts module and calculate the routing weights of the expert network through a gating mechanism.

[0042] In the specific implementation, the initial feature vector is input into the mixture-of-experts module, and the Softmax function is used to calculate the probability distribution of the expert network adaptation score to determine the routing weight vector of the expert network; The specific formula used is as follows: g ( x ) = softmax( Wgx ) Among them, x is the initial feature vector, Wg is the gating network weight matrix, g ( x ) is the probability distribution of the expert network fitness score; in g ( x ) g ( x ) i represents the routing weight vector of the i-th expert network. It should be noted that the core architecture of the present invention is based on Vision MoE (V-MoE), combining Vision Transformer (ViT) and Mixture of Experts (MoE) mechanisms. The MoE module contains multiple expert networks (Experts), and each expert network is an independent Transformer layer responsible for processing specific types of features. The gating mechanism dynamically assigns weights according to the input features and selects the most suitable expert network for the current input for processing. Through the sparse activation strategy, only some expert networks are activated, thus significantly reducing the computational complexity. In this way, the information of each input does not need to involve all the weights of the entire model in the calculation, thus ensuring the inference performance of the model.

[0043] Among them, the gating mechanism is the core component in the MoE architecture, used to dynamically select the most suitable expert network for processing the current input. It calculates the fitness of each expert for the input data and selects the most suitable expert to process the input according to these weights.

[0044] The softMax function is a commonly used activation function in machine learning and deep learning, especially when dealing with multi-classification problems. It converts a vector or a set of real numbers into a probability distribution, so that the value of each element is between 0 and 1, and the sum of all elements is 1. This makes the softMax function very suitable for the output layer of the model, converting the original output of the model into a probability prediction.

[0045] S4: Based on the calculated routing weights, adopt the sparse activation mechanism to activate the expert network for processing data.

[0046] In this step, sparse activation means that in the MoE architecture, only some of the expert networks are activated to process the input data, rather than having all experts participate in the calculation. This mechanism significantly reduces the computational cost while maintaining the performance of the model. A gating network is designed through a specific activation function to make its output sparse. Dynamically select the k expert networks with the highest weights for activation, and the outputs of the remaining expert networks are zero.

[0047] Exemplarily, the sparse activation mechanism can be expressed by the following formula:

[0048] This formula is the result generated during inference. y is the generated result matrix, is the gating weight of the i-th expert, is the output of the i-th expert. This formula is to sum the results calculated by different activated expert networks weighted by their weights to obtain the comprehensive result of all activated experts.

[0049] S5: Based on the activated expert networks, use the Mixture of Experts (MoE) layer of the hybrid expert module to output the feature vector according to the initial feature vector.

[0050] In the specific implementation, the initial feature vector is calculated by the following formula x of the feature vector MoE( x ):

[0051] where, is the output of the i-th expert network, E is the total number of expert networks, only the k expert networks with the highest weights are activated, and k < E. This formula shows how the input initial feature vector x is allocated to different expert networks, and the output of each expert network is obtained by weighted summation according to its routing weight.

[0052] S6: Aggregate and process the feature vector through the multi-head self-attention mechanism and the feed-forward network, and use the classifier to generate the classification result of the image to be processed.

[0053] In the specific implementation, the feature vector output by the MoE module is first further processed by the multi-head self-attention mechanism (Multi-head Self-Attention) and the feed-forward network (Feed-Forward Network) to be aggregated into a global feature representation to capture high-level semantic associations.

[0054] Then, the global feature is input into the Softmax classifier to calculate the probability distribution of each image category and output the final classification label.

[0055] In addition, it should be noted that this method essentially constructs an image recognition and classification model through the MoE module, in combination with the multi-head self-attention mechanism, the feed-forward network, and the Softmax classifier. Further, the present invention also discloses the training and optimization strategies of this model, which are as follows: (1) Pre-training and fine-tuning: The model is first pre-trained on a large-scale image dataset (such as ImageNet) to learn general image feature representations. During the pre-training process, data augmentation techniques (such as random cropping, color jittering, etc.) are used to improve the generalization ability of the model. On the specific task dataset, the model is fine-tuned to optimize the classification performance.

[0056] (2) Adaptive learning rate adjustment: During the training process, an adaptive learning rate adjustment mechanism is introduced. The learning rate is dynamically adjusted according to the training progress. A relatively high learning rate is adopted in the initial stage to achieve rapid convergence, and then the learning rate is gradually decreased to stabilize the training. In addition, the cosine annealing strategy is used to further optimize the learning rate adjustment.

[0057] (3) Sparse activation strategy: Through the sparse activation strategy, only part of the expert networks are activated to process the current input. The sparse activation ratio can be dynamically adjusted according to the task requirements. For example, more expert networks are activated in complex scenarios, while the number of activated networks is reduced in simple scenarios. This strategy significantly reduces the consumption of computing resources while maintaining the high performance of the model.

[0058] To further improve the robustness and accuracy of the model, the present invention also supports multi-modal information fusion. In addition to the pixel information of the image, the following multi-modal data can also be introduced: Depth information: The depth information of the image is obtained through a depth sensor or a depth estimation model and fused with the pixel information. The depth information provides additional three-dimensional spatial information for the model, which helps to improve the understanding of complex scenarios.

[0059] Semantic segmentation information: A pre-trained semantic segmentation model is used to obtain the semantic segmentation map of the image, and the segmentation map is combined with the pixel information to provide richer semantic information for the model.

[0060] Text description: The text descriptions related to the image (such as titles, labels, or descriptive texts provided by users) are collected, embedded as text feature vectors through natural language processing techniques, and fused with the image features. The text description provides additional context information for the model, which helps to improve the classification performance.

[0061] In this embodiment, an image recognition and classification method based on Vision MoE (V-MoE) is disclosed. The core objective of this method is to significantly improve the accuracy and efficiency of image recognition by optimizing the model architecture and training strategy, while greatly reducing the consumption of computing resources. Specifically, this method combines Vision Transformer (ViT) and Mixture of Experts (MoE) mechanism, and significantly improves the flexibility and adaptability of the model by dynamically selecting the expert network most suitable for the current task.

[0062] In terms of the model architecture, the present invention adopts an innovative design concept. First, the input image is segmented into multiple small patches, which can be regarded as local feature representations of the image. Subsequently, these patches are input into the MoE module. The MoE module is the core component of the present invention, which contains multiple expert networks, and each expert network is responsible for processing specific types of image features. For example, some expert networks may be good at processing edge information, while others may focus on color or texture features. Through the gating mechanism, the model can dynamically select the most suitable expert network according to the characteristics of the input features. The key to this design is that only part of the expert networks will be activated to process the current input image, thus significantly reducing the computational complexity. Compared with traditional dense networks, this method not only improves the computational efficiency but also enhances the model's adaptability to different image features.

[0063] In addition, the present invention also introduces a sparse activation strategy and an adaptive learning rate adjustment mechanism to further optimize the training process. The sparse activation strategy ensures that only the expert networks most relevant to the current task are activated, avoiding unnecessary computational overhead. The adaptive learning rate adjustment mechanism can dynamically adjust the learning rate according to the training progress of the model and the task difficulty, thus accelerating the convergence speed and improving the training efficiency of the model.

[0064] In the image recognition and classification process, after the input image is preprocessed, it is segmented into multiple small patches and input into the Vision MoE model. The gating mechanism dynamically selects the expert network for feature extraction according to the input features, and the extracted features are further aggregated into a global feature representation, and finally the class label of the image is output through the classifier.

[0065] The present invention significantly improves the generalization ability and recognition performance of the model through pre-training on a large-scale dataset and fine-tuning for specific tasks. Experimental results show that this method significantly reduces the consumption of computing resources (such as FLOPs) while maintaining high accuracy, especially performing well in processing large-scale image data. In addition, this method shows good adaptability in different types of image tasks, providing an efficient and flexible solution for the field of image recognition and classification.

[0066] As shown Figure 2 below, the following is an embodiment of an image recognition and classification system provided by an embodiment of the present disclosure. This system and the image recognition and classification methods of the above embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the image recognition and classification system, reference may be made to the embodiments of the above image recognition and classification methods.

[0067] An image recognition and classification system includes: an image preprocessing module, an image vectorization module, a routing weight calculation module, a network activation module, a feature recognition module, and a classification module.

[0068] The image preprocessing module is configured to obtain an image to be processed, perform normalization and scaling processing on the image, and generate a target image.

[0069] The image vectorization module is configured to perform block processing and vectorization processing on the target image to generate a plurality of image blocks containing one-dimensional vectors as initial feature vectors.

[0070] The routing weight calculation module is configured to input the initial feature vectors into a mixture of experts module and calculate the routing weights of the expert networks through a gating mechanism.

[0071] The network activation module is configured to output feature vectors based on the initial feature vectors by using the MoE layer of the mixture of experts module based on the activated expert networks.

[0072] The feature recognition module is configured to output feature vectors based on the one-dimensional vectors by using the MoE layer of the mixture of experts module based on the activated expert networks.

[0073] The classification module is configured to perform aggregation processing on the feature vectors through a multi-head self-attention mechanism and a feed-forward network, and generate a classification result of the image to be processed by using a classifier.

[0074] The image recognition and classification system provided in this embodiment generates a standard target image by performing normalization and scaling processing on the image to be processed, achieving the beneficial effects of reducing image data differences, providing a stable input for subsequent processing, and enhancing the generalization and robustness of the model; generating an initial feature vector by performing block and vectorization processing on the target image and attaching position encoding, achieving the beneficial effects of effectively extracting local features of the image, capturing spatial structure information, and enhancing the accuracy of feature expression; inputting the initial feature vector into the mixture-of-experts module, calculating routing weights using the gating mechanism and activating some expert networks using the sparse activation mechanism, achieving the beneficial effects of avoiding all expert networks from processing data simultaneously, reducing the computational amount and memory occupation, and improving the utilization efficiency of computing resources; generating a richer and more comprehensive feature expression by using the mixture-of-experts module to output a feature vector based on the activated expert network, enhancing the model's ability to learn and represent image features; aggregating the feature vector through the multi-head self-attention mechanism and the feed-forward network, and then calculating the probability distribution to output the classification label using the Softmax classifier, achieving the beneficial effects of fully mining the global information of the image, and enhancing the accuracy and reliability of classification.

[0075] Figure 3 Schematic diagram of the hardware structure of an electronic device for implementing each embodiment of the present invention.

[0076] The image recognition and classification method provided in the embodiments of the present application can be applied to an electronic device. Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.

[0077] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, keys, a camera, a display screen, and a subscriber identity module (SIM) card interface, etc.

[0078] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0079] Among them, the processor may be the nerve center and command center of the electronic device. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0080] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can save the instructions or data just used or recycled by the processor. If the processor needs to use the instruction or data again, it can be directly called from this memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the system efficiency.

[0081] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor through the external memory interface to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.

[0082] The internal memory can be used to store computer-executable program code, and the computer-executable program code includes instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory may include a program storage area and a data storage area. The internal memory may include a high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0083] The wireless communication function of the electronic device can be implemented through an antenna, a wireless communication module, a modem processor, a baseband processor, etc.

[0084] The wireless communication module can provide solutions for wireless communications applied to electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.

[0085] The electronic device can implement audio functions through an audio module, speaker, receiver, microphone, headphone jack, application processor, etc.

[0086] The electronic device can implement a shooting function through an ISP, camera, video codec, GPU, display screen, and application processor, etc.

[0087] The electronic device can implement a display function through a GPU, display screen, and application processor, etc.

[0088] The GPU is a microprocessor for image processing, connecting the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs that execute program instructions to generate or change display information.

[0089] The display screen is used to display images, videos, etc. The display screen includes a display panel.

[0090] The above-mentioned electronic device implements the image recognition and classification method of the present application. By normalizing and scaling the image, partitioning and vectorizing it, inputting the initial feature vector into the mixture of experts module and calculating the routing weights using the gating mechanism, activating some expert networks with the sparse activation mechanism, outputting the feature vector based on the activated network, aggregating the features through the multi-head self-attention mechanism and the feed-forward network, and finally obtaining the classification result with a classifier, it achieves the beneficial effects of reducing the difference in image data, effectively extracting image features, reasonably utilizing computing resources, enriching feature expressions, and improving the accuracy and reliability of image classification.

[0091] In the storage medium provided by the present application, there is a program product capable of implementing the image recognition and classification method.

[0092] The image recognition and classification method includes: Obtain an image to be processed, perform normalization and scaling processing on the image to generate a target image; The target image is segmented and vectorized to generate multiple image patches containing one-dimensional vectors as initial feature vectors; The initial feature vectors are input into the Mixture of Experts (MoE) module, and the routing weights of the expert networks are calculated through a gating mechanism; Based on the calculated routing weights, a sparse activation mechanism is used to activate the expert networks for processing data; Based on the activated expert networks, the MoE layer of the MoE module outputs feature vectors according to the initial feature vectors; The feature vectors are aggregated through a multi-head self-attention mechanism and a feed-forward network, and a classifier is used to generate the classification result of the image to be processed.

[0093] In some possible embodiments, the image recognition and classification method of the present disclosure may be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above in this specification.

[0094] The storage medium of the present disclosure may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0095] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. An image recognition and classification method, characterized in that, Including: Obtain the image to be processed, perform normalization and scaling processing on the image, and generate a target image; Perform block processing and vectorization processing on the target image to generate multiple image blocks containing one-dimensional vectors as initial feature vectors; Input the initial feature vectors into the Mixture of Experts (MoE) module, and calculate the routing weights of the expert networks through a gating mechanism; Based on the calculated routing weights, activate the expert networks for processing data using a sparse activation mechanism; Based on the activated expert networks, use the MoE layer of the MoE module to output feature vectors according to the initial feature vectors; Perform aggregation processing on the feature vectors through a multi-head self-attention mechanism and a feed-forward network, and use a classifier to generate the classification result of the image to be processed.

2. The image recognition and classification method according to claim 1, wherein The obtaining the image to be processed, performing normalization and scaling processing on the image, and generating a target image includes: Perform normalization processing on the image to be processed to map the image pixel value range to a specific interval; Scale the image to be processed to a unified image size to generate a target image.

3. The image recognition and classification method according to claim 2, wherein The performing block processing and vectorization processing on the target image to generate multiple image blocks containing one-dimensional vectors as initial feature vectors includes: Segment the target image into multiple image blocks of a fixed size, and regard each image block as a three-dimensional matrix; Through matrix transformation, flatten each three-dimensional matrix into a one-dimensional vector containing position information and append position encoding to obtain the initial feature vectors.

4. The image recognition and classification method according to claim 3, wherein The inputting the initial feature vectors into the MoE module and calculating the routing weights of the expert networks through a gating mechanism includes: Input the initial feature vectors into the MoE module, and use the Softmax function to calculate the probability distribution of the expert network fitness scores to determine the routing weight vector of the expert networks; The specific formula used is as follows: g ( x ) = softmax( Wgx ) Among them, x is the initial feature vector, Wg is the gating network weight matrix, g ( x ) is the probability distribution of the expert network fitness score; in g ( x ), g ( x ) i represents the routing weight vector of the i-th expert network.

5. The image recognition and classification method according to claim 4, wherein The activating the expert networks for processing data using a sparse activation mechanism based on the calculated routing weights includes: Sort the routing weight vectors of each expert network in descending order, select the k expert networks with the highest weights, and set the routing weight vectors of the remaining expert networks to zero.

6. The image recognition and classification method according to claim 5, characterized in that, The outputting feature vectors according to the initial feature vectors based on the activated expert networks using the MoE layer of the MoE module includes: The initial feature vector is calculated by the following formula x The feature vector MoE( x ): Among them, is the output of the i-th expert network, E is the total number of expert networks, and only the k expert networks with the highest activation weights are activated, where k < E.

7. The image recognition and classification method according to claim 6, characterized in that The performing aggregation processing on the feature vectors through a multi-head self-attention mechanism and a feed-forward network, and using a classifier to generate the classification result of the image to be processed includes: Perform aggregation processing on the feature vectors through a multi-head self-attention mechanism and a feed-forward network to aggregate them into global features; Input the global feature data into a Softmax classifier, calculate the probability distribution of each image category of the image to be processed, and output the final classification label.

8. An image recognition and classification system, characterized in that, The system adopts the image recognition and classification method according to any one of claims 1 to 7; The system includes: An image preprocessing module for obtaining the image to be processed, performing normalization and scaling processing on the image, and generating a target image; An image vectorization module for performing block processing and vectorization processing on the target image to generate multiple image blocks containing one-dimensional vectors as initial feature vectors; A routing weight calculation module, configured to input the initial feature vector into a mixture-of-experts module, and calculate the routing weights of an expert network through a gating mechanism; A network activation module, configured to output a feature vector based on the initial feature vector by using an MoE layer of the mixture-of-experts module based on the activated expert network; A feature recognition module, configured to output a feature vector based on a one-dimensional vector by using an MoE layer of the mixture-of-experts module based on the activated expert network; A classification module, configured to perform an aggregation process on the feature vector through a multi-head self-attention mechanism and a feed-forward network, and generate a classification result of the image to be processed by using a classifier.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, the steps of the image recognition and classification method according to any one of claims 1 to 7 are implemented.

10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the image recognition and classification method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • User security feature recognition method based on behavior pattern analysis

    CN120804894A

  • Block chain-based credible distributed hybrid expert model edge computing system

    CN121168529A

  • Model training method, vehicle control method, device, equipment and medium

    CN121214384A

  • Signal modulation identification method fused to Transform and hybrid expert mechanism

    CN121367632A

  • Remote sensing tree species identification method and device

    CN121564553A