Mineral substance identification method based on multi-modal architecture search and channel pruning technology
Through multimodal architecture search and channel pruning technology, the neural network model is optimized, and the bottlenecks of high-performance models running on resource-constrained platforms and the problem of identification accuracy deviation in complex environments is solved, achieving efficient and stable mineral recognition.
Patent Information
- Application Number
- CN202510018866.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-13
AI Technical Summary
When the existing high-performance neural network model runs on a resource-constrained platform, it has a large amount of computing and high storage requirements, making it difficult to provide real-time performance and stability. At the same time, the identification accuracy deviation is large in complex environments.
The mineral recognition method based on multimodal architecture search and channel pruning technology is adopted. Each operation on the hypernetwork model is assigned weights through a microscopic neural architecture search, and the operation that contributes the most to the model is found. During the search process, the channel pruning technology is used to reduce the number of model channels to obtain a high-performance and lightweight mineral recognition model.
It realizes a high-performance mineral identification model that operates efficiently on resource-constrained platforms, improving the recognition accuracy and model stability and reliability in complex environments.
Smart Images

Figure CN119992523A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a mineral identification method based on multimodal architecture search and channel pruning technology. Background Art
[0002] In the mining process of coal mines, accurate and rapid identification of mineral types is the key to ensuring safe production, promoting efficient resource utilization and optimizing mining processes. Traditional mineral identification technology usually relies on manual operation, which is not only time-consuming and laborious, but also often difficult to ensure the accuracy of mineral type identification in the complex and changeable coal mine environment. In addition, as coal mining gradually develops towards intelligence and automation, the demand for real-time and automatic identification of mineral types has become more urgent. Therefore, benefiting from the continuous development of deep learning technology, it has brought new possibilities for mineral identification.
[0003] Deep Neural Network (DNN) has demonstrated excellent performance in image processing, natural language processing, pattern recognition and other fields. However, designing a high-performance and efficient neural network architecture for the special environment of coal mines requires not only considering the diversity of minerals, but also dealing with complex factors such as environmental noise and changes in lighting conditions, which poses a huge challenge to the manual design of network architecture. Therefore, the emergence of Neural Architecture Search (NAS) technology provides an effective solution to this problem. NAS can automatically search and optimize the neural network architecture to find the network model that best suits a specific task, thereby achieving high-precision mineral identification in complex environments. The network architecture generated by NAS can not only significantly improve the recognition accuracy, but also reduce the limitations brought by human design.
[0004] However, these high-performance neural network models usually have large computational workloads and high storage requirements. In coal mining scenarios, most equipment is located underground or in remote areas, and is limited by space, energy, and computing power. This is a significant bottleneck for the limited computing resources and storage capacity at coal mine sites, limiting their deployment on resource-constrained platforms and the stability of actual operations.
[0005] That is, high-performance neural network models usually have complex architectures and a large number of parameters, which leads to very high requirements in terms of computing and storage. These requirements make it difficult to deploy the model on resource-constrained platforms. In coal mining environments, the computing power, memory, and storage resources of on-site computing equipment are often limited, so a model that can run efficiently under these constraints is needed. However, existing high-performance neural network models may not provide the required real-time performance and stability when running on resource-constrained platforms, limiting the feasibility of their practical applications.
[0006] Also, sensitivity to environmental and weather conditions: Currently, most mineral identification methods rely solely on RGB images for identification. In a coal mine environment, lighting conditions may be unstable, and environmental factors such as dust and smoke may also interfere with image quality. These factors cause models trained using only RGB images to deviate from the final accuracy of mineral prediction. In addition, weather changes (such as haze or rain and snow) can also cause changes in the color and contrast of images, making it difficult for RGB image-based identification systems to maintain stable performance under various environmental conditions. Summary of the invention
[0007] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0008] To this end, the purpose of this application is to propose a mineral identification method based on multimodal architecture search and channel pruning technology, which can provide a high-performance, lightweight mineral identification model for resource-constrained platforms.
[0009] To achieve the above objectives, the present application embodiment proposes a mineral identification method based on multimodal architecture search and channel pruning technology, including:
[0010] Acquire a mineral multimodal data set, and perform feature extraction on the mineral multimodal data set to obtain multi-scale multimodal features;
[0011] According to the multi-scale multi-modal features, a search space including different operations is defined and a hypernetwork based on operation units is constructed; wherein the search space uses a search method based on the operation units to perform neural architecture search, and the different operations are all provided with architecture parameters;
[0012] The network parameter weights and architecture parameters of the hypernetwork are alternately optimized using a differentiable neural architecture search technique, and the number of channels of the hypernetwork is reduced by a channel pruning technique based on channel shielding parameters during the alternate optimization process to obtain an optimized hypernetwork;
[0013] An initial mineral identification model is obtained according to the network parameter weights and architecture parameters of the optimized supernetwork; and the initial mineral identification model is pruned according to the optimal pruning channel corresponding to the channel shielding parameters of the optimized supernetwork to obtain a mineral identification model, so as to perform mineral identification through the mineral identification model.
[0014] In some implementations, the step of acquiring a mineral multimodal dataset and performing feature extraction on the mineral multimodal dataset to obtain multi-scale multimodal features includes:
[0015] Obtain mineral RGB image dataset and mineral depth image dataset;
[0016] Perform feature extraction on the mineral RGB image dataset through a VGG network to obtain multi-scale features of the RGB image;
[0017] Performing feature extraction on the mineral depth image dataset through the ResNeXt-101 network to obtain multi-scale features of the depth image;
[0018] The multi-scale features of the RGB image and the multi-scale features of the depth image are fused to obtain multi-scale multi-modal features.
[0019] In some implementations, the operation unit includes two input nodes, four intermediate nodes and one output node, and the input nodes, the intermediate nodes and the output nodes are connected in sequence through edges to form a directed acyclic graph; the different operations include multiple edge operations on each edge and multiple fusion operations on each of the intermediate nodes.
[0020] In some implementations, the super network includes an RGB image unit structure, a depth image unit structure, and a multimodal fusion unit structure, and the internal structures of the RGB image unit structure, the depth image unit structure, and the multimodal fusion unit structure are all operation units, and the operation unit includes two input nodes, four intermediate nodes, and one output node, and the input nodes, the intermediate nodes, and the output nodes are connected in sequence through edges to form a directed acyclic graph; the different operations include multiple edge operations on each edge and multiple fusion operations on each intermediate node; each of the edge operations and each of the fusion operations is set with an architecture parameter, the sum of the architecture parameters of the multiple edge operations on each edge is 1, and the sum of the architecture parameters of the multiple fusion operations on each intermediate node is 1; the architecture parameters of the edge operation are used to characterize the impact of the edge operation on the classification accuracy of the entire super network; the architecture parameters of the fusion operation of the intermediate node are used to determine the optimal fusion operation of the features input to the current intermediate node; wherein the architecture parameters are used as the architecture parameters of the super network.
[0021] In some implementations, the alternately optimizing the network parameter weights and the architecture parameters of the hypernetwork using a differentiable neural architecture search technique includes:
[0022] Initializing network parameter weights of the hypernetwork and the architecture parameters;
[0023] The network parameter weights and architecture parameters of the super network are updated alternately, and the alternating update includes fixing the architecture parameters and updating the network parameter weights based on a first loss function; and fixing the network parameter weights and updating the architecture parameters based on a second loss function.
[0024] In some implementations, the channel masking parameters include a first channel masking parameter for determining the number of channel masking operations on the edge of the RGB image unit structure, a second channel masking parameter for determining the number of channel masking operations on the edge of the depth image unit structure, and a third channel masking parameter for determining the number of channel masking operations on the edge of the multimodal fusion unit structure.
[0025] In some implementations, the alternating updating of network parameter weights and architecture parameters of the supernetwork includes:
[0026] Initialize the value of the channel shielding parameter, and perform the following steps until the value of the channel shielding parameter reaches a parameter threshold;
[0027] Alternatingly updating network parameter weights and architecture parameters of the hypernetwork until the hypernetwork converges;
[0028] According to the first updating rule, the value of the current channel shielding parameter is updated.
[0029] In some implementations, the first updating rule is to double the value of the current channel masking parameter.
[0030] In some implementations, obtaining an initial mineral identification model according to the network parameter weights and architecture parameters of the optimized supernetwork includes:
[0031] According to the network parameter weights and architecture parameter values of the optimized supernetwork, the operations with the largest weight system in the edge operations and the fusion operations are retained to obtain an initial mineral identification model.
[0032] In some implementations, the multiple edge operations include skip connections, 3x3 depthwise separable convolutions, 5x5 depthwise separable convolutions, 7x7 depthwise separable convolutions, 3x3 dilated convolutions, 5x5 dilated convolutions, 7x7 dilated convolutions, average pooling, maximum pooling, and None operations, and the multiple fusion operations include cascade operations and summation operations.
[0033] The mineral identification method based on multimodal architecture search and channel pruning technology provided by the present application assigns a weight to each operation on the hypernetwork model through differentiable neural architecture search, so as to find the operation that contributes the most to the model during the model training process, and apply channel pruning technology to reduce the number of model channels during the search process to reduce the amount of calculation and memory usage, so as to provide a high-performance, lightweight mineral identification model for resource-constrained platforms, and the model finally searched can be deployed on platforms with different hardware facilities. At the same time, the mineral identification model is obtained by jointly training the multiple modal features of minerals, so that the model can still have a high mineral identification accuracy rate under various environmental influences such as light and weather, so as to improve the stability and reliability of the model in complex environments.
[0034] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0036] Figure 1 A schematic diagram of a flow chart of a mineral identification method based on multimodal architecture search and channel pruning technology provided in an embodiment of the present application;
[0037] Figure 2 A schematic diagram of the internal structure of each operating unit provided in the embodiment of the present application;
[0038] Figure 3 A schematic diagram of the structure of a hypernetwork provided in an embodiment of the present application. DETAILED DESCRIPTION
[0039] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0040] The following describes a mineral identification method based on multimodal architecture search and channel pruning technology according to an embodiment of the present application with reference to the accompanying drawings.
[0041] Figure 1 A flowchart of a mineral identification method based on multimodal architecture search and channel pruning technology provided in an embodiment of the present application. Figure 1As shown, the mineral identification method based on multimodal architecture search and channel pruning technology includes the following steps:
[0042] Step S101, obtaining a mineral multimodal data set, and performing feature extraction on the mineral multimodal data set to obtain multi-scale multimodal features.
[0043] As an implementation method, a mineral multimodal dataset is obtained, and feature extraction is performed on the mineral multimodal dataset to obtain multi-scale multimodal features; including: obtaining a mineral RGB image dataset and a mineral depth image dataset; extracting features from the mineral RGB image dataset through a VGG network to obtain multi-scale features of the RGB image; extracting features from the mineral depth image dataset through a ResNeXt-101 network to obtain multi-scale features of the depth image; fusing the multi-scale features of the RGB image and the multi-scale features of the depth image to obtain multi-scale multimodal features.
[0044] As an implementation method, the VGG network is used as the backbone network for extracting multi-scale mineral RGB image features. The relevant formula is shown in Formula 1:
[0045] I i =VGG net_i (X RGB ) (1)
[0046] Among them, X RGB represents the RGB image input by the VGG backbone network, net_i represents the i-th neural network layer of the VGG backbone network, and I i Represents the output features of the i-th neural network layer of the VGG backbone network.
[0047] As an implementation method, the ResNeXt-101 network is used as the backbone network for extracting multi-scale mineral depth image features. The relevant formula is shown in Formula 2:
[0048] D i =ResNeXt net_i (X Depth ) (2)
[0049] Among them, X Depth Depth image of the ResNeXt-101 backbone network input, D i is the output feature of the i-th layer of the ResNeXt-101 backbone network.
[0050] It can be understood that in order to solve the problem of deviation in the final prediction accuracy of minerals in the model trained only with RGB images, the present invention adopts a mineral RGB image dataset and a mineral depth image dataset, and uses the VGG network and the ResNeXt-101 network as the backbone networks for extracting mineral RGB image features and mineral depth image features, respectively, aiming to extract multimodal features of different scales to improve the recognition accuracy of the mineral recognition model obtained by subsequent training.
[0051] Step S102, based on the multi-scale multi-modal features, define a search space including different operations and construct a hypernetwork based on operation units; wherein the search space uses a search method based on operation units to perform neural architecture search, and different operations are set with architecture parameters.
[0052] In some embodiments, the search space uses an operation unit-based search method to perform neural architecture search, each operation unit includes two input nodes, four intermediate nodes and one output node, and the input nodes, intermediate nodes and output nodes are connected in sequence through edges to form a directed acyclic graph; different operations include multiple edge operations on each edge and multiple fusion operations on each intermediate node.
[0053] As an implementation, the search space of DARTS is predetermined, such as Figure 2As shown, a neural architecture search is performed using an operation unit-based search method, where each operation unit consists of two input nodes, four intermediate nodes and one output node, and the edge operations on the edges and the fusion operations of the intermediate nodes are to be determined. Exemplarily, the different operations of the operation unit include 10 edge operations on each edge and 2 fusion operations on each intermediate node, wherein the 10 edge operations include skip connection, 3x3 depthwise separable convolution (3*3DepthwiseSeparable Convolution, 3_DSC), 5x5 depthwise separable convolution (5*5Depthwise SeparableConvolution, 5_DSC), 7x7 depthwise separable convolution (7*7Depthwise Separable Convolution, 7_DSC), 3x3 dilated convolution (3*3Dilated Convolution, 3_DC), 5x5 dilated convolution (5*5DilatedConvolution, 5_DC), 7x7 dilated convolution (7*7Dilated Convolution, 3_DC), average pooling (AP), maximum pooling (MP) and None operation, and the 2 fusion operations include a cascade operation (Concatenation) and a summation operation (Addition).
[0054] After pre-defining the search space, a super network needs to be constructed to find a subnetwork with lightweight, high efficiency and high performance, that is, a lightweight mineral identification network, that is, a mineral identification model.
[0055] In some embodiments, Figure 3 As shown in the figure, the super network includes RGB image unit structure, depth image unit structure and multimodal fusion unit structure. RGB ), depth image unit structure (Cell depth ) and multimodal fusion unit structure (Cell fusion)'s internal structure is an operation unit, which includes two input nodes, four intermediate nodes and one output node, and the input nodes, intermediate nodes and output nodes are connected in sequence through edges to form a directed acyclic graph; different operations of the operation unit include multiple edge operations on each edge and multiple fusion operations on each intermediate node; each edge operation and each fusion operation are set with architecture parameters, the sum of the architecture parameters of the multiple edge operations on each edge is 1, and the sum of the architecture parameters of the multiple fusion operations on each intermediate node is 1; the architecture parameters of the edge operation are used to characterize the impact of the edge operation on the classification accuracy of the entire hypernetwork; the architecture parameters of the fusion operation of the intermediate node are used to determine the optimal fusion operation of the features input to the current intermediate node; wherein, the architecture parameter of each edge operation is represented as ɑ and the architecture parameter of each fusion operation is represented as ε.
[0056] It can be understood that the internal structures of the RGB image unit structure, the depth image unit structure and the multimodal fusion unit structure are all as follows: Figure 2 shown.
[0057] Therefore, the present invention predefines a search space with different operations and constructs a super network based on operation units, so that in subsequent steps, the discrete relaxation strategy of differentiable architecture search (DARTS) is used to search for the best edge operation on each operation unit edge and the best fusion operation on the intermediate nodes, so as to search for a high-performance neural network architecture, i.e., a subnet.
[0058] Step S103, using differentiable neural architecture search technology to alternately optimize the network parameter weights and architecture parameters of the hypernetwork, and in the process of alternating optimization, reducing the number of channels of the hypernetwork through channel pruning technology based on channel shielding parameters to obtain an optimized hypernetwork.
[0059] As an implementation method, a differentiable neural architecture search technique is used to alternately optimize the network parameter weights and architecture parameters of a hypernetwork, including: initializing the network parameter weights and architecture parameters of the hypernetwork; alternately updating the network parameter weights and architecture parameters of the hypernetwork, wherein the alternating update includes fixing the architecture parameters and updating the network parameter weights based on a first loss function; and fixing the network parameter weights and updating the architecture parameters based on a second loss function.
[0060] The following is a further detailed description of the super network update process in combination with the formula. The super network update process includes the following steps:
[0061] Step S201, initializing the network parameter weight w and architecture parameters ɑ and ε of the hypernetwork model;
[0062] Step S202, fix the architecture parameters ɑ and ε of the hypernetwork model, and update the network parameter weight w of the hypernetwork model based on the first loss function, the first loss function is shown in Formula 3:
[0063] w * = arg min loss (w,(a,ε)) (3)
[0064] Among them, (a, ε) represents the fixed architecture parameters, w represents the network parameter weight that has not been updated, and w * Indicates the updated network parameter weight, arg min loss Represents the value of minimizing the loss function after fixing the architecture parameters (a, ε) to obtain the new network parameter weight w * .
[0065] Step S203, fix the network parameter weight w of the hypernetwork model, and update the architecture parameters ɑ and ε of the hypernetwork model based on the second loss function, the second loss function is shown in Formula 4:
[0066] (a, ε) * = arg min loss (w * , (a, ε)) (4)
[0067] Among them, w * represents the fixed network parameter weights, (a, ε) represents the unupdated architecture parameters, and (a, ε) * Indicates the updated architecture parameters, arg min loss Represents the fixed network parameter weight w * Then minimize the value of the loss function to obtain the new architecture parameters (α, ε) * .
[0068] Step S202 and step S203 are performed alternately until the optimization target is reached.
[0069] In the process of alternating optimization and updating of the above-mentioned super network, in order to reduce the amount of calculation, the present invention adds channel pruning technology to reduce the number of channels, so that the structure of the searched subnet has the characteristics of lightweight, high efficiency and high performance.
[0070] In some implementations, the channel masking parameters include a first channel masking parameter for determining the number of channel masking operations on the edge of an RGB image unit structure, a second channel masking parameter for determining the number of channel masking operations on the edge of a depth image unit structure, and a third channel masking parameter for determining the number of channel masking operations on the edge of a multimodal fusion unit structure.
[0071] In some embodiments, the network parameter weights and architecture parameters of the hypernetwork are updated alternately, including:
[0072] Initialize the value of the channel masking parameter, and perform the following steps until the value of the channel masking parameter reaches the parameter threshold; wherein the channel masking parameter includes a first channel masking parameter of the edge operation of the RGB image unit structure, a second channel masking parameter of the edge operation of the depth image unit structure, and a third channel masking parameter of the edge operation of the multimodal fusion unit structure;
[0073] The network parameter weights and architecture parameters of the hypernetwork are updated alternately until the hypernetwork converges;
[0074] According to the first updating rule, the value of the current channel shielding parameter is updated.
[0075] In some embodiments, the first updating rule is to double the value of the current channel masking parameter.
[0076] The following is a detailed description of the implementation process of reducing the number of channels for various operations predefined in the search space through channel pruning technology based on channel shielding parameters during the hypernetwork alternating optimization process, in conjunction with the formula, that is, the pruning process includes the following steps.
[0077] Step S301, by setting the first channel shielding parameter Z 1 To determine the number of channel shielding operations on the edge of the RGB image unit structure, the relevant formula is shown in formula (5-7):
[0078]
[0079] Among them, Op (i,j) (I i ) is the operation on the edge between node i and node j, I i represents the input feature (RGB image), O represents the 10 different edge operations defined on the edge, Op represents a type of edge operation to be selected, Op candidate operation architecture parameters, Represents an input feature (i.e. Cell) that participates in the operation calculation in the search space RGB input features). Represents an input feature that does not participate in the operation calculation in the search space. is a mask parameter matrix (whose matrix dimensions are the same as I i Consistent, consisting of 0 and 1), is designed to utilize Z 1 To determine the number of channel shielding operations on the edge, select 1 / Z 1The value of the number is 1, and the others are 0; 1 means executing the current operation, and 0 means not executing the current operation. Among them, the number of channel shielding refers to the number of channels that are "shielded" or "disabled" through the screening mechanism during the optimization process of the deep learning model. The shielded channels no longer participate in subsequent calculations, which plays a role in reducing the computing cost. For example, if a feature is a 20-dimensional * 20-dimensional two-dimensional matrix, then this feature has 400 eigenvalues. Assume that Z 1 is equal to 4, then 1 / Z 1 It can be understood as In this two-dimensional matrix, 1 / 4 of the values are 1, that is, 100 values are 1 (1 / 4*400=100), and the other 300 values are 0.
[0080] Step S302: Similarly, the second channel shielding parameter Z of the edge operation of the depth image unit structure and the multimodal fusion unit structure is 2 and the third channel shielding parameter Z 3 The related formulas are similar to step S301. Set the parameters Z 2 and Z 3 To determine the number of channel shielding operations on the edge of the depth image unit structure and the multimodal fusion unit structure, the relevant formula is shown in formula (8-13):
[0081]
[0082] Among them, formula (8-10) is the calculation formula of the depth image unit structure, Represents an input feature that participates in the operation calculation in the search space (i.e., the input feature of the deep image unit structure). Represents an input feature that does not participate in the operation calculation in the search space. Is a mask parameter matrix, whose matrix dimension is the same as D i Consistent, composed of 0 and 1, using Z 2 To determine the number of channel shielding operations on the edge. Formula (11-13) is the calculation formula for the multimodal fusion unit structure. Represents an input feature that participates in the operation calculation in the search space (i.e., the input feature of the multimodal fusion unit structure). Represents an input feature that does not participate in the operation calculation in the search space. is a mask parameter matrix, whose matrix dimension is the same as F i Consistent, composed of 0 and 1, using Z 3 To determine the number of channel shielding for edge operations.
[0083] Step S303, after determining the number of shielded channels on the edge, the fusion operation of the intermediate node of each operation unit is also calculated to update the architecture parameter ε inside the intermediate node, but the channels of the intermediate node are not shielded. The specific calculation formula is as follows;
[0084]
[0085] Among them, y1 and y2 represent the input features of the intermediate nodes of the unit structure of each hypernetwork (such as RGB image unit structure, deep image unit structure and multimodal fusion unit structure), f0 represents the set of predefined candidate operations on the intermediate nodes, and f represents a candidate operation. Schema parameters representing candidate operations, is the output feature of the intermediate node.
[0086] Step S304: After knowing the calculation principle of edge operation, node internal fusion operation and super network alternating update, in the super network optimization process, first initialize Z 1 , Z 2 and Z 3 The value of is 1. Then, the network parameter weight w and the architecture parameters ɑ and ε are optimized alternately to make the model converge. After the model converges, Z 1 , Z 2 and Z 3 Doubling to 2 makes the model converge again. Then Z 1 , Z 2 and Z 3 Double to 4 and continue training the model until Z 1 , Z 2 and Z 3 When the value of doubles to the set threshold, the search is stopped. For example, when it doubles to 16, the update optimization of the super network is stopped.
[0087] Step S104, obtaining an initial mineral identification model according to the network parameter weights and architecture parameters of the optimized supernetwork; and pruning the initial mineral identification model according to the optimal pruning channel corresponding to the channel shielding parameters of the optimized supernetwork to obtain a mineral identification model, so as to perform mineral identification through the mineral identification model.
[0088] As an implementation method, an initial mineral identification model is obtained according to the network parameter weights and architecture parameters of the optimized super network; including: according to the values of the network parameter weights and architecture parameters of the optimized super network, the operations with the largest weight system in the edge operations and fusion operations are retained to obtain the initial mineral identification model.
[0089] It can be understood that after stopping the search, the edge operation and the node fusion operation are discretized according to the value of the architecture parameter (ɑ, ε), that is, the operation with the largest architecture parameter of the edge operation and the fusion operation on the intermediate node is retained as the final subnet structure searched. That is, the DARTS strategy is used to search for the best edge operation on each operation unit and the best fusion operation on the intermediate node. Then, the channel shielding parameter Z found in the search process is 1 , Z 2 and Z 3 Applied to the final searched model structure, it aims to select the best pruning channel to perform feature operation calculations, thereby improving the computational efficiency of the model and reducing memory overhead, so that the searched model can be suitable for resource-constrained target platforms.
[0090] The mineral identification method based on multimodal architecture search and channel pruning technology in the embodiment of the present application assigns a weight to each operation on the hypernetwork model through differentiable neural architecture search, so as to find the operation that contributes the most to the model during the model training process, and apply channel pruning technology to reduce the number of model channels during the search process to reduce the amount of calculation and memory usage, so as to provide a high-performance, lightweight mineral identification model for resource-constrained platforms, and the model finally searched can be deployed on platforms with different hardware facilities. At the same time, the mineral identification model is obtained by jointly training the multiple modal features of minerals, so that the model can still have a high mineral identification accuracy rate under various environmental influences such as light and weather, so as to improve the stability and reliability of the model in complex environments.
[0091] In the description of the aforementioned embodiments, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0092] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of this application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0093] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0094] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute the instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0095] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0096] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0097] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0098] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A mineral identification method based on multimodal architecture search and channel pruning technology, characterized in that: The following steps are involved: Acquire a mineral multimodal data set, and perform feature extraction on the mineral multimodal data set to obtain multi-scale multimodal features; According to the multi-scale multi-modal features, a search space including different operations is defined and a hypernetwork based on operation units is constructed; wherein the search space uses a search method based on the operation units to perform neural architecture search, and the different operations are all provided with architecture parameters; The network parameter weights and architecture parameters of the hypernetwork are alternately optimized using a differentiable neural architecture search technique, and the number of channels of the hypernetwork is reduced by a channel pruning technique based on channel shielding parameters during the alternate optimization process to obtain an optimized hypernetwork; An initial mineral identification model is obtained according to the network parameter weights and architecture parameters of the optimized supernetwork; and the initial mineral identification model is pruned according to the optimal pruning channel corresponding to the channel shielding parameters of the optimized supernetwork to obtain a mineral identification model, so as to perform mineral identification through the mineral identification model.
2. The method according to claim 1, characterized in that The step of obtaining a mineral multimodal data set and performing feature extraction on the mineral multimodal data set to obtain multi-scale multimodal features comprises: Obtain mineral RGB image dataset and mineral depth image dataset; Perform feature extraction on the mineral RGB image dataset through a VGG network to obtain multi-scale features of the RGB image; Performing feature extraction on the mineral depth image dataset through the ResNeXt-101 network to obtain multi-scale features of the depth image; The multi-scale features of the RGB image and the multi-scale features of the depth image are fused to obtain multi-scale multi-modal features.
3. The method according to claim 1, characterized in that The operation unit includes two input nodes, four intermediate nodes and one output node, and the input nodes, the intermediate nodes and the output nodes are connected in sequence through edges to form a directed acyclic graph; the different operations include multiple edge operations on each edge and multiple fusion operations on each of the intermediate nodes.
4. The method according to claim 2, characterized in that: The super network includes an RGB image unit structure, a depth image unit structure and a multimodal fusion unit structure. The internal structures of the RGB image unit structure, the depth image unit structure and the multimodal fusion unit structure are all operation units. The operation unit includes two input nodes, four intermediate nodes and one output node, and the input nodes, the intermediate nodes and the output nodes are connected in sequence through edges to form a directed acyclic graph; the different operations include multiple edge operations on each edge and multiple fusion operations on each intermediate node; each of the edge operations and each of the fusion operations is set with an architecture parameter, the sum of the architecture parameters of the multiple edge operations on each edge is 1, and the sum of the architecture parameters of the multiple fusion operations on each intermediate node is 1; the architecture parameters of the edge operation are used to characterize the impact of the edge operation on the classification accuracy of the entire super network; the architecture parameters of the fusion operation of the intermediate node are used to determine the optimal fusion operation of the features input to the current intermediate node; wherein the architecture parameters are used as the architecture parameters of the super network.
5. The method according to claim 4, characterized in that The method of alternately optimizing the network parameter weights and the architecture parameters of the hypernetwork using the differentiable neural architecture search technology includes: Initializing network parameter weights of the hypernetwork and the architecture parameters; The network parameter weights and architecture parameters of the super network are updated alternately, and the alternating update includes fixing the architecture parameters and updating the network parameter weights based on a first loss function; and fixing the network parameter weights and updating the architecture parameters based on a second loss function.
6. The method according to claim 5, characterized in that The channel masking parameters include a first channel masking parameter for determining the number of channel masking operations on the edge of the RGB image unit structure, a second channel masking parameter for determining the number of channel masking operations on the edge of the depth image unit structure, and a third channel masking parameter for determining the number of channel masking operations on the edge of the multimodal fusion unit structure.
7. The method according to claim 6, characterized in that The alternately updating the network parameter weights and architecture parameters of the super network includes: Initialize the value of the channel shielding parameter, and perform the following steps until the value of the channel shielding parameter reaches a parameter threshold; Alternatingly updating network parameter weights and architecture parameters of the hypernetwork until the hypernetwork converges; According to the first updating rule, the value of the current channel shielding parameter is updated.
8. The method according to claim 7, characterized in that The first updating rule is to double the value of the current channel shielding parameter.
9. The method according to claim 7, characterized in that: The method of obtaining an initial mineral identification model according to the network parameter weights and architecture parameters of the optimized supernetwork comprises: According to the network parameter weights and architecture parameter values of the optimized supernetwork, the operations with the largest weight system in the edge operations and the fusion operations are retained to obtain an initial mineral identification model.
10. The method according to claim 3, characterized in that The multiple edge operations include skip connections, 3x3 depthwise separable convolutions, 5x5 depthwise separable convolutions, 7x7 depthwise separable convolutions, 3x3 dilated convolutions, 5x5 dilated convolutions, 7x7 dilated convolutions, average pooling, maximum pooling, and None operations, and the multiple fusion operations include cascade operations and summation operations.