Dynamic and Efficient Network Training Method, Device, Computer Equipment and Storage Medium Based on Category Hierarchy

Through a dynamic and efficient network training method based on category levels, using image enhancement and clustering technology to dynamically adjust the calculation path of neural networks, the problem of poor performance of neural networks in image classification is solved, and higher recognition accuracy and speed are achieved.

CN116071591BActive Publication Date: 2025-07-22CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310123092.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-16
Publication Date
2025-07-22
Estimated Expiration
2043-02-16

AI Technical Summary

Technical Problem

The improved neural networks in the prior art have poor performance in image classification tasks, which are limited by the defects in the training data.

Method used

Through dynamic and efficient network training methods based on category levels, including image enhancement processing, clustering, coarse recognition and fine recognition, the sub-network is selected using the path selection mask, dynamically adjusts the calculation path, and builds a loss function to update the neural network parameters.

Benefits of technology

It improves the recognition accuracy and recognition rate of neural networks, breaks the traditional static reasoning and fixed feedforward computing mode, and improves computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116071591B_ABST
    Figure CN116071591B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of deep learning, and specifically relates to a dynamic and efficient network training method, device, computer device and storage medium based on class hierarchy. The method includes: obtaining a sample image and its class label; invoking a first classification module of a neural network to identify the clustering result of the sample image, and obtaining the first predicted class of each sample image; invoking a second classification module of the neural network, and based on the corresponding relationship between the first predicted class and the sub-network, identifying each sample image based on the corresponding sub-network to obtain the second predicted class of each sample image; updating the neural network according to the first predicted class, the second predicted class and the class label of the sample image until a preset condition is met. The present invention starts from the data characteristics and further improves the network performance by using the image class relationship; focuses on learning the partial similar features of similar sample images, improving the recognition accuracy and recognition rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a dynamic and efficient network training method, device, computer device and storage medium based on class hierarchy. Background Art

[0002] Deep learning is a new technical field generated during the research process of machine learning. Specifically, deep learning is a method of deep representation learning of data in machine learning. Deep learning interprets data by establishing a neural network that simulates the human brain for analysis and learning. Different neural networks can be applicable to different scenarios (for example: classification) or provide different effects when used in the same scenario.

[0003] Currently, in order to optimize the performance of neural networks for image classification, in related technologies, the network structure of neural networks is improved to improve the classification accuracy of the improved neural networks. However, limited by the defects in the training data, the improved neural networks still have problems with poor performance. Summary of the Invention

[0004] To solve the above technical problems, the present invention proposes a dynamic and efficient network training method, device, computer device and storage medium based on class hierarchy.

[0005] In a first aspect, the present invention provides a dynamic and efficient network training method based on class hierarchy, including the following steps:

[0006] S1: Obtain n sample images and the class labels of the sample images; where n is an integer greater than 0;

[0007] The class labels of the sample images include: a first label class and a second label class, and the first label class and the second label class have an inclusion relationship, and a single first label class includes one or more second label classes;

[0008] S2: Perform image enhancement processing on the obtained n sample images, and cluster the n sample images after image enhancement according to the correlation between the sample images;

[0009] S3: Invoke the first classification module of the neural network to roughly identify the clustering results of the n sample images to obtain the first predicted classes of the n sample images respectively;

[0010] S4: For the n sample images, invoke the second classification module of the neural network, and perform fine identification on each sample image through the routing mask of the second classification module to obtain the second predicted classes of the n sample images respectively;

[0011] S5: Calculate the first loss value of the neural network based on the first predicted category and the category label of each of the n sample images, calculate the second loss value of the neural network based on the second predicted category and the category label of each of the n sample images, construct a loss function of the neural network based on the first loss value and the second loss value, and update the parameters of the neural network when the loss function is minimized;

[0012] S6: Repeat the above steps until the preset condition is satisfied, and use the neural network updated most recently as the target neural network.

[0013] In a second aspect, the present invention provides a dynamic and efficient network training device based on a category hierarchy, and the device includes: a data acquisition module, a clustering module, a first classification module, a second classification module, and an update module;

[0014] The data acquisition module is configured to acquire n sample images and the category labels of the sample images;

[0015] The clustering module is configured to perform image enhancement and clustering processing on the acquired sample images;

[0016] The first classification module is configured to perform rough category recognition on the clustered result to obtain the first predicted category of the sample image;

[0017] The second classification module is configured to perform fine category recognition on the sample image to obtain the second predicted category of the sample image;

[0018] The update module is configured to update the parameters of the target network according to the first predicted category and the second predicted category.

[0019] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the steps of the network training method are implemented.

[0020] In a fourth aspect, the present invention provides a computer storage medium, and the computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the network training method are implemented.

[0021] Advantages of the present invention: Starting from data characteristics, the present invention further improves network performance by utilizing the relationship between image categories; during the training process, it can focus on learning partial similar features of similar sample images, improving the recognition accuracy and recognition rate; it breaks the traditional static inference and fixed feed-forward calculation mode, selects a sub-network through a routing mask, makes the use of the neural network more reasonable, and improves the calculation efficiency. Description of the Drawings

[0022] Figure 1 Schematic diagram of the dynamic and efficient network training method based on category hierarchy of the present invention;

[0023] Figure 2 Schematic diagram of the sample image and its category label of the present invention;

[0024] Figure 3 Schematic diagram of channel selection of the present invention. Specific embodiments

[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0026] With the development of deep learning technology, neural networks are widely used in image classification scenarios. The inventor found in the research that in the face of a large amount of image data, the similarity between image data will affect the recognition accuracy of the neural network. In view of this, based on the similarity of training data, the present application proposes a dynamic and efficient network training method, device, computer device and storage medium based on category hierarchy to improve the performance of the neural network. The following will make a detailed description of the technical solutions such as the dynamic and efficient network training method, device, computer device and storage medium based on category hierarchy provided by the present application in conjunction with specific embodiments.

[0027] As Figure 1 shown, a dynamic and efficient network training method based on category hierarchy. This neural network training method can be applied to a terminal device or a server, or can be jointly executed by the terminal device and the server, and includes the following steps:

[0028] S1: Obtain n sample images and the category labels of the sample images; where n is an integer greater than 0;

[0029] S2: Perform image enhancement processing on the obtained n sample images, and cluster the n sample images after image enhancement according to the correlation between the sample images;

[0030] S3: Invoke the first classification module of the neural network to perform rough recognition on the clustering results of the n sample images to obtain the first predicted category of each of the n sample images;

[0031] S4: For the n sample images, invoke the second classification module of the neural network, and perform fine recognition on each sample image through the routing mask of the second classification module to obtain the second predicted category of each of the n sample images;

[0032] S5: Calculate the first loss value of the neural network according to the first predicted category and the category label of each of the n sample images, calculate the second loss value of the neural network according to the second predicted category and the category label of each of the n sample images, construct a loss function of the neural network based on the first loss value and the second loss value, and update the parameters of the neural network when the loss function is minimized;

[0033] S6: Repeat the above steps until the preset condition is satisfied, and use the neural network updated most recently as the target neural network.

[0034] n (n is an integer greater than 0) sample images and their category labels are used as the training data of the neural network. The sample labels include a first label category and a second label category. The first label category and the second label category have an inclusion relationship, and a single first label category includes one or more second label categories.

[0035] As Figure 2 shown, it is a schematic diagram of the sample images and their category labels of the embodiment of the present application; among them, 3 sample images are listed: the first category labels of sample image 1 and sample image 2 are the same as vehicle, and their second category labels are sedan and truck respectively. The first category label of sample image 3 is human, and the second category label is man.

[0036] Perform image enhancement processing on the obtained n sample images, including: based on the difference between the object area and the non-object area in the sample image, increase the attention to the object area and reduce the attention to the non-object area, so that the features of the sample image are more effective and clear, which can improve the accuracy of the clustering process, and further improve the classification accuracy of the first classification module.

[0037] Cluster the n sample images after image enhancement according to the correlation between the sample images, including: construct a two-layer coarse-to-fine category hierarchy by clustering the image features to obtain v fine categories c = {c1, c2,..., c v} and u coarse categories C = {C1, C2,..., C u}, and divide the clusters with similar visual features of the fine categories and the coarse categories into the same cluster. A cluster is a coarse category, and m coarse categories are obtained.

[0038] The classification principle of the first classification module is to solve the problem of similar images existing in the n sample images based on the m coarse categories obtained by the clustering process, which is beneficial to improving the accuracy and efficiency of the first classification module for identifying and classifying each coarse category, and then using the coarse category of each coarse category as the first predicted category of the corresponding sample image.

[0039] Invoke the first classification module of the neural network to perform rough recognition on the clustering results of the n sample images, including:

[0040] Based on the clustering results, perform rough recognition on each sample image in the m rough classes respectively, to obtain the rough categories corresponding to the m rough classes respectively;

[0041] For the m rough classes, use the rough category corresponding to each rough class as the first predicted category of the sample images in each rough class, to obtain the first predicted category of each of the n sample images.

[0042] Perform routing masking on the second classification module, and perform fine recognition on each sample image through the routed second classification module, including:

[0043] Perform routing masking according to the correspondence between the first predicted category and its sub-network parameters, determine the sub-network parameters of a single sample image, use the determined sub-network parameters of the single sample image to fix the network parameters of the shared network in the second classification module, and perform fine recognition on the object in the single sample image through the shared network with fixed network parameters, to obtain the second predicted category of the single sample image.

[0044] In the shared network, on the one hand, since different convolutional kernels have different importance for different first predicted categories, that is, different first predicted categories select different activated or discarded neural kernels, and on the other hand, since different iterative channels also have different importance for the first predicted categories, that is, different first predicted categories also select different channels, the above method realizes the reuse of features of large categories (first predicted categories), and improves the performance of the neural network.

[0045] For example, although the first predicted category vehicle can include multiple second predicted categories such as train, passenger car, freight car, sedan car, etc., all vehicles have certain common features. It is understandable that if a single sub-network is iteratively trained based on the sample images with these common features, the single sub-network can be more sensitive, accurate and fast in recognizing vehicle-related images, that is, the performance of the single sub-network in recognizing vehicle-related images can be significantly improved.

[0046] The shared network parameters are the network parameters of all determined first predicted categories. This reuse idea helps to reduce the amount of calculation. Even if the number of required sub-networks increases, there will be no problem of increased calculation amount due to the additional number of sub-networks generated.

[0047] The shared network is a directed graph including many nodes and their connection relationships. The generation of a sub-network or sub-network parameters can be regarded as determining a target path in the directed graph by node selection. It can be understood that the sizes of sub-networks are different, so the granularities of the selected paths are also different. In this regard, the embodiments of the present application also propose three different granularity routing methods for further training the neural network, including: weight selection, channel selection, and residual block selection;

[0048] The weight selection: The weight selection is pruning at the single weight granularity, that is, selecting an optimal combination from numerous weight parameters to minimize the cost function loss of the pruned target model (such as: sub-network); The pruning method introduced here is the pruning method based on the absolute value. In this method, the importance of the weight is measured according to the absolute value of the weight;

[0049] Update the weight mask for each layer of the sub-network corresponding to the first predicted category, including:

[0050]

[0051] where m i,j,h,w represents the weight mask for each layer of the sub-network corresponding to the first predicted category, W s i,j,h,w represents the weight of the sub-network corresponding to the first predicted category, λ represents the mask threshold. When the absolute value of the weight is greater than or equal to the threshold λ, the threshold is set to 1, indicating that the weight is retained. When the absolute value of the weight is less than the threshold λ, the threshold is set to 0, indicating that the weight is pruned;

[0052] The channel selection: The channel refers to the output channel in the convolutional layer. The channel selection is to select the output channel, that is, to select the convolution kernel and cut off the unimportant convolution kernels. The channel selection method used in the embodiments of the present application is the channel selection method based on the L2 norm. Among the convolution kernels in the same convolutional layer, the convolution kernels with smaller weights tend to have weakly activated feature maps, and the L2 norm is a good criterion for channel selection without data;

[0053] S41: Calculate the L2 norm of each convolution kernel :

[0054] S42: Sort the convolution kernels according to the size of s j ;

[0055] S43: Select a value in s j as the threshold according to the number of convolution kernels to be retained, and set the mask corresponding to the convolution kernels smaller than the threshold in s j to 0, indicating deleting the convolution kernel, s jThe mask corresponding to the convolution kernel greater than or equal to the threshold is set to 1, indicating that the convolution kernel is retained;

[0056] This channel selection method is hardware-friendly and can completely delete the unnecessary weights instead of using 0 to occupy the position. As Figure 3 shown, the feature map output by the previous layer is used as the input of the next layer. In the figure, the dark color in the solid-line convolution layer represents the convolution kernel deleted in the current layer, and the dark color in the dotted-line feature map layer represents the feature map that disappears due to the pruning of the convolution kernel. The light color represents the influence of the pruning of the upper-layer convolution kernel on the current layer. It is not difficult to find after observation that either all the convolution kernels in each layer are pruned after channel selection, or the same input channels of each convolution kernel are pruned. Therefore, after the training and pruning processes are completed, by copying these weights to a compact new network according to the corresponding positions, the weights that have been set to 0 can be completely discarded.

[0057] The selection of the residual block is as follows: The residual block can be a structural module in the ResNet. The ResNet is relatively deep and has many such residual blocks, and the association between blocks is not large. Pruning the residual block is a pruning method with a larger granularity. Before selecting the residual block, a mask is added to each block for selection, and the process of selecting the residual block is the update of the mask. Different from the hard mask introduced in the above non-structural pruning and channel selection methods, a soft mask is used here, and the mask is updated together with the loss function of the target network. During the training process, the mask has more than just two values of 0 and 1. After the training is completed, the residual block is selected according to the size of the mask.

[0058] Updating the mask together with the loss function of the target network includes:

[0059]

[0060] where L s represents the total loss function for updating the weights of the sub-network corresponding to the first predicted class, that is, the sum of the generator loss, the intermediate loss, and the output loss. R(m) represents the sparse regularization term for the soft mask, μ is the trade-off factor, m represents the routing mask, and m * represents the optimal routing mask.

[0061] The above are three dynamic routing methods proposed in the embodiments of the present application. They are mainly designed according to the requirements of different sizes of training data (n sample images) and the target neural network, and according to different routing granularities. The idea of iterative pruning is used in the routing method, and during the training process, the convolution of the shared network is selectively activated and discarded according to the first predicted class, so that the calculation path of the first predicted class can be dynamically adjusted according to the clustering result.

[0062] For the logical z calculated for each category by the Softmax function in the first classification module and the second classification module i compare it with other categories and convert it into a probability q i , select the one with the maximum probability and output its predicted category;

[0063] The probability calculation includes:

[0064]

[0065] where z j represents the logic of the j-th category, and there are a total of j categories; T represents a hyperparameter called temperature, which is usually set to 1. The higher the temperature parameter is set, the softer the probability distribution will be over the classes. That is, as the T parameter increases, the probability distribution output by the Softmax function will become more uniform.

[0066] The preset conditions, which are the training termination conditions of the neural network, can include one or more combinations of the following: the current neural network reaches the set accuracy requirement; the number of training iterations of the current neural network reaches the set maximum number of iteration requirements; among them, the accuracy requirement is set for the first prediction label and the second prediction label, and n sample images' first prediction labels and second prediction labels can be obtained in a single training iteration. Of course, the preset conditions can also be set according to the actual situation. For example: the training time reaches the set maximum training time requirement, etc., which are not specifically limited here.

[0067] Calculate the first loss value of the neural network:

[0068]

[0069] where L c represents the first loss value, f() represents the cross-entropy loss function, q c represents the normalization result for the first predicted categories of n sample images, y c represents the category labels of n sample images respectively, represents the normalization result for the first predicted category of the i-th sample image, represents the category label of the i-th sample image;

[0070] Calculate the second loss value of the neural network:

[0071]

[0072] where L f represents the second loss value, f() represents the cross-entropy loss function, q fRepresents the normalization result for the second predicted category of each of the n sample images, y f Represents the class label of each of the n sample images, Represents the normalization result for the second predicted category of the i-th sample image, Represents the class label of the i-th sample image;

[0073] Construct the loss function of the neural network based on the first loss value and the second loss value:

[0074]

[0075] where, L all Represents the target loss function of the neural network, α, β represent the first and second trade-off factors, γ represents the weight decay coefficient, W c Represents the weight of the first classification module, W f Represents the weight of the second classification module, L c Represents the first loss value, L f Represents the second loss value.

[0076] The first updated weight and the second updated weight of the above neural network can be calculated for the target loss function, and the calculation formula is as follows:

[0077]

[0078] where, L all is the target loss function, W c is the first weight of the first classification module, W f is the first updated weight of the first classification module, is the second weight of the second classification module, is the second updated weight of the second classification module.

[0079] As a more optimal implementation, based on different granularity routing methods (pruning methods), a loss term for mask usage can also be introduced, that is, calculate the routing loss value to obtain the optimal routing mask, that is, better sub-network parameters, and improve the performance of the sub-network.

[0080] Specifically, if a loss term is introduced into the target loss function, the target loss function can be expressed as:

[0081]

[0082] where, L all is the target loss function, W c is the first weight of the first classification module, W f is the first updated weight of the first classification module, is the second weight of the second classification module, is the second updated weight of the second classification module, m * is the sub-network parameter (i.e., the routing mask), and h(m) is the loss term used for the mask in the corresponding pruning method.

[0083] The present invention also provides a dynamic and efficient network training device based on category hierarchy, including: a data acquisition module, a clustering module, a first classification module, a second classification module, and an update module;

[0084] The data acquisition module is used to acquire n sample images and the category labels of the sample images;

[0085] The clustering module is used to perform image enhancement and clustering processing on the acquired sample images;

[0086] The first classification module is used to perform rough category recognition on the clustered result to obtain the first predicted category of the sample image;

[0087] The second classification module is used to perform fine category recognition on the sample image to obtain the second predicted category of the sample image;

[0088] The update module is used to update the parameters of the target network according to the first predicted category and the second predicted category.

[0089] On the one hand, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the network training method are implemented.

[0090] On the one hand, the present invention provides a computer storage medium. The computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the network training method are implemented.

[0091] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A dynamic and efficient network training method based on category hierarchy, characterized in that, Including: S1: Obtain n sample images and the class labels of the sample images; where n is an integer greater than 0; The class labels of the sample images include: a first label category and a second label category, and the first label category and the second label category have an inclusion relationship, and a single first label category includes one or more second label categories; S2: Perform image enhancement processing on the obtained n sample images, and cluster the n sample images after image enhancement according to the correlation between the sample images; Clustering the n sample images after image enhancement according to the correlation between the sample images, including: constructing a two-layer coarse-to-fine category hierarchy by clustering image features to obtain v fine categories c = {c1, c2,..., c v}, and u coarse categories C = {C1, C2,..., C u}, dividing the clusters with similar visual features of the fine categories and the coarse categories into the same cluster, and one cluster is a coarse category, obtaining m coarse categories; S3: Invoke the first classification module of the neural network to perform rough recognition on the clustering results of the n sample images to obtain the first predicted categories of the n sample images respectively; Invoking the first classification module of the neural network to perform rough recognition on the clustering results of the n sample images includes: Based on the clustering results, perform rough recognition on each sample image in m rough classes respectively to obtain the rough categories corresponding to the m rough classes respectively; For the m rough classes, use the rough category corresponding to each rough class as the first predicted category of the sample images in each rough class to obtain the first predicted categories of the n sample images respectively; S4: For the n sample images, invoke the second classification module of the neural network, and perform fine recognition on each sample image through the routing mask of the second classification module to obtain the second predicted categories of the n sample images respectively; Performing fine recognition on each sample image through the routing mask of the second classification module includes: Perform routing masking according to the correspondence between the first predicted category and its sub-network parameters to determine the sub-network parameters of a single sample image, use the determined sub-network parameters of the single sample image to fix the network parameters of the shared network in the second classification module, and perform fine recognition on the object in the single sample image through the shared network with fixed network parameters to obtain the second predicted category of the single sample image; The routing mask includes: weight selection, channel selection, and residual block selection; S5: Calculate the first loss value of the neural network according to the first predicted categories and class labels of the n sample images respectively, calculate the second loss value of the neural network according to the second predicted categories and class labels of the n sample images respectively, construct a loss function of the neural network based on the first loss value and the second loss value, and update the parameters of the neural network when the loss function is minimized; S6: Repeat the above steps until a preset condition is met, and use the neural network updated most recently as the target neural network.

2. The dynamic and efficient network training method based on category hierarchy according to claim 1, characterized in that, Performing image enhancement processing on the obtained n sample images includes: based on the difference between the object region and the non-object region in the sample image, increasing the attention to the object region and decreasing the attention to the non-object region.

3. A dynamic and efficient network training method based on category hierarchy according to claim 1, characterized in that, The weight selection: Update the mask of the weight of each layer of the sub-network corresponding to the first predicted category; Updating the mask of the weight of each layer of the sub-network corresponding to the first predicted category includes: Among them, m i,j,h,w represents the weight mask for each layer of the sub-network corresponding to the first predicted category, and W s i,j,h,w represents the weight of the sub-network corresponding to the first predicted category, and λ represents the mask threshold; The channel selection: Perform mask selection of the convolution kernel on the output channels in the convolution layer of the sub-network corresponding to the first predicted category; S41: Calculate the L2 norm of each convolutional kernel W s i,j : s j = ∑ i ‖W s i,j ||2; S42: Sort the convolutional kernels according to the size of s j ; S43: Select a value from s as the threshold according to the number of remaining convolutional kernels, and set the mask corresponding to the convolutional kernel with s j less than the threshold to 0, indicating that the convolutional kernel is deleted, and s j Set the mask corresponding to the convolutional kernel greater than or equal to the threshold to 1, indicating that the convolutional kernel is retained; j ​ The selection of the residual block: Add a selectable soft mask to each residual block of the sub-network corresponding to the first predicted category, and update the mask together with the loss function of the sub-network corresponding to the first predicted category; Updating the mask together with the loss function of the sub-network corresponding to the first predicted category includes: Among them, L s represents the total loss function for updating the sub-network weights corresponding to the first predicted category, that is, the sum of three losses: the generator loss, the intermediate loss, and the output loss. R(m) represents the sparse regularization term for the soft mask, μ is the trade-off factor, m represents the routing mask, and m * represents the optimal routing mask.

4. A dynamic and efficient network training method based on category hierarchy according to claim 1, characterized in that The specific content of S5 includes: Calculate the first loss value of the neural network: Among them, L c represents the first loss value, f() represents the cross-entropy loss function, and q c represents the normalization result for the first predicted category of n sample images, and y c represents the class labels of the n sample images respectively, represents the normalization result for the first predicted category of the i-th sample image, represents the class label of the i-th sample image; Calculate the second loss value of the neural network: Among them, L f represents the second loss value, f() represents the cross-entropy loss function, and q f represents the normalization result for the second predicted class of each of the n sample images, and y f represents the class label of each of the n sample images, represents the normalization result for the second predicted class of the i-th sample image, represents the class label of the i-th sample image; Construct the loss function of the neural network based on the first loss value and the second loss value: Among them, L all represents the neural network objective loss function, α and β represent the first and second trade-off factors, γ represents the weight decay coefficient, and W c represents the weight of the first classification module, and W f represents the weight of the second classification module, L c represents the first loss value, and L f represents the second loss value.

5. A dynamic and efficient network training device based on class hierarchy, which is used to implement a dynamic and efficient network training method based on class hierarchy as described in any one of claims 1-4, and is characterized in that, Including: Data acquisition module, clustering module, first classification module, second classification module, update module; The data acquisition module is used to acquire n sample images and the category labels of the sample images; The clustering module is used to perform image enhancement and clustering processing on the acquired sample images; The first classification module is used to perform rough category recognition on the clustered results to obtain the first predicted category of the sample images; The second classification module is used to perform fine category recognition on the sample images to obtain the second predicted category of the sample images; The update module is used to update the parameters of the target network according to the first predicted category and the second predicted category.

6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the network training method according to any one of claims 1 to 4 are implemented.

7. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the network training method according to any one of claims 1 to 4 are implemented.