Image analysis method and device, computer equipment and storage medium
By constructing a student model and incorporating a dynamic exit model, the problem of high resource consumption in existing high-precision image analysis methods is solved, achieving efficient image analysis.
Patent Information
- Application Number
- CN202511067425.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-11-07
AI Technical Summary
Existing image analysis methods struggle to balance high accuracy and low resource consumption, resulting in low efficiency and high resource consumption.
By constructing a first student model and updating the total loss function using a high-precision teacher model, and adding a dynamic exit model for training, an image analysis model is obtained, reducing computational load and achieving dynamic exit.
While maintaining high accuracy, the computational load and system consumption of the model were reduced, and the computational speed and efficiency were improved.
Smart Images

Figure CN120913035A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image analysis, and in particular to an image analysis method and device, computer equipment and a storage medium. BACKGROUND
[0002] With the development of information technology and the Internet, the demand for image analysis methods is increasing.
[0003] In the prior art, a large-scale image analysis model can achieve high-precision image analysis, but the large-scale image analysis model is low in efficiency and high in resource consumption. The model with less resource consumption is low in accuracy.
[0004] It can be seen that the image analysis method in the prior art cannot balance high precision and low resource consumption. SUMMARY
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present application provides an image analysis method, device, computer equipment and storage medium.
[0006] In a first aspect, the present application provides an image analysis method, comprising:
[0007] constructing a first student model according to a high-precision teacher model;
[0008] updating a total loss function of the first student model according to the high-precision teacher model and training data to obtain a second student model;
[0009] adding a dynamic exit model to the second student model and training to obtain an image analysis model.
[0010] Optionally, the first student model is a bidirectional state space model, and the first student model comprises a plurality of bidirectional state space layers, each of the bidirectional state space layers being calculated in the following manner:
[0011] z l =z l-1 +Vim(z l-1 ),l=1…L
[0012] Y=head(z L )
[0013] wherein z L is the intermediate feature output by the last layer of the first student model, Y is a prediction vector processed by a classification head, L is the total number of layers of the first student model, and Vim represents a bidirectional state space layer.
[0014] Optionally, the updating, according to the high-precision teacher model and training data, of the total loss function of the first student model to obtain a second student model comprises:
[0015] obtaining a self-loss function of the first student model;
[0016] obtaining, according to the high-precision teacher model, a distillation loss function of the first student model and a normalization loss function of the first student model;
[0017] obtaining, according to the self-loss function, the distillation loss function and the normalization loss function, a total loss function of the first student model;
[0018] obtaining, according to the total loss function, the second student model.
[0019] Optionally, the self-loss function of the first student model is obtained in the following manner:
[0020]
[0021] wherein Y is a prediction vector output by the first student model, y is a true label of the training data, (s) is a cross-entropy loss function;
[0022] the distillation loss function is obtained in the following manner:
[0023]
[0024] wherein, is a distillation loss function, is an inter-class relationship loss function, is an intra-class relationship loss function, D is a batch size during training of the first student model, is a Pearson distance between and , wherein is an i-th prediction vector of the first student model in one batch, is an i-th prediction vector of the high-precision teacher model in one batch, C is a number of classes of the training data, is a j-th column of a prediction matrix Y s of the first student model, is a j-th column of a prediction matrix Y (t) of the high-precision teacher model;
[0025] the normalization loss function is obtained in the following manner:
[0026]
[0027] wherein, is a normalization loss function, is an index set of training data belonging to the k-th class, and C is the number of classes of the training data, is the intermediate feature of the last layer output of the first student model is the feature vector of the i-th row of is the intermediate feature of the last layer output of the teacher network is the feature vector of the i-th of k is the intermediate parameter of the teacher network, e k is the unit vector in the direction of EMB k .
[0028] Optionally, the total loss function of the first student model is obtained according to the self-loss function, the distillation loss function and the normalization loss function in the following manner:
[0029]
[0030] wherein, is a total loss function, is a self-loss function of the first student model, is a distillation loss function, is a normalization loss function, and a and β are parameters for balancing the loss.
[0031] Optionally, the second student model is trained by adding a dynamic exit model to obtain an image analysis model, which comprises:
[0032] In the second student model, a plurality of dynamic exit points are obtained, and the dynamic exit module is added to each dynamic exit point;
[0033] According to the real label, the loss value of each dynamic exit module is obtained;
[0034] The minimum loss value in the loss value is obtained, and a pseudo label is obtained according to the minimum loss value;
[0035] The training data is evenly divided into first data and second data;
[0036] After freezing all parameters in the second student model, a training loss function of a gating module is obtained according to the pseudo label and the first data to update the dynamic exit module;
[0037] According to the real label and the second data, a training loss function of an intermediate classification head is obtained to update the dynamic exit module;
[0038] According to the real label, update the loss value of each dynamic exit module, and update the first data and the second data to update the dynamic exit module until a convergence condition is met, and obtain the image analysis network.
[0039] Optionally, according to the real label, the loss value of each dynamic exit module is obtained in the following manner:
[0040]
[0041] is the loss value of the b-th layer of the second student model when exiting, IC b represents the inference cost of the b-th layer of the second student model when exiting, Y b is the prediction vector of the b-th layer of the second student model, y is the real label of the training data, and λ is a weight parameter;
[0042] The training loss function of the gating module is obtained in the following manner:
[0043]
[0044] represents the value of i that minimizes is the training loss function of the gating module, g b represents the output value of the gating module when the b-th layer of the second student model exits, is a binary cross-entropy loss function, exit b is a pseudo label of the first data;
[0045] The training loss function of the intermediate classification head is obtained in the following manner:
[0046]
[0047] is the training loss function of the intermediate classification head, Y b is the prediction vector of the b-th layer of the second student model when exiting, y' is the real label of the second data, is a set of insertion points, is a cross-entropy loss function, w b represents the probability distribution of the training data when the b-th layer of the second student model exits.
[0048] In a second aspect, an image analysis device is provided, comprising:
[0049] a student model unit configured to construct a first student model according to the high-precision teacher model;
[0050] a training unit configured to update a total loss function of the first student model according to the high-precision teacher model and training data, to obtain a second student model;
[0051] a quitting unit configured to add a dynamic quitting model to the second student model and train the second student model, to obtain an image analysis model.
[0052] In a third aspect, a computer device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method according to any one of the preceding aspects when executing the computer program.
[0053] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program is executable on a processor to implement the method according to any one of the preceding aspects.
[0054] The present application provides an image analysis method, device, computer device, and storage medium. The method includes: constructing a first student model according to a high-precision teacher model; updating a total loss function of the first student model according to the high-precision teacher model and training data, to obtain a second student model; and adding a dynamic quitting model to the second student model and training the second student model, to obtain an image analysis model. The high-precision teacher model can be a high-precision model. The first student model is constructed according to the high-precision teacher model, and the loss function of the first student model is updated according to the high-precision teacher model and training data to obtain the second student model. The calculation amount of the second student model can be reduced on the basis of retaining the high precision of the high-precision teacher model, thereby reducing system consumption and improving the calculation speed of the model. In addition, the dynamic quitting model is added to the second student model in the present application, so that the image analysis model of the present application can quit dynamically without executing all models to quit, thereby further reducing system consumption and improving the calculation speed of the model. BRIEF DESCRIPTION OF DRAWINGS
[0055] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles behind the application.
[0056] To make the technical solutions of the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or the prior art description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained from these drawings without any creative effort.
[0057] Figure 1 An application environment diagram of the image analysis method of the embodiments of the present application is shown.
[0058] Figure 2 Fig. 1 shows a flowchart of an image analysis method according to an embodiment of the present application;
[0059] Figure 3 Fig. 1 shows a flowchart of an image analysis method according to an embodiment of the present application;
[0060] Figure 4 Fig. 2 shows a structural diagram of an image analysis model according to an embodiment of the present application;
[0061] Figure 5 Fig. 3 shows a structural block diagram of an image analysis device according to an embodiment of the present application;
[0062] Figure 6 Fig. 4 shows an internal structural diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION
[0063] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0064] Figure 1 Fig. 1 shows an application environment diagram of an image analysis method according to an embodiment. Referring to Fig. 1, the image analysis method is applied to an image analysis system. The image analysis method includes a terminal 110 and / or a server 120. The terminal 110 and the server 120 are connected through a network. The terminal 110 can be specifically a desktop terminal or a mobile terminal, and the mobile terminal can be specifically at least one of a mobile phone, a tablet computer, a notebook computer and the like. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. Figure 1 The image analysis method of the present application is applied to the terminal 110 and / or the server 120.
[0065]
[0066] Fig. 1 shows an image analysis method according to an embodiment of the present application. The method includes the following steps. Figure 2 In step 210, a first student model is constructed according to a high-precision teacher model.
[0067] In step 220, a total loss function of the first student model is updated according to the high-precision teacher model and training data, to obtain a second student model.
[0068]
[0069] Step 230, adding a dynamic exit model in the second student model and training to obtain an image analysis model.
[0070] In the embodiment of the present application, the high-precision teacher model can be a high-precision model, a first student model is constructed according to the high-precision high-precision teacher model, and a second student model is obtained by updating the loss function of the first student model according to the high-precision teacher model and training data. The calculation amount of the second student model can be reduced on the basis of preserving the high precision of the high-precision teacher model, thereby reducing system consumption and improving the calculation speed of the model. In addition, in the embodiment of the present application, a dynamic exit model is added to the second student model, so that the image analysis model of the embodiment of the present application can realize dynamic exit without executing all models to exit, thereby further reducing system consumption and improving the calculation speed of the model.
[0071] In the embodiment of the present application, the first student model is a bidirectional state space model, and the first student model includes a plurality of bidirectional state space layers, each of which is calculated in the following manner:
[0072] z l =z l-1 +Vim(z l-1 ),l=1…L
[0073] Y=head(z L )
[0074] Wherein, z L is the intermediate feature output by the last layer of the first student model, Y is the prediction vector after classification head processing, L is the total number of layers of the first student model, and Vim represents a bidirectional state space layer.
[0075] In the embodiment of the present application, the first student model is trained according to the high-precision teacher model and training data to obtain a second student model, comprising:
[0076] Obtaining the self-loss function of the first student model;
[0077] According to the high-precision teacher model, obtaining the distillation loss function of the first student model and the normalization loss function of the first student model;
[0078] According to the self-loss function, the distillation loss function and the normalization loss function, obtaining the total loss function of the first student model;
[0079] According to the total loss function, obtaining the second student model.
[0080] In the embodiment of the present application, the self-loss function of the first student model is obtained in the following manner:
[0081]
[0082] where Y (s) is the prediction vector output by the first student model, y is the true label of the training data, is the cross-entropy loss function;
[0083] The distillation loss function is obtained in the following manner:
[0084]
[0085] where, is the distillation loss function, is the inter-class relationship loss function, is the intra-class relationship loss function, D is the batch size during the training of the first student model, is the Pearson distance between and is the i-th prediction vector of the first student model in a batch, is the i-th prediction vector of the high-precision teacher model in a batch, C is the number of classes of the training data, is the j-th column of the prediction matrix Y s of the first student model, is the j-th column of the prediction matrix Y (t) of the high-precision teacher model;
[0086] The normalization loss function is obtained in the following manner:
[0087]
[0088] where, is the normalization loss function, is the index set of the training data belonging to the k-th class, C is the number of classes of the training data, is the i-th row of the feature vector of the intermediate feature output by the last layer of the first student model, is the i-th row of the feature vector of the intermediate feature output by the last layer of the teacher network, EMB k is the intermediate parameter of the teacher network, e k is the unit vector in the direction of EMB k .
[0089] In the embodiment of the present application, the total loss function of the first student model is updated according to the high-precision teacher model and the training data to obtain a second student model, which is a cyclic / iterative process, and a convergence condition can be set to exit the cyclic / iterative process, which will not be described here.
[0090] The Pearson distance can be obtained in the following way:
[0091]
[0092] d p (u,v)=1-ρ p (u,v)
[0093] ρ p (u,v) is the Pearson correlation coefficient between random variables u and v. Here, u and v are input variables, and are the mean of u and v, respectively, and Std(u) and Std(v) are the standard deviations of u and v. Cov(u,v) is the covariance of u and v. u i and v i are the i-th vectors of u and v, respectively, and there are K vectors in total, and d p (u,v) is the Pearson distance.
[0094] In the embodiment of the present application, the total loss function of the first student model is obtained according to the self-loss function, the distillation loss function and the normalization loss function in the following way:
[0095]
[0096] wherein, is the total loss function, is the self-loss function of the first student model, is the distillation loss function, is the normalization loss function, and alpha and beta are parameters for balancing the loss.
[0097] In the embodiment of the present application, the distillation loss function improves the performance in the distillation process, which helps to reduce the model consumption while maintaining high model accuracy. In addition, the total loss function of the embodiment of the present application is beneficial to reduce the model consumption, so that the model is more suitable for deployment to resource-constrained edge devices, while greatly reducing the inference energy consumption.
[0098] In the embodiment of the present application, the dynamic exit model is added to the second student model and trained to obtain an image analysis model, and the image analysis network is obtained by adding a dynamic exit module to the second student model, which comprises:
[0099] In the second student model, a plurality of dynamic exit points are obtained, and the dynamic exit module is added at each dynamic exit point.
[0100] According to the real label, a loss value of each dynamic exit module is obtained.
[0101] The minimum loss value in the loss value is obtained, and a pseudo label is obtained according to the minimum loss value.
[0102] The training data is evenly divided into first data and second data.
[0103] After freezing all parameters in the second student model, a training loss function of the gating module is obtained according to the pseudo label and the first data to update the dynamic exit module.
[0104] According to the real label and the second data, a training loss function of an intermediate classification head is obtained to update the dynamic exit module.
[0105] According to the real label, the loss value of each dynamic exit module is updated, and the first data and the second data are updated to update the dynamic exit module until a convergence condition is met, and the image analysis network is obtained.
[0106] Figure 3 The flowchart of adding a dynamic exit module in the embodiment of the application is shown as follows: Figure 3 As shown, the adding of the dynamic exit module comprises:
[0107] In step 310, in the second student model, a plurality of dynamic exit points are obtained, and the dynamic exit module is added at each dynamic exit point.
[0108] In the embodiment of the application, the dynamic exit points are usually set from the middle layer of the second student model, and each dynamic exit point is spaced by M layers.
[0109]
[0110] Wherein b=1, 2, 3, …, L-1, is a dynamic exit point; L is the total number of layers of the second student model.
[0111] In step 320, according to the real label, a loss value of each dynamic exit module is obtained.
[0112] In step 330, the minimum loss value in the loss value is obtained, and a pseudo label is obtained according to the minimum loss value.
[0113] In step 340, the training data is evenly divided into first data and second data.
[0114] Step 350, after freezing all parameters in the second student model, the dynamic exit module is updated according to the pseudo label and the training loss function of the first data acquisition gating module.
[0115] In the embodiment of the application, freezing all parameters in the second student model means that all parameters in the second student model are not updated during the training process.
[0116] Step 360, the training loss function of the intermediate classification head is obtained according to the real label and the second data, and the dynamic exit module is updated.
[0117] Step 370, it is judged whether the convergence condition is met, if yes, it is turned to step 380, and if no, it is turned to step 320.
[0118] Step 380, the second student network including the dynamic exit model is taken as an image analysis model.
[0119] In the embodiment of the application, the training data needs to be evenly divided into first data and second data in each cycle / iteration, but the data is evenly divided randomly.
[0120] In the embodiment of the application, the loss value of each dynamic exit module is obtained according to the real label in the following manner:
[0121]
[0122] The loss value of the b-th layer exit of the second student model is IC b Y represents the inference cost of the b-th layer exit selected by the second student model b y is the prediction vector of the b-th layer selected by the second student model, y is the real label of the training data, and λ is a weight parameter.
[0123] In the embodiment of the application, the training loss function of the gating module is obtained in the following manner:
[0124]
[0125] Y represents the inference cost of the b-th layer exit selected by the second student model Y represents the inference cost of the b-th layer exit selected by the second student model The training loss function of the gating module is g b Y represents the inference cost of the b-th layer exit selected by the second student model exit is a binary cross-entropy loss function b The pseudo label of the first data is.
[0126] In one embodiment of the present application, if there are a total of 5 dynamic exit points, a total of 5 dynamic exit models are added, and the loss values of the 5 dynamic exit models are calculated, wherein the loss value of the third dynamic exit model is the smallest, and the pseudo labels of the first and second dynamic exit models are 1, and the pseudo labels of the third, fourth and fifth dynamic exit models are 0.
[0127] In the embodiment of the present application, the training loss function of the intermediate classification head is obtained in the following manner:
[0128]
[0129] The training loss function of the intermediate classification head is Y b The prediction vector when the second student model exits at the b-th layer is y', and the true label of the second data is y, is a set of insertion points, is a cross-entropy loss function, and w b represents the probability distribution of the training data when the second student model exits at the b-th layer.
[0130] In the embodiment of the present application, the image analysis model of the dynamic exit module allows the sample to exit early if it meets the exit condition, which can reduce the inference / computation delay, and also helps to reduce the inference / computation power consumption, and realizes the dynamic inference / computation mechanism of the network. In addition, by reasonably planning the position of the exit point in the model, the sample can automatically select the best exit point in the inference stage, further improving the dynamic exit effect of the model.
[0131] Dividing the training data into first data and second data also becomes an alternating training strategy, and updating the parameters of the gating module and the intermediate classification head using the alternating training strategy can make the training stable, and further ensure the accuracy when the sample exits.
[0132] Figure 4 As shown in the figure, the image analysis model includes a plurality of bidirectional state space layers, and also includes a plurality of dynamic exit models. Figure 4 As shown in the figure, the image analysis model includes a plurality of bidirectional state space layers, and also includes a plurality of dynamic exit models. Among them, the dynamic exit model shown by the dashed line represents that the dynamic exit is unsuccessful, and the dynamic exit model shown by the solid line represents that the dynamic exit is successful. The bidirectional state space layer shown in the figure represents execution, and the bidirectional state space layer shown by the dashed line represents that it is no longer executed after the dynamic exit is successful.
[0133] Figure 5 As shown in the figure, the image analysis device includes a plurality of bidirectional state space layers, and also includes a plurality of dynamic exit models. Figure 5 As shown in the figure, the image analysis device includes:
[0134] The student model unit 510 is configured to construct a first student model according to the high-precision teacher model.
[0135] The training unit 520 is configured to update a total loss function of the first student model according to the high-precision teacher model and training data, so as to obtain a second student model.
[0136] The exit unit 530 is configured to add a dynamic exit model to the second student model and train the second student model, so as to obtain an image analysis model.
[0137] In the embodiment of the present application, the first student model is a bidirectional state space model, and the first student model comprises a plurality of bidirectional state space layers, and each bidirectional state space layer is calculated in the following manner:
[0138] z l =z l-1 +Vim(z l-1 ),l=1…L
[0139] Y=head(z L )
[0140] wherein z L is an intermediate feature output by the last layer of the first student model, Y is a prediction vector processed by a classification head, L is the total number of layers of the first student model, and Vim represents a bidirectional state space layer.
[0141] In the embodiment of the present application, the training unit 520 is further configured to:
[0142] obtain a self-loss function of the first student model;
[0143] obtain a distillation loss function of the first student model and a normalization loss function of the first student model according to the high-precision teacher model;
[0144] obtain a total loss function of the first student model according to the self-loss function, the distillation loss function and the normalization loss function;
[0145] obtain the second student model according to the total loss function.
[0146] The training unit 520 is further configured to obtain the self-loss function of the first student model in the following manner:
[0147]
[0148] wherein Y (s) is a prediction vector output by the first student model, y is a true label of training data, is a cross-entropy loss function;
[0149] The distillation loss function is obtained in the following manner:
[0150]
[0151]
[0152] wherein, is a distillation loss function, is an inter-class relation loss function, is an intra-relation loss function, D is a batch size during training of the first student model, is and a Pearson distance between, is an i-th prediction vector of the first student model in a batch, is an i-th prediction vector of the high-precision teacher model in a batch, C is a number of classes of the training data, is a j-th column of a prediction matrix Y s of the first student model, is a j-th column of a prediction matrix Y (t) of the high-precision teacher model;
[0153] The training unit 520 is further configured to obtain the normalization loss function in the following manner:
[0154]
[0155] wherein, is a normalization loss function, is an index set of the training data belonging to the k-th class, C is a number of classes of the training data, is an i-th row of a feature vector of an intermediate feature output by the last layer of the first student model, is an i-th row of a feature vector of an intermediate feature output by the last layer of the teacher network, EMB k is an intermediate parameter of the teacher network, e k is a unit vector in the direction of EMB k .
[0156] In the embodiment of the present application, the training unit 520 is further configured to obtain the total loss function of the first student model in the following manner:
[0157]
[0158] wherein, is a total loss function, is a self-owned loss function of the first student model, is a distillation loss function, is a normalization loss function, and a and β are parameters for balancing the loss.
[0159] In the embodiments of the present application, the exit unit 530 is further configured to:
[0160] In the second student model, a plurality of dynamic exit points are obtained, and the dynamic exit module is added at each dynamic exit point;
[0161] According to the real label, the loss value of each dynamic exit module is obtained;
[0162] The minimum loss value in the loss value is obtained, and the pseudo label is obtained according to the minimum loss value;
[0163] The training data is evenly divided into first data and second data;
[0164] After freezing all parameters in the second student model, the training loss function of the gating module is obtained according to the pseudo label and the first data to update the dynamic exit module;
[0165] According to the real label and the second data, the training loss function of the intermediate classification head is obtained to update the dynamic exit module;
[0166] According to the real label, the loss value of each dynamic exit module is updated, and the first data and the second data are updated to update the dynamic exit module until the convergence condition is met, and the image analysis network is obtained.
[0167] In the embodiments of the present application, the exit unit 530 is further configured to obtain the loss value of each dynamic exit module according to the real label in the following manner:
[0168]
[0169] is the loss value when the bth layer of the second student model exits, IC b represents the inference cost when the bth layer of the second student model exits, Y b is the prediction vector of the bth layer selected by the second student model, y is the real label of the training data, and λ is a weight parameter;
[0170] The exit unit 530 is further configured to obtain the training loss function of the gating module in the following manner:
[0171]
[0172] represents the value of i that minimizes , is a training loss function of the gating module, g b represents an output value of the gating module when the b-th layer of the second student model exits, is a binary cross-entropy loss function, exit b is a pseudo label of the first data;
[0173] The exit unit 530 is further configured to obtain a training loss function of the intermediate classification head in the following manner:
[0174]
[0175] is a training loss function of the intermediate classification head, Y b is a prediction vector when the b-th layer of the second student model exits, y' is a true label of the second data, is a set of insertion points, is a cross-entropy loss function, w b represents a probability distribution of the training data when the b-th layer of the second student model exits.
[0176] The embodiment of the present application can maintain high accuracy while reducing model consumption.
[0177] The embodiment of the present application also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the method as follows when executing the computer program: constructing a first student model according to a high-precision teacher model; updating a total loss function of the first student model according to the high-precision teacher model and training data, to obtain a second student model; adding a dynamic exit model to the second student model and training, to obtain an image analysis model.
[0178] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the method as follows: constructing a first student model according to a high-precision teacher model; updating a total loss function of the first student model according to the high-precision teacher model and training data, to obtain a second student model; adding a dynamic exit model to the second student model and training, to obtain an image analysis model.
[0179] The above image analysis method can solve the technical problems in the background art and achieve beneficial effects.
[0180] Figures 2 to 3 is a flowchart of the image analysis method in one embodiment. It should be understood that, although Figures 2 to 3The steps in the flowcharts are shown in sequence according to the arrows, but the steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the steps are not strictly limited in sequence, and the steps can be executed in other sequences. Moreover, Figures 2 to 3 At least part of the steps in the flowcharts can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of the sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with other steps or at least part of the sub-steps or stages of other steps.
[0181] Figure 6 An internal structure diagram of a computer device in an embodiment is shown. The computer device can be specifically a server 120 in Figure 1 As shown in Figure 6 , the computer device includes a processor, a memory, a network interface, an input device and a display screen connected through a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system, and can also store a computer program, which, when executed by the processor, can enable the processor to implement the image analysis method. The internal memory can also store a computer program, which, when executed by the processor, can enable the processor to execute the image analysis method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer overlaid on the display screen, or can be a key, trackball or touchpad arranged on the shell of the computer device, or can be an external keyboard, touchpad or mouse, etc.
[0182] Those skilled in the art can understand that Figure 6 the structure shown in the above is only a block diagram of part of the structure related to the present application scheme, and does not constitute a limitation on the computer device to which the present application scheme is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0183] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer readable storage medium, and when the program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, databases, or other media in this application refers to both non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0184] It should be noted that, in this document, the terms "first" and "second" and the like are used merely to distinguish one entity or action from another, and do not necessarily require or imply any actual such relationship or order between such entities or actions. Also, the terms "comprises", "comprising", or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements recited, but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element defined by an indefinite article "a" or "an" does not exclude the existence of additional identical elements in the process, method, article, or apparatus including the element.
[0185] The above description is only a specific implementation of the present application, which enables those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Accordingly, the present application is not limited to the embodiments shown herein, but is intended to conform to the principles and novel features set forth in the following claims.
[0186] The spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is intended to conform to the principles and novel features set forth in the following claims.
[0187] The spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is intended to conform to the principles and novel features set forth in the following claims.
[0188] the widest range of features consistent with the application.
Claims
1. An image analysis method characterized by, The method comprises: According to the high-precision teacher model, a first student model is constructed; According to the high-precision teacher model and training data, the total loss function of the first student model is updated to obtain a second student model; A dynamic exit model is added to the second student model and trained to obtain an image analysis model.
2. The method of claim 1, wherein, The first student model is a bidirectional state space model, and the first student model comprises a plurality of bidirectional state space layers, each of which is calculated in the following manner: z l = z l-1 + Vm(z l-1 ), l = 1...L Y = head(z L ) wherein z L is the intermediate feature of the last layer output of the first student model, Y is the prediction vector after the classification head processing, L is the total number of layers of the first student model, and Vim represents the bidirectional state space layer.
3. The method of claim 1, wherein, According to the high-precision teacher model and training data, the total loss function of the first student model is updated to obtain a second student model, comprising: Obtaining the self-loss function of the first student model; According to the high-precision teacher model, obtaining the distillation loss function of the first student model and the normalization loss function of the first student model; According to the self-loss function, the distillation loss function and the normalization loss function, the total loss function of the first student model is obtained; According to the total loss function, the second student model is obtained.
4. The method of claim 3, wherein, The self-loss function of the first student model is obtained in the following manner: wherein Y (s) is a prediction vector output by the first student model, y is a true label of the training data, is a cross-entropy loss function; The distillation loss function is obtained in the following manner: wherein, is a distillation loss function, is an inter-class relation loss function, is an intra-relation loss function, D is a batch size during training of the first student model, is and a Pearson distance between is the i-th prediction vector of the first student model in a batch, is the i-th prediction vector of the high-precision teacher model in a batch, C is a number of classes of the training data, is the j-th column of the prediction matrix Y s of the first student model, is the j-th column of the prediction matrix Y (t) of the high-precision teacher model. The normalization loss function is obtained in the following manner: wherein, is a normalized loss function, is an index set of training data belonging to the k-th class, C is the number of classes of the training data, is an intermediate feature of the first student model last layer output is the i-th row of the feature vector of is an intermediate feature of the teacher network last layer output is the i-th of the feature vector of EMB k is an intermediate parameter of the teacher network, e k is the EMB k is a unit vector in the direction of e.
5. The method of claim 4, wherein, According to the self-loss function, the distillation loss function and the normalization loss function, the total loss function of the first student model is obtained in the following manner: wherein, is the total loss function, is the self loss function of the first student model, is the distillation loss function, is the normalization loss function, and a and β are parameters balancing the losses.
6. The method of claim 1, wherein, The dynamic exit module is added to the second student model to obtain an image analysis network, comprising: In the second student model, a plurality of dynamic exit points are obtained, and the dynamic exit module is added to each dynamic exit point; According to the real label, the loss value of each dynamic exit module is obtained; The minimum loss value in the loss value is obtained, and the pseudo label is obtained according to the minimum loss value; The training data is evenly divided into first data and second data; After freezing all parameters in the second student model, the training loss function of the gating module is obtained according to the pseudo label and the first data to update the dynamic exit module; According to the real label and the second data, the training loss function of the intermediate classification head is obtained to update the dynamic exit module; According to the real label, the loss value of each dynamic exit module is updated, and the first data and second data are updated to update the dynamic exit module until the convergence condition is met, and the image analysis network is obtained.
7. The method of claim 6, wherein, According to the real label, the loss value of each dynamic exit module is obtained in the following manner: Lb is the loss value at the exit of the b-th layer of the second student model, IC b Yb represents the inference cost at the exit of the b-th layer of the second student model selection, Y b y is the true label of the training data, and λ is a weight parameter; The training loss function of the gating module is obtained in the following manner: denotes the value of i that minimizes is a training loss function of the gating module, g b denotes the output value of the gating module when the second student model exits at the b-th layer, is a binary cross-entropy loss function, exit b is a pseudo label of the first data; The training loss function of the intermediate classification head is obtained in the following manner: Y is a training loss function for the intermediate classification head, b y′ is a prediction vector at the exit of the b-th layer of the second student model, is a set of insertion points, w is a cross-entropy loss function, b represents a probability distribution of the training data at the exit of the b-th layer of the second student model.
8. An image analysis apparatus characterized by comprising: The device comprises: A student model unit for constructing a first student model according to a high-precision teacher model; A training unit for updating the total loss function of the first student model according to the high-precision teacher model and training data to obtain a second student model; A quitting unit is configured to add a dynamic quitting model to the second student model and train the second student model to obtain an image analysis model.
9. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program is executed by the processor to implement the method in any one of claims 1 to 7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the method in any one of claims 1 to 7.