Image recognition method and device, equipment and storage medium
Through the adaptive update of weights and coding coefficients, combined with the adaptive weighted Huber constrained sparse coding model, the problem of accurate recognition of face recognition in complex environments is solved, the recognition rate and robustness are improved, and the weight overfitting is avoided.
Patent Information
- Application Number
- CN202510041854.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, face recognition has problems with accurate recognition in complex environments, especially due to information interference caused by expression changes, lighting conditions, real disguise, continuous occlusion and pixel corrosion, resulting in weight overfitting and reduced universality.
By adaptively and alternately updating the adjustment weights and coding coefficients, maintain the consistency of coding coefficients and coding residuals, increase the difference between changes between classes and intra-class changes, and adopt an adaptively weighted Huber constrained sparse coding model to solve the weight overfitting phenomenon and improve the recognition rate.
It improves the robustness and recognition rate of face recognition in complex environments, enhances the universality and interpretability of weights, and avoids overfitting.
Smart Images

Figure CN119964217A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and in particular to an image recognition method, device, equipment and storage medium. Background Art
[0002] In intelligent driving, the accuracy and reliability of face recognition technology are still the key to practical applications. However, different facial expression changes, lighting conditions, real camouflage, continuous occlusion, pixel corrosion, etc. in the driving environment will reduce the valuable information of the face image and interfere with the recognition of the face. Face recognition technology still has the problem of accurate recognition. How to improve the robustness of face recognition in various complex and changing environments requires further research.
[0003] At present, sparse coding based on regression analysis is increasingly favored by researchers in the field of face recognition. At the same time, many researchers have noticed the relationship between different fidelity terms and the distribution of the coding residual. Most existing regression methods essentially use a single fidelity term of L1 loss or L2 loss to indicate that the coding residual conforms to the Gaussian distribution or Laplace distribution, but usually noisy samples do not follow a single distribution.
[0004] Reference patent CN108509843A "A face recognition method based on weighted Huber constrained sparse coding" proposes weighted Huber constrained sparse coding to improve the robustness of face recognition in occluded environments. Through a weighted sparse coding, the Huber loss is used to determine whether the fidelity term is L1 loss or L2 loss to optimize the intra-class and inter-class changes in face image classification. However, the weights proposed by this method use conventional logical functions. During the iteration process, the weights are prone to overfitting, and the weight calculation formula sets two constants, which reduces the versatility and needs to be solved urgently. Summary of the invention
[0005] The present invention provides an image recognition method to solve the weight overfitting phenomenon in the prior art; secondly, provides an image recognition device; secondly, provides an electronic device; thirdly, provides a computer storage medium; and finally, provides a computer program product.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] An image recognition method, the image recognition method comprising:
[0008] Obtain a query image and a set of training images of multiple categories;
[0009] Processing the query image and the training image sets of the multiple categories using an image recognition model to obtain the category of the query image;
[0010] Among them, the image recognition model is determined based on a first model, a second model and a third model, the first model is established based on the Huber loss function and weights, the coding residual in the Huber loss function is the coding residual between the query image and the target training image set, the weight is the weight set for the coding residual, the second model is established based on the weight function, and the third model is established based on the coding coefficient function.
[0011] According to the above technical means, based on the target training image sets of different categories, the weights and coding coefficients are adaptively and alternately updated and adjusted to maintain the consistency of the coding coefficients and the coding residuals, increase the difference between inter-class changes and intra-class changes, improve the recognition rate, and solve the overfitting phenomenon.
[0012] Further, the query image and the training image sets of the multiple categories are processed using the image recognition model to obtain the category of the query image, including: substituting the query image, the target training image set and the initial weight into a first relationship formula to obtain a first coding coefficient; wherein the first relationship formula is a relationship formula between the coding coefficient and the weight determined based on a fourth model, and the fourth model is determined based on the first model and the third model; substituting the query image, the target training image set and the first coding coefficient into a second relationship formula to obtain an updated first weight; wherein the second relationship formula is a relationship formula between the weight and the coding coefficient determined based on a fifth model, and the fifth model is determined based on the first model and the second model; performing an iterative process until an iteration condition is met, then determining the residual results of the query image in multiple categories based on the second coding coefficients and second weights corresponding to each of the acquired training image sets of the multiple categories; and taking the category corresponding to the smallest residual result as the category of the query image.
[0013] Furthermore, the residual results of the query image in multiple categories are determined based on the second coding coefficients and second weights corresponding to each of the acquired training image sets of the multiple categories, including: determining a first coding residual based on the query image, the target training image set and the corresponding second coding coefficient; when the first result and the second result meet a preset condition, determining the residual result of the query image in the target category based on the first coding residual corresponding to the target training image set; wherein the first result is determined based on the first coding residual, and the second result is determined based on the second weight and the Huber threshold.
[0014] Furthermore, the image recognition method also includes: when the first result and the second result do not meet the preset condition, determining the residual result of the query image in the target category based on the first encoding residual, the second weight and the Huber threshold corresponding to the target training image set.
[0015] Furthermore, the preset condition includes: the first result is less than or equal to the second result.
[0016] Furthermore, the image recognition method also includes: constructing a first Lagrangian function based on the first function and the first constraint condition corresponding to the fourth model; solving the optimal solution of the first Lagrangian function based on the alternating direction multiplier algorithm to obtain a first relationship between the coding coefficient and the weight.
[0017] Furthermore, the image recognition method also includes: constructing a second Lagrangian function based on the second function and the second constraint condition corresponding to the fifth model; solving the optimal solution of the second Lagrangian function based on a preset value to obtain a second relationship between the weight and the coding coefficient; wherein the preset value is the numerical value of the zero element included in the optimal solution.
[0018] Furthermore, the iteration condition includes one of the following:
[0019] The square of the L2 norm of the first coding coefficient corresponding to the current iteration number and the first coding coefficient corresponding to the previous iteration number is less than or equal to the first threshold;
[0020] The current number of iterations is greater than or equal to a second threshold.
[0021] Furthermore, the initial weight is one-tenth of the total number of pixels in an image.
[0022] An image recognition device, comprising:
[0023] An acquisition unit, used for acquiring a query image and a training image set of multiple categories;
[0024] A processing unit, configured to process the query image and the training image sets of the plurality of categories by using an image recognition model to obtain the category of the query image;
[0025] Among them, the image recognition model is determined based on a first model, a second model and a third model, the first model is established based on the Huber loss function and weights, the coding residual in the Huber loss function is the coding residual between the query image and the target training image set, the weight is the weight set for the coding residual, the second model is established based on the weight function, and the third model is established based on the coding coefficient function.
[0026] An electronic device comprises: a processor and a memory configured to store a computer program that can be run on the processor, wherein the processor is configured to execute the steps of the above method when running the computer program.
[0027] A computer-readable storage medium stores a computer program, which implements the steps of the above method when executed by a processor.
[0028] A computer program product comprises a computer program or instructions, wherein when the computer program or instructions are executed by a processor, the steps of the aforementioned method are implemented.
[0029] Beneficial effects of the present invention:
[0030] (1) The present invention adaptively and alternately updates and adjusts weights and coding coefficients according to target training image sets of different categories, maintains the consistency of coding coefficients and coding residuals, increases the difference between inter-class changes and intra-class changes, improves recognition rate, and solves the overfitting phenomenon;
[0031] (2) Based on the weighted Huber constrained sparse coding, the present invention introduces adaptive weights and coding coefficients, seeks the maximum likelihood estimate (MLE) of sparse coding by alternately updating the adaptive weights and coding coefficients, and obtains an adaptive weighted Huber constrained sparse coding model, so that the weights have stronger versatility and interpretability. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 FIG. 1 is a flow chart of an image recognition method according to an embodiment of the present invention. Figure 1 ;
[0033] Figure 2 It is a schematic diagram of Huber loss function;
[0034] Figure 3 (a) is the training image;
[0035] Figure 3 (b) is the query image with 50% Gaussian noise;
[0036] Figure 3 (c) is the fitted image with 50% Gaussian noise;
[0037] Figure 3 (d) is the query image with 30% black block occlusion;
[0038] Figure 3 (e) is the fitted image with 30% black block occlusion;
[0039] Figure 3 (f) is the query image with 30% white block occlusion;
[0040] Figure 3 (g) is the fitted image with 30% white block occlusion;
[0041] Figure 4 (a) is the encoding residual image of the L1 norm fidelity item with the same parameters in 40% black block occlusion;
[0042] Figure 4 (b) is the encoding residual image of the Huber loss fidelity item with the same parameters in 40% black block occlusion;
[0043] Figure 4 (c) is the encoding residual image of the L2 norm fidelity item with the same parameters in 40% black block occlusion;
[0044] Figure 5 FIG. 1 is a flow chart of an image recognition method according to an embodiment of the present invention. Figure 2 ;
[0045] Figure 6 FIG. 1 is a flow chart of an image recognition method according to an embodiment of the present invention. Figure 3 ;
[0046] Figure 7 FIG. 1 is a flow chart of an image recognition method according to an embodiment of the present invention. Figure 4 ;
[0047] Figure 8 FIG. 1 is a flow chart of an image recognition method according to an embodiment of the present invention. Figure 5 ;
[0048] Fig. 9 It is the ratio of the difference between inter-class variation and intra-class variation of ExYaleB face images to the intra-class variation;
[0049] Fig.10 A schematic diagram of the composition structure of an image recognition device in an embodiment of the present invention;
[0050] Fig.11 Schematic diagram of the composition structure of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] In order to more thoroughly understand the features and technical contents of the embodiments of the present invention, the implementation of the embodiments of the present invention is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not intended to limit the embodiments of the present invention.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein are only for the purpose of describing the present embodiment and are not intended to limit the present invention.
[0053] In the following description, references to “some embodiments,” “this embodiment,” “this embodiment,” and examples, etc., describe a subset of all possible embodiments, but it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.
[0054] If similar descriptions of "first / second" appear in the application documents, the following instructions are added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the present embodiment described here can be implemented in an order other than that illustrated or described here.
[0055] In this embodiment, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, object A and / or object B may represent three situations: object A exists alone, object A and object B exist at the same time, and object B exists alone.
[0056] The embodiment of the present invention provides an image recognition method. Figure 1 The process diagram of the image recognition method in the embodiment of the present invention is as follows Figure 1 ,like Figure 1 As shown, the image recognition method includes the following steps:
[0057] S101: Obtain a query image and a set of training images of multiple categories.
[0058] In the embodiment of the present invention, the query image is an image to be recognized. The multiple categories of training image sets include, but are not limited to: training image sets of different human face images, training image sets of different cat face images. Among them, the training image set of the same human face image includes multiple training images of the same human face.
[0059] S102: Using an image recognition model to process a query image and a plurality of categories of training image sets to obtain the category of the query image; wherein the image recognition model is determined based on a first model, a second model and a third model, the first model is established based on a Huber loss function and weights, the coding residual in the Huber loss function is the coding residual between the query image and the target training image set, the weight is a weight set for the coding residual, the second model is established based on a weight function, and the third model is established based on a coding coefficient function.
[0060] In the embodiment of the present invention, based on a training image set of multiple categories, an image recognition model is used to perform image recognition on a query image to determine which category of the multiple categories the query image belongs to.
[0061] In the embodiment of the present invention, the image recognition model can be understood as the sum of the first model, the second model and the third model.
[0062] In an embodiment of the present invention, the target training image set is a training image set of any category. The difference between the query image and the product of the target training image set and the coding coefficient is the coding residual in the Huber loss function. Specifically, the product of the pixel value of each pixel point of any target training image in the target training image set and the corresponding coding coefficient is determined, and then the average is calculated for the corresponding pixel points. The difference between the pixel value of each pixel point in the query image and the value after the average processing is the coding residual. The weight is the weight set for the coding residual as a whole.
[0063] It should be noted that the sparse coding model in the regression classifier is used as the basis, and the Huber loss function is introduced to replace the L1 fidelity term or the L2 fidelity term, and weight constraints are set for both the L1 fidelity term and the L2 fidelity term to obtain the first model, which can also be called a weighted Huber constrained sparse coding model.
[0064] Research on the human visual perception mechanism shows that in the human visual system, there is a series of cells from the retina to the cerebral cortex, which are described in the "receptive field" mode. Under the primary visual cortex, a single neuron only shows a strong response to information in a certain frequency band, and its spatial receptive field is described as a signal coding filter with local, directional and bandpass characteristics. Each neuron uses the sparse coding principle to express these stimuli, and describes the characteristics of the image in terms of edges, endpoints, stripes, etc. in the form of sparse coding. Since the l0 norm is a nondeterministic polynomial-time (NP) problem, the l1 norm is generally used as the most approximate solution to the l0 norm minimization problem. When the coding residual follows a Gaussian distribution, the sparse coding model can be expressed as:
[0065]
[0066] Among them, λ is the penalty coefficient of the l1 regular constraint term (or penalty term), y is the query image, X is the training image machine, Xθ refers to the mean of the training image set of the same category, and θ is the encoding coefficient.
[0067] When the coding residual follows the Laplace distribution, the sparse coding model can be expressed as:
[0068] minθ ‖y-Xθ‖1+λ‖α‖1 st α=θ (2)
[0069] Sparse coding models are mainly concerned with two issues: First, when the signal has noise or outliers, the fidelity term ( Or whether ‖y-Xθ‖1) can effectively describe the fidelity of the signal. Secondly, whether the regular constraint term ‖α‖1 makes the signal sparse enough.
[0070] For the first question, from the perspective of maximum a posteriori probability, using the l1 or l2 norm to define the fidelity term actually assumes that the coded residual follows a Gaussian or Laplace distribution. But in practice, a single assumption that the residual follows a certain distribution may not be very good, especially when the facial image is occluded, disguised, or corrupted. Therefore, the fidelity term that uses the l1 or l2 norm alone in the sparse coding model may not be robust enough in these cases.
[0071] From the perspective of statistical learning, the Huber loss function is a robust regression loss function. Compared with the mean square error, it is insensitive to outliers and is often used in classification problems. The Huber loss function can be expressed as:
[0072]
[0073] Where Z is the encoded residual and η is the Huber threshold.
[0074] Huber loss uses a mixture of l1 norm and l2 norm, such as Figure 2 As shown. Among them, Figure 2 The dotted broken line in the middle is the Huber loss function, the arc is the l2 loss function, and the other solid broken line is the l1 loss function.
[0075] In order to make the Huber loss smoothly connected at Z = η, a constant η is subtracted from the l1 norm 2 / 2. The Huber loss automatically matches the encoding residual through the threshold to conform to the l1 norm and l2 norm.
[0076] The sparse Huber model can improve the sparsity of the encoding coefficients in the Huber loss function. The Huber loss function corresponds to the standard form of the Alternating Direction Method of Multipliers (ADMM) model minf(θ)+g(z), which can be expressed as:
[0077] min θ g(z) st z=Xθ-y (4)
[0078] Where f(θ)=0,
[0079] Further changes can be obtained,
[0080] min θ g(z)+λ||α||1
[0081] st z=Xθ-y,α=θ (5)
[0082] Where g(z) is the Huber loss function, and z=y-Xθ constrains θ. λ||α||1 is the l1 norm penalty term, and α=θ constrains it. Within a certain range, the larger the λ value, the sparser θ.
[0083] In actual classification, we believe that noise is completely interference and should be eliminated. When comparing the query image with the training images of different categories, the noise pixels are assigned a very small non-negative weight coefficient. Therefore, the purpose of the weight is to eliminate the contribution of the noise area to the calculation of the coding residual.
[0084] In Huber regression, weight constraints are set for both the l1 fidelity term and the l2 fidelity term, expressed as:
[0085]
[0086] Among them, K is the unknown constant for smoothing the piecewise function, w=[w1,w2,…,w m ]∈R m×1 , I is a column vector of all 1s, w is the weight, m is the total number of pixels in an image, i is greater than or equal to 1 and less than or equal to m, and η is the Huber threshold. Formula (6) here is the first model mentioned above.
[0087] Based on this, Figure 3 (c) Figure 3 (e) Figure 3 (g) is the linear fitting image with weight constraints. We can observe that Figure 3 (c) Figure 3 (e) Figure 3 (g) It looks like Figure 3 (a) is a face image with valid pixels retained after noise is removed. Figure 3 (g) looks special because the weight coefficient of white block noise (i.e., white noise) is zero, making the grayscale value of these pixels zero, corresponding to the black in the [0, 255] grayscale image. Figure 3 (d) and fitted image Figure 3 (e), the weight coefficient of the black block area is zero, and any gray value of the linear fitting image in this area becomes zero, and the corresponding coding residual is also zero. Figure 3 (f) and Figure 3(g) will not be affected by the white block area.
[0088] It should be noted that the above formula (6) may be overfitting, that is, only a small part of the feature weights have values, while the other feature weights are 0. This can be transformed into the following model that does not contain the encoding residual information:
[0089]
[0090] Formula (7) here is the second model. The optimal solution of formula (7) is that all features are given the same weight And formula (7) can be regarded as the Gaussian prior of (6) when determining θ. Combining (6) and (7) we get the following objective function:
[0091]
[0092] swt T I=1,w i ≥0
[0093] After derivation, formula (8) can be transformed into:
[0094]
[0095] Here, a⊙b represents the multiplication of corresponding elements of a and b.
[0096] but It is the Huber function form of the objective function (8).
[0097] We try to select more effective training samples to fit the query samples and avoid the interference of noise samples. Therefore, the optimal coding coefficient is actually a sparse solution with sample selection function. The sparse coding of formula (9) is constructed by sparse constraints on the coding coefficients.
[0098]
[0099] w is the image weight. η is a constant threshold greater than zero, which determines whether the fidelity term uses the l1 norm or the l2 norm. There are different methods for determining the η threshold in many papers. This paper uses a threshold combined with weights, namely wη. wη makes the threshold more consistent with the distribution of training samples constrained by weight w.
[0100] Here, λ‖α‖1 (i.e., penalty term) in formula (10) is the third model. Formula (10) is the image recognition model.
[0101] When the coded residual in a [0, 255] grayscale image is zero, the image is black. Figure 4 (a) Figure 4(b) Figure 4 It can be observed in (c) that Figure 4 (b) The darkest, that is, the smallest coding residual, has the highest fit with the query sample. Therefore, Huber regression can obtain coding coefficients that better fit the query sample distribution.
[0102] In some embodiments, S102 may include the following steps S501 to S503:
[0103] S501: Substitute the query image, the target training image set and the initial weight into the first relationship formula to obtain a first coding coefficient; wherein the first relationship formula is a relationship formula between the coding coefficient and the weight determined based on the fourth model, and the fourth model is determined based on the first model and the third model.
[0104] In the embodiment of the present invention, the target training image set is a training image set of any one of multiple categories.
[0105] It should be noted that choosing a suitable initial value can make the algorithm easier to approach the optimal solution of the model. From formula (9), it can be observed that adaptive weight learning not only searches for effective pixels in the image, but also hopes to distribute weight values to effective pixels as evenly as possible. According to the optimal solution of formula (7), we set the initial weight of each pixel to
[0106] In the embodiment of the present invention, the first model is added to the third model to obtain a fourth model. A first relationship between the coding coefficient and the weight is determined based on the fourth model. Further, the query image, the target training image set and the initial weight are substituted into the first relationship to obtain a first coding coefficient for each pixel.
[0107] It should be noted that, for any type of training image set, the initial coding coefficient is the first coding coefficient.
[0108] S502: Substitute the query image, the target training image set and the first coding coefficient into the second relationship formula to obtain the updated first weight; wherein the second relationship formula is a relationship formula between the weight and the coding coefficient determined based on the fifth model, and the fifth model is determined based on the first model and the second model.
[0109] In the embodiment of the present invention, the first model is added to the second model to obtain a fifth model. A second relational expression between weights and coding coefficients is determined based on the fifth model. Further, the query image, the target training image set, and the first coding coefficient are substituted into the second relational expression to obtain a first weight of each pixel. That is, the initial weight is updated.
[0110] It should be noted that the updated initial weight, i.e., the first weight, is adaptively changed based on training image sets of different categories.
[0111] S503: Execute an iterative process until an iteration condition is met, then determine the residual results of the query image in multiple categories based on the second coding coefficients and second weights corresponding to the acquired training image sets of multiple categories; and take the category corresponding to the smallest residual result as the category of the query image.
[0112] The above execution process can be understood as the first iteration process.
[0113] Next, when the iteration condition is not met, an iteration process is performed based on S501 to S503 until the iteration condition is met.
[0114] Assuming that the iteration condition is satisfied after the Qth iteration is completed, the first coding coefficient and the first weight obtained after the Qth iteration are used as the second coding coefficient and the second weight.
[0115] Furthermore, the image recognition model determines the residual results of the query image in multiple categories based on the second coding coefficient and the second weight. The smaller the residual result, the closer the category of the query image is to the category of the corresponding residual result. Therefore, the category corresponding to the smallest residual result is taken as the category of the query image.
[0116] In the embodiment of the present invention, the adaptive weight and the coding coefficient are updated alternately. The weight is first fixed and the coding coefficient is calculated; then the coding coefficient is fixed and the weight is calculated.
[0117] In some embodiments, the image recognition method further includes the following steps S601 to S602:
[0118] S601: Constructing a first Lagrangian function based on a first function and a first constraint condition corresponding to the fourth model.
[0119] S602: Based on the alternating direction multiplier algorithm, an optimal solution of the first Lagrangian function is solved to obtain a first relationship between the coding coefficient and the weight.
[0120] In the embodiment of the present invention, the coding coefficient is solved by fixing the weight.
[0121] The fourth model can be expressed as:
[0122]
[0123] Among them, the left side of formula (11) is the first function, and the right side is the first constraint condition.
[0124] The Lagrangian expression of formula (11) is:
[0125]
[0126] ADMM is an algorithm that aims to combine the decomposability of the dual ascent method with the upper bound convergence property of the multiplier method. Since the Lagrangian expression requires the function to be strongly convex, the augmented Lagrangian expression is introduced to reduce the strong convexity constraint on the function. The augmented Lagrangian expression of formula (12) is as follows:
[0127]
[0128] Among them, formula (13) is the first Lagrangian function constructed.
[0129] Among them, ρ1 and ρ2 are greater than 0. The ADMM (alternating direction multiplier algorithm) iteration consists of:
[0130]
[0131] Where t is the number of iterations.
[0132] Substituting formula (13) into formulas (14) to (18), the iterative formula of ADMM is:
[0133]
[0134] Where u is The replacement variable of . By solving formulas (19) to (23), we can get:
[0135]
[0136] in is a diagonal matrix. The S operator is defined as:
[0137]
[0138] Next, substitute formulas (24) to (28) into formula (12) to find the optimal solution for the coding coefficient and obtain the first relationship between the coding coefficient and the weight.
[0139] In some embodiments, the image recognition method further includes the following steps S701 to S702:
[0140] S701: Construct a second Lagrangian function based on the second function and the second constraint condition corresponding to the fifth model.
[0141] S702: Based on a preset value, solve the optimal solution of the second Lagrangian function to obtain a second relationship between the weight and the coding coefficient; wherein the preset value is the value of the zero element included in the optimal solution.
[0142] In the embodiment of the present invention, the weight is solved by fixing the coding coefficient.
[0143] The default value here is the l that appears below.
[0144] The fifth model can be expressed as:
[0145]
[0146] swt T I=1,w i ≥0
[0147] Where, e = Xθ-y. The second function is the left side of formula (30), and the second constraint condition includes the right side and the bottom of formula (30). Transforming formula (30) and eliminating the constant term yields
[0148]
[0149] swt T I=1,w i ≥0
[0150] in,
[0151] The second Lagrangian function that minimizes formula (31) is
[0152]
[0153] Where κ and β ≥ 0 are Lagrange multiplier variables. And according to the Karush-Kuhn-Tucker conditions (KKT conditions), the optimal solution can be obtained.
[0154]
[0155] In order to reduce the impact of noise, it has been shown in the previous article that noisy pixels can be eliminated by learning a sparse weight vector. Therefore, the adaptive weight considers the weight coefficient of the noise pixel exceeding the threshold to be zero. m ] are sorted from small to large. Assuming that the optimal solution of w has l>0 zero elements, then according to formula (33) we can get w m-l >0 and w m-l+1 =0 (it can be understood that the weight of the first l is greater than 0, and the weight of l+1 to m = 0), that is:
[0156]
[0157] In addition, according to the constraint w T I=1 to get
[0158]
[0159] Substituting formula (35) into (34) we obtain
[0160]
[0161] Using the derived κ and γ, the optimal solution for w is
[0162]
[0163] The only parameter derived from formula (37) is l. l is more interpretable and its purpose is to determine the amount of noise in the sample. At the same time, formula (34) shows that w contains l zero elements, that is, the weight coefficient is a sparse solution.
[0164] The second relationship between the weight and the coding residual is formula (37).
[0165] In some embodiments, S504 may include the following steps S801 to S802:
[0166] S801: Determine a first coding residual based on a query image, a target training image set and corresponding second coding coefficients.
[0167] Specifically, the product of the pixel value of each pixel point of any target training image in the target training image set and the corresponding second coding coefficient is determined, and then the average of the corresponding pixel points is calculated, and the difference between the pixel value of each pixel point in the query image and the value after the average processing is obtained to obtain the first coding residual of each pixel point in the query image and the target training image set.
[0168] S802: When the first result and the second result meet a preset condition, determine the residual result of the query image in the target category based on the first coded residual corresponding to the target training image set; wherein the first result is determined based on the first coded residual, and the second result is determined based on the second weight and the Huber threshold.
[0169] In the embodiment of the present invention, the first result is determined based on the first coding residual, and the second result is determined based on the second weight and the Huber threshold. The preset condition includes: the first result is less than or equal to the second result.
[0170] That is, when the first result is less than or equal to the second result, the residual result of the query image in the target category is determined based on the first encoded residual corresponding to the target training image set.
[0171] S803: When the first result and the second result do not satisfy a preset condition, determine a residual result of the query image in the target category based on the first coded residual, the second weight and the Huber threshold corresponding to the target training image set.
[0172] Correspondingly, when the first result is greater than the second result, the residual result of the query image in the target category is determined based on the first encoded residual, the second weight and the Huber threshold corresponding to the target training image set.
[0173] Exemplarily, the encoding coefficients of the query image in each category are calculated as θ = [θ1, θ2, ..., θ i ,…,θ c ]∈R m×c , where θ i =[θ i1 ,θ i2 ,…,θ im ]∈R m×1 c is the total number of categories. Then the residual result of the query image in the i-th category is
[0174] e i =e i1 +e i2 , (38)
[0175] in, And the minimum e i The category to which it belongs is used as the category of the query image. The preset conditions include: The first result is |z|, and the second result is
[0176] In some embodiments, the iteration condition includes one of the following:
[0177] The square of the L2 norm of the first coding coefficient corresponding to the current iteration number and the first coding coefficient corresponding to the previous iteration number is less than or equal to the first threshold;
[0178] The current number of iterations is greater than or equal to a second threshold.
[0179] In the embodiment of the present invention, as the number of iterations increases, the loss value of AWHCSC gradually decreases. When the difference between θ in two iterations is small enough, AWHCSC reaches convergence and the iteration is terminated. The iteration conditions are as follows:
[0180]
[0181] Among them, ∈ is a sufficiently small positive number, namely the first threshold, and t is the number of iterations.
[0182] In the embodiment of the present invention, the second threshold is a preset number of iterations.
[0183] Based on the above embodiment, the embodiment of the present invention provides an implementation process of adaptive weighted Huber constrained sparse coding, including the following steps:
[0184] Step S1: Set the test image y, the training image set X, and initialize the weights Parameter l, threshold η, initialization is the mean of the corresponding training image subset mean(X i θ i );
[0185] Step S2: Set o∈[1,2,…,c] for distinguishing different face categories; and set the number of iterations t=1;
[0186] Step S3: Setting the initial coding residual
[0187] Step S4: Update weight coefficient;
[0188]
[0189] Step S5: Use ADMM to solve the subproblem min θ g(z)+λ‖a‖1, get the updated i-th coding coefficient;
[0190]
[0191] Step S6: Update the predicted value
[0192]
[0193] Step S7: Return to step S3 until the i-th type coding coefficient is updated;
[0194] Step S8: until all category coding coefficients are updated, output in If the iteration condition is met Or if t reaches the preset number of iterations, the training is completed; otherwise, t is updated to t+1 and the process returns to step S2 to continue.
[0195] Step S9: Using the updated θ * to classify.
[0196] The residual result of the query image in the i-th category is
[0197] e i =e i1 +e i2 (38)
[0198] in, And the minimum e i The category it belongs to is taken as the category of the query image.
[0199] The present invention proposes an adaptive weighted Huber constrained sparse coding (AWHCSC) and establishes an adaptive weighted Huber constrained sparse model. AWHCSC seeks the maximum likelihood estimation of sparse coding to solve the problem and has stronger robustness in the face of noise values (such as occlusion, corrosion, camouflage, etc.). In detail, the present invention (1) proposes an adaptive weighted Huber constrained model. The evolution process of the Huber regression model and the adaptive weighted Huber constrained model is derived, and the essence of the adaptive weight is analyzed to eliminate the contribution of the noise area to the coding residual calculation. (2) An adaptive weighted Huber constrained sparse coding model is proposed. First, the negative correlation between the adaptive weight and the coding residual is analyzed. Secondly, the adaptive weight update formula is derived, and the sparsity of the weight coefficient is analyzed. On the other hand, the alternating direction multiplier algorithm (ADMM) is used to solve the l_1 norm minimization problem in AWHCSC. (3) The coding residual calculation method of AWHCSC is proposed. Keeping the consistency of the coding coefficient and coding residual of AWHCSC and increasing the difference between inter-class changes and intra-class changes is more conducive to classification.
[0200] Fig. 9 is the ratio of the difference between the inter-class variation and the intra-class variation of the ExYaleB face image to the intra-class variation. The horizontal axis is the image category, and the vertical axis is the residual ratio. The essence of the encoding residual between the query image and the training image of the same category is the intra-class variation e s , and the essence of the encoding residual between the query image and the training images of different categories is the inter-class change e ns The residual weight can be expressed as Fig. 9 The query image is the first face image in the first category. The l1 curve represents the encoding residual e = ‖w⊙(Xθ-y)‖1, and the l2 curve represents the encoding residual We can observe that the residual proportion in AWHCSC is higher than that in both l1 and l2, which indicates that the inter-class variation in AWHCSC is more different than the intra-class variation and is more conducive to classification.
[0201] To implement the method of the embodiment of the present invention, based on the same inventive concept, the embodiment of the present invention further provides an image recognition device. Fig.10 FIG. 1 is a schematic diagram of the structure of the image recognition device in an embodiment of the present invention. Fig.10 As shown, the image recognition device 100 includes:
[0202] An acquisition unit 1001 is used to acquire a query image and a set of training images of multiple categories;
[0203] A processing unit 1002 is used to process the query image and the training image sets of the multiple categories using an image recognition model to obtain the category of the query image;
[0204] Among them, the image recognition model is determined based on a first model, a second model and a third model, the first model is established based on the Huber loss function and weights, the coding residual in the Huber loss function is the coding residual between the query image and the target training image set, the weight is the weight set for the coding residual, the second model is established based on the weight function, and the third model is established based on the coding coefficient function.
[0205] In the embodiment of the present invention, according to the target training image sets of different categories, the weights and coding coefficients are adaptively and alternately updated and adjusted to maintain the consistency of the coding coefficients and the coding residuals, increase the difference between the inter-class changes and the intra-class changes, improve the recognition rate, and solve the overfitting phenomenon.
[0206] In some embodiments, the processing unit 1002 is also used to substitute the query image, the target training image set and the initial weight into a first relationship formula to obtain a first coding coefficient; wherein the first relationship formula is a relationship formula between the coding coefficient and the weight determined based on a fourth model, and the fourth model is determined based on the first model and the third model; substitute the query image, the target training image set and the first coding coefficient into a second relationship formula to obtain an updated first weight; wherein the second relationship formula is a relationship formula between the weight and the coding coefficient determined based on a fifth model, and the fifth model is determined based on the first model and the second model; perform an iterative process until an iteration condition is met, then determine the residual results of the query image in multiple categories based on the second coding coefficients and second weights corresponding to each of the multiple categories of training image sets obtained; and take the category corresponding to the smallest residual result as the category of the query image.
[0207] In some embodiments, the processing unit 1002 is also used to determine a first coding residual based on the query image, the target training image set and the corresponding second coding coefficient; when the first result and the second result meet a preset condition, the residual result of the query image in the target category is determined based on the first coding residual corresponding to the target training image set; wherein the first result is determined based on the first coding residual, and the second result is determined based on the second weight and the Huber threshold.
[0208] In some embodiments, the processing unit 1002 is also used to determine the residual result of the query image in the target category based on the first encoded residual, the second weight and the Huber threshold corresponding to the target training image set when the first result and the second result do not meet the preset condition.
[0209] In some embodiments, the preset condition includes: the first result is less than or equal to the second result.
[0210] In some embodiments, the processing unit 1002 is further used to construct a first Lagrangian function based on the first function and the first constraint condition corresponding to the fourth model; based on the alternating direction multiplier algorithm, solve the minimum coding coefficient of the first Lagrangian function to obtain a first relationship between the coding coefficient and the weight.
[0211] In some embodiments, the processing unit 1002 is also used to construct a second Lagrangian function based on the second function and the second constraint corresponding to the fifth model; based on a preset value, solve the minimum weight value of the second Lagrangian function to obtain a second relationship between the weight and the coding coefficient; wherein the preset value is the numerical value of the zero element included in the optimal solution.
[0212] In some embodiments, the iteration condition includes one of the following:
[0213] The square of the L2 norm of the first coding coefficient corresponding to the current iteration number and the first coding coefficient corresponding to the previous iteration number is less than or equal to the first threshold;
[0214] The current number of iterations is greater than or equal to a second threshold.
[0215] In some embodiments, the initial weight is one-tenth of the total number of pixels in an image.
[0216] The embodiment of the present invention further provides another electronic device, Fig.11 FIG. 1 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Fig.11 As shown, the electronic device 110 includes: a processor 1101 and a memory 1102 configured to store a computer program that can be run on the processor;
[0217] The processor 1101 is configured to execute the method steps in the aforementioned embodiment when running a computer program.
[0218] Of course, in practical applications, Fig.11As shown, the various components in the electronic device 110 are coupled together via a bus system 1103. It is understood that the bus system 1103 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 1103 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig.11 Various buses are labeled as bus system 1103.
[0219] In practical applications, the processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processing device (DSPD), a programmable logic device (PLD), a field-programmable gate array (FPGA), a controller, a microcontroller, and a microprocessor. It is understandable that for different devices, the electronic device used to implement the functions of the processor may also be other, and the embodiments of the present invention do not specifically limit this.
[0220] The above-mentioned memory can be a volatile memory (volatile memory), such as a random access memory (RAM); or a non-volatile memory (non-volatile memory), such as a read-only memory (ROM), a flash memory, a hard disk (HDD) or a solid-state drive (SSD); or a combination of the above-mentioned types of memory, and provide instructions and data to the processor.
[0221] In an exemplary embodiment, an embodiment of the present invention further provides a computer-readable storage medium for storing a computer program.
[0222] Optionally, the computer-readable storage medium can be applied to any one of the methods in the embodiments of the present invention, and the computer program enables the computer to execute the corresponding processes implemented by the processor in each method in the embodiments of the present invention. For the sake of brevity, they are not described here.
[0223] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0224] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0225] In addition, all functional units in the embodiments of the present invention may be integrated into one processing module, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units. A person of ordinary skill in the art may understand that all or part of the steps of implementing the above method embodiments may be completed by hardware related to program instructions, and the above program may be stored in a computer-readable storage medium, which, when executed, executes the steps of the above method embodiments; and the above storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.
[0226] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.
[0227] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.
[0228] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0229] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. An image recognition method, characterized in that: The image recognition method comprises: Obtain a query image and a set of training images of multiple categories; Processing the query image and the training image sets of the multiple categories using an image recognition model to obtain the category of the query image; Among them, the image recognition model is determined based on a first model, a second model and a third model, the first model is established based on the Huber loss function and weights, the coding residual in the Huber loss function is the coding residual between the query image and the target training image set, the weight is the weight set for the coding residual, the second model is established based on the weight function, and the third model is established based on the coding coefficient function.
2. The image recognition method according to claim 1, characterized in that: The using the image recognition model to process the query image and the training image sets of the multiple categories to obtain the category of the query image includes: Substituting the query image, the target training image set and the initial weight into a first relational expression to obtain a first coding coefficient; wherein the first relational expression is a relational expression between the coding coefficient and the weight determined based on a fourth model, and the fourth model is determined based on the first model and the third model; Substituting the query image, the target training image set and the first coding coefficient into a second relational expression to obtain an updated first weight; wherein the second relational expression is a relational expression between a weight and a coding coefficient determined based on a fifth model, and the fifth model is determined based on the first model and the second model; An iterative process is performed until an iterative condition is met, and then residual results of the query image in multiple categories are determined based on the second coding coefficients and second weights corresponding to each of the multiple categories of training image sets obtained; and the category corresponding to the smallest residual result is used as the category of the query image.
3. The image recognition method according to claim 2, characterized in that: The determining of the residual results of the query image in multiple categories based on the second coding coefficients and second weights corresponding to the respective acquired training image sets of the multiple categories includes: Determining a first coding residual based on the query image, the target training image set, and corresponding second coding coefficients; In the case where the first result and the second result meet a preset condition, determining a residual result of the query image in the target category based on a first encoded residual corresponding to the target training image set; The first result is determined based on the first coding residual, and the second result is determined based on the second weight and the Huber threshold.
4. The image recognition method according to claim 3, characterized in that: The image recognition method further comprises: When the first result and the second result do not satisfy the preset condition, a residual result of the query image in the target category is determined based on the first encoding residual corresponding to the target training image set, the second weight and the Huber threshold.
5. The image recognition method according to claim 3 or 4, characterized in that: The preset condition includes: the first result is less than or equal to the second result.
6. The image recognition method according to claim 2, characterized in that: The image recognition method further comprises: Constructing a first Lagrangian function based on the first function and the first constraint condition corresponding to the fourth model; Based on the alternating direction multiplier algorithm, the optimal solution of the first Lagrangian function is solved to obtain a first relationship between the coding coefficient and the weight.
7. The image recognition method according to claim 2, characterized in that: The image recognition method further comprises: constructing a second Lagrangian function based on the second function and the second constraint condition corresponding to the fifth model; Based on a preset value, an optimal solution of the second Lagrangian function is solved to obtain a second relationship between the weight and the coding coefficient; wherein the preset value is the numerical value of the zero element included in the optimal solution.
8. The image recognition method according to claim 2, characterized in that: The iteration condition includes one of the following: The square of the L2 norm of the first coding coefficient corresponding to the current iteration number and the first coding coefficient corresponding to the previous iteration number is less than or equal to the first threshold; The current number of iterations is greater than or equal to a second threshold.
9. The image recognition method according to claim 2, characterized in that: The initial weight is one-tenth of the total number of pixels in an image.
10. An image recognition device, characterized in that: The image recognition device comprises: An acquisition unit, used for acquiring a query image and a training image set of multiple categories; A processing unit, configured to process the query image and the training image sets of the plurality of categories by using an image recognition model to obtain the category of the query image; Among them, the image recognition model is determined based on a first model, a second model and a third model, the first model is established based on the Huber loss function and weights, the coding residual in the Huber loss function is the coding residual between the query image and the target training image set, the weight is the weight set for the coding residual, the second model is established based on the weight function, and the third model is established based on the coding coefficient function.
11. An electronic device, characterized in that: The electronic device comprises: a processor and a memory configured to store a computer program capable of running on the processor, Wherein, the processor is configured to execute the steps of the image recognition method according to any one of claims 1 to 9 when running the computer program.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image recognition method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Weighted Huber constraint sparse coding-based face recognition method
CN108509843A