Robust classification method and device for multi-view binary ordered image data

By designing a multi-view binary ordered data classification method in image recognition, using robust loss function and viewing weight for deep fusion, the balance problem of overfitting and empirical risks in multi-view image data is solved, and the robustness and classification accuracy of image recognition are significantly improved.

CN120047716AActive Publication Date: 2025-05-27NANJING UNIV OF INFORMATION SCI & TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411963412.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

In the field of image recognition, how to effectively handle overfitting and empirical risks in multi-view binary ordered image data, improve the robustness of image recognition and meet practical application needs.

Method used

A multi-view binary ordered data classification method is designed. By extracting multi-view visual features and constructing a robust loss function, integrating feature distributions from different perspectives, using viewing angle weights and mapping matrices for deep fusion, avoiding overfitting and improving model generalization performance.

Benefits of technology

It significantly improves the classification accuracy of multi-view binary ordered image data, achieves a balance between minimizing empirical risks and avoiding overfitting, and has broad application prospects in the fields of computer vision, machine learning and pattern recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047716A_ABST
    Figure CN120047716A_ABST
Patent Text Reader

Abstract

The invention discloses a robust classification method and device for multi-view binary ordered image data. The method comprises the following steps: extracting multi-view visual features of training set binary ordered image data; the extracted features are sent into a multi-view binary ordered data classification model, a weight coefficient and a mapping matrix on each view are obtained through training, and images to be recognized are classified according to the weight coefficient and the mapping matrix; wherein the binary ordered data classification model applies different punishment to a binary ordered data sample with positive loss and a binary ordered data sample with negative loss, and adaptively learns a weight coefficient between visual angles by using consistency and complementarity information in image multi-visual-angle visual features. According to the method, from the microscopic perspective, the balance between empirical risk minimization and overfitting avoidance is well realized by mining the connotation of unbiased risk estimation of binary ordered data classification, the multi-view visual information of the image data is fully utilized, and the classification precision of the multi-view binary ordered image data is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and more particularly to a robust classification method and device for multi-view binary ordered image data. Background Art

[0002] In recent years, supervised learning methods based on accurately labeled samples have made significant progress in the field of pattern recognition. However, in many practical scenarios, accurately labeling data often requires a large amount of time and labor costs. In addition, directly labeling private or confidential data may have legal or ethical restrictions. Against this background, incomplete label learning has attracted much attention due to its lower labeling cost and broad application potential. Binary ordered data learning is an emerging incomplete label learning method, and its training samples are composed of feature pairs, in which the positive likelihood of the former in each pair of features is higher than that of the latter. Compared with traditional accurate labeling, obtaining the sorting relationship between data is usually more economical and practical. For binary ordered image data, how to properly handle the relationship between overfitting and empirical risk and effectively utilize the multi-view visual features of images remains an urgent problem to be solved. Therefore, it is necessary to design a new type of multi-view binary ordered data classifier to more accurately mine and utilize the ordered relationship between multi-view features, so as to improve the robustness of image recognition and meet the actual application requirements. Summary of the Invention

[0003] Object of the Invention: The object of the present invention is to provide a robust classification method and device for multi-view binary ordered image data, which uses the multi-view features of image samples to train a binary ordered data classifier, so as to improve the generalization ability of the classifier and greatly improve the classification accuracy of multi-view binary ordered image data.

[0004] Technical Solution: In a first aspect, the present invention provides a robust classification method for multi-view binary ordered image data, including the following steps:

[0005] Extract the multi-view visual features of the binary ordered image data in the training set;

[0006] Send the extracted features into a multi-view binary ordered data classification model, and train to obtain the weight coefficients and mapping matrices on each view. The binary ordered data classification model is expressed as the following optimization problem:

[0007] Problem P1:

[0008]

[0009] Constraints:

[0010] γ (v) ≥0

[0011]

[0012]

[0013]

[0014] In the formula, represents the visual feature of the i-th binary ordered image sample at the v-th perspective, where and are the visual feature vectors extracted from the first image and the second image respectively, and the first image has a greater positive likelihood than the second image; n represents the number of binary ordered image samples used for training; m represents the number of different perspectives; φ(·) represents the high-dimensional feature mapping function of each perspective; π + and π- are the probabilities representing the positive class and the negative class respectively; γ(v) represents the weight of the v-th perspective; w(v) is expressed as is the feature matrix, is the mapping matrix, and γ (v) are variables to be optimized; and are intermediate variables, a>0, b>0, c>0, d>0 and λ>0 are hyperparameters;

[0015] Based on the learned weight coefficients and mapping matrix, the image to be recognized is classified and recognized.

[0016] In a second aspect, the present invention further provides a robust classification device for multi-perspective binary ordered image data, including:

[0017] A data preprocessing module for extracting multi-perspective visual features of the binary ordered image data in the training set;

[0018] A model training module for sending the extracted features into a multi-perspective binary ordered data classification model to train the weight coefficients and mapping matrixes at each perspective. The binary ordered data classification model is expressed as the following optimization problem:

[0019] Problem P1:

[0020]

[0021] Constraints:

[0022] γ (v) ≥0

[0023]

[0024]

[0025] In the formula, denotes the visual feature of the \(i\)-th binary ordered image sample from the \(v\)-th perspective, where and are the visual feature vectors extracted from the first image and the second image respectively, and the first image has a greater positive likelihood than the second image; \(n\) represents the number of binary ordered image samples for training; \(m\) represents the number of different perspectives; \(\varphi(\cdot)\) represents the high-dimensional feature mapping function for each perspective; \(\pi\) + and \(\pi\) - are the probabilities representing the positive class and the negative class respectively; \(\gamma\) (v) represents the weight of the \(v\)-th perspective; \(w\) (v) is represented as is the feature matrix, is the mapping matrix, and \(\gamma\) (v) are variables to be optimized; and are intermediate variables, \(a > 0\), \(b > 0\), \(c > 0\), \(d > 0\) and \(\lambda>0\) are hyperparameters;

[0026] An image classification module, configured to classify and identify the image to be recognized based on the learned weight coefficients and the mapping matrix.

[0027] In a third aspect, the present invention further provides a computer device, including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and when the program is executed by the processor, the steps of the robust classification method for multi-view binary ordered image data as described in the first aspect of the present invention are implemented.

[0028] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the robust classification method for multi-view binary ordered image data as described in the first aspect of the present invention are implemented.

[0029] Advantageous Effects: The present invention proposes a robust classification method and device for multi-view binary ordered image data, designs a binary ordered data classification model based on a robust loss function, ensures the classification accuracy of the model on the training set by imposing a larger penalty on binary ordered data samples with positive losses, and imposes a smaller penalty on binary ordered data samples with negative losses, thereby effectively avoiding overfitting; in addition, by utilizing the consistency and complementarity information in the multi-view visual features of images, the model adaptively learns the weight coefficients between perspectives to achieve deep fusion between multi-view mapping models, effectively improving the generalization performance of the model. From a microscopic perspective, the present invention well balances empirical risk minimization and overfitting avoidance by exploring the connotation of unbiased risk estimation for binary ordered data classification, and fully utilizes the multi-view visual information of image data to greatly improve the classification accuracy of multi-view binary ordered image data. This method is simple and efficient, showing broad application prospects in the fields of computer vision, machine learning, pattern recognition, and data mining. Brief Description of the Drawings

[0030] Figure 1 It is a flowchart of the robust classification method for multi-view binary ordered image data according to an embodiment of the present invention;

[0031] Figure 2 It is a flowchart of solving the optimization problem by the alternating direction method of multipliers according to an embodiment of the present invention. Detailed Embodiments

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings.

[0033] The present invention proposes a multi-view framework for learning binary ordered data, and its core idea is to construct a robust loss function suitable for binary ordered data classification, integrate the feature distributions of different perspectives, and guide by the principles of consistency and complementarity to achieve a more efficient classification learning task. As Figure 1 shown, a robust classification method for multi-view binary ordered image data of the present invention includes the following steps:

[0034] Step S1, extract the multi-view visual features of the binary ordered image data in the training set.

[0035] A binary ordered image refers to taking two images as a sample, which is the meaning of "binary"; and the first image in the two images has a greater positive likelihood than the second image, which is the meaning of "ordered". That is to say, the features extracted from a binary ordered image contain the visual information extracted from the two images respectively. For the extraction of the visual features of an image sample, it is extracted from multiple perspectives. In the embodiments of the present invention, for black and white image data, 512-dimensional Histogram of Oriented Gradients (HOG) and 81-dimensional GIST descriptor are used as two different feature perspectives; for color image data, 144-dimensional HOG and 100-dimensional DenseHue descriptor are used as two feature perspectives. These features construct a multi-perspective representation based on specific visual features in each dataset, which helps to comprehensively mine the local and global feature information of the data.

[0036] Step S2: Feed the multi-perspective features extracted in S1 into the established multi-perspective binary ordered data classification model to train the weight coefficients and mapping matrices on each perspective.

[0037] The binary ordered data classification model in the present invention is expressed as the following optimization problem:

[0038] Problem P1:

[0039]

[0040] Constraints:

[0041] (1.1)γ (v) ≥0

[0042] (1.2)

[0043] (1.3)

[0044] (1.4)

[0045] In fact, a binary ordered image sample contains the visual information extracted from two images respectively. Let represent the visual feature of the i-th binary ordered image sample on the v-th perspective, where and are the visual feature vectors extracted from the first image and the second image respectively, and the first image has a greater positive likelihood than the second image; n represents the number of binary ordered samples used for training; m represents the number of different perspectives; φ(·) represents the high-dimensional feature mapping of each perspective; π + and π - are the probabilities representing the positive class and the negative class respectively; γ (v) represents the weight of the v-th perspective; Let w (v) be expressed as

[0046] is the feature matrix, is the mapping matrix, and γ (v) are variables to be optimized; and are intermediate variables, where a > 0, b > 0, c > 0, d > 0, and λ > 0 are hyperparameters.

[0047] Constraint (1.1) is a non-negativity constraint, ensuring that each perspective can provide a positive contribution to the final result; Constraint (1.2) is a normalization constraint, reasonably maintaining the balance between perspectives during the optimization process and avoiding the weight of a certain perspective being too large and ignoring other perspectives; Constraints (1.3)-(1.4) are soft constraint conditions, exerting an indirect influence through the constraint form of the optimization problem. The input variables of Problem P1 are The variables to be solved are w (v) and γ (v) . The present invention constructs a loss mechanism applicable to binary ordered data classification through the above constraint conditions, establishes collaborative constraints of different perspectives on the robust loss function, and promotes their consistency and complementarity.

[0048] After establishing the optimization problem P1, according to the representer theorem, we have:

[0049] w = φ(x)α + φ(x′)α′

[0050] α and α′ represent the mapping matrices corresponding to the binary ordered image samples;

[0051] Define the block vector:

[0052]

[0053] Define the block feature matrix:

[0054]

[0055] Map the extended input matrix to the feature space to obtain:

[0056]

[0057] Thus, define the block form of the kernel matrix as:

[0058]

[0059] where the kernel function k(·) is the inner product of the high-dimensional feature mapping function φ(·), and K(v), K′(v), K′′(v) are respectively composed of the kernel function and The composed kernel matrix.

[0060] Therefore, through the representer theorem, it can be further deduced that:

[0061]

[0062] According to the representer theorem, the kernel model is transformed.

[0063] After establishing the optimization problem P1, according to the representer theorem, the model transformed from the binary ordered data classification model is as follows:

[0064] Problem P2:

[0065]

[0066] Constraints:

[0067] γ (v) ≥0

[0068]

[0069] Among them, is the block form of the kernel matrix on the v-th perspective; γ(v) represents the weight of the v-th perspective; and γ(v) are variables to be optimized; and are intermediate variables, a>0, b>0, c>0, d>0 and λ>0 are hyperparameters, and v and μ represent different perspectives.

[0070] For problem P2, in order to quickly find the optimal solution, the present invention adopts the alternating direction multiplier method instead of the conventional method for solving quadratic programming problems, combines the momentum strategy related to gradient descent and the number of iterations, and gradually optimizes the objective function to achieve fast convergence. The algorithm inputs the number of perspectives m, the positive class sample probability π + , the negative class sample probability π-, the training set data parameters a, b, c, d and λ, and the step size parameter μ (i.e., the penalty parameter) of the alternating direction multiplier method. Referring to Figure 2 , the solution process of problem P2 specifically includes:

[0071] a) Initialization Let the iteration number l = 0, and determine the convergence threshold ε;

[0072] b) In problem P2, take Use the alternating direction multiplier method to solve the optimization problem regarding γ (v) to obtain

[0073] c) In problem P2, take Use the gradient descent method based on the momentum strategy to seek The optimal solution, denoted as

[0074] d) If ||γ (l+1) -γ (l) || > ε or Let l = l + 1, and return to step b); otherwise, output the optimal solution where

[0075] where, in step b), the method for updating γ (v) is as follows:

[0076] Define a vector representing the complexity of each perspective classifier model:

[0077] Define the perspective weight vector:

[0078]

[0079] For γ (v) , problem P2 can be transformed into problem P3:

[0080] Problem P3:

[0081]

[0082] Constraints:

[0083] τ ≥ 0

[0084] where p is the Lagrange multiplier vector, F = (1 T , I m ) (m+1)×m , H = (0 T , -I m×m ) (m+1)×m , 1 is the vector with all components being 1, I m is the m×m identity matrix, 0 is the vector with all components being 0, and μ is the penalty parameter of the alternating direction multiplier method.

[0085] Solve problem P3 through the following steps:

[0086] i) Initialize τ and p, and determine the convergence threshold;

[0087] ii) Update γ through the following formula:

[0088]

[0089] where I m is the m×m identity matrix.

[0090] iii) Update τ through the following formula:

[0091]

[0092] τ = max(τ, 0)

[0093] iv) Update p by the following formula:

[0094] p = p + μ(Fγ + Hτ - d)

[0095] v) Dynamically adjust the parameter μ:

[0096] μ = μ * β, where β is an adjustment coefficient greater than 1. In the embodiment, β takes the value of 1.05.

[0097] vi) If ||Fγ + Hτ - d|| 2 and the difference in the changes of γ and τ between two adjacent iterations is less than the set threshold, then obtain γ (l+1) = γ; otherwise, return to step ii).

[0098] The update in the above step c) is as follows: The method is as follows:

[0099] For Problem P2 can be transformed into problem P4:

[0100] Problem P4:

[0101]

[0102] Where:

[0103]

[0104] Solve problem P4 through the following steps:

[0105] i) Initialize q(v) = 0, the number of iterations s = 0, and determine the convergence threshold;

[0106] ii) Calculate the gradient of the objective function of problem P4 :

[0107]

[0108] Where:

[0109]

[0110]

[0111]

[0112]

[0113] iii) Dynamically modify the step size 1 / L according to the gradient norm (v) , to adapt to the change speed of the objective function, where:

[0114]

[0115] D = abcn

[0116] iv) Calculate the cumulative gradient:

[0117]

[0118] v) Update through the following formula

[0119]

[0120] vi) If the difference in changes between two adjacent iterations is less than the set threshold, then obtain Otherwise, let s = s + 1 and return to step ii).

[0121] Step S3, classify the image to be recognized through the learned weight coefficient γ* and the mapping matrix :

[0122] Classify using the classifier of the v-th perspective, where represents the kernel function:

[0123]

[0124] Weighted sum the classification results f(v) of each perspective according to the weight to obtain the final prediction result:

[0125]

[0126] Among them, is the visual feature of the image to be recognized from the v-th perspective.

[0127] In order to verify the effect and performance of the recognition method proposed in the present invention, a comparative experiment was carried out. In the experiment, the black and white image datasets MNIST, Kuzushiji, Fashion and the color image dataset CIFAR-10 were used. For the black and white image data, 512-dimensional Histogram of Oriented Gradients (HOG) and 81-dimensional GIST descriptor were used as two different feature perspectives; for the color image data, 144-dimensional HOG and 100-dimensional DenseHue descriptor were used as two feature perspectives. These features constructed a multi-perspective representation based on specific visual features in each dataset, which helped to comprehensively mine the local and global feature information of the data.

[0128] Table 1-4 shows the experimental results of five binary ordered data classification methods under different π + settings on different image datasets. The data in the table are the average classification accuracy and its standard deviation obtained by running each method ten times. It can be seen from the experimental results that, compared with Methods 1-4, the method of the present invention has the highest accuracy on all datasets (the last column), exceeding the second-best method (in bold) by up to 17% (MNIST dataset π + = 0.2). This fully demonstrates the excellent performance of the method of the present invention in aspects such as balancing overfitting and empirical risk, and fusing multi-view visual features of images. In addition, the standard deviation of the accuracy of the method of the present invention is generally about one order of magnitude lower than that of other methods, showing high training stability.

[0129] Table 1 Comparison of recognition results of five methods on the MNIST dataset under different π + settings

[0130]

[0131]

[0132] Table 2 Comparison of recognition results of five methods on the Kuzushiji dataset under different π + settings

[0133] Class_prior Pcomp-ABS Pcomp-ReLU Pcomp-Unbiased Pcomp-Teacher The method of the present invention <![CDATA[π + = 0.2]]> 0.8101±0.0096 0.8121±0.0105 0.8103±0.0138 0.7557±0.0936 0.862±0.0055 <![CDATA[π + = 0.5]]> 0.8073±0.0185 0.806±0.0119 0.8196±0.0137 0.5972±0.2304 0.8443±0.0016 <![CDATA[π + = 0.8]]> 0.8121±0.0087 0.8074±0.0123 0.8064±0.005 0.778±0.1128 0.8855±0.0058

[0134] Table 3 Comparison of recognition results of five methods on the Fashion dataset under different π + settings

[0135] Class_prior Pcomp-ABS Pcomp-ReLU Pcomp-Unbiased Pcomp-Teacher The method of the present invention <![CDATA[π + = 0.2]]> 0.8029±0.0035 0.8039±0.0093 0.8018±0.0035 0.8046±0.2485 0.9462±0.0148 <![CDATA[π + = 0.5]]> 0.9417±0.0278 0.9376±0.0193 0.9517±0.0143 0.704±0.3953 0.9662±0.001 <![CDATA[π + = 0.8]]> 0.8078±0.0092 0.8102±0.0113 0.806±0.0105 0.7914±0.2135 0.9693±0.0016

[0136] Table 4 Comparison of recognition results of five methods on the CIFAR-10 dataset under different π + settings

[0137] Class_prior Pcomp-ABS Pcomp-ReLU Pcomp-Unbiased Pcomp-Teacher The method of the present invention <![CDATA[π + = 0.2]]> 0.8005±0.0008 0.8016±0.004 0.8008±0.0014 0.6918±0.1311 0.8803±0.0008 <![CDATA[π + = 0.5]]> 0.8159±0.0183 0.8147±0.0117 0.8192±0.0095 0.6349±0.1779 0.8375±0.0008 <![CDATA[π + = 0.8]]> 0.8017±0.0038 0.8012±0.0027 0.8024±0.0038 0.7134±0.1156 0.848±0.0121

[0138] Based on the same technical concept as the method embodiment, the present invention also provides a robust classification device for multi-view binary ordered image data, including:

[0139] A data preprocessing module for extracting multi-view visual features of the binary ordered image data in the training set;

[0140] A model training module for sending the extracted features into a multi-view binary ordered data classification model to train the weight coefficients and mapping matrices on each view. The binary ordered data classification model is expressed as the following optimization problem:

[0141] Problem P1:

[0142]

[0143] Constraints:

[0144] γ (v) ≥0

[0145]

[0146]

[0147] Wherein, represents the visual feature of the i-th binary ordered image sample at the v-th perspective, where and are the visual feature vectors extracted from the first image and the second image respectively, and the first image has a greater positive likelihood than the second image; n represents the number of binary ordered image samples used for training; m represents the number of different perspectives; φ(·) represents the high-dimensional feature mapping function of each perspective; π + and π - are the probabilities representing the positive class and the negative class respectively; γ (v) represents the weight of the v-th perspective; Let w (v) be expressed as is the feature matrix, is the mapping matrix, and γ (v) are variables to be optimized; and are intermediate variables, a>0, b>0, c>0, d>0 and λ>0 are hyperparameters;

[0148] An image classification module, configured to classify and recognize the image to be recognized based on the learned weight coefficients and mapping matrix.

[0149] It should be understood that the robust classification device for multi-view binary ordered image data in the embodiments of the present invention can implement all the technical solutions in the above method embodiments, and the functions of its respective functional modules can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can refer to the relevant descriptions in the above embodiments and will not be elaborated here.

[0150] The present invention also provides a computer device, including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and when the program is executed by the processor, it implements the steps of the above-mentioned robust classification method for multi-view binary ordered image data.

[0151] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the robust classification method for multi-view binary ordered image data as described above are implemented.

[0152] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0153] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0154] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

Claims

1. A robust classification method for multi-view binary ordered image data, characterized in that: The following steps are involved: Extract multi-view visual features of binary ordered image data of training set; The extracted features are sent to the multi-view binary ordered data classification model to train the weight coefficients and mapping matrices at each view. The binary ordered data classification model is expressed as the following optimization problem: Question P1: Constraints: c (v) ≥0 In the formula, represents the visual features of the i-th binary ordered image sample at the v-th viewing angle, where and are the visual feature vectors extracted from the first image and the second image, respectively, and the first image has a greater positive likelihood than the second image; n represents the number of binary ordered image samples used for training; m represents the number of different viewing angles; φ(·) represents the high-dimensional feature mapping function for each viewing angle; π + and π - They represent the probabilities of positive and negative classes respectively; γ (v) represents the weight of the vth perspective; w (v) Expressed as is the feature matrix, is the mapping matrix, and γ (v) is the variable to be optimized; and are intermediate variables, a>0, b>0, c>0, d>0 and λ>0 are hyperparameters; Based on the learned weight coefficients and mapping matrix, the image to be identified is classified and identified.

2. The method according to claim 1, characterized in that: Extract multi-view visual features of binary ordered image data in the training set, including: For black and white image data, 512-dimensional histogram of oriented gradients and 81-dimensional GIST descriptors are used as visual features of two different perspectives; for color image data, 144-dimensional HOG and 100-dimensional DenseHue descriptors are used as visual features of two perspectives.

3. The method according to claim 1, characterized in that The weight coefficients and mapping matrices at each perspective are obtained through training, including: According to the representation theorem, the problem P1 of the binary ordered data classification model is converted into a kernel model as follows: Question P2: Constraints: c (v) ≥0 in, is the block form of the kernel matrix at the vth viewing angle. The block form of the kernel matrix is The kernel function k(·) is the inner product of the high-dimensional feature mapping function φ(·), K (v) , K′ (v) , K″ (v) The kernel functions are and The nuclear matrix composed of and γ (v) is the variable to be optimized; For problem P2, the alternating direction multiplier method is used, combined with gradient descent and momentum strategy related to the number of iterations, to gradually optimize the objective function to achieve fast convergence.

4. The method according to claim 3, characterized in that The solution to problem P2 includes the following steps: S201, input the number of viewing angles m, the probability of positive samples π + , the probability of negative class samples π - , training set data Parameters a, b, c, d and λ, step size parameter μ and convergence threshold ε of the alternating direction multiplier method; initialization Let the number of iterations l = 0; S202, in question P2 Using the alternating direction multiplier method, we can solve the problem about γ (v) The optimization problem is S203, in question P2 Using the gradient descent method based on the momentum strategy to find The optimal solution of S204, if ||γ (l+1) -γ (l) ||>ε or Let l = l + 1, return to step S202; otherwise, output the optimal solution in 5. The method according to claim 4, characterized in that In step S202, update γ (v) The method is as follows: Define a vector representing the complexity of the classifier model for each view Define the view weight vector For γ (v) For example, problem P2 is transformed into problem P3: Constraints: τ ≥ 0 Where p is the Lagrange multiplier vector, F = (1 T , I m ) (m+1)×m , H=(0 T , -I m×m ) (m+1)×m , 1 is a vector whose components are all 1, I m is the m×m identity matrix, and 0 is a vector whose components are all 0; Solve problem P3 by following the steps below: i) Initialize τ and p and determine the convergence threshold; ii) Update γ by the following formula: iii) Update τ by the following formula: iv) Update p by the following formula: p = p + μ (Fγ + Hτ - d); v) Dynamic adjustment parameter μ: μ = μ*β, β is an adjustment coefficient greater than 1; vi) If ||Fγ+Hτ-d||2 and the difference between the changes of γ and τ in two adjacent iterations are less than the set threshold, then γ is obtained. (l+1) =γ; otherwise, return to step ii).

6. The method according to claim 4, characterized in that In step S203, update The method is as follows: for For example, problem P2 is transformed into problem P4: in: Solve problem P4 by following the steps below: i) Initialization q (v) =0, iteration number s=0, determine the convergence threshold; ii) Calculate the objective function of problem P4 The gradient is: in: iii) Dynamically modify the step size 1 / L according to the gradient norm (v) , to adapt to the changing speed of the objective function, where: D=abcn iv) Calculate the accumulated gradient: v) Update by vi) If When the difference between the changes in two adjacent iterations is less than the set threshold, we get Otherwise, set s=s+1 and return to step ii).

7. The method according to claim 1, characterized in that Based on the learned weight coefficients and mapping matrix, the image to be identified is classified and identified, including: Use the classifier of the vth perspective for classification, where Represents the kernel function: The classification results of each view are f (v) By weight γ (v)* The weighted summation gives the final prediction result: in, It is the visual feature of the image to be identified at the v viewing angle.

8. A robust classification device for multi-view binary ordered image data, characterized in that: include: A data preprocessing module is used to extract multi-view visual features of binary ordered image data in the training set; The model training module is used to send the extracted features into the multi-view binary ordered data classification model to obtain the weight coefficients and mapping matrices at each view. The binary ordered data classification model is expressed as the following optimization problem: Question P1: Constraints: c (v) ≥0 In the formula, represents the visual features of the i-th binary ordered image sample at the v-th viewing angle, where and are the visual feature vectors extracted from the first image and the second image, respectively, and the first image has a greater positive likelihood than the second image; n represents the number of binary ordered image samples used for training; m represents the number of different viewing angles; φ(·) represents the high-dimensional feature mapping function for each viewing angle; π + and π - They represent the probabilities of positive and negative classes respectively; γ (v) represents the weight of the vth perspective; w (v) Expressed as is the feature matrix, is the mapping matrix, and γ (v) is the variable to be optimized; and are intermediate variables, a>0, b>0, c>0, d>0 and λ>0 are hyperparameters; The image classification module is used to classify and identify the image to be identified based on the learned weight coefficients and mapping matrix.

9. A computer device, characterized in that: include: one or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and when the programs are executed by the processors, the steps of the robust classification method for multi-view binary ordered image data as described in any one of claims 1-7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the robust classification method for multi-view binary ordered image data as described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Secure item identification and authentication system and method based on unclonable features

    CN102037676A

  • Airborne laser radar point cloud semantic annotation method under multi-view feature joint learning

    CN111275077A

  • Picture identification method and device based on multi-view comparison confidence

    CN117237748A

  • Picture classification method and system based on multi-view complementary labels

    CN117274726A

  • Multi-view feature fusion image classification method and system based on tensor learning

    CN117671333A