An artificial intelligence visual inspection method based on linear discrimination

By adopting deep convolutional neural networks, graph neural networks and dynamic topological adaptive linear discriminant analysis models in artificial intelligence vision detection, combining collaborative reinforcement learning and deep reinforcement learning, dynamically adjusting feature relationship graphs and optimizing model parameters, the problem of redundancy in feature expression and poor model adaptability is solved, and higher robustness and target recognition accuracy are achieved.

CN118968163BActive Publication Date: 2025-05-13贯文信息技术(苏州)有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411029861.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-05-13
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

The prior art ignores the dynamic correlation and weight adjustment between features in artificial intelligence visual detection, resulting in inaccurate feature expression redundancy and importance evaluation, and lacks adaptability to dynamically changing visual data, reducing the robustness and adaptability of the model.

Method used

Using an artificial intelligence vision detection method based on linear discrimination, advanced features are extracted through deep convolutional neural networks, and the graph neural network dynamically adjusts node connections and their weights, constructs a dynamic feature relationship diagram, and uses a dynamic topological adaptive linear discriminant analysis model to perform feature dimensionality reduction. Combining the collaborative reinforcement learning framework and deep reinforcement learning algorithm, we optimize the dynamic topological adaptive linear discriminant analysis model to improve visual detection performance.

Benefits of technology

It significantly enhances the feature processing capabilities in complex visual scenarios, improves the robustness of the detection system and target recognition accuracy, and improves the performance and accuracy in the face of variable visual environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968163B_ABST
    Figure CN118968163B_ABST
Patent Text Reader

Abstract

The present invention discloses an artificial intelligence visual detection method based on linear discriminant, which relates to the field of artificial intelligence visual detection technology, including receiving raw visual data, and preprocessing, using a deep convolutional neural network to extract high-level features in the raw visual data, and generating a feature matrix; according to the feature matrix, dynamically adjusting node connections and their weights through a graph neural network, and constructing a dynamic feature relationship graph; applying optimized dynamic topology adaptive linear discriminant analysis model parameters to perform feature dimension reduction and classification detection processing on real-time visual data, and realizing efficient recognition and classification of targets through a classifier. The present invention dynamically optimizes feature extraction and dimension reduction by integrating deep learning and graph neural network technology, and continuously improves performance with the help of reinforcement learning, thereby realizing efficient recognition and classification of visual data and enhancing the accuracy and adaptability of detection in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence visual detection, in particular to an artificial intelligence visual detection method based on linear discrimination. Background Art

[0002] The rapid development of artificial intelligence technology has greatly promoted the innovation in the field of visual inspection. As one of the core applications of machine vision, visual inspection is widely used in industrial manufacturing, autonomous driving, medical diagnosis and other fields. Its accuracy and efficiency are directly related to the reliability and response speed of the system. With the rise of deep learning, especially convolutional neural networks, computer vision systems have made significant progress in image recognition and target detection, and can process more complex image data and extract high-order features. In this context, linear discriminant analysis, as a classic data dimensionality reduction and classification technology, has also begun to be combined with deep learning, in order to maintain the expressiveness of features while enhancing the generalization ability and computational efficiency of the model.

[0003] In the field of artificial intelligence visual inspection technology, traditional methods often ignore the dynamic association and weight adjustment between features, resulting in redundancy in feature expression and inaccurate importance assessment. On the other hand, existing technologies lack adaptability to dynamically changing visual data during feature dimensionality reduction. Fixed topological structures and weight distribution strategies are difficult to fully cope with complex and changing inspection scenarios, reducing the robustness and adaptability of the model. In addition, feature selection and dimensionality reduction strategies are mostly based on static assumptions, and fail to fully utilize advanced methods such as reinforcement learning for online optimization, limiting further improvements in detection accuracy. Summary of the invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an artificial intelligence visual inspection method based on linear discrimination to solve the problems of insufficient feature representation, poor model adaptability and limited classification accuracy in visual inspection.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0007] In the first aspect, an embodiment of the present invention provides an artificial intelligence visual detection method based on linear discriminant, which includes receiving raw visual data and performing preprocessing, using a deep convolutional neural network to extract high-level features in the raw visual data, and generating a feature matrix; according to the feature matrix, dynamically adjusting node connections and their weights through a graph neural network to construct a dynamic feature relationship graph; using a dynamic topology adaptive linear discriminant analysis model to perform feature dimension reduction on the feature relationship weights in the dynamic feature relationship graph, and output a reduced dimension feature vector; inputting the reduced dimension feature vector into a principal component analysis model to form a collaborative reinforcement learning framework; based on the collaborative reinforcement learning framework, using a deep reinforcement learning algorithm to optimize the dynamic topology adaptive linear discriminant analysis model to improve visual detection performance; using the optimized dynamic topology adaptive linear discriminant analysis model parameters to perform feature dimension reduction and classification detection processing on real-time visual data, and realizing efficient recognition and classification of targets through a classifier.

[0008] As a preferred solution of the artificial intelligence visual inspection method based on linear discrimination described in the present invention, the preprocessing includes graying, scaling, denoising and standardization processing.

[0009] As a preferred solution of the artificial intelligence visual detection method based on linear discrimination described in the present invention, the steps of extracting high-level features from the original visual data using a deep convolutional neural network and generating a feature matrix are as follows:

[0010] The residual network of the deep convolutional neural network with 50 layers is used as the feature extraction network, which includes multiple residual modules, each of which consists of multiple convolutional layers;

[0011] The normalized image is sent as input to the fifty-layer residual network, and forward propagated through the convolution layer, BN layer, and ReLU activation function;

[0012] Before the last convolutional layer of the fifty-layer residual network, the feature map is extracted to generate the feature matrix.

[0013] As a preferred solution of the artificial intelligence visual inspection method based on linear discrimination described in the present invention, wherein: according to the feature matrix, the node connections and their weights are dynamically adjusted through the graph neural network to construct a dynamic feature relationship graph, and the specific steps are as follows:

[0014] Assume the feature matrix F is expressed as:

[0015] F∈R n×d

[0016] Among them, n represents the number of samples, d represents the feature dimension, and R is a set of real numbers;

[0017] Based on the feature matrix F, the attention similarity weight function A is used to calculate each pair of nodes vi With v j The connection probability A ij The expression is:

[0018] A ij =σ(α·cos(F i , F j )-β·|F i -F j | p )

[0019] Among them, α is a hyperparameter that adjusts the effect of cosine similarity in calculating connection probability, β is a hyperparameter that controls the effect of feature difference in calculating connection probability, and cos(F i , F j ) is the cosine similarity between the feature vectors of node i and node j, p is the norm symbol, indicating the p-order norm of the vector, |F i -F j | p is the p-order norm distance between the feature vectors of node i and node j, σ is the logistic function;

[0020] Based on the feature matrix F, the weight function W is transferred through the feature correlation to calculate each pair of nodes v i and v j The weight W between ij The expression is:

[0021] W ij =tanh(γ·cos(F i , F j )+δ)

[0022] Among them, tanh is the hyperbolic tangent function, γ is a hyperparameter that controls the influence of cosine similarity in weight allocation, and δ is an offset hyperparameter;

[0023] The attention similarity weight function A, the feature correlation transfer weight function W and the feature matrix F are used to update the node features through the graph convolution operation to obtain the new feature matrix H(1) expression:

[0024]

[0025] Among them, σ′ is a nonlinear activation function, is the degree matrix The square root inverse of is the self-loop adjacency matrix, I n is the n×n identity matrix, where n is the number of nodes in the graph;

[0026] The new feature matrix H (1) Replace the feature matrix F and based on H (1) Recalculate Aij and W ij The expression is:

[0027]

[0028] in, and Update the feature matrices F to H through graph convolution operations respectively (1) After that, the new connection probability and new feature transfer weight between node i and node j are recalculated;

[0029] based on and Perform graph convolution operations multiple times to generate hidden layer features H (k) , where k is the index variable, representing the kth iteration number. Multiple rounds of iterations are performed in this way until convergence, completing the construction of the dynamic feature relationship graph.

[0030] As a preferred solution of the artificial intelligence visual inspection method based on linear discriminant described in the present invention, wherein: the dynamic topology adaptive linear discriminant analysis model is used to perform feature dimension reduction on the feature relationship weights in the dynamic feature relationship graph, and the dimension reduction feature vector is output, and the specific steps are as follows:

[0031] The final feature matrix H obtained based on iteration (k) , calculate the dynamic connection probability between node i and node j in the dynamic feature relationship graph The expression is:

[0032]

[0033] Among them, dc is the dynamic connection, α final To adjust the optimal hyperparameters for the effect of cosine similarity in calculating connection probability, and are nodes i and j in the final feature matrix H (k) The eigenvector, β final Optimal hyperparameters to control the role of feature differences in computing connection probabilities;

[0034] The final feature matrix H obtained based on iteration (k) , calculate the weight between node i and node j in the dynamic feature relationship graph The expression is:

[0035]

[0036] Among them, δ final is the optimal hyperparameter for the offset, γ final is the hyperparameter for adjusting the cosine similarity, λ is the reinforcement coefficient for intra-category connections, is the indicator function;

[0037] H (k) , and Integration, construct the objective function expression as follows:

[0038]

[0039] Where W′ is the dimension reduction projection matrix, is the maximum value operation for the dimension reduction projection matrix W′, T is the transpose, μ c is the mean of the feature vector of category c, is the feature matrix H (k) The feature vector of the i-th sample in;

[0040] The gradient descent optimization algorithm is used to solve the above objective function and generate the feature dimension reduction matrix Z after dimension reduction: Z = H (k) W′;

[0041] Based on the feature dimension reduction matrix Z after dimension reduction, each column represents the dimension reduction feature vector of a sample.

[0042] As a preferred solution of the artificial intelligence visual inspection method based on linear discrimination of the present invention, wherein: the dimension reduction feature vector is input into the principal component analysis model to form a collaborative reinforcement learning framework, and the specific steps are:

[0043] Centralize the reduced eigenvector matrix Z and output the centralized eigenvector matrix Z cd ;

[0044] Based on the centralized feature matrix Z cd , calculate the covariance matrix C′ between the internal features;

[0045] Perform eigenvalue decomposition on the covariance matrix C′ to obtain the eigenvector V and eigenvalue A′;

[0046] According to the required dimension reduction dimension d′, select the eigenvectors corresponding to the first d′ largest eigenvalues ​​to form the matrix V d′ ;

[0047] Using the matrix V d 'Convert the reduced dimension feature matrix Z to the principal component space in the principal component analysis model to generate a low-dimensional feature vector matrix Z' PCA The expression is: Z′ PCA =ZV d′ ;

[0048] The eigenvector matrix Z′ PCAEach row of is a component of the environment state, constructing a state space where each state s t They all correspond to the specific state configuration of the environment at time step t;

[0049] Constructing an action space based on the state space, focusing on the management of feature dimensions through the action space, including the decision of increasing and decreasing the number of principal components and fine-tuning the dimensionality reduction parameters of the principal component analysis model;

[0050] The incentive mechanism is designed based on the key performance indicators of target detection task accuracy and F1 score, and the reward function expression is designed as follows:

[0051] R(s t , a)

[0052] Among them, R is the reward value, a is the action variable;

[0053] The deep Q learning network is used to implement collaborative reinforcement learning. The neural network model expression is constructed as follows:

[0054] Q(s,a|θ)

[0055] Among them, θ is the real-time network parameter; s is the state variable;

[0056] By interacting with the environment, the dimension reduction parameters in the principal component analysis model are continuously adjusted to optimize feature selection, and the best strategy of the deep Q learning network is used to guide this process, thus forming the final collaborative reinforcement learning framework.

[0057] As a preferred solution of the artificial intelligence visual inspection method based on linear discrimination described in the present invention, wherein: based on the collaborative reinforcement learning framework, the deep reinforcement learning algorithm is used to optimize the dynamic topology adaptive linear discriminant analysis model to improve the visual inspection performance, and the specific steps are:

[0058] Based on storing historical states t 、Action a t , Instant Reward R t+1 and the next state s t+1 Build a buffer to store quads (s t , a t , R t+1 ,s t+1 );

[0059] The experience replay mechanism is used to extract data from the buffer, and the network parameter θ is updated through the Bellman equation to optimize the strategy. The update expression of the Bellman equation is:

[0060]

[0061] Where L(θ) is the loss function, U(D) is uniformly sampled from the experience replay buffer D, γ is the discount factor, a′ is the best action, and θ - is the target network parameter, Q is the Q value function;

[0062] Ensemble learning idea, train multiple deep Q network models, each model focuses on a different feature subset;

[0063] Feedback the detection performance to the deep Q-network model to dynamically adjust the learning focus of the model;

[0064] According to the decision results of the deep Q-network model, the optimal hyperparameters in the dynamic topology adaptive linear discriminant analysis model are dynamically adjusted, and the weight distribution of the feature relationship is optimized.

[0065] As a preferred solution of the artificial intelligence visual inspection method based on linear discrimination described in the present invention, wherein: the efficient recognition and classification of the target is achieved by the classifier, and the specific steps are:

[0066] Map the reduced-dimensional feature vector to the feature space and perform normalization to generate a reduced-dimensional standard feature space;

[0067] By integrating support vector machines, random forests, and neural network classifiers, we train the ensemble learning strategy of voting, weighted averaging, and stacking on different feature subsets in the standard feature space after dimensionality reduction.

[0068] For each frame of data input in real time, the feature vector obtained after dimensionality reduction processing is immediately classified and predicted by the integrated classifier to output the best matching category label.

[0069] In a second aspect, an embodiment of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the artificial intelligence visual inspection method based on linear discrimination as described in the first aspect of the present invention is implemented.

[0070] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the artificial intelligence visual inspection method based on linear discrimination as described in the first aspect of the present invention.

[0071] The beneficial effects of the present invention are as follows: the present invention integrates a comprehensive framework of deep convolutional networks, graph neural networks, dynamic topology adaptive linear discriminant analysis and collaborative reinforcement learning; uses deep convolutional neural networks to perform feature extraction process, which greatly improves the quality and extraction efficiency of features; dynamically constructs and adjusts feature relationship graphs, and uses adaptive linear discriminant analysis to effectively perform feature dimensionality reduction, thereby overcoming the limitations of traditional methods in feature expression, model adaptability and classification accuracy, significantly enhancing the feature processing capabilities in complex visual scenes, and improving the robustness of the detection system and target recognition accuracy, especially in the face of changing visual environments, showing higher performance and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0073] Figure 1 This is a flow chart of the artificial intelligence visual inspection method based on linear discrimination in Example 1.

[0074] Figure 2 This is a flow chart for achieving efficient identification and classification of targets based on a classifier in Example 1. DETAILED DESCRIPTION

[0075] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0076] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0077] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0078] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, provides an artificial intelligence visual detection method based on linear discrimination, comprising the following steps:

[0079] S1. Receive raw visual data, perform preprocessing, and use a deep convolutional neural network to extract high-level features from the raw visual data to generate a feature matrix.

[0080] Furthermore, preprocessing includes grayscale, scaling, denoising and normalization;

[0081] Grayscale conversion usually uses a weighted average method to calculate the grayscale value of each pixel based on the different contribution ratios of the three channels of red, green, and blue (RGB). Its purpose is to convert a color image into a grayscale image, reduce the data dimension, remove color information, and focus on basic information such as shape and texture. This is sufficient for many visual processing tasks, and can also simplify computational complexity and reduce storage requirements;

[0082] Scaling usually uses the nearest neighbor interpolation, bilinear interpolation, or bicubic interpolation. The nearest neighbor is fast but may introduce pixelation. Bilinear and bicubic interpolation are slow but can maintain good image smoothness.

[0083] Denoising mainly uses median, Gaussian and bilateral filtering to eliminate random and irregular interference in the image, such as Gaussian noise, salt and pepper noise, etc., to improve image quality and make subsequent feature extraction more accurate;

[0084] The standardization processing method mainly uses Z-score standardization, which can ensure that all input images have a uniform scale and distribution, and avoid inconsistent feature expression caused by factors such as lighting changes and equipment differences.

[0085] Furthermore, the fifty layers of the residual network in the deep convolutional neural network are used as the feature extraction network, which includes multiple residual modules, each of which consists of multiple convolutional layers;

[0086] The normalized image is sent as input to the fifty-layer residual network, and forward propagated through the convolution layer, BN layer, and ReLU activation function;

[0087] Before the last convolutional layer of the fifty-layer residual network, the feature map is extracted to generate the feature matrix.

[0088] It should be noted that each module usually contains multiple convolutional layers, accompanied by BN layers and ReLU activation functions. The BN layer is used to accelerate the training process and improve the generalization ability of the model. By standardizing the input of each layer, the network is less sensitive to initialization. The ReLU activation function can introduce nonlinearity and enhance the learning ability of the network.

[0089] It should also be noted that the preprocessed standardized image is input into the starting layer of ResNet-50. The image data is gradually forward propagated through multiple convolutional layers, BN layers, and ReLU layers as the depth of the network increases. In this process, the network learns and gradually refines more and more abstract features. Finally, the feature map extracted by the penultimate convolutional layer of ResNet-50 (usually before the last convolutional layer, because the last layer is often a global average pooling layer or a fully connected layer for classification tasks) is regarded as a highly compressed and information-rich feature representation. These feature maps are flattened into one-dimensional vectors and combined into a feature matrix. Each sample corresponds to a row in the matrix, which provides rich input for subsequent feature relationship graph construction, dimensionality reduction, and classification.

[0090] S2. Based on the feature matrix, the node connections and their weights are dynamically adjusted through the graph neural network to construct a dynamic feature relationship graph.

[0091] Furthermore, let the feature matrix F be expressed as:

[0092] F∈R n×d

[0093] Among them, n represents the number of samples, d represents the feature dimension, and R is a set of real numbers;

[0094] Based on the feature matrix F, the attention similarity weight function A is used to calculate each pair of nodes v i With v j The connection probability A ij The expression is:

[0095] A ij =σ(α·cos(F i , F j )-β·|F i -F j | p )

[0096] Among them, α is a hyperparameter that adjusts the effect of cosine similarity in calculating connection probability, β is a hyperparameter that controls the effect of feature difference in calculating connection probability, and cos(F i , F j ) is the cosine similarity between the feature vectors of node i and node j, p is the norm symbol, indicating the p-order norm of the vector, |F i -F j | p is the p-order norm distance between the feature vectors of node i and node j, σ is the logistic function;

[0097] Preferably, by combining cosine similarity and p-order norm distance, a more delicate metric is created to evaluate the strength of the relationship between nodes. The settings of hyperparameters α and β enable the model to balance the contribution of similarity and difference in connection construction. The logistic function σ ensures that the output value is in the interval (0, 1), which is suitable as a probability value;

[0098] It should be noted that the above-mentioned logical function σ usually refers to the sigmoid function or the activation function in a more general sense, which is used to map any real value to between 0 and 1, which can be interpreted as a probability value. The mathematical expression of the sigmoid function is:

[0099]

[0100] This function has an S-shaped curve, with the function value approaching 0 as the input x approaches negative infinity and approaching 1 as x approaches positive infinity. Near 0, the function grows rapidly, which makes it a good output for a binary classifier, or here as a calculation of the probability of connection between nodes.

[0101] In the context of visual inspection, the logistic function σ is used to convert the calculated similarity or other metrics between nodes into connection probabilities. For example, if the cosine similarity between two nodes is high, then after conversion by the logistic function σ, the connection probability between the two nodes will also be high, and vice versa. This conversion helps to establish dynamic connections based on feature similarity in the network, thereby constructing a dynamic feature relationship graph.

[0102] Therefore, "σ is a logical function" means that when calculating the connection probability between node i and node j, a logical function (such as the sigmoid function) is used to ensure that the output value is in the interval (0, 1) and is suitable as a probability value. This conversion is crucial for subsequent graph neural network operations because it determines which nodes will form valid edges in the graph, which in turn affects the propagation and processing of information on the graph.

[0103] Based on the feature matrix F, the weight function W is transferred through the feature correlation to calculate each pair of nodes v i and v j The weight W between ij The expression is:

[0104] W ij =tanh(γ·cos(F i , F j )+δ)

[0105] Among them, tanh is a hyperbolic tangent function, which is used to map the linear combination of cosine similarity and an offset hyperparameter δ. This design allows the model to emphasize those highly correlated node pairs during feature transfer while maintaining nonlinear characteristics, thereby enhancing the expressiveness of the model. γ is a hyperparameter that controls the influence of cosine similarity in weight distribution, and δ is an offset hyperparameter.

[0106] The attention similarity weight function A, the feature correlation transfer weight function W and the feature matrix F are used to update the node features through the graph convolution operation to obtain the new feature matrix H(1) expression:

[0107]

[0108] Among them, σ′ is a nonlinear activation function, is the degree matrix The square root inverse of is the self-loop adjacency matrix, I n is the n×n identity matrix, where n is the number of nodes in the graph;

[0109] Preferably, this process not only integrates the structural information reflected by the adjacency matrix A, but also incorporates the feature correlation reflected by the feature correlation transfer weight function W and the content information of the initial feature matrix F. In particular, the addition of self-loop I n It ensures that each node can introspect its characteristics, and the degree matrix The inverse square root operation of achieves normalization and ensures the smoothness of information transmission;

[0110] The new feature matrix H (1) Replace the feature matrix F and based on H (1) Recalculate A ij and W ij The expression is:

[0111]

[0112] in, and They are the new connection probability and new feature transfer weight between node i and node j recalculated after updating the feature matrix F to H(1) through graph convolution operation;

[0113] Preferably, the new connection probability is calculated again by using the updated feature matrix H(1) And the new feature transfer weight This process is repeated in subsequent iterations, forming a dynamic evolution process. This mechanism enables the model to gradually refine the connections between nodes and optimize feature representation in multiple iterations until some form of convergence is reached or the predetermined stopping condition is met.

[0114] based on and Perform graph convolution operations multiple times to generate hidden layer features H (k) , where k is the index variable, representing the kth iteration number. Multiple rounds of iterations are performed in this way until convergence, completing the construction of the dynamic feature relationship graph.

[0115] It should be noted that by iteratively fusing the similarities, differences and structural information of features, the connection structure and feature weights between nodes are continuously optimized, and finally a graph model that can more accurately capture the intrinsic correlation of data is formed. This method has significant advantages in the understanding and classification of complex data structures, especially in the field of visual data processing, and can effectively improve the recognition accuracy and generalization ability of the model.

[0116] S3. Use the dynamic topology adaptive linear discriminant analysis model to perform feature dimensionality reduction on the feature relationship weights in the dynamic feature relationship graph and output the reduced dimensionality feature vector.

[0117] Furthermore, based on the final feature matrix H obtained by iteration (k) , calculate the dynamic connection probability between node i and node j in the dynamic feature relationship graph The expression is:

[0118]

[0119] Among them, dc is the dynamic connection, α final To adjust the optimal hyperparameters for the effect of cosine similarity in calculating connection probability, and are nodes i and j in the final feature matrix H (k) The eigenvector, β final Optimal hyperparameters to control the role of feature differences in computing connection probabilities;

[0120] The final feature matrix H obtained based on iteration (k) , calculate the weight between node i and node j in the dynamic feature relationship graph The expression is:

[0121]

[0122] Among them, δ final is the optimal hyperparameter for the offset, used to adjust the overall level of weights, γ final is a hyperparameter for adjusting cosine similarity, which is used to control the contribution of cosine similarity in weight calculation. λ is the reinforcement coefficient of intra-category connection. is the indicator function, tanh is the hyperbolic tangent function, which is used for smoothing;

[0123] H (k) , and Integration, construct the objective function expression as follows:

[0124]

[0125] Where W′ is the dimension reduction projection matrix, is the maximum value operation for the dimension reduction projection matrix W′, T is the transpose, μ c is the mean of the feature vector of category c, is the feature matrix H (k) The feature vector of the i-th sample in;

[0126] Preferably, the objective function is constructed to maximize the dispersion between categories while minimizing the dispersion within categories. This is the core idea of ​​linear discriminant analysis. c The optimal dimensionality reduction projection matrix W is found by minimizing the divergence in the dimensionality reduction space and the reconstruction error of each sample in the space. This ensures that the features after dimensionality reduction can maintain the separability between categories and retain the information of the original data as much as possible.

[0127] It should be noted that when we talk about maximizing the center points of categories, we actually mean that in dimensionality reduction or classification tasks, we hope to make the distance between these center points of different categories as large as possible through algorithm design, while minimizing the distance between the members within each category and the center point of its category. Doing so can enhance the separability between categories and improve the accuracy of classification. Therefore, the "maximization" involved in the optimization process actually refers to the purpose of maximizing the distance between category center points through algorithm adjustment (such as parameter optimization in the dynamic topology adaptive linear discriminant analysis model (DTALDA) model mentioned above), and the "mean of the eigenvector of category c" is the specific mathematical object used to quantify and operate in this process.

[0128] It should also be noted that the "DTALDA model" refers to the "dynamic topology adaptive linear discriminant analysis model". The DTALDA model is a method for feature dimensionality reduction based on the dynamic feature relationship graph. Its goal is to optimize feature dimensionality reduction by maximizing the dispersion between categories while minimizing the dispersion within categories, which is the core idea of ​​linear discriminant analysis (LDA); specifically, the DTALDA model attempts to adjust the distance between the center points of the categories to be as large as possible, while reducing the distance from the internal members of each category to the center point of its category, so as to enhance the separability between categories and improve the accuracy of classification; in the process of optimizing the DTALDA model, parameter adjustment is involved, such as the selection of the best hyperparameters, which will affect the calculation of dynamic connection probability and weight, thereby affecting the effect of feature dimensionality reduction. By using optimization algorithms such as gradient descent, the model will iteratively update these parameters until a local optimal solution is found, that is, the reduced-dimensional features can maintain the separability between categories and retain the information of the original data as much as possible. Therefore, the "DTALDA model" is an important component of the entire visual detection method. It is responsible for the task of feature dimensionality reduction, and its optimization process is to obtain better visual detection performance.

[0129] The gradient descent optimization algorithm is used to solve the above objective function and generate the feature dimension reduction matrix Z after dimension reduction: Z = H (k) W′;

[0130] Preferably, the gradient descent optimization algorithm is an effective method for solving the objective function, which gradually approaches the maximum value of the objective function by iteratively updating W'. This process involves calculating the gradient of the objective function with respect to W' and adjusting the value of W' accordingly until it converges to a local optimal solution.

[0131] Based on the feature dimension reduction matrix Z after dimensionality reduction, each column represents the dimension reduction feature vector of a sample. These feature vectors carry the key classification information of the original data and significantly reduce the dimension, which is beneficial to subsequent classification and recognition tasks.

[0132] S4. Input the dimension-reduced feature vector into the principal component analysis model to form a collaborative reinforcement learning framework.

[0133] Furthermore, the reduced eigenvector matrix Z is centralized and the centralized eigenvector matrix Z is outputted. cd ;

[0134] Based on the centralized feature matrix Z cd , calculate the covariance matrix C′ between the internal features;

[0135] Perform eigenvalue decomposition on the covariance matrix C′ to obtain the eigenvector V and eigenvalue A′;

[0136] It should be noted that eigenvalue decomposition is an important tool for matrix analysis, mainly used to understand the properties of matrices and simplify calculations. For a given square matrix (such as the covariance matrix C′ here), eigenvalue decomposition can reveal the inherent structure of the matrix. The following are the specific steps for eigenvalue decomposition:

[0137] The eigenvalues ​​are calculated based on the covariance matrix C′. The eigenvalues ​​are scalars λ that satisfy the following equation, expressed as:

[0138] C′v=λv;

[0139] Where v is the eigenvector and λ is the eigenvalue. This equation shows that when the matrix C′ acts on the eigenvector v, it will not change the direction of v except for possibly changing its length. The eigenvalue is actually the zero point of the matrix C′-λI, where I is the identity matrix.

[0140] Once the eigenvalue λ is obtained, the corresponding eigenvector v can be found by solving the homogeneous linear equation system C′v-λv=0 or C′v=λv. Usually, the eigenvector is normalized to make its length 1 (unit vector);

[0141] Typically, we sort the eigenvalues ​​by their magnitude. Larger eigenvalues ​​correspond to more important eigenvectors because they explain more of the variance in the data.

[0142] According to the required dimension reduction dimension d′, select the eigenvectors corresponding to the first d′ largest eigenvalues ​​to form the matrix V d′ ,In this step, in addition to manually selecting the dimension reduction dimension d′, a mechanism can also be designed to allow the algorithm to dynamically adjust the number of principal components and automatically find the optimal dimension reduction level based on the feedback during the learning process. This can be achieved by setting a reward function to encourage the number of dimensions that can improve the performance of the detection task;

[0143] Using the matrix V d′ Convert the reduced dimension feature matrix Z to the principal component space in the principal component analysis model to generate a low-dimensional feature vector matrix Z′ PCA The expression is: Z′ PCA =ZV d′ ;

[0144] The eigenvector matrix Z′ PCA Each row of is a component of the environment state, constructing a state space where each state s t They all correspond to the specific state configuration of the environment at time step t. Consider building state s t When , it not only contains the current low-dimensional feature vector matrix Z′PCA ,It is also possible to add historical behavior sequences, time series characteristics, or other aspects of the environment to provide a more comprehensive state description and enhance the learning ability of the model;

[0145] Constructing an action space based on the state space, focusing on the management of feature dimensions through the action space, including the decision of increasing and decreasing the number of principal components and fine-tuning the dimensionality reduction parameters of the principal component analysis model to explore the best feature representation;

[0146] The incentive mechanism is designed to balance these goals by weighted summation or other combinations based on the key performance indicators of target detection task accuracy and F1 score, but not limited to (other key indicators such as model generalization ability, computational efficiency or memory usage). The reward function expression is designed as:

[0147] R(s t , a)

[0148] Among them, R is the reward value, a is the action variable;

[0149] The deep Q learning network is used to implement collaborative reinforcement learning. The neural network model expression is constructed as follows:

[0150] Q(s,a|θ)

[0151] Among them, θ is the real-time network parameter; s is the state variable;

[0152] By interacting with the environment, the dimension reduction parameters in the principal component analysis model are continuously adjusted to optimize feature selection, and the best strategy of the deep Q learning network is used to guide this process, thus forming the final collaborative reinforcement learning framework.

[0153] It should be noted that by applying the reduced-dimensional feature vector to principal component analysis and integrating it into the collaborative reinforcement learning framework, dynamic optimization of feature representation and automatic adjustment of model parameters are achieved, aiming to automatically and efficiently identify the most discriminative feature subsets to improve the performance of object detection and other related tasks.

[0154] S5. Based on the collaborative reinforcement learning framework, the deep reinforcement learning algorithm is used to optimize the dynamic topology adaptive linear discriminant analysis model to improve the visual inspection performance.

[0155] Furthermore, based on storing historical state s t 、Action a t , Instant Reward R t+1 and the next state s t+1 Build a buffer to store quads (s t , a t , R t+1 ,s t+1), which helps to break the temporal correlation between data and improve learning efficiency and stability. This method allows the algorithm to learn from past experiences rather than relying only on recent data, increasing the diversity and generalization ability of learning;

[0156] The experience replay mechanism is used to extract data from the buffer, and the network parameter θ is updated through the Bellman equation to optimize the strategy. The update expression of the Bellman equation is:

[0157]

[0158] Where L(θ) is the loss function, U(D) is uniformly sampled from the experience replay buffer D, γ is the discount factor, a′ is the best action, and θ - is the target network parameter, Q is the Q value function, R t+1 represents the immediate reward at time step t+1;

[0159] It should be noted that in the update step of the reinforcement learning algorithm, the immediate reward is a measure of the direct benefit obtained after taking a specific action in a specific state. It is one of the goals that the learning algorithm tries to maximize, especially from the perspective of long-term cumulative rewards (also called returns); t It is usually combined with the estimated value of the state value function or action value function (such as the Q function) to update the parameters of the model through the Bellman equation in order to make better decisions in the future; specifically, the update rule takes into account the current immediate reward and the expectation of future rewards, the latter of which is weighted by the discount factor γ to reflect the importance of long-term benefits.

[0160] Preferably, by using the data in the experience replay, the network parameters are updated according to the Bellman equation. The core is to reduce the gap between the predicted Q value and the maximum expected Q value calculated according to the next state, so as to guide the optimization of the strategy. The loss function L(θ) emphasizes the balance between immediate rewards and future rewards (adjusted by the discount factor γ), and encourages the algorithm to make long-term beneficial action choices.

[0161] Ensemble learning ideas, train multiple deep Q network models, each model focuses on a different feature subset, and further improves the robustness of decision making by voting or averaging results between models;

[0162] Feedback the detection performance to the deep Q-network model to dynamically adjust the learning focus of the model;

[0163] According to the decision results of the deep Q-network model, the optimal hyperparameters in the dynamic topology adaptive linear discriminant analysis model are dynamically adjusted, and the weight distribution of the feature relationship is optimized to ensure that the model can self-optimize as the visual data changes, thereby improving the model's adaptability and visual detection performance.

[0164] S6. Apply the optimized dynamic topology adaptive linear discriminant analysis model parameters to perform feature dimension reduction and classification detection processing on real-time visual data, and achieve efficient recognition and classification of targets through the classifier.

[0165] Furthermore, the feature vector after dimensionality reduction is mapped to the feature space and normalized to ensure that all features are on the same scale, eliminate the influence caused by different dimensions, and generate a standard feature space after dimensionality reduction. This step is crucial for subsequent classification. The normalized feature space becomes the basis for the subsequent machine learning algorithm input, which improves the algorithm's stability and convergence speed.

[0166] By integrating support vector machines, random forests, and neural network classifiers, an integrated learning strategy of voting, weighted averaging, and stacking is trained on different feature subsets in the standard feature space after dimensionality reduction to improve the accuracy and robustness of classification;

[0167] Preferably, different types of classifiers such as support vector machines, random forests, and neural networks are combined to form a powerful classification system, each of which is based on different assumptions and learning mechanisms and can capture patterns in the data from different perspectives;

[0168] For each frame of data input in real time, the feature vector obtained after dimensionality reduction processing is immediately classified and predicted by the integrated classifier, and the best matching category label is output, so as to efficiently and accurately identify and classify the target object;

[0169] It should be noted that this comprehensive strategy not only improves the efficiency of real-time visual data processing, but also significantly enhances the accuracy and robustness of classification through integrated learning and dynamic optimization mechanisms. It is very suitable for target recognition and classification tasks in complex scenes.

[0170] This embodiment also provides a computer device, which is suitable for the case of an artificial intelligence visual inspection method based on linear discrimination, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the artificial intelligence visual inspection method based on linear discrimination proposed in the above embodiment.

[0171] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0172] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the artificial intelligence visual detection method based on linear discrimination as proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, disk or optical disk.

[0173] In summary, the present invention constructs an end-to-end solution from raw visual data preprocessing to efficient target recognition and classification by integrating advanced algorithms such as deep convolutional neural networks, graph neural networks, dynamic topology adaptive linear discriminant analysis, principal component analysis and deep reinforcement learning, which realizes efficient recognition and classification of visual data and enhances the accuracy and adaptability of detection in complex environments.

[0174] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. An artificial intelligence visual inspection method based on linear discrimination, characterized in that: include, Receive raw visual data, perform preprocessing, use deep convolutional neural network to extract high-level features in the raw visual data, and generate a feature matrix; According to the feature matrix, the node connections and their weights are dynamically adjusted through the graph neural network to construct a dynamic feature relationship graph; Using the dynamic topology adaptive linear discriminant analysis model, the feature relationship weights in the dynamic feature relationship graph are reduced in dimension and the reduced dimension feature vector is output; The dimension-reduced feature vectors are input into the principal component analysis model to form a collaborative reinforcement learning framework; Based on the collaborative reinforcement learning framework, the deep reinforcement learning algorithm is used to optimize the dynamic topology adaptive linear discriminant analysis model to improve the visual inspection performance; The optimized dynamic topology adaptive linear discriminant analysis model parameters are used to perform feature dimension reduction and classification detection processing on real-time visual data, and the classifier is used to achieve efficient recognition and classification of targets; The specific steps of achieving efficient identification and classification of targets through a classifier are as follows: Map the reduced-dimensional feature vector to the feature space and perform normalization to generate a reduced-dimensional standard feature space; By integrating support vector machines, random forests, and neural network classifiers, we train the ensemble learning strategy of voting, weighted averaging, and stacking on different feature subsets in the standard feature space after dimensionality reduction. For each frame of data input in real time, the feature vector obtained after dimensionality reduction processing is immediately classified and predicted by the integrated classifier to output the best matching category label.

2. The artificial intelligence visual inspection method based on linear discrimination as claimed in claim 1, characterized in that: The preprocessing includes grayscale, scaling, denoising and standardization.

3. The artificial intelligence visual inspection method based on linear discrimination as claimed in claim 2, characterized in that: The steps of using a deep convolutional neural network to extract high-level features from raw visual data and generate a feature matrix are as follows: The residual network of the deep convolutional neural network with 50 layers is used as the feature extraction network, which includes multiple residual modules, each of which consists of multiple convolutional layers. The normalized image is sent as input to the fifty-layer residual network, and forward propagated through the convolution layer, BN layer, and ReLU activation function; Before the last convolutional layer of the fifty-layer residual network, the feature map is extracted to generate the feature matrix.

4. The artificial intelligence visual inspection method based on linear discrimination as claimed in claim 3, characterized in that: According to the feature matrix, the node connections and their weights are dynamically adjusted through the graph neural network to construct a dynamic feature relationship graph. The specific steps are as follows: Assume the feature matrix F is expressed as: F∈R n×d Among them, n represents the number of samples, d represents the feature dimension, and R is a set of real numbers; Based on the feature matrix F, the attention similarity weight function A is used to calculate each pair of nodes v i With v j The connection probability A ij The expression is: A ij =σ(α·cos(F i ,F j )-β·|F i -F j | p ) Among them, α is a hyperparameter that adjusts the effect of cosine similarity in calculating connection probability, β is a hyperparameter that controls the effect of feature difference in calculating connection probability, and cos(F i ,F j ) is the cosine similarity between the feature vectors of node i and node j, p is the norm symbol, indicating the p-order norm of the vector, |F i -F j | p is the p-order norm distance between the feature vectors of node i and node j, σ is the logistic function; Based on the feature matrix F, the weight function W is transferred through the feature correlation to calculate each pair of nodes v i and v j The weight W between ij The expression is: W ij =tanh(γ·cos(F i ,F j )+δ) Among them, tanh is the hyperbolic tangent function, γ is a hyperparameter that controls the influence of cosine similarity in weight allocation, and δ is an offset hyperparameter; The attention similarity weight function A, the feature correlation transfer weight function W and the feature matrix F are used to update the node features through the graph convolution operation to obtain the new feature matrix H (1) The expression is: Among them, σ' is a nonlinear activation function, is the degree matrix The square root inverse of is the self-loop adjacency matrix, I n is the n×n identity matrix, where n is the number of nodes in the graph; The new feature matrix H (1) Replace the feature matrix F and based on H (1) Recalculate A ij and W ij The expression is: in, and Update the feature matrices F to H through graph convolution operations respectively (1) After that, the new connection probability and new feature transfer weight between node i and node j are recalculated; based on and Perform graph convolution operations multiple times to generate hidden layer features H (k) , where k is the index variable, representing the kth iteration number. Multiple rounds of iterations are performed in this way until convergence, completing the construction of the dynamic feature relationship graph.

5. The artificial intelligence visual inspection method based on linear discrimination as claimed in claim 4, characterized in that: The reduced-dimensional feature vector is input into the principal component analysis model to form a collaborative reinforcement learning framework. The specific steps are: Centralize the reduced eigenvector matrix Z and output the centralized eigenvector matrix Z cd ; Based on the centralized feature matrix Z cd , calculate the covariance matrix C' between the internal features; Perform eigenvalue decomposition on the covariance matrix C' to obtain the eigenvector V and eigenvalue A'; According to the required dimension reduction dimension d', select the eigenvectors corresponding to the first d' largest eigenvalues ​​to form the matrix V d' ; Using the matrix V d' Convert the reduced dimension feature matrix Z to the principal component space in the principal component analysis model to generate a low-dimensional feature vector matrix Z' PCA The expression is: Z' PCA =ZV d' ; The eigenvector matrix Z' PCA Each row of is a component of the environment state, constructing a state space where each state s t They all correspond to the specific state configuration of the environment at time step t; Constructing action space based on state space, focusing on the management of feature dimensions through action space, including decision-making on increasing and decreasing the number of principal components and fine-tuning the dimensionality reduction parameters of the principal component analysis model; The incentive mechanism is designed based on the key performance indicators of target detection task accuracy and F1 score, and the reward function expression is designed as follows: R(s t ,a) Among them, R is the reward value, a is the action variable; The deep Q learning network is used to implement collaborative reinforcement learning. The neural network model expression is constructed as follows: Q(s,a|θ) Among them, θ is the real-time network parameter; s is the state variable; By interacting with the environment, the dimension reduction parameters in the principal component analysis model are continuously adjusted to optimize feature selection, and the best strategy of the deep Q learning network is used to guide this process, thus forming the final collaborative reinforcement learning framework.

6. The artificial intelligence visual inspection method based on linear discrimination as claimed in claim 5, characterized in that: Based on the collaborative reinforcement learning framework, the deep reinforcement learning algorithm is used to optimize the dynamic topology adaptive linear discriminant analysis model to improve the visual detection performance. The specific steps are as follows: Based on storing historical states t 、Action a t , Instant Reward R t+1 and the next state s t+1 Build a buffer to store quads (s t ,a t ,R t+1 ,s t+1 ); The experience replay mechanism is used to extract data from the buffer and the network parameters θ are updated through the Bellman equation to optimize the strategy. Ensemble learning idea, train multiple deep Q network models, each model focuses on a different feature subset; Feedback the detection performance to the deep Q-network model to dynamically adjust the learning focus of the model; According to the decision results of the deep Q-network model, the optimal hyperparameters in the dynamic topology adaptive linear discriminant analysis model are dynamically adjusted, and the weight distribution of the feature relationship is optimized.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the artificial intelligence visual inspection method based on linear discrimination described in any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the artificial intelligence visual inspection method based on linear discrimination described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Brain vision detection and analysis device and method based on neural feedback

    CN110585591A

  • Full-view visual inspection method and device for semiconductor manufacturing industry

    CN115035325A