Facial expression recognition method and system based on improved neural network architecture search
Through improved neural network architecture search and optimization methods, including path position identification coding, CNN-LSTM evaluation, quantum computing and biometric simulation, combined with genetic algorithms and life cycle selection strategies, the problems of low efficiency and complex optimization in the field of facial expression recognition are solved, and efficient and accurate neural network architecture search and optimization are achieved.
Patent Information
- Application Number
- CN202411922428.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-06-03
AI Technical Summary
The existing neural network architecture search methods have problems such as low evaluation efficiency, complex optimization process, and easy to fall into local optimal solutions in the field of facial expression recognition.
Improved neural network architecture search and optimization methods are adopted, including generating path position identification coding (PIPE), rapid CNN-LSTM evaluation, introducing quantum computing and biometric simulation methods, combining genetic algorithms and life cycle selection strategies, multi-objective optimization and dynamic threshold screening.
It significantly improves the efficiency and performance of neural network architecture search, improves the model performance and computing efficiency of facial expression recognition tasks, and can quickly filter out potentially excellent candidate architectures and find the global optimal neural network architecture.
Smart Images

Figure CN120088820A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural network architecture search (NAS) and optimization, and particularly to a facial expression recognition method and system based on improved neural network architecture search and optimization. Background Art
[0002] In the early stage, the neural network architectures for facial expression recognition tasks were all manually designed, which had problems of low efficiency and poor accuracy. Prior knowledge was still required to design a network architecture with good performance. Therefore, it has become an urgent task to apply NAS in the field of facial expression recognition. Existing neural network architecture search methods have problems such as low evaluation efficiency and complex optimization processes. Traditional methods consume a large amount of resources in the evaluation and screening process of the initial population, and are prone to falling into local optimal solutions in the subsequent optimization process, making it difficult to obtain the globally optimal neural network architecture. Summary of the Invention
[0003] In view of the problems existing in the above-mentioned prior art, the present invention is proposed.
[0004] To solve the above technical problems, the present invention provides the following technical solutions. A facial expression recognition method based on improved neural architecture search and optimization includes: S1. Generate path position identification encoding (PIPE), generate path position identification encoding vectors for the input-to-output paths of the neural architecture, and arrange them in ascending order of path length; traverse all paths, record the one-hot encoding and activation function feature vectors of each operation node, and arrange them in ascending order of operation index; concatenate the path position identification encoding vectors to form a complete encoding representation, ensuring encoding consistency and efficiency; add timestamp information and feature extraction step characteristics to the encoding vectors to record the execution order of each operation node and the specific facial expression feature extraction process, and record the response of a specific convolutional layer to the eye region; use quantum computing technology to perform parallel processing on the path position identification encoding to improve the encoding speed and accuracy; introduce a biological feature simulation method into the path position identification encoding to imitate the way of the biological neural network in processing facial expressions, including enhancing the sensitivity to subtle differences such as smiling and frowning expressions, and improving the naturalness and effectiveness of the encoding; S2. Conduct CNN-LSTM fast evaluation, input the encoded network structure sequence into the convolutional layer to extract facial expression features, and perform dimensionality reduction through the pooling layer; input the feature sequence into the LSTM layer to capture the long-term dependence relationship of the expression features, and at the same time prevent overfitting through the Dropout layer.Add a self-attention mechanism to the LSTM layer to enhance the ability to capture subtle facial expression changes; add an adversarial training module after the Dropout layer to improve the robustness and anti-interference ability of the model by introducing adversarial samples; introduce a multi-task learning method to not only predict facial expression classification but also other related performance metrics, such as expression intensity or confidence, to improve the comprehensive prediction ability of the model; add an adaptation module for hardware accelerators in the model to optimize the performance of the model on different hardware platforms; S3, EvoXBench performance evaluation; the selected candidate architectures will be evaluated for multi-objective optimization on different search spaces and datasets on the EvoXBench platform, and specific evaluation metrics will be used to evaluate the candidate architectures for facial expression recognition; S4, set dynamic thresholds based on performance distribution, mark the network structures with comprehensive scores higher than the thresholds as candidates, and manage these candidate structures using a queue; the queue is dynamically adjusted according to the latest performance data to ensure that each candidate network structure enters the comprehensive evaluation and optimization in the second stage in sequence; introduce a deep learning prediction model in the threshold setting to dynamically predict and adjust the threshold to improve the intelligence of screening; adopt a hierarchical queue management method to hierarchically manage candidate structures according to different task types and performance metrics; add an automated performance evaluation and adjustment module to evaluate the latest performance of candidate structures in real time and dynamically adjust the queue order; add an adaptation module for hardware accelerators to ensure the efficient operation of the architecture on different hardware platforms; S5, add a fitness decay mechanism to the life cycle selection strategy to prevent premature convergence caused by excessive improvement of individual fitness within the life cycle; introduce a hybrid mode of multiple life cycle selection strategies to dynamically adjust the selection strategy according to different task requirements; add an automated recognition and adjustment module for life cycle stages to dynamically adjust the life cycle stage of an individual according to its performance; integrate the simulated annealing algorithm and combine it with the life cycle selection strategy to improve the global search ability and optimization effect; S6, perform genetic algorithm optimization and initialize the population; perform selection, crossover, and mutation operations through a life cycle-based selection strategy; introduce a diversity mutation mechanism of biological evolution in the mutation operation to simulate the diverse evolution process in nature and improve the mutation effect; add a mutation operation selection module based on reinforcement learning to automatically select the optimal mutation operation through a reinforcement learning algorithm; introduce a fuzzy logic control system to infer the best mutation strategy based on the current state of an individual through fuzzy reasoning; add a multi-objective optimization mechanism to the mutation operation to optimize multiple performance metrics simultaneously and enhance the comprehensive effect of the mutation operation; S7, based on the crowding distance selection strategy, when selecting individuals, calculate the crowding distance of each individual and preferentially select those individuals located in non-crowded areas; S8, after screening out the optimal architecture, perform a final performance evaluation and fine-tuning on it to ensure its best performance in the facial expression recognition task. Apply the fine-tuned architecture to specific expression recognition problems as the core basis of the algorithm to achieve efficient and accurate expression classification and analysis.
[0005] As a preferred solution of the facial expression recognition method based on improved neural network architecture search and optimization according to the present invention, wherein: the steps of generating the path position identification code include:
[0006] S21. Generate a feature path encoding vector for each operation node and arrange them in ascending order of operation index, which is used to represent the feature extraction operations of each layer in the facial expression recognition model;
[0007] S22. Add the temperature characteristic information of the operation node to the feature path encoding, which is used to record the performance of the neural network under different environmental conditions (such as different lighting or temperature conditions), so as to optimize the accuracy and robustness of facial expression recognition;
[0008] S23. Introduce a biometric simulation method into the feature path encoding, imitate the feature extraction path of the biological neural network when processing facial expressions, improve the naturalness and effectiveness of the encoding, and thus enhance the model's ability to capture subtle expression changes.
[0009] As a preferred solution of the facial expression recognition method based on improved neural network architecture search and optimization according to the present invention, wherein:
[0010] S31. Add a self-attention mechanism to the LSTM layer to improve the ability to capture long-term dependence relationships and improve the recognition accuracy of subtle expression changes;
[0011] S32. Add an adversarial training module after the Dropout layer. By introducing adversarial samples, improve the robustness and anti-interference ability of the model, and enhance the reliability of facial expression recognition;
[0012] S33. Introduce a multi-task learning method to predict multiple performance indicators of the network structure simultaneously and improve the comprehensive prediction ability of the model.
[0013] As a preferred solution of the facial expression recognition method based on improved neural network architecture search and optimization according to the present invention, wherein: the threshold setting and automatic transition steps include:
[0014] S51. Set a dynamic threshold based on the performance distribution, mark the network structures with a comprehensive score higher than the threshold as candidates, and manage these candidate structures using a queue;
[0015] S52. The queue is dynamically adjusted according to the latest performance data to ensure that each candidate network structure enters the comprehensive evaluation and optimization in the second stage in sequence;
[0016] S53. Introduce a deep learning prediction model into the threshold setting to dynamically predict and adjust the threshold and improve the intelligence level of screening;
[0017] S54. Adopt a hierarchical queue management method to hierarchically manage candidate structures according to different task types and performance metrics;
[0018] S55. Add an automated performance evaluation and adjustment module to evaluate the latest performance of candidate structures in real time and dynamically adjust the queue order.
[0019] As a preferred solution of the facial expression recognition method based on improved neural network architecture search and optimization according to the present invention, wherein:
[0020] S71. Incorporate a fitness decay mechanism into the life cycle selection strategy to prevent premature convergence caused by excessive improvement of individual fitness within the life cycle;
[0021] S72. Introduce a hybrid mode of multiple life cycle selection strategies and dynamically adjust the selection strategy according to different task requirements;
[0022] S73. Add an automated recognition and adjustment module for life cycle stages to dynamically adjust the life cycle stage of an individual according to its performance;
[0023] S74. Integrate the simulated annealing algorithm and combine it with the life cycle selection strategy to improve the global search ability and optimization effect.
[0024] As a preferred solution of the facial expression recognition method based on improved neural network architecture search and optimization according to the present invention, after screening out the optimal architecture, the following steps are carried out:
[0025] S101. Conduct a final performance evaluation on the optimal architecture to verify its effectiveness in practical applications;
[0026] S102. Further fine-tune the optimal architecture according to the final evaluation results to ensure its stability and efficiency in different facial expression recognition application scenarios, including the ability to handle different lighting conditions and expression changes;
[0027] S103. Apply the finally determined optimal architecture to specific facial expression recognition problems as the core algorithm basis to solve the expression classification and analysis problems in practical applications.
[0028] As a preferred solution of the facial expression recognition system based on improved neural network architecture search and optimization according to the present invention, it includes an architecture combination definition module, a candidate architecture screening module, a multi-objective optimization evaluation module, a dynamic threshold screening module, and an optimal architecture verification module.
[0029] A computer device includes a memory and a processor. The memory stores a computer program. It is characterized in that when the processor executes the computer program, the steps of any one of the facial expression recognition methods based on improved neural network architecture search and optimization are implemented.
[0030] A computer-readable storage medium stores a computer program thereon. It is characterized in that when the computer program is executed by a processor, the steps of any one of the facial expression recognition methods based on improved neural network architecture search and optimization are implemented.
[0031] Advantages of the present invention: The present invention can significantly improve the efficiency and performance of neural network architecture search, and enhance the model performance and computational efficiency of facial expression recognition tasks. Specifically, through a two-stage progressive method, potential excellent candidate architectures can be quickly screened out; through an improved genetic algorithm, a globally optimal neural network architecture can be found in a relatively short time; through a life-cycle-based selection strategy, the selection probability can be dynamically adjusted to prevent premature convergence; by introducing advanced technologies such as quantum computing and biometric simulation, the accuracy and effectiveness of encoding and optimization are further improved. While solving the problems of the prior art, the present invention provides an efficient and reliable solution for the search and optimization of neural network architectures for facial expression recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for description in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0033] Figure 1 It is a schematic diagram of the algorithm flow of the facial expression recognition method based on improved neural network architecture search and optimization provided by an embodiment of the present invention.
[0034] Figure 2 It is a schematic diagram of the network structure encoding of the facial expression recognition method based on improved neural network architecture search and optimization provided by an embodiment of the present invention.
[0035] Figure 3 It is a schematic diagram of the LSTM-CNN local feature extraction design of the facial expression recognition method based on improved neural network architecture search and optimization provided by an embodiment of the present invention.
[0036] Figure 4 It is a schematic diagram of the rough evaluation predefined architecture of the facial expression recognition method based on improved neural network architecture search and optimization provided by an embodiment of the present invention.
[0037] Figure 5 It is a graph showing the accuracy change of each model during the training process of the embodiment.
[0038] Figure 6 It is a graph showing the loss value change of each model during the training process of the embodiment.
[0039] Figure 7 It is a comparison graph of the training time required for each model of the embodiment under different accuracy targets. Detailed implementation manners
[0040] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0041] Embodiment 1
[0042] Referring to Figures 1 - 4 , which is the first embodiment of the present invention. This embodiment provides a facial expression recognition method based on improved neural network architecture search and optimization, including:
[0043] The present invention relates to a facial expression recognition method based on improved neural network architecture search and optimization. This method effectively improves the efficiency and performance of neural network architecture search through a phased optimization strategy, path position identification coding, a CNN-LSTM model, and a genetic algorithm based on a life cycle selection strategy. The specific embodiments are as follows:
[0044] S1: Define a network architecture combination suitable for facial expression recognition;
[0045] In the embodiments of the present application, when defining the network architecture combination suitable for facial expression recognition, path position identification coding is used to encode the neural network architecture to capture the multi-path information of the network architecture;
[0046] In the embodiments of the present application, it should be noted that the NAS task is divided into two stages: in the first stage, path position identification coding and a first model are used to quickly evaluate and screen the initial population; in the second stage, the selected candidate architectures are comprehensively trained and optimized, and the performance is improved through a genetic algorithm based on a life cycle selection strategy. As Figure 1 shown, the algorithm flow of the facial expression recognition method based on improved neural network architecture search and optimization can be roughly divided into several modules such as network structure coding, model design, rough evaluation, threshold setting, and genetic optimization.
[0047] In the embodiment of the present application, the path location identification code includes assigning a unique index to each operation node of the neural network. If an operation node is located in the current input-to-output path, a corresponding one-hot vector is generated according to the operation node type; if it is not in the path, it is represented by a vector of all zeros. An encoding vector of the operation position is generated for each path from input to output, and the encoding vectors of each path are concatenated to form a complete representation. The encoding vectors of all paths are concatenated to form an overall path encoding representation of the neural network architecture.
[0048] In the embodiment of the present application, it should be noted that the path location identification code (PIPE) is generated, as Figure 2 shown, generating a path location identification code vector for the input-to-output path of the neural architecture and arranging them in ascending order of path length; traversing all paths, recording the one-hot encoding and activation function feature vector of each operation node, and arranging them in ascending order of operation index; concatenating the path location identification code vectors to form a complete encoding representation, ensuring encoding consistency and efficiency; adding timestamp information and feature extraction step characteristics to the encoding vector to record the execution order of each operation node and the specific facial expression feature extraction process, and recording the response of a specific convolutional layer to the eye region; using quantum computing technology to perform parallel processing on the path location identification code to improve the encoding speed and accuracy; introducing a biometric simulation method into the path location identification code to imitate the way of the biological neural network to process facial expressions, including enhancing the sensitivity to subtle differences such as smiling and frowning expressions, and improving the naturalness and effectiveness of the encoding.
[0049] S2: Input the facial expression image data into the first model to identify and screen out candidate architectures that can process facial expression information;
[0050] In the embodiment of the present application, most preferably, the first model is a CNN-LSTM model;
[0051] It should be noted that a CNN-LSTM fast evaluation is performed, as Figure 3 shown, inputting the encoded network structure sequence into the convolutional layer to extract facial expression features and performing dimensionality reduction through the pooling layer; inputting the feature sequence into the LSTM layer to capture the long-term dependence relationship of the expression features, and at the same time preventing overfitting through the Dropout layer. Adding a self-attention mechanism to the LSTM layer to enhance the ability to capture subtle facial expression changes; adding an adversarial training module after the Dropout layer to improve the robustness and anti-interference ability of the model by introducing adversarial samples; introducing a multi-task learning method to not only predict facial expression classification but also other related performance indicators, such as expression intensity or confidence, to improve the comprehensive prediction ability of the model; adding an adaptation module for hardware accelerators to the model to optimize the performance of the model on different hardware platforms.
[0052] Furthermore, generate path position identification coding vectors for the input-to-output paths of the neural architecture and sort them in ascending order of path length. The neural network architecture is defined by a directed acyclic graph (DAG), where nodes represent operations in the architecture, the adjacency matrix is A, and the operation nodes are O i . The path position identification coding P i = PE(A, O i ), where PE represents the path position identification coding function.
[0053] Furthermore, traverse all paths, record the one-hot coding and activation function feature vectors of each operation node, and sort them in ascending order of operation index. For example, for each path, the one-hot coding represents the operation type T, and the activation function feature vector represents the activation function F of that node. The operations for the one-hot coding and activation function feature vectors are: One-hot(O i ) = [0, 1, 0,..., 0], Activation(O i ) = [0, 0, 1, 0,..., 0]
[0054] Furthermore, concatenate the path position identification coding vectors to form a complete coding representation, ensuring coding consistency and efficiency. Assume the path position identification coding vectors are Pi, and the concatenated complete coding is C = contact(P 1 , P 2 ,..., P n ), where concat represents the concatenation operation.
[0055] Furthermore, add timestamp information Ts(O i ) = t i to the coding vector, which is used to record the execution time of each operation node in the path.
[0056] Furthermore, add the temperature feature of the operation node to the path position identification coding. The temperature feature Temp(O i ) = T i .
[0057] Furthermore, use quantum algorithms to perform parallel processing on the path position identification coding. The quantum computing technology is represented as follows: where α i is the amplitude coefficient and ψ i > is the ground state.
[0058] Furthermore, introduce a biometric simulation method into the path position identification coding to mimic the path characteristics of the biological neural network. The biometric simulation method is realized by simulating the path characteristics of the biological neural network, which improves the naturalness and effectiveness of the coding.
[0059] Perform genetic algorithm optimization and initialize the population; perform selection, crossover, and mutation operations through a life-cycle-based selection strategy; add a fitness decay mechanism to the life-cycle selection strategy to prevent premature convergence caused by excessive improvement of individual fitness within the life cycle; introduce a hybrid mode of multiple life-cycle selection strategies to dynamically adjust the selection strategy according to different task requirements; add an automated recognition and adjustment module for the life-cycle stage to dynamically adjust the life-cycle stage of individuals according to their performance; integrate the simulated annealing algorithm and combine it with the life-cycle selection strategy to improve the global search ability and optimization effect; introduce a diversity mutation mechanism of biological evolution in the mutation operation to simulate the diverse evolution process in nature and improve the mutation effect; add a mutation operation selection module based on reinforcement learning to automatically select the optimal mutation operation through the reinforcement learning algorithm; introduce a fuzzy logic control system to fuzzy-infer the best mutation strategy according to the current state of individuals; add a multi-objective optimization mechanism to the mutation operation to optimize multiple performance indicators simultaneously and enhance the comprehensive effect of the mutation operation;
[0060] Input the encoded network structure sequence C into a one-dimensional convolutional layer to extract features and reduce the dimension through a pooling layer. Assume the weight of the one-dimensional convolutional layer is W, the bias is b, and the pooling operation is pool: Conv1D(C) = W·C + b, Pooling(C) = pool(C)
[0061] Input the feature sequence H into the LSTM layer to capture long-term dependency relationships and prevent overfitting through the Dropout layer. The hidden state of the LSTM layer is h, the cell state is c, and the Dropout probability is p: h t ,c t = LSTM(H t ,h t-1 ,c t-1 )h' t = Dropout(h t ,p)
[0062] The LSTM output ht′ calculates the final prediction score ŷ through the average layer and optimizes the model parameters using a combination of mean squared error and mean absolute error. The formula is as follows:
[0063] Add a self-attention mechanism to the LSTM layer. The self-attention mechanism is used to enhance the model's ability to capture long-term dependency relationships and improve the prediction accuracy.
[0064] Add an adversarial training module after the Dropout layer. The adversarial training module improves the robustness and anti-interference ability of the model by introducing adversarial samples. The adversarial training operation is: AdvTr(H,δ) = H + δ, where δ is the adversarial noise.
[0065] Introduce the multi-task learning method to predict multiple performance indicators of the network structure simultaneously. The multi-task learning method predicts multiple performance indicators simultaneously in one model, and the operation is: MultiTask(H) = [f 1 (H), f 2 (H),..., f n (H)], where f i represents the prediction function of the i-th task.
[0066] Add an adaptation module for the hardware accelerator in the model. The hardware accelerator adaptation module improves the adaptability and efficiency of the model by optimizing the performance of the model on different hardware platforms. The operation of the hardware accelerator adaptation module is: HardwareAdapt(H,A) = A(H), where A is the hardware accelerator adaptation function.
[0067] In an alternative embodiment, the first model can also be an RNN model;
[0068] In an alternative embodiment, the first model can also be a Transformer model;
[0069] S3: The selected candidate architectures will be evaluated by multi-objective optimization based on different search spaces and datasets on the first platform to verify their adaptability in various facial expression recognition scenarios.
[0070] In the embodiment of the present application, the most preferred first platform is EvoXBench.
[0071] Furthermore, after the initial screening, EvoXBench is used for further performance evaluation. The candidate architectures {Ai} selected in the initial screening stage are further evaluated for performance on EvoXBench, and the candidate architecture set is {A 1 , A 2 ,…, A n}.
[0072] It should be noted that, through a multi-layer perceptron (MLP) surrogate model or a database query method, as Figure 4 shown, the prediction error, model complexity, and test performance of the network architecture are evaluated.
[0073] Furthermore, the optimal architecture is selected through Pareto sorting, and the evaluation results are fed back to the evolutionary algorithm to optimize the search strategy. The optimal architecture with the best performance is selected using Pareto sorting, and the evaluation results are fed back to the evolutionary algorithm to optimize the search strategy. The Pareto sorting operation is as follows: Indicates whether the individual xi is on the Pareto front.
[0074] It should be noted that dedicated evaluation modules for different domain tasks are added in EvoXBench. The dedicated evaluation modules for different domain tasks are carried out by evaluating performance metrics for specific task requirements.
[0075] Furthermore, an intelligent data cleaning and preprocessing module is introduced. The intelligent data cleaning and preprocessing module improves data quality by automatically cleaning and preprocessing data.
[0076] Furthermore, a real-time performance monitoring and feedback mechanism is added. The real-time performance monitoring and feedback mechanism dynamically adjusts by monitoring performance metrics in real time and feeding the monitoring results back into the system.
[0077] Furthermore, a combined evaluation function integrating multiple optimization algorithms is integrated. The combined evaluation function adaptively evaluates different types of network architectures by integrating multiple optimization algorithms.
[0078] In an alternative embodiment, the first platform may also be NASBench;
[0079] In an alternative embodiment, the first platform may also be bench-ENAS;
[0080] S4: According to the architecture performance distribution evaluated by the first platform, dynamically set a screening threshold to screen architectures with good performance, and use the first algorithm to improve the performance of the screened architectures.
[0081] By setting a dynamic threshold based on the performance distribution, the comprehensive score S i Network structures with a score higher than the threshold are marked as candidates, and a queue is used to manage these candidate structures. The comprehensive score calculation operation is: S i = αf e (x i ) + βf c (x i ) + γf H (x i ) where f e (x i ) represents the prediction error, f c (x i ) represents the model complexity, f H (x i ) represents the test performance, and α, β, γ are weight coefficients.
[0082] The queue is dynamically adjusted according to the latest performance data to ensure that each candidate network structure enters the comprehensive evaluation and optimization in the second stage in sequence. The queue management method is as follows: Let the queue be Q. For the network structures with comprehensive scores higher than the threshold, add them to the tail of Q; when a candidate network structure passes the evaluation and optimization in the second stage, remove this network structure from the head of Q.
[0083] S5.3 In the dynamic threshold setting, dynamically predict and adjust the threshold to improve the intelligence of screening. The deep learning prediction model learns historical data to dynamically predict and adjust the threshold. The operation of the prediction model is: Predict(x) = W·x + b, where W is the weight matrix and b is the bias term.
[0084] S5.4 Adopt a hierarchical queue management method to hierarchically manage candidate structures according to different task types and performance metrics. The hierarchical queue management method classifies and manages candidate structures according to task types and performance metrics to improve management efficiency. The hierarchical queue management method is as follows: Hierarchically manage candidate structures according to different task types Ti and performance metrics Pi:
[0085] S5.5 Add an automated performance evaluation and adjustment module to evaluate the latest performance of candidate structures in real time and dynamically adjust the queue order. The automated performance evaluation and adjustment module dynamically adjusts the queue order by evaluating the latest performance of candidate structures in real time to ensure the real-time nature of the evaluation. The evaluation function is expressed as: AutoEval(A j ) = f e (A j ), f c (A j ), f H (A j )
[0086] In the embodiment of the present application, the first algorithm proposes an improved genetic algorithm, which effectively improves the efficiency and performance of neural network architecture search through innovative methods such as phased optimization strategies, path position identification coding, life cycle selection strategies, and regularization genetic operations.
[0087] It should be noted that a life cycle-based selection strategy is introduced, and different selection probabilities are designed according to the life cycle stages of individuals (such as young, middle-aged, and old) to generate the optimal network structure. The operation of calculating the selection probability is: Among them, the fitness decay function is defined as:
[0088] Furthermore, a fitness decay mechanism is added to the life cycle selection strategy to prevent premature convergence caused by excessive improvement of individual fitness within the life cycle. The fitness decay mechanism prevents premature convergence by reducing individual fitness, and the operation is as follows:
[0089] It should be noted that a hybrid mode of multiple life cycle selection strategies is introduced, and the selection strategy is dynamically adjusted according to different task requirements. The hybrid mode combines multiple selection strategies and dynamically adjusts the selection probability according to task requirements. The operation is as follows:
[0090] Furthermore, an automated recognition and adjustment module for the life cycle stage is added to dynamically adjust the life cycle stage of an individual according to its performance. The automated recognition and adjustment module dynamically adjusts the life cycle stage by monitoring the individual's performance.
[0091] It should be noted that the simulated annealing algorithm is integrated and combined with the life cycle selection strategy. The simulated annealing algorithm improves the global search ability by introducing a temperature parameter and combining with the life cycle selection strategy.
[0092] Crossover and Mutation: Regularized Genetic Optimization
[0093] When performing the mutation operation, it is judged whether to perform the mutation by generating a random number. If the condition is satisfied, a position is randomly selected and a mutation operation is selected from a predefined mutation list, such as adding a random pooling layer, adding a random block, deleting or changing a pooling layer, etc., to adjust the depth of the network architecture. The operation of calculating the mutation probability is as follows:
[0094] A diversity mutation mechanism of biological evolution is introduced in the mutation operation to simulate the diverse evolution process in nature and improve the mutation effect. The operation of the diversity mutation mechanism is: diverse_mutate(O i ) = O i + rand(δ), where δ is the mutation amplitude and rand(δ) represents a random perturbation.
[0095] A fuzzy logic control system is introduced to infer the best mutation strategy based on the current state of the individual through fuzzy reasoning. The fuzzy logic control system selects the best mutation strategy according to the current state of the individual through fuzzy reasoning. The operation of the fuzzy logic control is as follows: where μ i (x) is the fuzzy membership function and y i is the output of the fuzzy rule.
[0096] A multi-objective optimization mechanism is added to the mutation operation, and multiple performance indicators are optimized simultaneously: minF(x) = (f 1 (x), f 2 (x),..., fm (x)), where F(x) is the objective function vector and f i (x) is the i-th objective function.
[0097] Calculate the crowding distance of each individual, which is used to measure the sparsity of the individual in the feature space. The operation of calculating the crowding distance is as follows: where CD i is the crowding distance of the i-th individual, and are the values of the two individuals before and after on the m-th objective respectively, and are the maximum and minimum values on the m-th objective respectively.
[0098] The larger the crowding distance, the sparser the individual in the feature space and the higher the selection priority. Specifically, by comparing the crowding distances of each individual, select the individuals with larger crowding distances for the next operation.
[0099] Introduce a dynamic adjustment mechanism in the calculation of the crowding distance, dynamically adjust the calculation method of the crowding distance according to the individual performance, and improve the flexibility of the selection strategy. The operation of the dynamic adjustment mechanism is: adj(CD i ) = α·CD i = + β·per i where α and β are adjustment coefficients, and per i is the performance score of the ii-th individual.
[0100] Add a crowding distance calculation module for the multi-dimensional feature space to improve the accuracy and applicability of the crowding distance calculation. The multi-dimensional feature space calculation module operates by considering features in multiple dimensions: CD i = where N is the number of feature dimensions considered.
[0101] Introduce a crowding distance calculation method based on graph neural networks, extract features and perform calculations on individuals through graph neural networks, and improve the intelligence level of the selection strategy. The operation is as follows: where σ is the activation function, N(x) is the set of neighbor nodes of node x, c i,x is the normalization constant, W is the weight matrix, and h i is the feature of the neighbor nodes.
[0102] Integrate the crowding distance calculation mechanism for multi-objective optimization, consider the crowding situations of multiple objectives simultaneously, and improve the comprehensive effect of the selection strategy. The operation of the multi-objective optimization mechanism is: minF(x) = (f 1 (x), (f 2 (x),...,(f m(x)), where F(x) is the objective function vector and f i (x) is the i-th objective function.
[0103] In an alternative embodiment, the first algorithm can also be a genetic programming algorithm;
[0104] In an alternative embodiment, the first algorithm can also be a dung beetle optimization algorithm;
[0105] S5: Through comprehensive performance evaluation and optimization, select the optimal neural network architecture and verify it on the test set of the facial expression recognition task to ensure high efficiency and high accuracy.
[0106] It should be noted that after completing all genetic algorithm operations, select the neural network architecture with the highest comprehensive score as the optimal architecture. The comprehensive score calculation operation is: S i = αf e (A i ) + βf c (A i ) + γf H (A i ) where f e (A i ) represents the prediction error of the i-th architecture, f c (A i ) represents the model complexity of the iii-th architecture, f H (A i ) represents the test performance of the i-th architecture, and α, β, and γ are weight coefficients.
[0107] It should be noted that perform a final performance evaluation on the optimal architecture to verify its effectiveness in practical applications. The final performance evaluation metrics include but are not limited to prediction accuracy, computational efficiency, and resource consumption.
[0108] It should be noted that apply the finally determined optimal architecture to specific facial expression recognition problems as the core algorithm basis to solve the expression classification and analysis problems in practical applications.
[0109] It should be noted that in this application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0110] The above are only specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments shown herein, but rather will conform to the broadest scope consistent with the principles and novel features claimed herein.
[0111] Embodiment 2 provides a facial expression recognition method based on improved neural network architecture search and optimization. To verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.
[0112] It should be noted that when applying the finally determined optimal architecture to the facial expression recognition task, a comprehensive dataset containing 100,500 pictures is first constructed, which is sourced from FER2013, CK+, RAF-DB, and AffectNet respectively. In terms of data preprocessing, all images are grayscaled and the hues of the images are unified. Then, image normalization and histogram equalization are performed, and data augmentation methods are adopted, including random cropping, rotation, horizontal flipping, and color jittering to enhance the diversity of the data.
[0113] It should be noted that the initial population size is set to 100 individuals. The specific generation algorithm uses a two-layer LSTM, with each layer containing 128 hidden units, a sequence length of 20, and an encoding dimension of 32. In terms of training hyperparameter adjustment, the initial learning rate is set to 0.01, which is dynamically adjusted according to the validation accuracy during the training process. The learning rate is decayed every 10 epochs, with a decay rate of 0.1. The batch size is set to 64, and the optimizer used is Adam, with the initial parameters being: a learning rate of 0.01, β1 being 0.9, β2 being 0.999, and epsilon being 1e-8. The crossover probability of the genetic algorithm is 0.8, and the mutation probability is 0.1. The age attribute is implemented through an improved tournament algorithm. The young stage is the first 10 epochs, with a selection probability of 0.6; the middle-aged stage is from 10 to 20 epochs, with a selection probability of 0.3; the old-aged stage is more than 20 epochs, with a selection probability of 0.1. In terms of data preprocessing and augmentation, the image pixel values are normalized to the [0, 1] interval, and data standardization is performed so that the mean pixel value of each image is 0 and the standard deviation is 1. The early stopping strategy terminates the training in advance when the model validation accuracy no longer improves for 3 consecutive epochs to prevent overfitting.
[0114] It should be noted that the designed pre-training strategy combines modular design. In the initial pre-training stage, a basic shallow convolutional network is used, and independent convolutional layers, pooling layers, activation functions, and normalization layers are embedded for training. On this basis, pre-trained convolutional modules, residual modules, dense modules, and Inception modules are gradually added. The initial performance of each module in the facial expression recognition task is obtained through pre-training, the performance of the pre-trained modules is evaluated and sorted, the module with the best performance is selected to narrow the search space. In the narrowed search space, the pre-trained weights are used to accelerate the convergence of the NAS algorithm, and the pre-trained modules are combined to form a complete network structure. Finally, the best network structure is fine-tuned and verified to ensure the best performance of the model in the facial expression recognition task. Finally, a control experiment is designed to compare the accuracy, loss value, and other indicators of the method of the present invention with other methods. As Figure 5 shown, the accuracy of PGONAS rises from about 85% to around 87% within the first 20 epochs, which is better than other models, and breaks through 94% at 60 epochs. In terms of the change in the loss value, as Figure 6 shown, the loss value of PGONAS rapidly drops to around 0.35 within the first 50 epochs and approaches 0.1 at 100 epochs. As Figure 7As shown, the training times required for the method of the present invention to reach accuracies of 93%, 97%, and 97.8% are 13.76 hours, 17.48 hours, and 26.55 hours respectively, all significantly lower than those of other comparative models. In the comparison of inference times, the average inference time per image of the method of the present invention is 0.45 seconds, significantly faster than other models. At the same time, its number of parameters (16.75M) and FLOPs (2.69G) are also lower than those of other models.
[0115] It should be noted that in this application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0116] The above are only specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features claimed herein.
[0117] Embodiment 3
[0118] The third embodiment of the present invention provides a system for a facial expression recognition method based on improved neural network architecture search and optimization, characterized in that it includes an architecture combination definition module, a candidate architecture screening module, a multi-objective optimization evaluation module, a dynamic threshold screening module, and an optimal architecture verification module.
[0119] The architecture combination definition module defines the network architecture combination method suitable for facial expression recognition;
[0120] The candidate architecture screening module inputs facial expression image data into the first model to identify and screen candidate architectures capable of processing facial expression information;
[0121] Multi-objective optimization evaluation module. The selected candidate architectures will be evaluated for multi-objective optimization based on different search spaces and datasets on the first platform to verify their adaptability in various facial expression recognition scenarios;
[0122] Dynamic threshold screening module. According to the architecture performance distribution after evaluation on the first platform, dynamically set the screening threshold to screen architectures with good performance, and use the first algorithm to improve the performance of the screened architectures;
[0123] Optimal architecture verification module. Through comprehensive performance evaluation and optimization, select the optimal neural network architecture and verify it on the test set of the facial expression recognition task to ensure high efficiency and high accuracy. If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0124] Logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0125] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0126] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0127] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A facial expression recognition method based on improved neural network architecture search and optimization, characterized by: The steps include: Define the network architecture combination suitable for facial expression recognition; Inputting facial expression image data into the first model, identifying and screening candidate architectures capable of processing facial expression information; The selected candidate architectures will be evaluated on the first platform through multi-objective optimization based on different search spaces and datasets to verify their adaptability in various facial expression recognition scenarios. According to the architecture performance distribution evaluated by the first platform, a screening threshold is dynamically set to screen architectures with good performance, and the performance of the screened architectures is improved by using the first algorithm; Through comprehensive performance evaluation and optimization, the optimal neural network architecture is selected and verified on the test set of facial expression recognition tasks to ensure high efficiency and accuracy.
2. The facial expression recognition method based on improved neural network architecture search and optimization as claimed in claim 1, characterized in that: The network architecture combination method suitable for facial expression recognition is defined, and the neural network architecture is encoded using path position identification coding to capture multi-path information of the network architecture; The path location identification encoding includes assigning a unique index to each operation node of the neural network.
3. The facial expression recognition method based on improved neural network architecture search and optimization as claimed in claim 1, characterized in that: The dynamically setting screening threshold comprises: Collect the performance data of all candidate network architectures, normalize the performance data, randomly initialize the weight coefficients, calculate the comprehensive score and fitness value for each candidate network structure, and select the architecture based on the comprehensive score.
4. The facial expression recognition method based on improved neural network architecture search and optimization as claimed in claim 1, characterized in that: The first algorithm comprises: A life cycle-based selection strategy was adopted to screen individuals at different life cycle stages for crossover and mutation operations.
5. The facial expression recognition method based on improved neural network architecture search and optimization as claimed in claim 4, characterized in that, The life cycle-based selection strategy includes: According to the life cycle stage of the individual, a dynamic selection probability is set to generate a random number between 0 and 1. If the random number is not greater than the selection probability, the individual is selected for crossover and mutation operations.
6. The facial expression recognition method based on improved neural network architecture search and optimization as claimed in claim 4, characterized in that, The crossover operation includes: Two parent individuals are selected according to the fitness value and a random number is generated. If the random number is less than the predefined crossover probability, the parent network architecture is divided into two parts and the corresponding parts are exchanged to generate child individuals. Otherwise, the selected parent individuals are directly retained as offspring.
7. The facial expression recognition method based on improved neural network architecture search and optimization as claimed in claim 4, characterized in that: The mutation operation includes: A position is randomly selected from the individual's network structure, and a mutation operation is randomly selected for adjustment.
8. A system based on the facial expression recognition method based on improved neural network architecture search and optimization according to any one of claims 1 to 7, characterized in that: It includes architecture combination definition module, candidate architecture screening module, multi-objective optimization evaluation module, dynamic threshold screening module and optimal architecture verification module.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.