Intelligent image processing and real-time identification method for camera module

By combining multi-scale visual coding network, Transformer object detection network, improved ant lion optimization algorithm and taboo search algorithm, real-time preprocessing of camera module image data and multi-scale feature extraction are realized, and the system's robustness and response speed are improved through dynamic parameter tuning, solving the problem of insufficient recognition accuracy and response speed in the existing technology.

CN120219720AInactive Publication Date: 2025-06-27SHENZHEN ZHIYUANDAKE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371577.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In complex and changeable practical application scenarios, existing camera modules have problems such as low recognition accuracy, slow response speed and poor system robustness, especially under the influence of factors such as lighting, background interference, and motion blur.

Method used

A multi-scale visual coding network and a Transformer-based object detection network are adopted, combined with an improved Ant-Lion optimization algorithm and taboo search algorithm, the image data collected by the camera module is preprocessed in real time, multi-scale feature extraction and global feature modeling, and parameter adaptive tuning is performed through dynamic step generation and dynamic double-layer taboo mechanisms.

Benefits of technology

It realizes rapid preprocessing of camera module image data, multi-scale feature extraction, efficient global search of candidate parameters and local fine optimization, which significantly improves the real-time, accuracy and robustness of the system and adapts to changes in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219720A_ABST
    Figure CN120219720A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent image processing and real-time identification method for a camera module, and the method comprises the following steps: S1, collecting original image data, and carrying out the preprocessing of the original image data, and generating a preprocessed image; s2, performing multi-scale feature extraction on the preprocessed image to generate image feature data; s3, performing global feature modeling by using a target detection network based on Transform, and generating a target detection candidate result; s4, setting a parameter search space, randomly generating a candidate parameter combination, and performing fitness evaluation; s5, performing global search through an improved ant lion optimization algorithm; s6, performing local fine optimization by adopting a tabu search algorithm, and determining an optimal parameter combination; and S7, applying the optimal parameter combination to an image preprocessing and target detection module, and periodically executing S4 to S7 to realize dynamic optimization. The objective of the invention is to realize high-efficiency, low-delay and high-robustness image processing and real-time target identification of the camera module in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent image processing and target detection, and in particular, to an intelligent image processing and real-time recognition method for a camera module. Background Art

[0002] The existing image processing and real-time recognition technologies for camera modules mainly rely on traditional image preprocessing and object detection methods based on convolutional neural networks (CNNs). Current technologies usually use fixed image filtering, denoising, and enhancement algorithms to preprocess the acquired raw images to improve image quality, and then detect and identify objects in the images through deep learning networks. Such methods can achieve certain recognition accuracy in static environments and single scenarios, but in complex and changing actual application scenarios, due to the influence of factors such as lighting, background interference, and motion blur, their robustness and real-time performance often cannot be effectively guaranteed. Traditional image processing algorithms have problems such as slow processing speed, insufficient robustness to low-quality images, and difficulty in realizing multi-scale feature extraction. Although traditional CNN models have achieved remarkable results in feature extraction and object recognition, due to the fixed model structure and parameter adjustment relying on a large number of experiments, they often cannot meet the requirements of low latency and high accuracy in applications such as real-time monitoring, autonomous driving, and industrial inspection.

[0003] In addition, the existing technologies mainly use empirical tuning or gradient descent-based methods for model parameter optimization. These methods are prone to falling into local optima when facing complex non-convex optimization problems and cannot dynamically adapt to environmental changes. With the wide application of camera modules in various real-time recognition applications, higher requirements are put forward for the real-time performance and robustness of parameter optimization methods. Traditional parameter optimization methods often require a large number of parameter searches and trainings in an offline environment. Once deployed to an actual system, the fixity of the parameters makes it difficult for the system to adaptively adjust in the face of environmental changes, resulting in a decrease in recognition accuracy and an increase in response time. To address this problem, some swarm intelligence-based optimization algorithms, such as the ant lion optimization algorithm and the tabu search algorithm, have been proposed for global search and local fine-tuning of parameters. However, the existing ant lion optimization algorithm has problems such as excessive randomness and unstable search process in practical applications, while the tabu search algorithm has good refinement ability in local search but lacks the efficiency of global search. It is difficult to balance global search and local exploitation when using these two algorithms alone.

[0004] In the image preprocessing stage, existing technologies mostly use fixed-scale or multi-scale convolutional neural networks for feature extraction. However, this method is insufficient in responding to the diversity of image features in different scenarios and is prone to ignoring the changes in small but crucial information in the image. At the same time, due to the obvious differences in image content at different scales, the fixed multi-scale feature extraction method is difficult to adapt to the diversity and deformability of targets in complex scenarios. To improve the effectiveness of feature extraction, some studies have attempted to introduce the Transformer structure and capture long-range dependencies through the global self-attention mechanism. However, such methods are often limited by computational resources and real-time requirements in practical applications, and they are also highly sensitive to model parameters. Once the parameter settings are unreasonable, it will affect the accuracy and stability of object detection.

[0005] In the object detection and recognition link, existing Transformer-based object detection networks model image features through the global self-attention mechanism, which can better integrate the context information of different regions in the image, thereby improving the accuracy of object detection. However, there are still many deficiencies in parameter tuning for this method. Especially in real-time systems, due to the lack of an adaptive mechanism, traditional parameter tuning methods often cannot quickly respond to environmental changes, resulting in large fluctuations in detection latency and recognition accuracy. Based on this, some researchers have proposed using swarm intelligence optimization algorithms to adaptively adjust the key parameters of object detection networks. However, existing optimization strategies mostly use simple random walk and fixed-step update mechanisms, which makes it difficult for the search process to balance global exploration and local fine optimization, and is prone to falling into local optimal solutions or having an unstable search process.

[0006] Existing technologies have certain limitations in dynamic parameter tuning, mainly manifested in the lack of utilization of image multi-scale feature gradient information and insufficient monitoring of the diversity of candidate parameter combinations. Traditional optimization methods usually ignore the guiding role of image features in parameter updates and fail to incorporate subtle gradient changes in the image into the parameter adjustment mechanism, resulting in low search efficiency and long response latency. In addition, due to the lack of effective regulation of population diversity during the optimization process, candidate parameter combinations tend to concentrate at the initial stage of the search, thus affecting the global search effect and the final robustness of the system. To address this defect, it is necessary to propose a dynamic step size generation method that combines image multi-scale feature gradients, adaptive chaotic perturbation, and population diversity entropy control, and apply it to the parameter optimization of intelligent image processing and real-time object detection and recognition in camera modules to achieve a balance between global fast convergence and local fine search, thereby improving the real-time performance, accuracy, and robustness of the system.

[0007] Therefore, how to provide an intelligent image processing and real-time recognition method for camera modules is an urgent problem that needs to be solved by those skilled in the art. Summary of the Invention

[0008] An object of the present invention is to propose an intelligent image processing and real-time recognition method for a camera module. The present invention makes full use of multi-scale visual coding, a Transformer-based object detection network, and improved ant lion optimization algorithm and tabu search algorithm to adaptively optimize the key parameters of the system, and details the key technical steps of real-time image preprocessing, object detection and recognition, and dynamic parameter optimization, thereby realizing fast preprocessing of the image data collected by the camera, multi-scale feature extraction, global feature modeling, and efficient global search and local fine optimization of candidate parameters. This method has the advantages of high real-time performance, high recognition accuracy, strong adaptability to environmental changes, and excellent system robustness.

[0009] A method for intelligent image processing and real-time recognition of a camera module according to an embodiment of the present invention includes the following steps:

[0010] S1. Collect the original image data, preprocess the original image data to generate a preprocessed image;

[0011] S2. Extract multi-scale features from the preprocessed image to generate image feature data;

[0012] S3. Use a Transformer-based object detection network to perform global feature modeling on the image feature data to generate object detection candidate results;

[0013] S4. Set a parameter search space in the object detection candidate results, randomly generate multiple candidate parameter combinations, and evaluate the fitness of each candidate parameter combination;

[0014] S5. Use the improved ant lion optimization algorithm to perform a global search on the candidate parameter combinations, update the candidate parameter combinations to lock in the high-performance parameter region;

[0015] S6. Use the tabu search algorithm to construct a parameter neighborhood within the locked high-performance parameter region, and perform a local fine search within the neighborhood to determine the optimal parameter combination;

[0016] S7. Apply the optimal parameter combination to update the relevant parameters in the preprocessing of the original image data, the extraction of the image feature data, and the generation of the object detection candidate results, and periodically repeat steps S4 to S7 during the actual operation process.

[0017] Optionally, the S2 specifically includes:

[0018] S21. Perform size normalization processing on the preprocessed image, adjust the width and height of the image to the set standard input resolution, and unify the input scale;

[0019] S22. Input the normalized image into a multi-scale visual encoding network, and use the encoding modules at different levels in the multi-scale visual encoding network to extract the hierarchical features of the image. Among them, the shallow encoding module extracts low-level features such as edges and textures, the middle layer extracts intermediate-level features such as geometric structures and shape contours, and the deep encoding module extracts abstract semantic features related to categories, obtaining feature maps at all levels;

[0020] S23. Introduce a feature pyramid structure based on the feature maps at all levels, and generate multiple scale branches by setting different downsampling rates at different scales to obtain feature maps at each scale;

[0021] S24. Perform an operation to unify the number of channels on the feature maps at each scale, and compress or expand the channels of the feature maps at each scale through a linear transformation module, so that the feature maps at each scale are consistent in the channel dimension;

[0022] S25. Adopt a horizontal connection and progressive upsampling strategy to transfer the deep features to the shallow layer through upsampling operations, and fuse the feature maps at each scale. The fusion methods include weighted summation or feature splicing;

[0023] S26. Perform unified encoding processing on the fused feature maps at each scale, and output image feature data with multi-scale structural information and global context information.

[0024] Optionally, the specific steps of S3 are as follows:

[0025] S31. The image feature data is X ∈ R H×W×C , where H represents the height of the image, W represents the width of the image, C represents the number of channels, and R is the set of real numbers;

[0026] S32. Map the image feature data into a sequence form X' ∈ R N×d through an embedding layer, where N = H × W represents the sequence length and d represents the embedding dimension;

[0027] S33. Model the sequence form with a multi-head self-attention mechanism. For the i-th attention head, calculate the multi-head self-attention:

[0028]

[0029] where, Q i , K i and V i respectively represent the query, key, and value matrices of the i-th attention head, d k is the dimension of each attention head, φ i is the adaptive normalization factor, R i is the relative position encoding matrix of the i-th attention head, α i is the adaptive scaling factor of the i-th attention head, Ai denotes the output result calculated by the i-th attention head, and softmax represents the normalization operation;

[0030] S34. Concatenate the outputs of all attention heads and obtain the global feature representation F through a linear transformation;

[0031] S35. Add the position encoding PE to the global feature representation to obtain the encoded feature F′ = F + PE;

[0032] S36. Input the encoded feature into the target detector module, and respectively extract the coordinates and class probabilities of the target region through the regression and classification sub-modules to generate the target detection candidate results.

[0033] Optionally, the S4 specifically includes:

[0034] S41. Set the parameter search space Ω, where Ω = {θ∣θ = {θ1, θ2, θ3, θ4, θ5}}, where θ1 represents the filter size, θ2 represents the number of layers of the encoding module in the multi-scale visual encoding network, θ3 represents the number of attention heads in the Transformer-based target detection network, θ4 represents the detection box regression parameters in the target detector module, and θ5 represents the non-maximum suppression threshold;

[0035] S42. Randomly generate M candidate parameter combinations within the parameter search space Ω, denoted as θ (j) , where j = 1, 2, …, M;

[0036] S43. Calculate the fitness score η (j) for each candidate parameter combination θ (j) :

[0037]

[0038] where, Acc(θ (j) ) represents the recognition accuracy of the candidate parameter combination θ (j) in the target detection module, T(θ (j) ) represents the processing delay of the candidate parameter combination, T max represents the preset maximum allowable processing delay, Rob(θ (j) ) represents the robustness score of the candidate parameter combination θ (j) in different scenarios, and ω1, ω2, and ω3 are weight coefficients;

[0039] S44. Construct the fitness function G as a mapping function, and for each candidate parameter combination, G(θ (j) ) = η (j) ;

[0040] S45. For all candidate parameter combinations {θ(j)} Sort according to the value of the fitness function G(θ (j) ) to determine the priority of each candidate parameter combination;

[0041] S46. Output the sorted candidate parameter combinations and the corresponding fitness scores {θ (j) , η (j)}.

[0042] Optionally, the S5 specifically includes:

[0043] S51. Construct a set of candidate parameter combinations {θ (j) ∣ j = 1, 2, …, M}, where θ (j) represents the j-th candidate parameter combination, M is the total number of candidate parameter combinations, and the candidate parameter combinations are derived from the parameter search space Ω and include filter size, number of encoding module layers, number of attention heads, detection box regression parameters, and non-maximum suppression threshold;

[0044] S52. Based on the constructed fitness function G(θ (j) ), perform fitness evaluation on each candidate parameter combination to obtain a set of fitness scores {η (j)}, and select the candidate parameter combination with the highest fitness score as the elite individual in the current round of iteration, denoted as θ * :

[0045]

[0046] where, is the index operator for finding the maximum value;

[0047] S53. For each candidate parameter combination θ (j) , obtain the output result of the corresponding image multi-scale feature extraction module, calculate the gradient change amount of the candidate parameter combination in the image feature dimension, and construct a multi-scale feature gradient vector G(θ (j) (t)) as a measure of parameter sensitivity to provide semantic guidance for the step direction;

[0048] S54. Generate a global perturbation factor based on the chaotic map, the initial chaotic variable is c(0), and use the Logistic map function to calculate the chaotic factor c(t):

[0049] c(t) = 4c(t - 1)(1 - c(t - 1));

[0050] where, c(t - 1) represents the chaotic factor generated in the (t - 1)-th iteration;

[0051] S55. Calculate the diversity level of the current population as an adaptive adjustment factor for the search step size, and define the normalized Shannon entropy as the diversity metric function:

[0052]

[0053] where p j represents the normalized occurrence frequency of the j-th candidate parameter combination in the population, ln is the logarithmic function, and D(t) represents the diversity metric value of the entire candidate parameter combination population in the t-th iteration;

[0054] S56. For each candidate parameter combination, a dynamic step size generation method that fuses image multi-scale feature gradients, adaptive chaotic perturbations, and population diversity entropy control is used to calculate the global search step size Δθ (j) (t):

[0055] Δθ (j) (t) = λ·G(θ (j) (t)) + κ·c(t)·D(t)·(θ * - θ (j) (t)) + δ·(2·rand - 1);

[0056] where λ is the feature gradient regulation factor, κ is the chaotic perturbation intensity coefficient, δ is the Gaussian perturbation amplitude factor, rand is a random number obeying a uniform distribution, and θ (j) (t) represents the position of the j-th candidate parameter combination at the t-th iteration;

[0057] S57. According to the global search step size Δθ (j) (t), the position of the j-th candidate parameter combination is updated to obtain a new candidate parameter combination θ (j) (t + 1):

[0058] θ (j) (t + 1) = θ (j) (t) + Δθ (j) (t);

[0059] S58. The updated candidate parameter combination θ (j) (t + 1) is re-evaluated using the fitness function to obtain a new fitness score η (j) (t + 1). If η (j) (t + 1) > η (j) (t), then the updated candidate parameter combination is retained as the current solution; otherwise, the original solution is retained;

[0060] S59. Repeat S53 to S58 until the preset number of iteration rounds or convergence conditions are met, and then output the set of candidate parameter combinations in the current population whose fitness scores are higher than the preset threshold as the locked high-performance parameter region.

[0061] Optionally, the specific content of S6 includes:

[0062] S61. For each candidate parameter combination θ in the locked high-performance parameter region (j) Construct a local parameter set with a dynamically adaptive neighborhood where the radius of the neighborhood ∈ (j) Is dynamically determined according to the fitness change of the candidate parameter combination and the population diversity

[0063]

[0064] where η(θ (j) ) represents the fitness score of the candidate parameter combination θ (j) , represents the average fitness of the current candidate parameter set, μ and ν are preset proportionality coefficients, D(t) is the normalized Shannon entropy of the current candidate parameter set, and tanh is the hyperbolic tangent function;

[0065] S62. Generate local candidate parameter combinations within the local parameter set where r is a random vector uniformly distributed in the interval [-1, 1], and G(θ

[0066]

[0067] ) is the multi-scale image feature gradient information vector corresponding to the candidate parameter combination θ (j) , and ||G(θ (j) )|| is the Euclidean norm; (j) )|| is the Euclidean norm;

[0068] S63. Calculate the fitness score for each parameter combination in the local candidate parameter combination using the fitness function G(θ)

[0069] S64. Select the parameter combination with the highest fitness score in the local candidate parameter combination as the local optimal solution

[0070] S65. Control the local search process using a dynamic double-layer taboo mechanism, and establish a short-term taboo list T s and a long-term memory list T l :

[0071] When a candidate parameter combination θ (j) is selected as the local optimal solution after being updated by local search, it is added to the short-term taboo list T s , and the retention period is L s ;

[0072] For candidate parameter combinations that have been improved continuously but not selected as global elites, add them to the long-term memory list T l, with a retention period of L l ;

[0073] Define the dynamic taboo set as T = T s ∪T l , during the local search process, if the candidate parameter combination exists in the dynamic taboo set, it will be excluded within the retention period until the retention period expires;

[0074] S66. Repeat steps S61 to S65 until the preset local search termination condition is met, and output the candidate parameter combination with the highest fitness score during the local search process as the determined optimal parameter combination.

[0075] The beneficial effects of the present invention are as follows:

[0076] By deeply integrating a multi-scale visual coding network and a Transformer-based object detection network, and combining an improved ant lion optimization algorithm and a taboo search algorithm, the present invention realizes real-time preprocessing of image data collected by a camera module, fine multi-scale feature extraction, and dynamic adaptive tuning of global and local parameters. In the background art, due to the insufficient ability to adapt to image multi-scale information and dynamic environments, existing methods often result in low object detection accuracy, slow response speed, and poor system robustness. The present invention uses a dynamic step size generation method that combines multi-scale image feature gradient information, adaptive chaotic perturbation, and population diversity entropy control to effectively guide the parameter update direction. At the same time, a dynamic double-layer taboo mechanism is used to prevent the search process from falling into local optima, enabling global search and local fine search to work together, significantly improving the stability and response speed of the system in complex environments.

[0077] In addition, during the real-time image processing and object detection process of the present invention, by optimizing the network structure and parameter adjustment strategy, efficient image preprocessing, feature modeling, and candidate parameter combination update are realized, enabling the system to maintain high accuracy and low latency in a dynamic environment. The swarm intelligence optimization strategy adopted not only improves the efficiency of parameter tuning, but also effectively prevents premature convergence by dynamically monitoring population diversity, thereby enhancing the adaptability of the system to environmental changes. Generally speaking, the present invention has obvious advantages in terms of real-time performance, accuracy, robustness, and self-adaptability, overcomes the deficiencies existing in traditional methods, and provides an efficient, stable, and low-latency image processing and object detection solution for fields such as intelligent monitoring, autonomous driving, and industrial inspection. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0079] Figure 1 Flow chart of an intelligent image processing and real-time recognition method for a camera module proposed by the present invention;

[0080] Figure 2 Schematic diagram of the global and local search optimization of candidate parameter combinations by the ant lion optimization algorithm and the tabu search algorithm for an intelligent image processing and real-time recognition method for a camera module proposed by the present invention. Specific implementation manner

[0081] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0082] Refer to Figure 1 and Figure 2 , an intelligent image processing and real-time recognition method for a camera module, comprising the following steps:

[0083] S1. Collect original image data, preprocess the original image data, and generate a preprocessed image;

[0084] S2. Extract multi-scale features from the preprocessed image to generate image feature data;

[0085] S3. Use a Transformer-based object detection network to perform global feature modeling on the image feature data to generate object detection candidate results;

[0086] S4. Set a parameter search space in the object detection candidate results, randomly generate multiple candidate parameter combinations, and evaluate the fitness of each candidate parameter combination;

[0087] S5. Use an improved ant lion optimization algorithm to perform global search on the candidate parameter combinations, and update the candidate parameter combinations to lock in the high-performance parameter region;

[0088] S6. Use the tabu search algorithm to construct a parameter neighborhood within the locked high-performance parameter region, and perform local fine search within the neighborhood to determine the optimal parameter combination;

[0089] S7. Apply the optimal parameter combination to update the relevant parameters in the processes of preprocessing the original image data, extracting the image feature data, and generating the object detection candidate results, and periodically repeat steps S4 to S7 during the actual operation process.

[0090] The present invention integrates a Transformer-based object detection network with an improved swarm intelligence optimization strategy, achieving seamless connection of image preprocessing, feature extraction, and object detection. The improved ant lion optimization algorithm and tabu search algorithm are used to globally and locally dynamically adaptively tune key parameters, thereby significantly improving the real-time performance and recognition accuracy of the camera module in complex environments. This method can not only quickly process and extract multi-scale features of images to achieve global feature modeling, but also effectively reduce the response delay of the system through dynamic parameter optimization, ensuring high-precision object detection effects in harsh scenarios such as low light, strong reflection, and dynamic backgrounds. Compared with traditional fixed-parameter methods, the present invention can adaptively adjust model parameters, avoid falling into local optima, significantly enhance the robustness of the system, and at the same time improve the parameter update speed and local search accuracy, providing a more efficient, stable, and low-latency solution for fields such as intelligent monitoring, autonomous driving, and industrial inspection.

[0091] In this embodiment, step S2 specifically includes:

[0092] S21. Perform size normalization processing on the preprocessed image, adjust the width and height of the image to the set standard input resolution, and unify the input scale;

[0093] S22. Input the normalized image into a multi-scale visual coding network, and use coding modules at different levels in the multi-scale visual coding network to extract hierarchical features of the image. Among them, the shallow coding module extracts low-order features such as edges and textures, the middle layer extracts medium-order features such as geometric structures and shape contours, and the deep coding module extracts abstract semantic features related to categories, obtaining feature maps at all levels;

[0094] S23. Introduce a feature pyramid structure based on the feature maps at all levels, generate multiple scale branches by setting different downsampling rates at different scales, and obtain feature maps at each scale;

[0095] S24. Perform a unified operation on the number of channels of the feature maps at each scale, compress or expand the number of channels of the feature maps at each scale through a linear transformation module, and the feature maps at each scale are consistent in the channel dimension;

[0096] S25. Adopt a horizontal connection and progressive upsampling strategy, transfer the deep features to the shallow layer through upsampling operations, fuse the feature maps at each scale, and the fusion methods include weighted summation or feature splicing;

[0097] S26. Perform unified coding processing on the fused feature maps at each scale, and output image feature data with multi-scale structural information and global context information.

[0098] The multi-scale visual coding network of the present invention performs size normalization on the preprocessed image to ensure that the image has a unified scale before input, laying a solid foundation for subsequent multi-level feature extraction. The normalized image is processed by the shallow, middle, and deep coding modules in the multi-scale visual coding network to extract different levels of features such as edges, textures, geometric structures, shape contours, and abstract semantics, enabling the full decomposition and expression of image information. On this basis, a feature pyramid structure is introduced, and multiple scale branches are generated by setting different downsampling rates to effectively capture the feature information of targets of different sizes in the image. Subsequently, a linear transformation module is used to perform channel unification operations on the feature maps of each scale, ensuring that the multi-scale features are consistent in the channel dimension, thereby avoiding information loss and mismatch problems. Finally, through the horizontal connection and progressive upsampling strategies, the deep semantic information and shallow detail features are efficiently fused to output image feature data with rich multi-scale structural information and global context information. This method significantly improves the expression ability and robustness of image features, provides more accurate and stable data support for the subsequent target detection and recognition modules, greatly enhances the real-time performance and recognition accuracy of the system in complex environments, and provides efficient and reliable technical guarantees for fields such as intelligent monitoring, autonomous driving, and industrial inspection.

[0099] In this embodiment, the specific steps of S3 are as follows:

[0100] S31. The image feature data is X ∈ R H×W×C , where H represents the image height, W represents the image width, C represents the number of channels, and R is the set of real numbers;

[0101] S32. The image feature data is mapped into a sequence form X′ ∈ R N×d through an embedding layer, where N = H × W represents the sequence length and d represents the embedding dimension;

[0102] S33. Perform multi-head self-attention mechanism modeling on the sequence form. For the i-th attention head, calculate the multi-head self-attention:

[0103]

[0104] where, Q i , K i and V i respectively represent the query, key, and value matrices of the i-th attention head, d k is the dimension of each attention head, φ i is the adaptive normalization factor, R i is the relative position encoding matrix of the i-th attention head, α i is the adaptive scaling factor of the i-th attention head, and A idenotes the output result calculated by the i-th attention head, and softmax represents the normalization operation;

[0105] S34. Concatenate the outputs of all attention heads and obtain the global feature representation F through a linear transformation;

[0106] S35. Add the position encoding PE to the global feature representation to obtain the encoded feature F′ = F + PE;

[0107] S36. Input the encoded feature into the target detector module, and respectively extract the coordinates and class probabilities of the target region through the regression and classification sub-modules to generate target detection candidate results.

[0108] The present invention uses a Transformer-based object detection network to perform global feature modeling on image feature data. By converting high-dimensional image feature data into a sequence form and using the multi-head self-attention mechanism to capture key information in the image from multiple perspectives, it realizes the organic integration of local details and global context in the image. In this process, the image feature data is mapped into a sequence through the embedding layer, so that the spatial information and channel information in the original data are retained, and a high-quality input is provided for the subsequent operation of the attention mechanism. Each attention head calculates the query, key, and value matrices respectively, and introduces an adaptive normalization factor and a relative position encoding matrix, further enhancing the sensitivity of the model to spatial position information and structural information in the image, making it more robust when processing image features in complex scenarios. Subsequently, the outputs of all attention heads are concatenated, and a unified global feature representation is obtained through a linear transformation, and then the position encoding is added to fully supplement the position information in the feature. Finally, the coordinates and class probabilities of the target region are respectively extracted through the regression and classification sub-modules to generate target detection candidate results. This method not only greatly improves the accuracy and response speed of object detection, but also can still maintain high recognition stability and adaptive ability in complex backgrounds, multi-scale targets, and dynamically changing scenarios, providing an efficient, stable, and real-time solution for fields such as intelligent monitoring, autonomous driving, and industrial detection.

[0109] In this embodiment, the specific content of S4 includes:

[0110] S41. Set the parameter search space Ω, where Ω = {θ∣θ = {θ1, θ2, θ3, θ4, θ5}}, where θ1 represents the filter size, θ2 represents the number of layers of the encoding module in the multi-scale visual encoding network, θ3 represents the number of attention heads in the Transformer-based object detection network, θ4 represents the detection box regression parameters in the target detector module, and θ5 represents the non-maximum suppression threshold;

[0111] S42. Randomly generate M candidate parameter combinations within the parameter search space Ω, denoted as θ (j) , where j = 1, 2, …, M;

[0112] S43. For each candidate parameter combination θ (j) calculate the fitness score η (j) :

[0113]

[0114] where Acc(θ (j) ) represents the recognition accuracy of the candidate parameter combination θ (j) in the target detection module, T(θ (j) ) represents the processing delay of the candidate parameter combination, T max represents the preset maximum allowable processing delay, Rob(θ (j) ) represents the robustness score of the candidate parameter combination θ (j) under different scenarios, and ω1, ω2, and ω3 are weight coefficients;

[0115] S44. Construct the fitness function G as a mapping function, and for each candidate parameter combination, G(θ (j) ) = η (j) ;

[0116] S45. Sort all the candidate parameter combinations {θ (j)} according to the value of the fitness function G(θ (j) ) to determine the priority of each candidate parameter combination;

[0117] S46. Output the sorted candidate parameter combinations and their corresponding fitness scores {θ (j) , η (j)}.

[0118] By constructing a comprehensive fitness function, the present invention comprehensively evaluates randomly generated candidate parameter combinations within the parameter search space, significantly improving the performance of the target detection module. The fitness function simultaneously considers the recognition accuracy, processing delay, and robustness score of candidate parameter combinations in target detection, and achieves a comprehensive balance of multiple indicators through weighted summation, thereby ensuring that the system can quickly distinguish parameter combinations with better performance during the optimization process. Using this fitness function, the present invention can sort all candidate parameters, determine the priority of each combination, and provide a scientific basis for subsequent parameter update and global search. Through this precise parameter screening, the system can maintain efficient and stable recognition performance in complex and changing environments, significantly reduce the processing delay, improve the response speed, and enhance the system's adaptability to various interference conditions. Thus, the present invention not only solves the problem of fixed parameters and difficult self-adaptive adjustment in traditional methods, but also achieves a breakthrough improvement in target detection accuracy and system robustness, providing reliable and efficient technical support for fields such as intelligent monitoring, autonomous driving, and industrial inspection.

[0119] In this embodiment, step S5 specifically includes:

[0120] S51. Construct a set of candidate parameter combinations {θ (j) ∣j = 1, 2, …, M}, where θ (j) represents the j-th candidate parameter combination, M is the total number of candidate parameter combinations, and the candidate parameter combinations are derived from the parameter search space Ω and include filter size, number of encoding module layers, number of attention heads, detection box regression parameters, and non-maximum suppression threshold;

[0121] S52. Based on the constructed fitness function G(θ (j) ), perform fitness evaluation on each candidate parameter combination to obtain a set of fitness scores {η (j)}, and select the candidate parameter combination with the highest fitness score as the elite individual in the current round of iteration, denoted as θ * :

[0122]

[0123] where is the index operator for finding the maximum value;

[0124] S53. For each candidate parameter combination θ (j) , obtain the output result of the corresponding image multi-scale feature extraction module, calculate the gradient change amount of the candidate parameter combination in the image feature dimension, and construct a multi-scale feature gradient vector G(θ (j) (t)) as a measure of parameter sensitivity to provide semantic guidance for the step direction;

[0125] S54. Generate a global disturbance factor based on chaos mapping. The initial chaos variable is c(0). The Logistic mapping function is used to calculate the chaos factor c(t):

[0126] c(t)=4c(t-1)(1-c(t-1));

[0127] Where c(t-1) represents the chaotic factor generated in the t-1th iteration;

[0128] S55. Calculate the diversity level of the current population as an adaptive adjustment factor for the search step size, and define the standardized Shannon entropy as the diversity measurement function:

[0129]

[0130] Among them, p j represents the normalized frequency of occurrence of the jth candidate parameter combination in the population, ln is a logarithmic function, and D(t) represents the diversity measure of the entire candidate parameter combination population in the tth iteration;

[0131] S56. For each candidate parameter combination, the dynamic step size generation method integrating the image multi-scale feature gradient, adaptive chaotic perturbation and population diversity entropy control is used to calculate the global search step size Δθ (j) (t):

[0132] Δθ (j) (t) = λ·G(θ (j) (t))+κ·c(t)·D(t)·(θ * -θ (j) (t))+δ·(2·rand-1);

[0133] Among them, λ is the characteristic gradient control factor, κ is the chaotic perturbation intensity coefficient, δ is the Gaussian perturbation amplitude factor, rand is a random number that obeys uniform distribution, and θ (j) (t) represents the position of the jth candidate parameter combination at the tth iteration;

[0134] S57, according to the global search step Δθ (j) (t) Perform position update on the jth candidate parameter combination to obtain a new candidate parameter combination θ (j) (t+1):

[0135] θ (j) (t+1)=θ (j) (t)+Δθ (j) (t);

[0136] S58, update the candidate parameter combination θ (j)(t + 1) is re-evaluated using the fitness function to obtain a new fitness score η (j) (t + 1), if η (j) (t + 1) > η (j) (t), then the updated candidate parameter combination is retained as the current solution, otherwise the original solution is retained;

[0137] S59. Repeat steps S53 to S58 until the preset number of iterations or convergence condition is met, and then output the set of candidate parameter combinations in the current population whose fitness scores are higher than the preset threshold as the locked high-performance parameter region.

[0138] The present invention realizes the global search and local fine optimization of the parameter space by constructing a set of candidate parameter combinations and using the comprehensive fitness function to score each parameter combination. Specifically, the present invention first randomly generates multiple candidate parameter combinations within the parameter search space, and conducts a preliminary evaluation based on the recognition accuracy, processing delay, and robustness score in the target detection module, so as to select elite individuals. Subsequently, the system uses the image multi-scale feature gradient information as a measure of parameter sensitivity, and provides semantic guidance for the subsequent generation of dynamic step sizes by calculating the gradient change amount. At the same time, a chaotic map is introduced to generate a global perturbation factor, and the population diversity is quantified by combining the normalized Shannon entropy, effectively regulating the size of the search step. Using the dynamic step size generation method that combines image feature gradient, chaotic perturbation, and diversity entropy control, the system can adaptively adjust the position of each candidate parameter combination during the global search process, so as to quickly jump out of the local optimum. After repeated iterations, position updates, and fitness re-evaluations, the parameter combinations with fitness scores higher than the preset threshold are finally output to form a locked high-performance parameter region. Experimental data shows that this method can improve the search efficiency by about 30% in a complex environment, and significantly improve the target detection recognition accuracy and system robustness, providing an efficient and intelligent parameter optimization solution for real-time image processing and target detection.

[0139] In this embodiment, the specific steps of S6 are as follows:

[0140] S61. For each candidate parameter combination θ in the locked high-performance parameter region (j) Construct a local parameter set with a dynamically adaptive neighborhood where the radius ∈ of the neighborhood (j) is dynamically determined according to the fitness change of the candidate parameter combination and the population diversity:

[0141]

[0142] where η(θ (j) ) represents the fitness score of the candidate parameter combination θ (j) , It represents the average fitness of the current candidate parameter set, μ and ν are preset proportionality coefficients, D(t) is the normalized Shannon entropy of the current candidate parameter set, and tanh is the hyperbolic tangent function;

[0143] S62. Generate local candidate parameter combinations within the said local parameter set where r is a random vector uniformly distributed in the interval [-1, 1], and G(θ

[0144]

[0145] ) is the multi-scale image feature gradient information vector corresponding to the candidate parameter combination θ (j) ), and ||G(θ (j) )|| is the Euclidean norm; (j) )|| is the Euclidean norm;

[0146] S63. Calculate the fitness score for each parameter combination in the local candidate parameter combinations using the fitness function G(θ)

[0147] S64. Select the parameter combination with the highest fitness score within the local candidate parameter combinations as the local optimal solution

[0148] S65. Control the local search process using a dynamic double-layer taboo mechanism, and establish a short-term taboo list T s and a long-term memory list T l :

[0149] When a certain candidate parameter combination θ (j) is selected as the local optimal solution after being updated by local search, add it to the short-term taboo list T s , and the retention period is L s ;

[0150] For candidate parameter combinations that have been improved continuously but not selected as global elites, add them to the long-term memory list T l , and the retention period is L l ;

[0151] Define the dynamic taboo set as T = T s ∪T l , and during the local search process, if a candidate parameter combination exists in the dynamic taboo set, exclude it within the retention period until the retention period expires;

[0152] S66. Repeat steps S61 to S65 until the preset local search termination condition is met, and output the candidate parameter combination with the highest fitness score during the local search process as the determined optimal parameter combination.

[0153] The present invention adopts a dynamic double-layer taboo mechanism to strictly control the local search process, effectively solving the problem that candidate parameter combinations are prone to falling into local optima during the local optimization process. By constructing a dynamic adaptive neighborhood within the locked high-performance parameter region and dynamically determining the neighborhood radius according to the fitness change of candidate parameter combinations and population diversity, it can finely capture the subtle changes of parameters in multi-scale feature extraction of images. At the same time, multi-scale image feature gradient information is introduced during the generation of local candidate parameters, effectively providing semantic guidance for parameter update to ensure that the generated local candidate parameters better meet the requirements of the target detection network. Through fitness evaluation and sorting of local candidate parameters, the system can select the local optimal solution and establish a dynamic taboo set by combining short-term and long-term taboo lists, thereby preventing repeated searches and ineffective iterations of similar parameters. After repeated local searches and parameter updates, the parameter combination with the highest fitness score is finally output, significantly improving the robustness and response speed of the system in complex environments. Experimental data show that this local search optimization method not only effectively reduces the occurrence rate of local optimal traps, but also achieves higher search efficiency and accuracy during the global parameter tuning process, providing stable and accurate parameter support for real-time target detection and greatly improving the overall recognition performance.

[0154] Embodiment 1:

[0155] To verify the feasibility of the present invention in implementation, the present invention is applied to a modern commercial complex located in the core area of the city center, with a building area of more than 20,000 square meters, a complex internal structure, common multi-channel access, and dense crowds. Since the lighting conditions in this area are good during the day, but there are obvious low-light environments at night and in rainy weather, and the crowds come and go frequently with large background dynamic changes, extremely high requirements are put forward for the monitoring system in real-time image preprocessing, target detection and recognition. Traditional monitoring systems generally have problems such as poor real-time performance, high false alarm rates, and insufficient robustness in such scenarios, often resulting in untimely detection of abnormal situations, and using fixed parameters in parameter tuning, which cannot adapt to complex environmental changes, thus affecting the overall recognition effect and safety.

[0156] The present invention proposes an intelligent image processing and real-time recognition method for a camera module. This method uses a high-resolution camera terminal to collect image data in the monitoring area in real time. The collected original images are denoised, enhanced, and filtered through a depth vision preprocessing module to eliminate adverse factors such as low light, strong reflection, and motion blur. Subsequently, the preprocessed images are sent to a multi-scale vision coding network for feature extraction, and then a target detection network based on Transformer performs global feature modeling on the extracted image features to generate target detection candidate results. On this basis, the present invention uses an improved swarm intelligence optimization algorithm to dynamically and adaptively optimize the key parameters of the target detection network. This algorithm realizes the global search and local fine optimization of candidate parameter combinations by fusing image multi-scale feature gradients, adaptive chaotic perturbations, and a dynamic step size generation method controlled by population diversity entropy. The system effectively avoids the problem of local optimal solutions through a dynamic double-layer taboo mechanism and maintains extremely high recognition accuracy and response speed in complex environments. After actual deployment and long-term operation tests, this system can achieve high-precision and low-latency target detection and recognition in different scenarios, significantly reducing the false alarm rate and improving the overall security and robustness of the system.

[0157] To intuitively reflect the improvement effects of the method of the present invention and traditional methods in terms of parameter optimization, real-time response, and recognition accuracy, we conducted tests for three months in multiple monitoring areas of the complex (including the lobby, corridors, entrances and exits, and parking lots). The comparison data records are as follows. The test data cover well-lit environments, low-light environments, dynamic background environments, and high-density pedestrian flow environments. The following table is a comparison data table, recording key performance indicators such as the average detection delay, recognition accuracy, improvement rate of parameter update speed, local search improvement rate, and system robustness index in each test environment. The test locations include four main areas within the business complex, and a total of more than 1000 groups of data samples were collected. The data is true and valid.

[0158] Table 1 Comparison Data of the Performance of the Business Complex Monitoring System

[0159]

[0160]

[0161] As can be seen from Table 1, the method of the present invention exhibits excellent performance in various test environments. In the well-lit hall area, the average detection delay of the system is as low as 40 milliseconds, the recognition accuracy rate is as high as 99.3%, the parameter update speed is increased by 38%, the local search improvement rate reaches 30%, and the system robustness index is as high as 0.95; in the corridor area with poor low-light environment, although the detection delay increases slightly, the recognition accuracy rate still remains above 98.0%, and the improvement effects of parameter update and local search are stable; in the dynamic background environment of the entrance and exit, the system can effectively cope with the interference caused by the rapid movement of people and vehicles and maintain a recognition accuracy rate of 98.5%; in the parking lot entrance with high-density pedestrian flow, despite the large environmental interference and the system detection delay of 52 milliseconds, the overall recognition effect is still satisfactory, and the system robustness index performs excellently. Based on the data of multiple test areas, the average detection delay of the method of the present invention in the mixed environment is 45 milliseconds, the recognition accuracy rate reaches 98.8%, the parameter update speed is increased by 37%, the local search improvement rate is 29%, and the system robustness index reaches 0.94, all of which are significantly better than the traditional fixed-parameter tuning method.

[0162] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.

Claims

1. A camera module intelligent image processing and real-time recognition method, characterized in that: The steps include: S1, collecting original image data, preprocessing the original image data, and generating a preprocessed image; S2, extracting multi-scale features from the preprocessed image to generate image feature data; S3, using the Transformer-based object detection network to perform global feature modeling on image feature data and generate candidate object detection results; S4, setting a parameter search space in the target detection candidate results, and randomly generating multiple candidate parameter combinations, and performing fitness evaluation on each candidate parameter combination; S5. Use the improved ant lion optimization algorithm to perform a global search on the candidate parameter combinations, update the candidate parameter combinations and lock the high-performance parameter area; S6. Use the taboo search algorithm to construct a parameter neighborhood in the locked high-performance parameter region, and perform a local fine search in the neighborhood to determine the optimal parameter combination; S7, applying the optimal parameter combination to update the relevant parameters in the process of preprocessing the original image data, extracting the image feature data, and generating the target detection candidate results, and periodically repeating steps S4 to S7 during the actual operation.

2. The camera module intelligent image processing and real-time recognition method according to claim 1, characterized in that: The S2 specifically includes: S21, performing size normalization processing on the pre-processed image, adjusting the width and height of the image to a set standard input resolution, and unifying the input scale; S22, inputting the normalized image into a multi-scale visual coding network, and using coding modules at different levels in the multi-scale visual coding network to extract hierarchical features of the image, wherein the shallow coding module extracts low-order features of edges and textures, the middle layer extracts mid-order features of geometric structures and shape contours, and the deep coding module extracts abstract semantic features related to categories, thereby obtaining feature maps at all levels; S23, introducing a feature pyramid structure based on feature maps at all levels, generating multiple scale branches by setting downsampling rates of different scales, and obtaining feature maps of each scale; S24, performing a channel number unification operation on the feature maps of each scale, and compressing or expanding the channels of the feature maps of each scale through a linear transformation module, so that the feature maps of each scale are consistent in the channel dimension; S25, adopting the strategy of horizontal connection and step-by-step upsampling, transferring the deep features to the shallow layer through upsampling operation, and fusing the feature maps of each scale, and the fusion method includes weighted summation or feature splicing; S26, uniformly encoding the fused feature maps of each scale, and outputting them as image feature data with multi-scale structural information and global context information.

3. The camera module intelligent image processing and real-time recognition method according to claim 1, characterized in that: The S3 specifically includes: S31, image feature data is X∈R H×W×C , where H represents the image height, W represents the image width, C represents the number of channels, and R is a real number set; S32, map the image feature data into a sequence form X through the embedding layer ′ ∈R N×d , where N = H × W represents the sequence length and d represents the embedding dimension; S33. Model the multi-head self-attention mechanism for the sequence form, where for the i-th attention head, calculate the multi-head self-attention: Among them, Q i , K i and V i denote the query, key, and value matrices of the ith attention head, respectively, and d k is the dimension of each attention head, φ i is the adaptive normalization factor, R i is the relative position encoding matrix of the i-th attention head, α i is the adaptive scaling factor of the ith attention head, A i represents the output result calculated by the i-th attention head, and softmax represents the normalization operation; S34, concatenate the outputs of all attention heads and obtain the global feature representation F through linear transformation; S35. Add position encoding PE to the global feature representation to obtain the encoded feature F ′ =F+PE; S36. Input the encoded features into the target detector module, extract the coordinates and category probability of the target area through the regression and classification submodules respectively, and generate target detection candidate results.

4. The camera module intelligent image processing and real-time recognition method according to claim 1, characterized in that: The S4 specifically includes: S41. Set the parameter search space Ω, where Ω = {θ|θ = {θ1,θ2,θ3,θ4,θ5}}, where θ1 represents the filter size, θ2 represents the number of layers of the encoding module in the multi-scale visual encoding network, θ3 represents the number of attention heads in the Transformer-based object detection network, θ4 represents the detection box regression parameter in the object detector module, and θ5 represents the non-maximum suppression threshold; S42, randomly generate M candidate parameter combinations in the parameter search space Ω, denoted as θ (j) , where j = 1, 2, ..., M; S43, for each candidate parameter combination θ (j) Calculate the fitness score η (j) : Among them, Acc(θ (j) ) represents the candidate parameter combination θ (j) The recognition accuracy in the target detection module, T(θ (j) ) represents the processing delay of the candidate parameter combination, T max represents the preset maximum allowable processing delay, Rob(θ (j) ) represents the candidate parameter combination θ (j) Robustness scores in different scenarios, ω1, ω2 and ω3 are weight coefficients; S44, construct the fitness function G as a mapping function, for each candidate parameter combination, there is G(θ (j) ) = η (j) ; S45. For all candidate parameter combinations {θ (j) According to the fitness function G(θ (j) ) values ​​to determine the priority of each candidate parameter combination; S46, output the sorted candidate parameter combinations and the corresponding fitness scores {θ (j) ,η (j) }.

5. The camera module intelligent image processing and real-time recognition method according to claim 1, characterized in that: The S5 specifically includes: S51, construct a candidate parameter combination set {θ (j) |j=1,2,…,M}, where θ (j) represents the jth candidate parameter combination, M is the total number of candidate parameter combinations, and the candidate parameter combination comes from the parameter search space Ω and includes filter size, number of encoding module layers, number of attention heads, detection box regression parameters and non-maximum suppression threshold; S52, based on the constructed fitness function G(θ (j) ) performs fitness evaluation on each candidate parameter combination to obtain a fitness score set {η (j) }, and select the candidate parameter combination with the highest fitness score in the current iteration as the elite individual, denoted by θ * : in, Indexing operator for maximum value; S53, for each candidate parameter combination θ (j) Obtain the output results of the corresponding image multi-scale feature extraction module, calculate the gradient change of the candidate parameter combination in the image feature dimension, and construct the multi-scale feature gradient vector Gθ (j) (t), as a measure of parameter sensitivity, provides semantic guidance for the step direction; S54. Generate a global disturbance factor based on chaos mapping. The initial chaos variable is c(0). The Logistic mapping function is used to calculate the chaos factor c(t): c(t)=4c(t-1)(1-c(t-1)); Where c(t-1) represents the chaotic factor generated in the t-1th iteration; S55. Calculate the diversity level of the current population as an adaptive adjustment factor for the search step size, and define the standardized Shannon entropy as the diversity measurement function: Among them, p j represents the normalized frequency of occurrence of the jth candidate parameter combination in the population, ln is a logarithmic function, and D(t) represents the diversity measure of the entire candidate parameter combination population in the tth iteration; S56. For each candidate parameter combination, the dynamic step size generation method integrating the image multi-scale feature gradient, adaptive chaotic perturbation and population diversity entropy control is used to calculate the global search step size Δθ (j) (t): Dth (j) (t)=λ·G(θ (j) (t))+κ·c(t)·D(t)·(θ * -θ (j) (t))+δ∈·(2·rand-1); Among them, λ is the characteristic gradient control factor, κ is the chaotic perturbation intensity coefficient, δ is the Gaussian perturbation amplitude factor, rand is a random number that obeys uniform distribution, and θ (j) (t) represents the position of the jth candidate parameter combination at the tth iteration; S57, according to the global search step Δθ (j) (t) Perform position update on the jth candidate parameter combination to obtain a new candidate parameter combination θ (j) (t+1): i (j) (t+1)=θ (j) (t)+Δθ (j) (t); S58, update the candidate parameter combination θ (j) (t+1) Re-evaluate using the fitness function to obtain a new fitness score η (j) (t+1), if η is satisfied (j) (t+1)>η (j) (t), the updated candidate parameter combination is retained as the current solution, otherwise the original solution is retained; S59, repeatedly executing S53 to S58 until a preset number of iterations or convergence conditions are met, and outputting a set of candidate parameter combinations in the current population whose fitness scores are higher than a preset threshold as a locked high-performance parameter region.

6. The camera module intelligent image processing and real-time recognition method according to claim 1, characterized in that: The S6 specifically includes: S61, for each candidate parameter combination θ in the locked high performance parameter region (j) Constructing a local parameter set with a dynamic adaptive neighborhood where the radius of the neighborhood ∈ (j) Dynamic determination based on the fitness changes and population diversity of candidate parameter combinations: Among them, ηθ (j) represents the candidate parameter combination θ (j) The fitness score of the current candidate parameter set is η, μ and ν are preset proportional coefficients, D(t) is the normalized Shannon entropy of the current candidate parameter set, and tanh is the hyperbolic tangent function; S62, in the local parameter set Generate local candidate parameter combinations Among them, r is a random vector uniformly distributed in the interval [-1,1], Gθ (j) is the candidate parameter combination θ (j) The corresponding multi-scale image feature gradient information vector, |Gθ (j) | is the Euclidean norm; S63, using the fitness function G(θ) to calculate the fitness score for each parameter combination in the local candidate parameter combination S64. Select the parameter combination with the highest fitness score from the local candidate parameter combinations as the local optimal solution S65, use a dynamic double-layer taboo mechanism to control the local search process and establish a short-term taboo list T s and long-term memory list T l : When a candidate parameter combination θ (j) When it is selected as the local optimal solution after local search update, it is added to the short-term taboo list T s , the retention period is L s ; For candidate parameter combinations that have been improved multiple times but not selected as global elites, add them to the long-term memory list T l , the retention period is L l ; Define the dynamic taboo set as T = T s ∪T l ,During the local search process, if the candidate parameter combination exists in the ,dynamic taboo set, it will be excluded during the retention period until the ,retention period expires; S66. Repeat steps S61 to S65 until a preset local search termination condition is met, and output the candidate parameter combination with the highest fitness score in the local search process as the determined optimal parameter combination.