Image recognition system driven by artificial intelligence and recognition method thereof

Through the parallel fusion of multi-layer convolutional neural networks and model adaptation mechanism, the accuracy and robustness problems of existing image recognition systems in complex scenarios are solved, and efficient and secure image recognition is achieved. It is suitable for scenes with mixed multiple types of images and blurred boundaries and has privacy protection capabilities.

CN120747709AActive Publication Date: 2025-10-03SUZHOU RUIXIN INTELLIGENT TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510840306.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-03
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

Existing image recognition systems have problems with low recognition accuracy and poor robustness when dealing with mixed multi-category images, blurred boundaries or low-quality images. In addition, centralized data training can easily lead to privacy leaks and is difficult to meet the needs of complex scenarios.

Method used

An artificial intelligence-driven image recognition system was designed, which adopts a multi-layer convolutional neural network structure, combines visual geometry group network, residual neural network and efficient convolutional neural network to run in parallel, integrates weighted averaging strategy, introduces model adaptation module and federated learning mechanism, and supports edge computing and privacy protection.

Benefits of technology

It improves the accuracy and robustness of image recognition, enhances the system's adaptability and privacy protection, adapts to long-term performance stability in complex scenarios, and meets actual application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747709A_ABST
    Figure CN120747709A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, in particular to an artificial intelligence driven image recognition system and method, and the system comprises an image collection module, a preprocessing module, a feature extraction module, an image recognition module and a result output module. The image acquisition module is used for acquiring input images or video frames in real time; the preprocessing module performs normalization, noise reduction, enhancement and other operations on the acquired image data; the feature extraction module is used for coding and representing key regions of the image through a multilayer neural structure based on a deep convolutional neural network, and extracting multi-scale spatial features; the image recognition module fuses recognition results of a plurality of neural network models based on an integrated learning mechanism, and outputs final classification or recognition information; and the result output module is used for outputting the calculated result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition technology, and in particular to an artificial intelligence-driven image recognition system and a recognition method thereof. Background Art

[0002] With the rapid development of computer vision and artificial intelligence technologies, image recognition, as a core branch of AI, has been widely applied in various industries, including security surveillance, medical diagnosis, intelligent transportation, and industrial inspection. Traditional image recognition methods primarily rely on manually designed feature extractors, which have limitations when dealing with image deformation, noise, or complex background environments. In recent years, the rise of deep learning, particularly convolutional neural networks, has greatly improved the accuracy and generalization of image recognition. Multi-layer feature learning effectively extracts semantic information at different levels in an image, enabling efficient classification and recognition of image objects. However, existing systems mostly rely on a single neural network model, lacking flexibility and adaptability. They also exhibit bias in the recognition of edge cases and multi-class confusion areas, making them unable to meet the needs of diverse and complex real-world scenarios. Furthermore, the computational resource consumption of large models limits their deployment in resource-constrained environments, such as edge devices and mobile devices.

[0003] Therefore, there is an urgent need for a new image recognition system that not only integrates the advantages of multiple neural network models, improving recognition accuracy while also possessing greater robustness and versatility, but also supports optimized discrimination of low-quality images, complex backgrounds, and boundary samples. At the same time, with increasing user awareness of privacy protection, models trained on traditional centralized data are prone to data leakage risks, which has also driven research on privacy-preserving intelligent recognition technologies such as federated learning and adaptive update mechanisms. To achieve the evolution of image recognition systems from "static judgment" to "intelligent adaptation," the system must also possess the ability to self-learn online, optimizing model parameters through continuous feedback training to improve long-term performance stability in specific scenarios. Therefore, designing an AI-driven image recognition system that integrates multiple network structures, possesses high robustness, adaptability, and privacy protection mechanisms, has become a current technological development trend and research focus in the field of intelligent vision.

[0004] In view of the above situation, in order to overcome the above technical problems, the present invention designs an artificial intelligence-driven image recognition system and a recognition method thereof to solve the above technical problems. Summary of the Invention

[0005] The technical objective of this invention is to design an artificial intelligence-driven image recognition system and its recognition method, optimize model parameters through continuous feedback training, and improve long-term performance stability in specific scenarios.

[0006] In order to achieve the above technical objectives, the present invention provides the following technical solutions: This AI-driven image recognition system aims to improve image recognition accuracy, robustness, and adaptability. It is particularly suitable for applications with mixed images, blurred boundaries, or poor image quality. The system primarily comprises an image acquisition module, a preprocessing module, a feature extraction module, an image recognition module, a result output module, and a model adaptation module. These modules work together to form a complete intelligent image recognition process.

[0007] The image acquisition module acquires input image or video frame data in real time and supports a variety of image sources, including industrial cameras, webcams, mobile terminals, and drone acquisition devices. This module features stabilization and adaptive exposure control capabilities to ensure stable image quality. The acquired raw image data is fed into the preprocessing module for image standardization, where it performs normalization, noise reduction, and enhancement. Normalization maps image pixel values ​​to a fixed range to minimize the impact of illumination variations. Noise reduction removes interfering noise from the image, improving the effectiveness of subsequent feature extraction. Enhancement, including contrast adjustment and edge enhancement, enhances image detail.

[0008] The feature extraction module is one of the core modules in the system, employing a multi-layer neural architecture based on a deep convolutional neural network (CNN). This architecture extracts multi-scale spatial features from images, learning layer-by-layer information such as edges, texture, shape, and high-level semantics. By introducing residual connections or attention mechanisms into the network, the system can more accurately represent and encode key areas, providing rich and accurate feature support for subsequent recognition.

[0009] The image recognition module, based on an ensemble learning mechanism, fuses the recognition results of multiple neural network models to output more robust classification or recognition information. Specifically, the image recognition module consists of a primary discriminant network, an auxiliary correction network, and a confidence assessment unit. The primary discriminant network is responsible for making preliminary classification predictions, employing a normalized exponential function (Softmax) to probabilistically output image feature vectors. The auxiliary correction network, primarily targeting image samples with blurred boundaries or low confidence, employs a residual-structured neural network model to correct and optimize preliminary results. The confidence assessment unit, based on the outputs of multiple models, uses a weighted voting mechanism or a confidence fusion algorithm to modify the final recognition results, improving the overall system stability and classification credibility.

[0010] The system also features a model adaptation module that dynamically adjusts the recognition strategy based on the real-time distribution characteristics of the input image. This module automatically adjusts the discrimination threshold by analyzing the category weights and sample distribution of the image input. It then optimizes model parameters and adjusts the structure specifically for boundary samples or low-confidence categories, thereby enhancing the system's robustness and generalization capabilities.

[0011] The result output module is responsible for displaying or outputting the final recognition results to the terminal system, which may include text label output, coordinate position return or image segmentation result presentation, and supports access to the host computer system, database or other application platforms.

[0012] The preprocessing module includes an image normalization unit, a histogram equalization unit, and an edge enhancement unit. The three work together to optimize the quality of the original image and enhance key features, providing a good input basis for subsequent deep neural network processing. Among them, the image normalization unit is responsible for linearly mapping the pixel values ​​of the input image to a numerical range of [0,1] to reduce the dynamic differences in the image caused by changes in lighting, uneven exposure, or different imaging devices, and to improve the generalization ability of the model. The histogram equalization unit performs contrast enhancement processing on the image based on the grayscale distribution characteristics of the image. By adjusting the distribution of pixel values, the details of the dark and bright areas in the image are made clearer, making it easier for the subsequent network to extract more discriminative mid- and high-level semantic features. The edge enhancement unit uses the Laplace operator to extract the edges of the image, and strengthens the contour and structural information in the image by responding to the second-order derivative, effectively highlighting the boundaries and key contours of objects in the image. In particular, the edge-enhanced image will be input as a feature channel into the first layer of the deep convolutional neural network, and will participate in the initial convolution operation together with the original image channel, thereby enhancing the model's perception of local details and edge information, and helping to improve recognition effects in complex backgrounds or low-contrast scenes.

[0013] The main discriminant network in the image recognition module utilizes three mainstream deep neural network structures: the Visual Geometry Group Network (VGG), the Residual Neural Network (ResNet), and the Efficient Convolutional Neural Network (EfficientNet), running in parallel. This multi-model collaboration improves recognition accuracy and robustness. Each network independently receives the same image feature input and completes classification predictions. The system dynamically assigns weights to each network in the final fusion of results based on its accuracy performance in cross-validation during the training phase. The specific fusion method uses an integrated weighted average strategy, which weights and superimposes the probability distributions output by the three models to obtain more representative and stable classification probabilities. The final recognition output takes the category label with the largest weighted probability value as the system's recognition result, effectively alleviating the problem of a single model misjudging marginal samples and improving the overall system's adaptability and judgment capabilities in complex image environments.

[0014] The image recognition results obtained by the result output module have a wide range of application adaptability and can be directly applied to multiple practical scenarios such as medical image recognition (such as tumor detection and lesion localization), intelligent traffic monitoring (such as license plate recognition and violation analysis), and face recognition (such as identity authentication and behavior tracking). The system supports loading dedicated data sets built for different industries and fine-tuning the model in the field by combining transfer learning technology, thereby achieving rapid adaptation to the recognition requirements of specific tasks. At the same time, the module integrates a multi-task learning mechanism to perform target detection, image segmentation, and image classification tasks in parallel within a unified deep neural network structure. This not only improves overall computational efficiency, but also significantly enhances the system's multifunctional recognition capabilities and its ability to understand complex image scenes, providing a highly flexible and scalable image processing solution for practical deployment.

[0015] The auxiliary correction network uses the attention mechanism to assign weights to the intermediate feature maps to highlight the salient features of the target area. The attention mechanism is defined as: Among them, Q, K, V are query, key, and value vectors, and dk is the dimension; this mechanism is used to enhance the ability to distinguish boundary samples.

[0016] The confidence evaluation unit outputs a confidence score for each recognition result. The confidence calculation method is: Where zi is the output value of the i-th category. If the confidence of all categories is less than the set threshold T, the manual review interface is triggered.

[0017] The model adaptation module consists of a classifier dynamic adjustment unit and a training data resampling unit, which aims to solve the problems of category imbalance and overfitting. The classifier dynamic adjustment unit introduces a temperature scaling algorithm in the inference stage to adjust the output of the normalized exponential function (Softmax). By adjusting the temperature coefficient T, the model's excessive confidence in the main frequency category is reduced, thereby improving the ability to distinguish long-tail category samples and alleviating the prediction bias phenomenon. At the same time, the training data resampling unit analyzes the density distribution of samples in the feature space, upsamples low-frequency samples, downsamples or weighted samples high-frequency samples, and automatically generates new training batches to improve the model's generalization ability on multi-category data and effectively avoid the occurrence of overfitting. This module supports dynamic linkage between the training and inference stages, enhancing the system's adaptability and stability in the face of changes in data distribution in practical applications.

[0018] The entire system is deployed on edge computing devices to fully leverage the real-time and data security advantages of edge computing. Specifically, the image acquisition and preprocessing modules are integrated into the front-end acquisition devices, enabling real-time image acquisition and preliminary preprocessing, such as normalization, noise reduction, and enhancement, to be completed directly on the device side, reducing data transmission volume and improving data quality. The feature extraction and image recognition modules are deployed on the edge computing unit using lightweight neural network models. These models have been optimized through techniques such as pruning and quantization, and have low computing resource requirements, ensuring that the system can still operate efficiently under limited hardware conditions. This architecture supports autonomous recognition functions in offline environments while achieving low-latency real-time image processing, meeting the stringent requirements of fast response and local processing for applications such as intelligent monitoring, unmanned driving, and mobile devices. The distributed deployment of edge computing also effectively protects the privacy and security of user data, avoids uploading large amounts of sensitive image data to the cloud, and enhances the practicality and reliability of the system.

[0019] The system supports a model update mechanism based on federated learning to achieve distributed collaborative training and protect user privacy. Specifically, each terminal device independently performs model training locally, updates model parameters using its own collected image data, and uploads the gradient information or model weight differences calculated during the training process to the central server. The central server securely aggregates the gradient information from each terminal to generate the latest parameters of the global model, and then broadcasts this global model back to all terminal devices to achieve unified model updates. This process does not require uploading the original image data, effectively avoiding the risk of user privacy leakage and complying with data protection regulations. At the same time, the federated learning mechanism significantly improves the generalization and adaptability of the model by integrating the diverse data features of multiple terminals, ensuring that the system can maintain a high recognition accuracy in different application scenarios. This mechanism is particularly suitable for application in environments with dispersed and highly sensitive data, such as medical image recognition and intelligent monitoring, and takes into account the dual needs of privacy protection and model performance.

[0020] An artificial intelligence-driven image recognition method is used in conjunction with the artificial intelligence-driven image recognition system described above; the method is characterized in that the steps of the method are as follows: S1: Acquires raw image data or video frame data through the image acquisition module, supports static image and dynamic frame input, and image formats include standard formats such as JPEG, PNG or YUV; S2: normalize, grayscale equalize, and edge enhance the acquired image. The normalization method is linear mapping to the interval [0, 1]. The enhancement method includes the Laplace operator. S3: Use a deep convolutional neural network to perform multi-layer convolution and pooling operations on the preprocessed image to extract multi-scale semantic features. The convolution structure uses residual connections, and the output of the middle layer of the network is used as the input feature φ for subsequent recognition; S4: The extracted features are input into the main discriminant network and the auxiliary correction network for preliminary judgment and correction respectively. The main network uses softmax for classification prediction, and the auxiliary network uses the attention mechanism to optimize the boundary samples. Finally, the category label with the highest confidence is output through the integrated weighted average method. S5: The final recognition result and its confidence level are output. If the confidence level is lower than the preset threshold T, the review mechanism is triggered. At the same time, the input sample features are recorded for adaptive module training to dynamically update the model weights and achieve continuous optimization of recognition accuracy.

[0021] The beneficial effects of the present invention are as follows: (1) The present invention effectively improves the accuracy and robustness of image classification and recognition by introducing an image recognition system with parallel fusion of multiple network structures. The three mainstream deep neural networks, Visual Geometry Group Network (VGG), Residual Neural Network (ResNet) and Efficient Convolutional Neural Network (EfficientNet), are used to run in parallel, combined with a dynamic weighted fusion strategy to achieve the complementary advantages of different networks and significantly reduce the misjudgment and bias problems that are prone to occur in a single model. At the same time, the design of the auxiliary correction network and confidence assessment unit further optimizes the ability to distinguish boundary samples and low-confidence results, making the system more stable and adaptable when processing complex and diverse images. The system's multi-level, multi-angle feature extraction and discrimination method greatly improves the ability to capture and express image details, effectively improves the accuracy and reliability of recognition, and meets the demand for high-performance image recognition in practical applications.

[0022] (2) The present invention also has significant advantages in adaptability and privacy protection. The model adaptation module realizes dynamic adjustment of category distribution and sample structure, enhances the recognition effect of long-tail categories and edge samples, avoids overfitting, and improves the generalization ability and stability of the model. Combined with the federated learning mechanism, this system supports local training of terminal devices and model aggregation updates of central servers, effectively protecting the security of user privacy data and avoiding the risk of leakage of sensitive information. This distributed training architecture not only improves the real-time update capability of the system, but also adapts to the diverse needs of different terminals and scenarios, greatly enhancing the practicality and promotion value of the system. In summary, the present invention provides an artificial intelligence-driven image recognition solution with superior performance, flexibility and security, which has broad application prospects and important industrial promotion significance. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0024] The above and other aspects of the present invention will now be described, by way of example only, with reference to the accompanying drawings, in which: Figure 1 It is a schematic diagram of the system flow of the present invention; Figure 2 It is a step diagram of the method of the present invention. DETAILED DESCRIPTION

[0025] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0026] like Figure 1-2 As shown in Figure 1, an AI-driven image recognition system aims to improve image recognition accuracy, robustness, and adaptability. It is particularly suitable for applications with mixed images, blurred boundaries, or poor image quality. The system primarily includes an image acquisition module, a preprocessing module, a feature extraction module, an image recognition module, a result output module, and a model adaptation module. These modules work together to form a complete intelligent image recognition process.

[0027] The image acquisition module acquires input image or video frame data in real time and supports a variety of image sources, including industrial cameras, webcams, mobile terminals, and drone acquisition devices. This module features stabilization and adaptive exposure control capabilities to ensure stable image quality. The acquired raw image data is fed into the preprocessing module for image standardization, where it performs normalization, noise reduction, and enhancement. Normalization maps image pixel values ​​to a fixed range to minimize the impact of illumination variations. Noise reduction removes interfering noise from the image, improving the effectiveness of subsequent feature extraction. Enhancement, including contrast adjustment and edge enhancement, enhances image detail.

[0028] The feature extraction module is one of the core modules in the system, employing a multi-layer neural architecture based on a deep convolutional neural network (CNN). This architecture extracts multi-scale spatial features from images, learning layer-by-layer information such as edges, texture, shape, and high-level semantics. By introducing residual connections or attention mechanisms into the network, the system can more accurately represent and encode key areas, providing rich and accurate feature support for subsequent recognition.

[0029] The image recognition module, based on an ensemble learning mechanism, fuses the recognition results of multiple neural network models to output more robust classification or recognition information. Specifically, the image recognition module consists of a primary discriminant network, an auxiliary correction network, and a confidence assessment unit. The primary discriminant network is responsible for making preliminary classification predictions, employing a normalized exponential function (Softmax) to probabilistically output image feature vectors. The auxiliary correction network, primarily targeting image samples with blurred boundaries or low confidence, employs a residual-structured neural network model to correct and optimize preliminary results. The confidence assessment unit, based on the outputs of multiple models, uses a weighted voting mechanism or a confidence fusion algorithm to modify the final recognition results, improving the overall system stability and classification credibility.

[0030] The system also features a model adaptation module that dynamically adjusts the recognition strategy based on the real-time distribution characteristics of the input image. This module automatically adjusts the discrimination threshold by analyzing the category weights and sample distribution of the image input. It then optimizes model parameters and adjusts the structure specifically for boundary samples or low-confidence categories, thereby enhancing the system's robustness and generalization capabilities.

[0031] The result output module is responsible for displaying or outputting the final recognition results to the terminal system, which may include text label output, coordinate position return or image segmentation result presentation, and supports access to the host computer system, database or other application platforms.

[0032] The preprocessing module includes an image normalization unit, a histogram equalization unit, and an edge enhancement unit. The three work together to optimize the quality of the original image and enhance key features, providing a good input basis for subsequent deep neural network processing. Among them, the image normalization unit is responsible for linearly mapping the pixel values ​​of the input image to a numerical range of [0,1] to reduce the dynamic differences in the image caused by changes in lighting, uneven exposure, or different imaging devices, and to improve the generalization ability of the model. The histogram equalization unit performs contrast enhancement processing on the image based on the grayscale distribution characteristics of the image. By adjusting the distribution of pixel values, the details of the dark and bright areas in the image are made clearer, making it easier for the subsequent network to extract more discriminative mid- and high-level semantic features. The edge enhancement unit uses the Laplace operator to extract the edges of the image, and strengthens the contour and structural information in the image by responding to the second-order derivative, effectively highlighting the boundaries and key contours of objects in the image. In particular, the edge-enhanced image will be input as a feature channel into the first layer of the deep convolutional neural network, and will participate in the initial convolution operation together with the original image channel, thereby enhancing the model's perception of local details and edge information, and helping to improve recognition effects in complex backgrounds or low-contrast scenes.

[0033] The main discriminant network in the image recognition module utilizes three mainstream deep neural network structures: the Visual Geometry Group Network (VGG), the Residual Neural Network (ResNet), and the Efficient Convolutional Neural Network (EfficientNet), running in parallel. This multi-model collaboration improves recognition accuracy and robustness. Each network independently receives the same image feature input and completes classification predictions. The system dynamically assigns weights to each network in the final fusion of results based on its accuracy performance in cross-validation during the training phase. The specific fusion method uses an integrated weighted average strategy, which weights and superimposes the probability distributions output by the three models to obtain more representative and stable classification probabilities. The final recognition output takes the category label with the largest weighted probability value as the system's recognition result, effectively alleviating the problem of a single model misjudging marginal samples and improving the overall system's adaptability and judgment capabilities in complex image environments.

[0034] The image recognition results obtained by the result output module have a wide range of application adaptability and can be directly applied to multiple practical scenarios such as medical image recognition (such as tumor detection and lesion localization), intelligent traffic monitoring (such as license plate recognition and violation analysis), and face recognition (such as identity authentication and behavior tracking). The system supports loading dedicated data sets built for different industries and fine-tuning the model in the field by combining transfer learning technology, thereby achieving rapid adaptation to the recognition requirements of specific tasks. At the same time, the module integrates a multi-task learning mechanism to perform target detection, image segmentation, and image classification tasks in parallel within a unified deep neural network structure. This not only improves overall computational efficiency, but also significantly enhances the system's multifunctional recognition capabilities and its ability to understand complex image scenes, providing a highly flexible and scalable image processing solution for practical deployment.

[0035] The auxiliary correction network uses the attention mechanism to assign weights to the intermediate feature maps to highlight the salient features of the target area. The attention mechanism is defined as: Among them, Q, K, V are query, key, and value vectors, and dk is the dimension; this mechanism is used to enhance the ability to distinguish boundary samples.

[0036] The confidence evaluation unit outputs a confidence score for each recognition result. The confidence calculation method is: Where zi is the output value of the i-th category. If the confidence of all categories is less than the set threshold T, the manual review interface is triggered.

[0037] The model adaptation module consists of a classifier dynamic adjustment unit and a training data resampling unit, which aims to solve the problems of category imbalance and overfitting. The classifier dynamic adjustment unit introduces a temperature scaling algorithm in the inference stage to adjust the output of the normalized exponential function (Softmax). By adjusting the temperature coefficient T, the model's excessive confidence in the main frequency category is reduced, thereby improving the ability to distinguish long-tail category samples and alleviating the prediction bias phenomenon. At the same time, the training data resampling unit analyzes the density distribution of samples in the feature space, upsamples low-frequency samples, downsamples or weighted samples high-frequency samples, and automatically generates new training batches to improve the model's generalization ability on multi-category data and effectively avoid the occurrence of overfitting. This module supports dynamic linkage between the training and inference stages, enhancing the system's adaptability and stability in the face of changes in data distribution in practical applications.

[0038] The entire system is deployed on edge computing devices to fully leverage the real-time and data security advantages of edge computing. Specifically, the image acquisition and preprocessing modules are integrated into the front-end acquisition devices, enabling real-time image acquisition and preliminary preprocessing, such as normalization, noise reduction, and enhancement, to be completed directly on the device side, reducing data transmission volume and improving data quality. The feature extraction and image recognition modules are deployed on the edge computing unit using lightweight neural network models. These models have been optimized through techniques such as pruning and quantization, and have low computing resource requirements, ensuring that the system can still operate efficiently under limited hardware conditions. This architecture supports autonomous recognition functions in offline environments while achieving low-latency real-time image processing, meeting the stringent requirements of fast response and local processing for applications such as intelligent monitoring, unmanned driving, and mobile devices. The distributed deployment of edge computing also effectively protects the privacy and security of user data, avoids uploading large amounts of sensitive image data to the cloud, and enhances the practicality and reliability of the system.

[0039] The system supports a model update mechanism based on federated learning to achieve distributed collaborative training and protect user privacy. Specifically, each terminal device independently performs model training locally, updates model parameters using its own collected image data, and uploads the gradient information or model weight differences calculated during the training process to the central server. The central server securely aggregates the gradient information from each terminal to generate the latest parameters of the global model, and then broadcasts this global model back to all terminal devices to achieve unified model updates. This process does not require uploading the original image data, effectively avoiding the risk of user privacy leakage and complying with data protection regulations. At the same time, the federated learning mechanism significantly improves the generalization and adaptability of the model by integrating the diverse data features of multiple terminals, ensuring that the system can maintain a high recognition accuracy in different application scenarios. This mechanism is particularly suitable for application in environments with dispersed and highly sensitive data, such as medical image recognition and intelligent monitoring, and takes into account the dual needs of privacy protection and model performance.

[0040] An artificial intelligence-driven image recognition method is used in conjunction with the artificial intelligence-driven image recognition system described above; the method is characterized in that the steps of the method are as follows: S1: Acquires raw image data or video frame data through the image acquisition module, supports static image and dynamic frame input, and image formats include standard formats such as JPEG, PNG or YUV; S2: normalize, grayscale equalize, and edge enhance the acquired image. The normalization method is linear mapping to the interval [0, 1]. The enhancement method includes the Laplace operator. S3: Use a deep convolutional neural network to perform multi-layer convolution and pooling operations on the preprocessed image to extract multi-scale semantic features. The convolution structure uses residual connections, and the output of the middle layer of the network is used as the input feature φ for subsequent recognition; S4: The extracted features are input into the main discriminant network and the auxiliary correction network for preliminary judgment and correction respectively. The main network uses softmax for classification prediction, and the auxiliary network uses the attention mechanism to optimize the boundary samples. Finally, the category label with the highest confidence is output through the integrated weighted average method. S5: The final recognition result and its confidence level are output. If the confidence level is lower than the preset threshold T, the review mechanism is triggered. At the same time, the input sample features are recorded for adaptive module training to dynamically update the model weights and achieve continuous optimization of recognition accuracy.

[0041] Various modifications to the present disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but should be given the widest scope consistent with the principles and novel features disclosed herein. Although one or more exemplary embodiments of the present disclosure have been described with reference to the accompanying drawings, it will be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure as defined in the appended claims.

Claims

1. An artificial intelligence driven image recognition system, characterized in that: It includes image acquisition module, preprocessing module, feature extraction module, image recognition module and result output module; Image acquisition module: real-time acquisition of input images or video frames; The preprocessing module performs normalization, noise reduction, enhancement and other operations on the collected image data; Feature extraction module: This module encodes and represents key image regions through a multi-layer neural structure based on a deep convolutional neural network, extracting multi-scale spatial features. The image recognition module fuses the recognition results of multiple neural network models based on an integrated learning mechanism and outputs the final classification or recognition information; The image recognition module specifically includes: a main discriminant network, an auxiliary correction network, and a confidence assessment unit. The main discriminant network uses a normalized exponential function to perform preliminary classification predictions, the auxiliary correction network uses a residual structure to perform secondary judgments on edge samples, and the confidence assessment unit corrects the final output results based on a weighted voting mechanism. The system further includes a model adaptation module for dynamically adjusting the recognition threshold according to the category distribution of the input image to improve the recognition accuracy and stability of edge samples. Result output module: outputs the calculated results.

2. The artificial intelligence-driven image recognition system according to claim 1, characterized in that: The preprocessing module includes an image normalization unit, a histogram equalization unit, and an edge enhancement unit, wherein the image normalization unit maps pixel values ​​to the range of [0, 1], the equalization unit performs contrast enhancement based on the grayscale distribution of the image, and the edge enhancement unit uses the Laplacian operator to extract edge features; The edge enhancement processing result participates in the first layer convolution operation of the subsequent deep convolutional neural network to improve the recognition ability of local details.

3. The artificial intelligence-driven image recognition system according to claim 1, characterized in that: The main discriminant network in the image recognition module uses three structures: visual geometry group network, residual neural network and efficient convolutional neural network to run in parallel. Its classification output is fused through integrated weighted averaging. The fusion weight is dynamically adjusted according to the cross-validation accuracy during training. The final recognition output is the category with the maximum probability of weighted voting.

4. The artificial intelligence-driven image recognition system according to claim 1, characterized in that: The results obtained in the result output module can be applied to multiple scenarios such as medical image recognition, traffic monitoring, and face recognition. By configuring specific data sets for domain fine-tuning, it supports a multi-task learning mechanism and integrates target detection, image segmentation, and classification tasks into a unified network structure, thereby enhancing the system's multifunctional recognition capabilities.

5. The artificial intelligence-driven image recognition system according to claim 1, characterized in that: The auxiliary correction network uses the attention mechanism to assign weights to the intermediate feature maps to highlight the salient features of the target area. The attention mechanism is defined as: Where Q, K, V are query, key, and value vectors, d k This mechanism is used to enhance the discrimination ability of boundary samples.

6. The artificial intelligence-driven image recognition system according to claim 1, characterized in that: The confidence evaluation unit outputs a confidence score for each recognition result. The confidence calculation method is: where z i is the output value of the i-th category. If the confidence of all categories is less than the set threshold T, the manual review interface is triggered.

7. The artificial intelligence-driven image recognition system according to claim 1, characterized in that: The model adaptation module includes a classifier dynamic adjustment unit and a training data resampling unit. The dynamic adjustment unit uses a temperature scaling algorithm to correct the normalized exponential function output according to the category distribution to improve the discrimination of long-tail samples; the resampling unit generates new training batches according to the sample density to avoid overfitting of the model to high-frequency categories.

8. The artificial intelligence-driven image recognition system according to claim 1, characterized in that: The system is deployed on edge computing devices, the image acquisition module and preprocessing module are embedded in the front-end device, and the feature extraction and recognition modules are deployed on the edge computing unit through a lightweight model, supporting offline recognition and low-latency real-time processing.

9. The artificial intelligence-driven image recognition system according to claim 1, characterized in that: The system supports a model update mechanism based on federated learning. Each terminal device trains model parameters locally and uploads gradient information to the central server, which aggregates and broadcasts the global model. This mechanism protects user privacy and ensures model generalization capabilities.

10. An artificial intelligence-driven image recognition method, the method being used in conjunction with an artificial intelligence-driven image recognition system according to any one of claims 1 to 9; characterized in that: The steps of the method are as follows: S1: Acquires raw image data or video frame data through the image acquisition module, supports static image and dynamic frame input, and image formats include standard formats such as JPEG, PNG or YUV; S2: normalize, grayscale equalize, and edge enhance the acquired image. The normalization method is linear mapping to the interval [0, 1]. The enhancement method includes the Laplace operator. S3: Use a deep convolutional neural network to perform multi-layer convolution and pooling operations on the preprocessed image to extract multi-scale semantic features. The convolution structure uses residual connections, and the output of the middle layer of the network is used as the input feature φ for subsequent recognition; S4: The extracted features are input into the main discriminant network and the auxiliary correction network for preliminary judgment and correction respectively. The main network uses softmax for classification prediction, and the auxiliary network uses the attention mechanism to optimize the boundary samples. Finally, the category label with the highest confidence is output through the integrated weighted average method. S5: The final recognition result and its confidence level are output. If the confidence level is lower than the preset threshold T, the review mechanism is triggered. At the same time, the input sample features are recorded for adaptive module training to dynamically update the model weights and achieve continuous optimization of recognition accuracy.

Citation Information

Patent Citations

  • Unmanned aerial vehicle aerial image target detection method based on adaptive model integration

    CN113313058A

  • Defect detection method and device

    CN116385387A

  • Focus image segmentation method and device, electronic equipment and readable storage medium

    CN116883432A

  • Model training method and device, object recognition method and device and electronic equipment

    CN118781471A

  • Image processing method for enhancing image recognition

    CN119274044A