Image recognition system and method based on computer vision
By integrating a variety of deep learning technologies in the image recognition system, including ResNet architecture and multi-scale feature fusion, supporting multi-task learning and improved training algorithms, the problem of insufficient recognition performance and generalization capabilities of existing systems in complex scenarios is solved, and efficient, accurate and stable image recognition effects are achieved.
Patent Information
- Application Number
- CN202510295460.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-06-20
AI Technical Summary
The existing deep learning-based image recognition system has problems of lighting, angle changes and occlusion sensitivity when processing complex image data, and it is difficult to adapt to the training needs of large-scale data sets, the generalization ability is limited, it is difficult to maintain stable recognition performance in different scenarios, and it lacks support for multi-task learning.
A computer vision-based image recognition system is adopted, and by integrating a variety of advanced deep learning technologies, including image acquisition, preprocessing, feature extraction, multi-task learning, system processing and training enhancement modules, it realizes efficient, accurate and highly generalized image recognition functions. The system adopts the ResNet architecture combined with multi-scale feature fusion technology for feature extraction, which supports multi-task learning to achieve collaborative optimization through shared feature extraction layers and task-specific branch networks, and accelerates model convergence through improved hybrid accuracy training algorithms and adaptive learning rate adjustment strategies.
It significantly improves the training efficiency and recognition accuracy of the model, enhances the generalization ability of the system, maintains stable recognition performance in a variety of complex scenarios, and supports a variety of image recognition tasks, reduces the consumption of computing resources, and is suitable for a variety of practical application scenarios.
Smart Images

Figure CN120182706A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically provides an image recognition system and method based on computer vision. Background Art
[0002] With the rapid development of computer technology, image recognition technology has been widely used in various fields. Traditional image recognition methods mainly rely on manual feature extraction and simple classification algorithms, such as Support Vector Machine (SVM), decision tree, etc. However, these methods have many limitations when dealing with complex image data. For example, they are sensitive to factors such as image illumination, angle change, occlusion, etc., and it is difficult to meet the training requirements of large-scale data sets.
[0003] In recent years, the rise of deep learning technology has brought new breakthroughs to image recognition. Convolutional Neural Network (CNN) has become the mainstream technology in the field of image recognition due to its powerful feature extraction ability and adaptive learning ability. However, existing deep learning-based image recognition systems still have some problems: on the one hand, model training requires a large amount of labeled data, and the cost of data acquisition and annotation is high; on the other hand, the generalization ability of the model is limited, and it is difficult to maintain stable recognition performance in different scenarios.
[0004] In addition, most existing image recognition systems focus on single tasks, such as object detection or image classification, and lack support for multi-task learning. In practical applications, it is often necessary to complete multiple image recognition tasks simultaneously. For example, in security monitoring, it is necessary to identify the identity of personnel and detect abnormal behaviors at the same time, and existing systems are difficult to meet such complex requirements.
[0005] Therefore, an image recognition system and method based on computer vision are proposed to solve the above problems. Summary of the Invention
[0006] The present invention provides an image recognition system and method based on computer vision, which realizes efficient, accurate and strongly generalized image recognition functions by integrating a variety of advanced deep learning technologies, can support multiple image recognition tasks simultaneously, and maintains stable performance in different scenarios to solve the problems in the background art.
[0007] To achieve the above object, the present invention provides the following technical solution: An image recognition system based on computer vision, comprising: An image acquisition module: used to acquire image data to be recognized, supporting multiple resolutions and formats; An image preprocessing module: used to perform grayscale conversion, normalization, and denoising preprocessing operations on the image; Feature extraction module: used to automatically extract features of images, extract high-level feature representations of images based on deep convolutional neural networks, adopt the ResNet architecture, and combine multi-scale feature fusion technology; Multi-task learning module: supports simultaneous completion of multiple image recognition tasks, including object detection, image classification, and semantic segmentation, and realizes collaborative optimization by sharing feature extraction layers and introducing task-specific branch networks; System processing module: optimizes and corrects the recognition results, including non-maximum suppression and confidence threshold screening, to improve the accuracy and reliability of the recognition results; Training enhancement module, adopts an improved mixed-precision training algorithm, combines an adaptive learning rate adjustment strategy, accelerates model convergence, improves training efficiency, and introduces transfer learning technology to enhance the generalization ability of the model.
[0008] Furthermore, the processing process of the image acquisition module includes the following steps: Image acquisition: Obtain original image data through the image sensor sub-module; Data fusion: Integrate multi-sensor data through the multi-source data fusion sub-module; Real-time calibration: Correct image distortion and illumination deviation through the real-time calibration sub-module; Data caching and transmission: Send the processed image data to the preprocessing module through the data caching and transmission sub-module.
[0009] Furthermore, the image preprocessing module includes: Image input module: used to receive original image data, supporting JPEG, PNG, and BMP formats; Grayscale module: Convert color images to grayscale images; Normalization module: Adjust the pixel values of grayscale images to a specific range; Denoising module: Denoise the normalized images; Output module: Output the preprocessed images to subsequent processing modules or storage devices.
[0010] Furthermore, the ResNet architecture in the feature extraction module includes: Introduce multi-scale convolution operations in the residual modules of ResNet, use convolutional kernels of different sizes to process input feature maps in parallel, and capture image features at different scales; Through the feature pyramid network structure, fuse low-level high-resolution features with high-level semantic features to generate feature representations with rich semantic information and spatial details.
[0011] Furthermore, the feature extraction formula of the feature extraction module is: ; Wherein: F out is the output feature map, F in is the input feature map, W conv is the convolution kernel weight, W scale is the multi-scale convolution weight.
[0012] Furthermore, the multi-task learning module realizes multi-task collaborative optimization in the following way: Define a shared feature extraction network for extracting the general feature representation of the image; For each task, design a task-specific branch network, such as the bounding box regression network for the object detection task and the fully-connected classification network for the image classification task; Introduce a multi-task loss function to jointly optimize the performance of all tasks. The loss function is defined as follows: ; where L total is the total loss function, N is the number of tasks, λ i is the task weight, and L i is the loss function of the i-th task.
[0013] Furthermore, the image classification is based on the extracted features and uses a support vector machine for image classification; the objective function of the SVM classifier: ; Where: w is the weight vector, b is the bias term, C is the penalty parameter, and ξ i is the slack variable.
[0014] Furthermore, the processing process of the system processing module includes the following steps: Non-maximum suppression: Remove redundant detection boxes and retain the optimal results; Confidence threshold screening: Filter out the recognition results with low confidence; Boundary optimization: Optimize the recognition boundary using the CRF algorithm; Result output and visualization: Output the final result and generate a visualization image.
[0015] Furthermore, the training enhancement module adopts the following method: Use 16-bit floating-point numbers for forward and backward propagation calculations to reduce memory occupancy and calculation time. At the same time, introduce dynamic loss scaling technology to avoid gradient underflow; Adopt an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the loss change during the training process. The learning rate adjustment formula is as follows: ; Wherein: η t is the learning rate for the t th iteration, η 0 is the initial learning rate, T is the total number of iterations, p is the decay exponent.
[0016] The computer vision-based image recognition method includes the following steps: Use an image acquisition module to obtain image data to be recognized; Perform grayscale conversion, normalization, and denoising preprocessing operations on the image in the preprocessing module; In the feature extraction module, extract the high-level feature representation of the image through the ResNet architecture; In the multi-task learning module, run multiple image recognition tasks simultaneously, including object detection, image classification, and semantic segmentation; In the system processing module, perform non-maximum suppression (NMS) and confidence threshold screening optimization processing on the recognition results; In the training enhancement module, adopt an improved mixed-precision training algorithm and an adaptive learning rate adjustment strategy, and introduce transfer learning technology to improve the generalization ability of the model.
[0017] Compared with the prior art, the present invention provides a computer vision-based image recognition system and method, having the following beneficial effects: For the computer vision-based image recognition system and method, the present invention adopts a mixed-precision training algorithm and an adaptive learning rate adjustment strategy, significantly improving the training efficiency of the model. In the tests under various complex scenarios, the image recognition of the present invention is high. The multi-task learning module of the present invention can support multiple image recognition tasks simultaneously, and realizes collaborative optimization through a shared feature extraction layer and task-specific branch networks. By optimizing the structure and training process of the deep learning model, the accuracy and efficiency of image recognition are improved, while reducing the consumption of computing resources, being applicable to a variety of actual application scenarios, and having broad market prospects and application values. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 is a schematic flowchart of a computer vision-based image recognition system and method of the present invention; Figure 2Schematic diagram of the image acquisition module of an image recognition system and method based on computer vision according to the present invention; Figure 3 Schematic diagram of the image preprocessing module of an image recognition system and method based on computer vision according to the present invention.
[0020] Figure 4 Schematic diagram of the system processing module of an image recognition system and method based on computer vision according to the present invention. Detailed implementation manners
[0021] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention in conjunction with the accompanying drawings of the specification.
[0022] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0023] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selectively exclusive embodiment from other embodiments.
[0024] Please refer to Figures 1-4 , the present invention discloses an image recognition system based on computer vision, including: Image acquisition module: used to obtain image data to be recognized, supporting multiple resolutions and formats; The processing process of the image acquisition module includes the following steps: Image acquisition: Obtain raw image data through the image sensor sub-module; Data fusion: Integrate multi-sensor data through the multi-source data fusion sub-module; Real-time calibration: Correct image distortion and illumination deviation through the real-time calibration sub-module; Technical advantages: High compatibility: Support image input of multiple resolutions and formats; High robustness: Adapt to complex scenarios through multi-source data fusion and real-time calibration; High efficiency: Adopt cache and transmission technologies to ensure real-time performance; High precision: Improve image quality through geometric distortion correction and illumination correction.
[0025] Data caching and transmission: Send the processed image data to the preprocessing module through the data caching and transmission sub-module.
[0026] The image acquisition module needs to determine the type of data source, which may include but is not limited to the following: Camera devices: such as high-definition cameras, infrared cameras, panoramic cameras, etc., for real-time acquisition of dynamic scene images; Static image files: Read pre-shot image files (such as JPEG, PNG formats) from local storage devices (such as hard disks, USB drives) or network storage; Video stream frame extraction: Extract single-frame images as needed from video files or real-time video streams for subsequent processing; Medical imaging devices: such as X-ray machines, CT scanners, ultrasound devices, etc., for obtaining medical image data.
[0027] Image preprocessing module: Used to perform grayscale conversion, normalization, and denoising preprocessing operations on images to improve the accuracy of subsequent recognition. The image preprocessing module includes: Image input module: Used to receive raw image data, supporting JPEG, PNG, and BMP formats; Grayscale conversion module: Convert color images to grayscale images; Normalization module: Adjust the pixel values of grayscale images to a specific range; Normalization is the process of adjusting image pixel values to a specific range (such as [0, 1] or [-1, 1]), which helps to improve the efficiency and stability of subsequent processing.
[0028] Denoising module: Perform denoising processing on the normalized images, where Gaussian filtering algorithm is used for denoising: ; where σ is the standard deviation of the Gaussian kernel, which determines the filtering intensity, and the normalization operation scales the pixel values to the [0,1] interval.
[0029] Output module: Output the preprocessed images to subsequent processing modules or storage devices.
[0030] Feature extraction module: Used to automatically extract features of images, extract high-level feature representations of images based on deep convolutional neural network (CNN), adopt ResNet architecture, and combine multi-scale feature fusion technology to enhance the robustness of features.
[0031] The ResNet architecture in the feature extraction module includes: Introduce multi-scale convolution operations in the residual modules of ResNet, use convolutional kernels of different sizes to process the input feature maps in parallel, and capture image features at different scales; Through the Feature Pyramid Network (FPN) structure, fuse the low-level high-resolution features with the high-level semantic features to generate feature representations with rich semantic information and spatial details.
[0032] The feature extraction formula of the feature extraction module is: ; Wherein: F out is the output feature map, F in is the input feature map, W conv is the convolutional kernel weight, W scale is the multi-scale convolutional weight.
[0033] The multi-task learning module supports simultaneous completion of multiple image recognition tasks, including object detection, image classification, and semantic segmentation, and realizes collaborative optimization by sharing a feature extraction layer and introducing task-specific branch networks.
[0034] The multi-task learning module realizes multi-task collaborative optimization in the following ways: Define a shared feature extraction network for extracting the general feature representation of images; For each task, design a task-specific branch network, such as the bounding box regression network for object detection tasks and the fully connected classification network for image classification tasks; Introduce a multi-task loss function to jointly optimize the performance of all tasks. The loss function is defined as follows: ; where L total is the total loss function, N is the number of tasks, λ i is the task weight, and L i is the loss function of the i-th task.
[0035] The image classification is based on the extracted features and uses a support vector machine (SVM) for image classification; the objective function of the SVM classifier: ; Where: w is the weight vector, b is the bias term, C is the penalty parameter, and ξ i is the slack variable.
[0036] The system processing module: optimizes and corrects the recognition results, including non-maximum suppression (NMS) and confidence threshold screening, to improve the accuracy and reliability of the recognition results; The processing process of the system processing module includes the following steps: Non-maximum suppression (NMS): removes redundant detection boxes and retains the optimal results; Confidence threshold screening: filters out low-confidence recognition results; Boundary optimization: uses the CRF algorithm to optimize the recognition boundary and identifies the boundary through the energy function. The energy function is defined as: ; Where: is the unary potential function, representing the pixel point Category probability; Is a binary potential function, representing the pixel point and The spatial relationship between; The optimization process of CRF adopts the mean field approximation algorithm, the number of iterations is 5 times, and the convergence threshold is 0.01; Result output and visualization: Output the final result and generate a visualization image.
[0037] The training enhancement module adopts an improved mixed-precision training algorithm, combines an adaptive learning rate adjustment strategy, accelerates model convergence, improves training efficiency, and introduces transfer learning technology to enhance the generalization ability of the model.
[0038] The training enhancement module adopts the following method: Use 16-bit floating-point numbers (FP16) for forward and backward propagation calculations, reduce memory occupancy and calculation time, and at the same time introduce dynamic loss scaling technology to avoid gradient underflow; Adopt an adaptive learning rate adjustment strategy, dynamically adjust the learning rate according to the loss change during the training process, and the learning rate adjustment formula is as follows: ; Where: η t Is the learning rate of the t th iteration, η 0 is the initial learning rate, T Is the total number of iterations, p Is the decay exponent.
[0039] An image recognition method based on computer vision, including the following steps: Use the image acquisition module to obtain the image data to be recognized; Perform grayscale conversion, normalization, and denoising preprocessing operations on the image in the preprocessing module; In the feature extraction module, extract the high-level feature representation of the image through the ResNet architecture; In the multi-task learning module, run multiple image recognition tasks simultaneously, including object detection, image classification, and semantic segmentation; In the system processing module, perform non-maximum suppression (NMS) and confidence threshold screening optimization processing on the recognition results; In the training enhancement module, adopt an improved mixed-precision training algorithm and an adaptive learning rate adjustment strategy, and introduce transfer learning technology to enhance the generalization ability of the model.
[0040] This application mainly consists of an image acquisition module, an image preprocessing module, a feature extraction module, a multi-task learning module, a system processing module, and a training enhancement module. Among them, the image acquisition module obtains image data from sensors and processes it through sub-modules such as multi-source data fusion, real-time calibration, caching, and forwarding; the image preprocessing module performs grayscale conversion, normalization, and denoising operations on the images; the feature extraction module extracts multi-level image features based on the ResNet architecture, combined with the multi-task collaborative optimization mechanism of the multi-task learning module, and simultaneously supports multiple tasks such as object detection, image classification, and semantic segmentation; the system processing module further optimizes the recognition results to improve their accuracy and reliability; the training enhancement module adopts an improved mixed-precision training algorithm and an adaptive learning rate adjustment strategy to accelerate the model training process and improve the generalization ability.
[0041] This application significantly improves the efficiency and accuracy of image recognition, enhances the generalization ability of the system, can adapt to a variety of different image recognition tasks, and provides stable recognition quality in complex scenarios.
[0042] By integrating a variety of advanced technologies, it ensures efficient, accurate, and stable image recognition performance in various scenarios.
[0043] Specifically, through the design of multiple levels such as feature extraction, multi-task learning, mixed-precision training, adaptive learning rate adjustment, and transfer learning, the above technical effects are achieved. The feature extraction module uses the ResNet architecture and multi-scale feature fusion to ensure the effective extraction of high-level features; the multi-task learning module supports multiple image recognition tasks with collaborative optimization; the mixed-precision training and adaptive learning rate adjustment strategies improve the training efficiency and model accuracy; the transfer learning technology enhances the generalization ability.
[0044] This system can be applied to the following scenarios: security monitoring scenarios for personnel identity recognition and abnormal behavior detection; autonomous driving scenarios for traffic sign recognition and obstacle detection; industrial inspection scenarios for product quality inspection and defect recognition; medical imaging analysis scenarios for disease diagnosis and pathological analysis.
[0045] Embodiment 1: Security Monitoring Scenario In the security monitoring scenario, the image recognition system of the present invention can be used for personnel identity recognition and abnormal behavior detection. The specific implementation steps are as follows: Use a high-definition camera to collect image data of the monitoring area.
[0046] In the preprocessing module, perform grayscale conversion and normalization on the images to remove noise.
[0047] In the feature extraction module, extract the high-level feature representation of the images through an improved ResNet architecture.
[0048] In the multi-task learning module, the person identity recognition task and the abnormal behavior detection task are run simultaneously. For the person identity recognition task, a fully-connected classification network is used to output the identity information of the person; for the abnormal behavior detection task, a bounding box regression network is used to detect the area of the abnormal behavior.
[0049] In the system processing module, non-maximum suppression and confidence threshold screening are performed on the recognition results, and the final recognition results are output.
[0050] Embodiment 2: Autonomous driving scenario In the autonomous driving scenario, the image recognition system of the present invention can be used for traffic sign recognition and obstacle detection. The specific implementation steps are as follows: Use an in-vehicle camera to collect road image data.
[0051] In the preprocessing module, the image is de-fogged and normalized to improve the image quality.
[0052] In the feature extraction module, advanced feature representations of the image are extracted through an improved ResNet architecture.
[0053] In the multi-task learning module, the traffic sign recognition task and obstacle detection are run simultaneously.
[0054] The present invention provides an image recognition system and method based on computer vision. Its working principle is to achieve efficient, accurate, and strongly generalization-capable image recognition functions by integrating a variety of advanced deep learning technologies, including modules such as image acquisition, preprocessing, feature extraction, multi-task learning, post-processing, and training and optimization.
[0055] The system first obtains the image data to be recognized through the image acquisition module, and improves the image quality through the real-time calibration and multi-source data fusion sub-module; subsequently, the image preprocessing module performs grayscale conversion, normalization, and denoising on the image to ensure the stability of the input features; the feature extraction module uses the ResNet architecture and multi-scale feature fusion technology to automatically extract the advanced feature representations of the image, improving the recognition accuracy and spatial details of the model; the multi-task learning module supports simultaneous completion of diverse image recognition tasks such as object detection, image classification, and semantic segmentation, and realizes collaborative optimization through a shared feature extraction layer and task-specific branch networks, improving the overall performance of the model; the system processing module optimizes the recognition results using methods such as non-maximum suppression and confidence threshold screening, enhancing the accuracy and reliability of the results; Finally, the training enhancement module adopts an improved mixed-precision training algorithm and an adaptive learning rate adjustment strategy to accelerate model convergence, and introduces transfer learning technology to enhance the generalization ability of the model, ensuring that the system can maintain stable high recognition performance in various complex scenarios. Through the steps and methods of the above work, the present invention significantly improves the accuracy and real-time performance of image recognition, reduces the consumption of computing resources, is applicable to a variety of practical application scenarios, and has broad application value.
[0056] In summary, for the computer vision-based image recognition system and method, the present invention adopts a mixed-precision training algorithm and an adaptive learning rate adjustment strategy, significantly improving the training efficiency of the model. In the tests under various complex scenarios, the image recognition of the present invention is high. The multi-task learning module of the present invention can support multiple image recognition tasks simultaneously, and realizes collaborative optimization through a shared feature extraction layer and task-specific branch networks. By optimizing the structure and training process of the deep learning model, the accuracy and efficiency of image recognition are improved, while the consumption of computing resources is reduced. It is applicable to a variety of practical application scenarios and has broad market prospects and application value.
[0057] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0058] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0059] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with suitable combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0060] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. An image recognition system based on computer vision, characterized in that: include: Image acquisition module: used to obtain image data to be identified, supporting multiple resolutions and formats; Image preprocessing module: used to perform grayscale, normalization, and denoising preprocessing operations on images; Feature extraction module: used to automatically extract image features, extract high-level feature representation of images based on deep convolutional neural networks, using ResNet architecture and combined with multi-scale feature fusion technology; Multi-task learning module: supports multiple image recognition tasks at the same time, including object detection, image classification, and semantic segmentation, and achieves collaborative optimization by sharing feature extraction layers and introducing task-specific branch networks; System processing module: optimizes and corrects the recognition results, including non-maximum suppression and confidence threshold screening, to improve the accuracy and reliability of the recognition results; Training enhancement module: It adopts an improved mixed precision training algorithm and combines it with an adaptive learning rate adjustment strategy to accelerate model convergence and improve training efficiency. It also introduces transfer learning technology to enhance the generalization ability of the model.
2. The computer vision-based image recognition system according to claim 1, characterized in that: The processing process of the image acquisition module includes the following steps: Image acquisition: Obtain raw image data through the image sensor submodule; Data fusion: Integrate multi-sensor data through multi-source data fusion submodule; Real-time calibration: Correct image distortion and illumination deviation through the real-time calibration submodule; Data caching and transmission: The processed image data is sent to the preprocessing module through the data caching and transmission submodule.
3. The computer vision-based image recognition system according to claim 1, characterized in that: The image preprocessing module comprises: Image input module: used to receive raw image data, supporting JPEG, PNG and BMP formats; Grayscale module: convert color images into grayscale images; Normalization module: adjusts the pixel values of the grayscale image to a specific range; Denoising module: denoise the normalized image; Output module: outputs the preprocessed image to the subsequent processing module or storage device.
4. The computer vision-based image recognition system according to claim 1, characterized in that: The ResNet architecture in the feature extraction module includes: Introducing multi-scale convolution operations in the residual module of ResNet, using convolution kernels of different sizes to process input feature maps in parallel and capture image features of different scales; Through the feature pyramid network structure, the high-resolution features of the low layer are fused with the semantic features of the high layer to generate a feature representation with rich semantic information and spatial details.
5. The computer vision-based image recognition system according to claim 4, characterized in that: The feature extraction formula of the feature extraction module is: ; in: F out is the output feature map, F in is the input feature map, W conv is the convolution kernel weight, W scale is the multi-scale convolution weight.
6. The computer vision-based image recognition system according to claim 1, characterized in that: The multi-task learning module achieves multi-task collaborative optimization in the following ways: Define a shared feature extraction network to extract a common feature representation of the image; For each task, design a task-specific branch network, such as a bounding box regression network for object detection tasks and a fully connected classification network for image classification tasks; A multi-task loss function is introduced to jointly optimize the performance of all tasks. The loss function is defined as follows: Among them, L total is the total loss function, N is the number of tasks, λ i is the task weight, L i is the loss function of the i-th task.
7. The computer vision-based image recognition system according to claim 1, characterized in that: The image classification is based on extracting features and using a support vector machine for image classification; the objective function of the SVM classifier is: ; Where: w is the weight vector, b is the bias term, C is the penalty parameter, ξ i is the slack variable.
8. The computer vision-based image recognition system according to claim 1, characterized in that: The processing process of the system processing module includes the following steps: Non-maximum suppression: remove redundant detection boxes and retain the optimal result; Confidence threshold screening: filter low-confidence recognition results; Boundary optimization: CRF algorithm is used to optimize the recognition boundary; Result output and visualization: Output the final results and generate visualization images.
9. The computer vision-based image recognition system according to claim 1, characterized in that: The training enhancement module adopts the following method: Use 16-bit floating point numbers for forward propagation and back propagation calculations to reduce memory usage and calculation time, and introduce dynamic loss scaling technology to avoid gradient underflow; Adopt an adaptive learning rate adjustment strategy to dynamically adjust the learning rate according to the loss changes during training. The learning rate adjustment formula is as follows: ; in: η t For the t The learning rate of the iteration, η 0 is the initial learning rate, T is the total number of iterations, p is the decay index.
10. The computer vision-based image recognition method according to any one of claims 1 to 9, characterized in that: The following steps are involved: Use the image acquisition module to obtain image data to be identified; In the preprocessing module, the image is grayed, normalized, and denoised; In the feature extraction module, the high-level feature representation of the image is extracted through the ResNet architecture; In the multi-task learning module, multiple image recognition tasks are run simultaneously, including object detection, image classification, and semantic segmentation; In the system processing module, the recognition results are subjected to non-maximum suppression (NMS) and confidence threshold screening optimization processing; In the training enhancement module, an improved mixed precision training algorithm and adaptive learning rate adjustment strategy are used, and transfer learning technology is introduced to improve the generalization ability of the model.
Citation Information
Cited By
Image recognition system driven by artificial intelligence and recognition method thereof
CN120747709A