Image recognition system and method based on deep learning
Through the deep learning image recognition system, combined with data reception, segmentation and recognition modules, the problem of low image recognition accuracy in complex environments is solved, and high-precision target object recognition is achieved.
Patent Information
- Application Number
- CN202510342279.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In complex environments, such as when lighting conditions are drastically changed, occlusion or image blur, existing image recognition technologies based on deep learning are difficult to accurately identify target objects, and the recognition accuracy is difficult to meet actual needs.
The image recognition system based on deep learning is adopted, including a data reception module, an image segmentation module and an image recognition module. Through data preprocessing, deep learning semantic segmentation algorithm and hierarchical fusion model, high-level semantic features are extracted and high-precision recognition results are generated.
It improves the accuracy and efficiency of image recognition, and can accurately identify target objects in complex environments. It is suitable for scenarios such as intelligent security monitoring and medical diagnosis.
Smart Images

Figure CN120259662A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning, and particularly relates to an image recognition system and method based on deep learning. Background Art
[0002] In today's digital age, deep learning technology is widely used in the field of image recognition, driving the intelligent transformation of many industries. At present, the image recognition technology methods based on deep learning mainly focus on convolutional neural networks (CNNs), recurrent neural networks (RNNs) and their variants, etc. However, in complex environments, such as when the lighting conditions change drastically, there are occlusions or the images are blurred, the accuracy of image recognition will drop significantly. In surveillance scenarios at night or under direct strong light, it is very difficult for a CNN-based object detection model to accurately identify the target object; current methods often require complex calculations and a large amount of training data, and the recognition accuracy is difficult to meet the actual needs. Summary of the Invention
[0003] Based on this, it is necessary to provide an image recognition system based on deep learning for the above technical problems, which can understand the semantic information of images and improve the accuracy of image recognition.
[0004] In a first aspect, the present application provides an image recognition system based on deep learning, including a data receiving module, an image segmentation module and an image recognition module;
[0005] The data receiving module is used to preprocess the acquired original image to obtain a target image to be recognized;
[0006] The image segmentation module is used to perform semantic segmentation on the target image based on a deep learning semantic segmentation algorithm to obtain multiple morphemes, where the attributes of the morphemes include coordinate information and class labels;
[0007] The image recognition module is used to obtain high-level semantic features based on the morphemes and generate an image recognition result based on the high-level semantic features.
[0008] In one embodiment, the image segmentation module includes an instance extraction unit and a hierarchical feature extraction unit:
[0009] The instance extraction unit is used to segment instances in various class labels based on the morphemes, and the attributes of the instances include attributes inherited from the morphemes and instance labels;
[0010] The hierarchical feature extraction unit is used to obtain primary features and intermediate features, perform primary feature extraction on each instance, and the primary features include category, location, and visual attributes; obtain intermediate features between multiple instances, and the intermediate features include interaction relationships and / or topological relationships; the primary features and intermediate features are used to extract high-level semantic features.
[0011] In one of the embodiments, the image recognition module includes an advanced semantic feature fusion unit:
[0012] The advanced semantic fusion unit is used to input the primary features and intermediate features into a hierarchical fusion model based on deep learning to obtain advanced semantic features; the advanced semantic features include at least one of object relationship features, scene semantic features, and action behavior features.
[0013] Furthermore, the hierarchical fusion model based on deep learning is a Transformer model under a hierarchical fusion strategy, and the Transformer model under the hierarchical fusion strategy is trained using the following method:
[0014] Establish a connection with the training image library and obtain training images, perform feature transformation on the training images to obtain sequence features, and perform hierarchical fusion calculation on the sequence features using the following formula:
[0015]
[0016] where X l is the input of the l-th layer of the Transformer, and l i is a preset layer index;
[0017] Perform multi-head self-attention calculation on the input X l that has fused hierarchical features, using the following formula:
[0018]
[0019] MHSA(Q, K, V) = Concat(head1, head2, …, head h )W 0
[0020]
[0021]
[0022] where is the output of the multi-head self-attention, head i is the attention head, Concat(head1, …, head h ) is the splicing function of multiple attention heads, W 0 is the splicing integration linear transformation matrix, is the projection matrix projected into the query space, is the projection matrix projected into the key space, is the projection matrix projected into the value space, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the feature dimension;
[0023] Using the following formula, the output of the multi-head self-attention is input into the feed-forward network to obtain the feed-forward network output:
[0024]
[0025] FFN(x) = ReLU(xW1 + b1)W2 + b2
[0026] where is the feed-forward network output, FFN is a two-layer fully connected network, W1 is the weight matrix of the first-layer fully connected network, b1 is the bias vector of the first-layer fully connected network, ReLU is the activation function, W2 is the weight matrix of the second-layer fully connected network, and b2 is the bias vector of the second-layer fully connected network;
[0027] Perform layer normalization on the feed-forward network output, and use the normalized result as the input of the next layer of the Transformer. Repeat the above steps for calculation until the preset model convergence condition is reached; where the calculation formula for layer normalization is as follows:
[0028]
[0029] where LayerNorm is the layer normalization function.
[0030] In a possible embodiment, the image segmentation module further includes a timing acquisition unit and a dynamic recognition unit:
[0031] The timing acquisition unit is used to acquire the time information of the target image, sort the target images according to the time sequence, and set multiple target images that are continuous and the interval time does not exceed the preset value as a continuous frame image group;
[0032] The dynamic recognition unit is used to continuously detect multiple target images in the continuous frame image group, obtain dynamic features and generate events based on the dynamic features, and the events are used to indicate the image recognition result.
[0033] Preferably, the data reception module includes a dataset expansion sub-module:
[0034] The dataset expansion sub-module is used to perform real-time generalization enhancement on the acquired original image based on the generative adversarial network architecture to obtain an enhanced image;
[0035] where the generative adversarial network architecture includes a generator and a discriminator. The generator is used to learn the feature distribution of the original image to generate a simulated image;
[0036] The discriminator is used to distinguish between the original image and the simulated image;
[0037] The simulated image is used for semantic segmentation.
[0038] Preferably, the dataset augmentation sub-module includes an image inpainting unit for:
[0039] Perform an integrity check on the acquired original image to obtain the incomplete region;
[0040] If the proportion of the incomplete region in the area of the whole image is within a preset threshold range, use a neural network to perform image inpainting on the original image to obtain a repaired image; wherein, the repaired image is used for semantic segmentation.
[0041] In a second aspect, the present application also provides an image recognition method based on deep learning, including:
[0042] Preprocess the acquired original image to obtain a target image to be recognized;
[0043] Perform semantic segmentation on the target image based on a deep learning semantic segmentation algorithm to obtain multiple morphemes, wherein the attributes of the morphemes include coordinate information and class labels;
[0044] Obtain high-level semantic features based on the morphemes, and generate an image recognition result based on the high-level semantic features.
[0045] In a third aspect, the present application also provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the above-mentioned image recognition method based on deep learning is implemented.
[0046] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method as described above is implemented.
[0047] The above-mentioned image recognition system and method based on deep learning integrate a data reception module, an image segmentation module and an image recognition module. The data reception module preprocesses the original image to make it a standardized target image to be recognized, providing a high-quality data basis for image processing. The image segmentation module uses a deep learning semantic segmentation algorithm to segment the target image into multiple morphemes with coordinate information and class labels, accurately parsing the image semantics. The image recognition module extracts high-level semantic features based on these morphemes, covering information such as object relationships and scene semantics, so as to generate a high-precision recognition result. The cooperation of each module improves the recognition efficiency and recognition accuracy of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 FIG. is a schematic structural diagram of an image recognition system based on deep learning provided by an embodiment of the present invention;
[0050] Figure 2 FIG. is a schematic flowchart of an image recognition method based on deep learning provided by an embodiment of the present invention. Detailed Embodiments
[0051] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0052] First, a brief introduction to the nouns involved in the embodiments of the present application is given.
[0053] Deep learning refers to machine learning based on deep neural network models and methods. It has evolved on the basis of algorithm models such as statistical machine learning and artificial neural networks, combined with the development of contemporary big data and high computing power. The most important technical feature of deep learning is its ability to automatically extract features, and the extracted features are also called deep features or deep feature representations. Compared with manually designed features, deep features have stronger and more robust representation capabilities.
[0054] Semantic segmentation is a computer vision task that can identify and understand the content of each pixel in an image through a deep learning model, and the annotation of its semantic regions is at the pixel level. The core of semantic segmentation lies in marking the semantic categories of each pixel in the image and classifying them into specific semantic categories.
[0055] Image Enhancement is an image processing technology whose main purpose is to improve the visual effect of an image or enhance the recognition ability of target features in the image through specific algorithms and methods.
[0056] According to the above-mentioned noun explanations, the implementation environment of an image recognition system based on deep learning provided by the embodiments of this application is described. Schematically, this implementation environment includes: an image acquisition device, a processor, and a storage device. Among them, the image acquisition device can be an industrial camera, a surveillance camera, a mobile phone camera, a CT scanner, an MRI (Magnetic Resonance Imaging) device, etc.; the processor includes but is not limited to a central processing unit (CPU), a graphics processing unit (GPU), an artificial intelligence chip, etc.; the storage device can be a distributed storage system or a centralized storage, which is not limited here.
[0057] Combined with the above-mentioned noun explanations and the implementation environment, the application scenarios of the embodiments of this application are described.
[0058] The image recognition system based on deep learning provided by the embodiments of this application can be applied to the following scenarios including but not limited to:
[0059] In the intelligent security monitoring scenario, the camera, as the data acquisition hardware, continuously collects the image data of the monitoring area for 24 hours and transmits it to the data receiving module. The image is preprocessed such as denoising and size adjustment to be transformed into a target image suitable for subsequent processing. The image segmentation module, based on the deep learning semantic segmentation algorithm, segments people, vehicles, suspicious objects, etc. in the target image into different morphemes, and marks the coordinate information and class labels. The image recognition module extracts the high-level semantic features of these morphemes, such as the behavior actions (running, wandering) of people, the driving direction and speed of vehicles, etc. The system can judge in real time whether there are abnormal behaviors. Once an abnormality is found, such as someone breaking into a restricted area or a vehicle going in the wrong direction, an alarm is immediately issued.
[0060] In the field of medical diagnosis, medical imaging equipment collects the medical images of patients. The data receiving module preprocesses the original medical images, such as denoising and normalization, to make them meet the subsequent processing standards. The image segmentation module segments organs, tissues, lesions, etc. in the medical image to form morphemes with coordinates and class labels. The image recognition module, by extracting high-level semantic features, assists doctors in judging whether the organs are normal, whether there are lesions, and the type and severity of the lesions. For example, in the recognition of lung CT images, the system can identify nodules in the lungs and analyze their size, shape, density and other features to provide a diagnostic reference for doctors.
[0061] Schematically, the image recognition system based on deep learning provided by the embodiments of this application can also be applied to other application scenarios. Only examples are given here, and the specific application scenarios are not limited.
[0062] In an exemplary embodiment, such as Figure 1As shown, an image recognition system 10 based on deep learning is provided. Taking the application of this system to the aforementioned processor as an example for illustration, it can be understood that this system can also be implemented through the interaction between the processor of the image acquisition device and other processors / controllers / servers. In this embodiment, the system includes the following modules:
[0063] A data receiving module 11, which is used to preprocess the acquired original image to obtain a target image to be recognized.
[0064] Specifically, the data receiving module 11 can obtain the original image from multiple data sources, such as real-time capture by a camera, reading from a storage device, or receiving through network transmission, and perform a series of preprocessing operations. For an image with noise, algorithms such as Gaussian filtering and median filtering are used to remove the noise and make the image clearer; for images with inconsistent resolutions, the image scaling algorithm is used to adjust them to a unified resolution to ensure the consistency of subsequent processing; if the brightness and contrast of the image are not good, methods such as histogram equalization and contrast stretching are used for optimization to enhance the detail information of the image. After these preprocessing steps, a target image to be recognized suitable for subsequent processing is obtained.
[0065] An image segmentation module 12, which is used to perform semantic segmentation on the target image based on the deep learning semantic segmentation algorithm to obtain multiple morphemes, where the attributes of the morphemes include coordinate information and class labels.
[0066] Specifically, the image segmentation module 12 takes the deep learning semantic segmentation algorithm as the core. In the training stage of the algorithm, a large number of labeled image data are used to train the semantic segmentation model. These labeled data contain the category information of each pixel. By learning these data, the model gradually masters the feature representations of different objects and scenes in the image. For example, semantic segmentation models such as U-Net and DeepLab series are used to automatically extract the features of the image using convolutional neural networks. In the inference stage, the target image preprocessed by the data receiving module is input into the trained semantic segmentation model. The model classifies each pixel in the image, divides the pixels with the same category into one region to form multiple morphemes, and assigns coordinate information and class labels to each morpheme to accurately identify different objects and regions in the image.
[0067] An image recognition module 13, which is used to obtain high-level semantic features based on the morphemes and generate an image recognition result based on the high-level semantic features.
[0068] Specifically, the image recognition module 13 extracts high-level semantic features based on the morphemes obtained by the image segmentation module 12. Based on these high-level semantic features, classification is performed through a classifier to generate the final image recognition result, and judge the category to which the image belongs or identify specific target objects, actions, or scenarios in the image.
[0069] The above deep learning-based image recognition system integrates a data reception module, an image segmentation module, and an image recognition module. The data reception module preprocesses the original image to make it a standardized target image to be recognized, providing a high-quality data basis for image processing. The image segmentation module uses a deep learning semantic segmentation algorithm to segment the target image into multiple morphemes with coordinate information and class labels, accurately parsing the image semantics. The image recognition module extracts high-level semantic features based on these morphemes, covering information such as object relationships and scene semantics, to generate high-precision recognition results. The cooperation of each module improves the recognition efficiency and accuracy of this system.
[0070] In one embodiment, the image segmentation module 12 may include an instance extraction unit 121 and a hierarchical feature extraction unit 122.
[0071] The instance extraction unit 121 is used to segment instances based on morphemes in various class labels. The attributes of the instances include attributes inherited from morphemes and instance labels.
[0072] Specifically, when processing an image, all morphemes and their corresponding class labels are traversed. For each class label, the instance extraction unit 121 analyzes the set of morphemes with the same class label. For example, in a natural scene image, if there are multiple morphemes labeled as "tree", the instance extraction unit will use the coordinate information of these morphemes and apply a clustering algorithm (such as the DBSCAN algorithm) or rules based on spatial proximity to merge adjacent morphemes belonging to the same class label into one instance. During the merging process, the instance inherits the coordinate information and class label of these morphemes as its own attributes, and assigns a unique instance label to each instance to distinguish different instances, thus completing the segmentation work from morphemes to instances.
[0073] The hierarchical feature extraction unit 122 is used to obtain primary features and intermediate features, extract primary features for each instance, where the primary features include category, location, and visual attributes; obtain intermediate features between multiple instances, where the intermediate features include interaction relationships and / or topological relationships; the primary features and intermediate features are used to extract high-level semantic features.
[0074] Specifically, for each instance, the hierarchical feature extraction unit extracts its primary features respectively. In terms of category feature extraction, the morpheme class labels inherited by the instance are used to determine the category of the instance; when extracting location features, the statistical information of all morpheme coordinates in the instance (such as the center coordinate, bounding box coordinate, etc.) is calculated to accurately represent the position of the instance in the image; for visual attribute features, a convolutional neural network (CNN) is used to extract features from the image region corresponding to the instance. The instance image region is input into a specific layer of a pre-trained CNN model to obtain the feature vectors output by this layer. These vectors contain visual information such as color and texture, which are used as the visual attribute features of the instance. When obtaining the intermediate features of multiple instances, for the interaction relationship, it is determined by analyzing the position and spatial relationship between different instances. For example, if an instance of "person" is spatially close to an instance of "car" and there is a relative movement trend, it may be judged that there is an interaction relationship of "approaching" or "avoiding" between them; for the topological relationship, a spatial graph model between instances is constructed, with instances as nodes and the spatial relationships between instances as edges, and graph algorithms (such as the shortest path algorithm, graph convolutional network, etc.) are used to mine the topological structure information between instances, such as whether an instance is surrounded by other instances, or whether there is a connected path between instances, so as to obtain intermediate features.
[0075] In one embodiment, the image recognition module 13 includes a high-level semantic feature fusion unit 131.
[0076] The high-level semantic fusion unit 131 is used to input the primary features and intermediate features into a hierarchical fusion model based on deep learning to obtain high-level semantic features; the high-level semantic features include at least one of object relationship features, scene semantic features, and action behavior features.
[0077] Specifically, the primary features and intermediate features are obtained from the hierarchical feature extraction unit of the image segmentation module 12 and input into a hierarchical fusion model based on deep learning, so that the generated high-level semantic features can more comprehensively and deeply reflect the semantic content of the image, providing a more accurate basis for image recognition.
[0078] Furthermore, the hierarchical fusion model based on deep learning is a Transformer model under a hierarchical fusion strategy, and the Transformer model under the hierarchical fusion strategy is trained using the following method:
[0079] Establish a connection with the training image library and obtain training images, perform feature transformation on the training images to obtain sequence features, and use the following formula to perform hierarchical fusion calculation on the sequence features.
[0080]
[0081] Where, X lis the input to the l-th layer of the Transformer, where l i is a pre-set layer index.
[0082] Perform multi-head self-attention calculation on the input X that fuses hierarchical features l using the following formula:
[0083]
[0084] MHSA(Q, K, V) = Concat(head1, head2, …, head h )W 0
[0085]
[0086] where, is the output of multi-head self-attention, head i is the attention head, Concat(head1, …, head h ) is the concatenation function of multiple attention heads, W 0 is the concatenation integration linear transformation matrix, is the projection matrix projected into the query space, is the projection matrix projected into the key space, is the projection matrix projected into the value space, Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the feature dimension.
[0087] Use the following formula to input the multi-head self-attention output into the feed-forward network to obtain the feed-forward network output:
[0088]
[0089] FFN(x) = ReLU(xW1 + b1)W2 + b2
[0090] where, is the feed-forward network output, FFN is a two-layer fully connected network, W1 is the weight matrix of the first layer of the fully connected network, b1 is the bias vector of the first layer of the fully connected network, ReLU is the activation function, W2 is the weight matrix of the second layer of the fully connected network, and b2 is the bias vector of the second layer of the fully connected network.
[0091] Perform layer normalization on the feed-forward network output and use the normalized result as the input to the next layer of the Transformer, repeating the above steps for calculation until the preset model convergence condition is reached; where, the calculation formula for layer normalization is as follows:
[0092]
[0093] Among them, LayerNorm is a layer normalization function.
[0094] Exemplarily, the hierarchical fusion strategy allows the model to fuse features at different levels. It can capture the local detailed features of the image through shallow fusion and integrate the global semantic features through deep fusion. When processing complex scene images, it can effectively combine different scales and types of information in the image, improve the model's understanding ability of the image content, and thus improve the accuracy of image recognition. Specifically, the multi-head self-attention mechanism enables the model to capture the dependencies between features in parallel from different subspaces. Through the collaborative work of multiple attention heads, it can simultaneously focus on different aspects of the image, such as the shape, position, texture, etc. of the object, and can capture the relationships between features more comprehensively and meticulously, enhancing the model's ability to model complex image structures. The layer normalization operation normalizes the output of the feed-forward network, making the distribution of the input data of each layer more stable during the training process of the model, alleviating the problems of gradient disappearance and gradient explosion, accelerating the model convergence, and improving the training efficiency. Even when the training data has large differences or the number of model layers is large, layer normalization can ensure the stable training of the model. By learning a large number of training images and combining data augmentation techniques to expand the dataset, the model can learn rich image features and variation rules. When facing unseen images, it can make reasonable inferences based on the learned knowledge, accurately identify the image content, and has strong generalization ability, suitable for various complex and changeable actual application scenarios.
[0095] In a possible embodiment, the image segmentation module 12 further includes a timing acquisition unit 123 and a dynamic recognition unit 124.
[0096] The timing acquisition unit 123 is configured to acquire the time information of the target image, sort the target images according to the time sequence, and set multiple target images that are continuous and the time interval between them does not exceed a preset value as a continuous frame image group.
[0097] Specifically, when the data receiving module 11 transfers the preprocessed target image to the image segmentation module 12, the timing acquisition unit extracts the time information from the metadata of the image. For example, if the image is captured in real time by a camera, the camera device usually adds a timestamp to each frame of the image, and this timestamp records the specific moment when the image is taken; for images read from a storage device, the relevant time information such as the creation or modification time of the image saved by the storage system is extracted, and the images are sorted according to the time sequence. Then, the sorted target images are traversed to check the time interval between adjacent images. If the time interval between adjacent images does not exceed the preset value, they are divided into the same continuous frame image group.
[0098] A dynamic recognition unit 124 is configured to continuously detect multiple target images within a group of consecutive frame images, obtain dynamic features, and generate events based on the dynamic features, where the events are used to indicate the image recognition results.
[0099] Specifically, for each group of consecutive frame images, the dynamic recognition unit 124 continuously detects multiple target images within the group, applies object detection algorithms such as YOLO, Faster R-CNN, etc., to identify target objects of interest in each frame image, and records information such as their positions and categories. By comparing the changes in the positions, sizes, postures, etc. of the same target object in consecutive frames, dynamic features are obtained. For example, for a moving vehicle, dynamic features such as its displacement, speed, and direction change between adjacent frames are calculated. Based on the obtained dynamic features, combined with preset rules or models, corresponding events are generated. Exemplarily, if it is detected that a person moves from one side of the screen to the other side in consecutive frames and at a relatively fast speed, an event of "person moving quickly" may be generated; if it is detected that a vehicle stays in a certain area for more than a preset duration, an event of "vehicle staying for a long time" is generated. These events are used to indicate the image recognition results and provide a basis for subsequent analysis and decision-making.
[0100] Preferably, the data receiving module 11 includes a data set expansion sub-module 111.
[0101] The data set expansion sub-module 111 is configured to perform real-time generalization enhancement on the obtained original images based on a generative adversarial network architecture to obtain enhanced images; wherein, the generative adversarial network architecture includes a generator and a discriminator, the generator is used to learn the feature distribution of the original images to generate simulated images; the discriminator is used to distinguish between the original images and the simulated images; and the simulated images are used for semantic segmentation.
[0102] Exemplarily, the dataset augmentation sub-module 111 constructs a generative adversarial network architecture including a generator and a discriminator. The generator is usually composed of a series of convolutional layers, deconvolutional layers, fully connected layers, etc., and is used to convert a random noise vector into a simulated image. The discriminator is a binary classifier, generally composed of convolutional layers and fully connected layers, and is used to judge whether the input image is an original image or a simulated image. The obtained original image is input into the generative adversarial network. At the same time, a random noise vector is provided for the generator, and the dimension and distribution of the noise vector are set according to the specific network design. The generator starts to learn the feature distribution of the original image. It attempts to generate a simulated image similar to the original image in appearance, texture, structure, etc. by performing a series of transformation operations on the random noise vector. After multiple iterations and adjustments, the generator outputs a simulated image and inputs it together with the original image into the discriminator. The discriminator classifies according to the input image and calculates the classification error. The generator and the discriminator perform adversarial training, and through continuous adversarial training, the performance of the generator and the discriminator is improved. The original image and the generated simulated image are merged to obtain an enhanced image for subsequent semantic segmentation.
[0103] Preferably, the dataset augmentation sub-module 111 includes an image inpainting unit 1110 for:
[0104] Performing an integrity check on the obtained original image to obtain the incomplete region.
[0105] If the proportion of the incomplete region in the area of the whole image is within the preset threshold range, use a neural network to perform image inpainting on the original image to obtain an inpainted image; wherein, the inpainted image is used for semantic segmentation.
[0106] Specifically, multiple methods are used to check the integrity of the original image. For example, the edge information of the image is detected through an edge detection algorithm. If the edges are discontinuous or missing, it may indicate that the image is incomplete. Or the features such as the color and texture of the image are used for analysis, and the feature distribution of the normal image is compared to determine whether there are abnormal regions. If an incomplete image is found, algorithms such as connected component analysis are used to locate the incomplete region. By marking and segmenting the parts that are significantly different from the normal region, the position and scope of the incomplete region are determined, and its coordinate information is recorded, and the proportion of the incomplete region in the total area of the whole image is calculated. The calculated proportion of the incomplete region in the total area of the whole image is compared with the preset threshold range. The threshold range is preset according to specific application scenarios and requirements, for example, set to 5%-30%. If the proportion of the incomplete region in the total area of the whole image is within this threshold range, it is considered that the image can be repaired; if the proportion is too small, it may have little impact on subsequent semantic segmentation and no repair is required; if the proportion is too large, the repair effect may be poor and no repair is carried out either. According to the characteristics of the image and the repair requirements, a suitable neural network model is selected. For example, a repair model based on a convolutional neural network (CNN), such as Context Encoder, which learns the features of the image through an encoder-decoder structure and can fill in the incomplete region according to the context information of the image. The repaired image can more completely present the object and scene information in the image. When the semantic segmentation model processes these images, it can more accurately identify and segment different objects and regions. In the semantic segmentation of natural scene images, the repaired image can avoid the situation of incorrect or missed segmentation of objects caused by the incomplete part, improving the accuracy and integrity of the segmentation result.
[0107] In summary, the image recognition system based on deep learning provided by the embodiments of the present application integrates a data receiving module, an image segmentation module, and an image recognition module. The data receiving module acquires the original image and performs image preprocessing. The dataset augmentation sub-module can perform real-time generalization and enhancement on the original image based on the generative adversarial network architecture. Its image repair unit can check the incomplete original image and repair it with a neural network when the proportion is appropriate to obtain the enhanced and repaired image for subsequent semantic segmentation. The image segmentation module uses a deep learning semantic segmentation algorithm to segment the target image into multiple morphemes with coordinate information and class labels. Among them, the instance extraction unit segments instances based on the morphemes, the hierarchical feature extraction unit obtains primary and intermediate features, the timing acquisition unit sorts the target image according to time and divides it into continuous frame image groups, and the dynamic recognition unit detects the continuous frame image groups to obtain dynamic features and generate events. The advanced semantic fusion unit of the image recognition module inputs the primary and intermediate features into the Transformer model under the hierarchical fusion strategy, and obtains advanced semantic features through operations such as feature transformation, hierarchical fusion, multi-head self-attention calculation, feed-forward network calculation, and layer normalization. Finally, an accurate image recognition result is generated based on this. Through the integrated operation of each module, the above solution can deepen the understanding of the image by the system by extracting advanced semantic features, thereby improving the accuracy and reliability of the image recognition result.
[0108] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0109] Based on the same inventive concept, the embodiments of the present application also provide a method for image recognition based on deep learning to implement the functions of the above-mentioned image recognition system based on deep learning. The implementation solutions provided by this method to solve problems are similar to the implementation solutions recorded in the above system. Therefore, the specific limitations in one or more embodiments of the method for image recognition based on deep learning provided below can refer to the limitations on the image recognition system based on deep learning in the above text, and will not be repeated here.
[0110] In an exemplary embodiment, as Figure 2 shown, a method for image recognition based on deep learning is provided, including:
[0111] Step 201: Preprocess the obtained original image to obtain the target image to be recognized.
[0112] Step 202: Perform semantic segmentation on the target image based on the deep learning semantic segmentation algorithm to obtain multiple morphemes, where the attributes of the morphemes include coordinate information and class labels.
[0113] Step 203: Obtain high-level semantic features based on the morphemes, and generate an image recognition result based on the high-level semantic features.
[0114] In one embodiment, the method may further include the following steps:
[0115] Step 2021: Segment instances in each class label based on the morphemes, where the attributes of the instances include attributes inherited from the morphemes and instance labels.
[0116] Step 2022: Obtain primary features and intermediate features, extract primary features for each instance, where the primary features include category, location, and visual attributes; obtain intermediate features between multiple instances, where the intermediate features include interaction relationships and / or topological relationships; the primary features and intermediate features are used to extract high-level semantic features.
[0117] In one embodiment, the method may further include the following steps:
[0118] Step 2031: Input the primary features and intermediate features into a deep learning-based hierarchical fusion model to obtain high-level semantic features; the high-level semantic features include at least one of object relationship features, scene semantic features, and action behavior features.
[0119] Furthermore, the deep learning-based hierarchical fusion model may be a Transformer model under a hierarchical fusion strategy, and the Transformer model under the hierarchical fusion strategy is trained using the following method:
[0120] Establish a connection with the training image library and obtain training images, perform feature transformation on the training images to obtain sequence features, and perform hierarchical fusion calculation on the sequence features using the following formula:
[0121]
[0122] where, X l is the input of the l-th layer of the Transformer, and l i is the preset layer index;
[0123] Perform multi-head self-attention calculation on the input X l after fusing hierarchical features, using the following formula:
[0124]
[0125] MHSA(Q, K, V) = Concat(head1, head2, …, head h )W 0
[0126]
[0127]
[0128] where is the output of the multi - head self - attention, head i is the attention head, Concat(head1, …, head h ) is the concatenation function of multiple attention heads, W 0 is the concatenation integration linear transformation matrix, is the projection matrix projected into the query space, is the projection matrix projected into the key space, is the projection matrix projected into the value space, Q is the query matrix, K is the key matrix, V is the value matrix, d k is the feature dimension;
[0129] Using the following formula, the output of the multi - head self - attention is input into the feed - forward network to obtain the output of the feed - forward network:
[0130]
[0131] FFN(x) = ReLU(xW1 + b1)W2 + b2
[0132] where is the output of the feed - forward network, FFN is a two - layer fully - connected network, W1 is the weight matrix of the first - layer fully - connected network, b1 is the bias vector of the first - layer fully - connected network, ReLU is the activation function, W2 is the weight matrix of the second - layer fully - connected network, and b2 is the bias vector of the second - layer fully - connected network;
[0133] Perform layer normalization on the output of the feed - forward network, and use the normalized result as the input of the next - layer Transformer. Repeat the above steps for calculation until the preset model convergence condition is reached; where the formula for layer normalization is as follows:
[0134]
[0135] where LayerNorm is the layer normalization function.
[0136] In one possible embodiment, the method further includes the following steps:
[0137] In step 2023, obtain the time information of the target images, sort the target images according to the chronological order, and set multiple target images that are consecutive and have an interval time not exceeding a preset value as a consecutive frame image group.
[0138] In step 2024, continuously detect multiple target images within the consecutive frame image group, obtain dynamic features, and generate an event based on the dynamic features, where the event is used to indicate the image recognition result.
[0139] Preferably, the method may include the following steps:
[0140] In step 2011, the dataset augmentation sub-module is used to perform real-time generalization enhancement on the obtained original images based on the generative adversarial network architecture to obtain enhanced images; wherein, the generative adversarial network architecture includes a generator and a discriminator. The generator is used to learn the feature distribution of the original images to generate simulated images; the discriminator is used to distinguish between the original images and the simulated images; and the simulated images are used for semantic segmentation.
[0141] Preferably, the method may include the following steps:
[0142] In step 2012, perform an integrity check on the obtained original images to obtain incomplete regions.
[0143] In step 2013, if the proportion of the incomplete region in the area of the entire image is within a preset threshold range, use a neural network to perform image repair on the original images to obtain repaired images; wherein, the repaired images are used for semantic segmentation.
[0144] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of a method for image recognition based on deep learning as described above are implemented.
[0145] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0146] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0147] The above-described embodiments merely represent several implementation manners of the embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the embodiments of the present application.
Claims
1. An image recognition system based on deep learning, characterized in that, The system includes: a data receiving module, an image segmentation module, and an image recognition module; The data receiving module is used to preprocess the acquired original image to obtain a target image to be recognized; The image segmentation module is used to perform semantic segmentation on the target image based on a deep learning semantic segmentation algorithm to obtain multiple morphemes, where the attributes of the morphemes include coordinate information and class labels; The image recognition module is used to obtain high-level semantic features based on the morphemes and generate an image recognition result based on the high-level semantic features.
2. The system according to claim 1, wherein The image segmentation module includes an instance extraction unit and a hierarchical feature extraction unit: The instance extraction unit is used to segment instances in each of the class labels based on the morphemes, and the attributes of the instances include attributes inherited from the morphemes and instance labels; The hierarchical feature extraction unit is used to obtain primary features and intermediate features, perform primary feature extraction on each of the instances, where the primary features include category, location, and visual attributes; obtain intermediate features between multiple of the instances, where the intermediate features include interaction relationships and / or topological relationships; the primary features and the intermediate features are used to extract the high-level semantic features.
3. The system according to claim 2, characterized in that, The image recognition module includes a high-level semantic feature fusion unit: The high-level semantic fusion unit is used to input the primary features and the intermediate features into a deep learning-based hierarchical fusion model to obtain high-level semantic features; the high-level semantic features include at least one of object relationship features, scene semantic features, and action behavior features.
4. The system according to claim 3, characterized in that The deep learning-based hierarchical fusion model is a Transformer model under a hierarchical fusion strategy, and the Transformer model under the hierarchical fusion strategy is trained using the following method: Establish a connection with a training image library and obtain training images, perform feature transformation on the training images to obtain sequence features, and perform hierarchical fusion calculation on the sequence features using the following formula: Among them, X l is the input of the l-th layer of the Transformer, where l i is a preset layer index; The input X that incorporates hierarchical features l is subjected to multi-head self-attention calculation using the following formula: MHSA(Q, K, V) = Concat(head1, head2, …, head h )W 0 Among them, is the output of the multi-head self-attention, head i is the attention head, Concat(head1,…,head h ) is the concatenation function of multiple attention heads, W 0 is the concatenation integration linear transformation matrix, is the projection matrix projected into the query space, is the projection matrix projected into the key space, is the projection matrix projected into the value space, Q is the query matrix, K is the key matrix, V is the value matrix, d k is the feature dimension; Use the following formula to output the multi-head self-attention Input it into the feed-forward network to obtain the output of the feed-forward network: FFN(x) = ReLU(xW1 + b1)W2 + b2 Among them, is the output of the feedforward network, FFN is a two-layer fully connected network, W1 is the weight matrix of the first-layer fully connected network, b1 is the bias vector of the first-layer fully connected network, ReLU is the activation function, W2 is the weight matrix of the second-layer fully connected network, and b2 is the bias vector of the second-layer fully connected network; Perform layer normalization on the output of the feedforward network and use the normalized result as the input of the next layer of the Transformer, and repeat the above steps for calculation until a preset model convergence condition is reached; where the calculation formula for layer normalization is as follows: where LayerNorm is the layer normalization function.
5. The system according to claim 2, wherein The image segmentation module further includes a timing acquisition unit and a dynamic recognition unit: The timing acquisition unit is used to obtain the time information of the target image, sort the target images in chronological order, and set multiple target images that are continuous and have an interval time not exceeding a preset value as a continuous frame image group; The dynamic recognition unit is used to continuously detect multiple target images in the continuous frame image group, obtain dynamic features and generate an event based on the dynamic features, and the event is used to indicate the image recognition result.
6. The system according to claim 2, wherein The data receiving module includes a dataset augmentation sub-module: The dataset augmentation sub-module is used to perform real-time generalization enhancement on the acquired original image based on a generative adversarial network architecture to obtain an enhanced image; The generative adversarial network architecture includes a generator and a discriminator, wherein the generator is used to learn the feature distribution of the original image to generate a simulated image; The discriminator is used to distinguish the original image from the simulated image; The simulated image is used for semantic segmentation.
7. The system according to claim 6, characterized in that, The data set expansion submodule includes an image restoration unit, which is used to: Performing an integrity check on the acquired original image to obtain a defective area; If the proportion of the incomplete area to the entire image area is within a preset threshold range, a neural network is used to perform image restoration on the original image to obtain a restored image; wherein the restored image is used for semantic segmentation.
8. An image recognition method based on deep learning, characterized in that, The method comprises: Preprocessing the acquired original image to obtain the target image to be identified; Performing semantic segmentation on the target image based on a deep learning semantic segmentation algorithm to obtain a plurality of morphemes, wherein the attributes of the morphemes include coordinate information and class labels; A high-level semantic feature is obtained according to the morphemes, and an image recognition result is generated based on the high-level semantic feature.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the method according to claim 8 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to claim 8 is implemented.