An image recognition method, system and medium based on an air-spectrum joint model
Through the cascading hierarchical residual structure and attention mechanism of the space-spectral joint model, the problem of low accuracy of recognition of multi-category target bodies is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411069132.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2044-08-06
AI Technical Summary
When existing image recognition technology recognizes multi-class target bodies with similar colors or colorless and transparent colors, the accuracy is low. Traditional algorithms such as Gaussian classifiers and convolutional neural networks have poor robustness, gradient disappearance or explosion, making it difficult to effectively distinguish multi-class target bodies.
The image recognition method based on the space-spectral joint model is adopted, and the cascaded hierarchical residual structure and attention mechanism are combined with spectral features and spatial feature extraction networks to perform preprocessing and training to improve the recognition accuracy of multiple target bodies.
The recognition accuracy of multiple types of target bodies in the image is improved, the recognition problem of similar colors or colorless transparent target bodies is solved, and the robustness of the model and the ability to obtain global information is enhanced.
Smart Images

Figure CN118918442B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition technology, and in particular, to an image recognition method, system and medium based on a spatial-spectral joint model. Background Art
[0002] With the continuous progress of image recognition technology, its applications in fields such as animal husbandry, agriculture, medicine, and industry have become more and more extensive. In animal husbandry, the recognition of pasture grasses can effectively prevent the invasion of foreign grass seeds, realize the detection of grassland degradation, protect local grass seeds and provide a good and sufficient feed source for animal husbandry; in agriculture, image recognition technology provides an efficient and non-destructive detection method for the classification and detection of pathogenic molds and the classification of crop seeds, and provides certain help for the control of pathogenic molds and the correct classification of crop seeds; in medicine, image recognition technology can provide effective technical support for the identification of the authenticity, quality and inferiority of medical cotton.
[0003] However, in the existing cotton processing plants, the identification and screening of impurities in unginned cotton, and the classification and recycling of transparent and white materials in landfills, the application of image recognition technology is not very extensive. This is because when facing multiple types of target objects with similar colors or colorless and transparent ones, due to their extremely small color differences, ordinary visual cameras simply cannot effectively distinguish multiple types of target objects. And traditional algorithms, such as Gaussian classifiers and support vector machines, have poor robustness and are not suitable for the recognition and classification of multiple types of target objects. For another example, the scale information obtained by convolutional neural networks in deep learning algorithms is relatively single, and it is difficult to obtain the global information of multiple types of target objects, resulting in a relatively low accuracy rate for identifying multiple types of target objects.
[0004] With the development of the times, many scholars have tried to obtain a larger receptive field by increasing the number of layers of convolutional neural networks, and then obtain the global information of the target object. However, with the sharp increase in the number of network layers, it will undoubtedly bring problems such as gradient disappearance or gradient explosion to the convolutional neural network, and ultimately will also lead to a significant reduction in the accuracy rate of identifying multiple types of target objects in the image. Summary of the Invention
[0005] The purpose of the present application is to provide an image recognition method, system and medium based on a spatial-spectral joint model, which improves the accuracy of identifying multiple types of target objects in an image.
[0006] To achieve the above purpose, the present application provides the following solutions.
[0007] First aspect, the present application provides an image recognition method based on an air-spectrum joint model. The image recognition method based on the air-spectrum joint model includes: obtaining a plurality of different sample images; each sample image contains multiple types of target objects; the target objects of each type have similar colors or are colorless and transparent; preprocessing all the sample images to obtain a plurality of sample slice images; the preprocessing at least includes: black and white correction, three-dimensional cropping, normalization, multiplicative scatter correction, and slicing; dividing all the sample slice images into a training set, a validation set, and a test set according to a ratio; constructing an image recognition basic model; the image recognition basic model at least includes: a spectral feature extraction network and a spatial feature extraction network; the spectral feature extraction network includes: a first hierarchical residual structure and a second hierarchical residual structure; the spatial feature extraction network includes: a third hierarchical residual structure and a fourth hierarchical residual structure; the first hierarchical residual structure and the second hierarchical residual structure are connected in a cascaded manner; the third hierarchical residual structure and the fourth hierarchical residual structure are connected in a cascaded manner; both the first hierarchical residual structure and the third hierarchical residual structure contain n layer structures, n the layer structures are connected in a cascaded manner; both the first hierarchical residual structure and the third hierarchical residual structure have n -1 layer structures containing convolutional blocks; both the second hierarchical residual structure and the fourth hierarchical residual structure contain m layer structures, m the layer structures are connected in a cascaded manner; both the second hierarchical residual structure and the fourth hierarchical residual structure have m -1 layer structures containing convolutional blocks; n , m are integers greater than 2 and satisfy n > m ; based on the training set and the label set, training the image recognition basic model until the average accuracy and the overall accuracy of the recognition of multiple types of target objects in the sample slice images are both greater than a set threshold; the label set contains a plurality of sample slice images annotating various target objects; using the test set and the validation set to test and validate the trained image recognition basic model, and determining the trained, tested, and validated image recognition basic model as the air-spectrum joint model; obtaining an image to be recognized, and inputting the image to be recognized into the air-spectrum joint model to recognize multiple types of target objects in the image to be recognized.
[0008] Optionally, the image recognition basic model further includes: a high-dimensional mapping module, a channel attention mechanism module, a spatial attention mechanism module, and an information fusion module; the high-dimensional mapping module is respectively connected to the spectral feature extraction network and the spatial feature extraction network; the spectral feature extraction network is connected to the channel attention mechanism module; the spatial feature extraction network is connected to the spatial attention mechanism module; the channel attention mechanism module and the spatial attention mechanism module are respectively connected to the information fusion module; the high-dimensional mapping module is used to increase the dimension of the sample slice image to obtain a hyperspectral sample slice image; the spectral feature extraction network is used to extract the spectral features of the hyperspectral sample slice image to obtain an initial spectral feature map; the spatial feature extraction network is used to extract the spatial features of the hyperspectral sample slice image to obtain an initial spatial feature map; the channel attention mechanism module is used to change the channel attention weights of the initial spectral feature map to obtain an optimized spectral feature map; the spatial attention mechanism module is used to change the spatial attention weights of the initial spatial feature map to obtain an optimized spatial feature map; the information fusion module is used to perform information fusion on the optimized spectral feature map and the optimized spatial feature map to obtain feature fusion information; the feature fusion information is used to identify multiple types of target bodies in the sample slice image.
[0009] Optionally, the internal operation relationship of the first hierarchical residual structure is as follows.
[0010] 。
[0011] In the formula, is the output of the i th layer structure in the first hierarchical residual structure, is the input of the n th layer structure in the first hierarchical residual structure, is the output of the n -1th layer structure in the first hierarchical residual structure, is the first convolution operator.
[0012] Optionally, the internal operation relationship of the second hierarchical residual structure is as follows.
[0013] 。
[0014] In the formula, is the output of the i th layer structure in the second hierarchical residual structure, is the input of the m th layer structure in the second hierarchical residual structure, is the output of the m -1th layer structure in the second hierarchical residual structure, is the first convolutional operator.
[0015] Optionally, the operation relationship inside the third hierarchical residual structure is as follows.
[0016] .
[0017] In the formula, is the output of the i th layer structure in the third hierarchical residual structure, is the input of the n th layer structure in the third hierarchical residual structure, is the output of the n -1th layer structure in the third hierarchical residual structure, is the second convolutional operator.
[0018] Optionally, the operation relationship inside the fourth hierarchical residual structure is as follows.
[0019] .
[0020] In the formula, is the output of the i th layer structure in the fourth hierarchical residual structure, is the input of the m th layer structure in the fourth hierarchical residual structure, is the output of the m -1th layer structure in the fourth hierarchical residual structure, is the second convolutional operator.
[0021] Optionally, the calculation formula of the average accuracy rate is as follows.
[0022] .
[0023] In the formula, is the average accuracy rate, is the number of samples of the j th type of target object correctly recognized, is the j th type of target object, is the total number of types of target objects;
[0024] The calculation formula of the overall accuracy rate is as follows.
[0025] .
[0026] In the formula, is the overall accuracy rate, is the total number of samples of target objects of all types, is the number of samples correctly recognized among target objects of all types.
[0027] In a second aspect, the present application also provides a computer system, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the image recognition method based on the spatial-spectral joint model.
[0028] In a third aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the image recognition method based on the spatial-spectral joint model.
[0029] According to the specific embodiments provided by the present application, the following technical effects are disclosed.
[0030] The present application considers how to improve the accuracy of multi-class target recognition in images from two aspects. In the first aspect, in the selection of training set images, the sample images selected by the present application have multi-class target bodies with similar colors or colorless and transparent, and sample images with large color differences are not selected, so as to ensure that the training set images are more targeted. Secondly, the present application also preprocesses the sample images, such as black and white correction, multiplicative scatter correction, slicing, etc., further improving the accuracy of the training set images. In the second aspect, in terms of the model structure, in order to obtain a larger receptive field and avoid problems such as gradient disappearance or explosion, the present application designs a hierarchical residual structure with front and rear cascades, that is, the first hierarchical residual structure and the second hierarchical residual structure in the spectral feature extraction network, and the third hierarchical residual structure and the fourth hierarchical residual structure in the spatial feature extraction network. Since each residual structure is composed of multiple layer structures, a larger receptive field can be obtained. At the same time, in order to avoid too many layers, a front and rear cascading method is used to separate multiple layer structures. Based on the above two considerations, the present application can not only ensure the effectiveness of model training, but also capture the global information of the target body by using the model structure, thus improving the accuracy of multi-class target recognition in images. Description of the Drawings
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0032] Figure 1 It is a flowchart of the image recognition method based on the spatial-spectral joint model provided by the embodiment of the present application.
[0033] Figure 2 It is a schematic structural diagram of the high-dimensional mapping module - spectral feature extraction network - channel attention mechanism module provided by the embodiment of the present application.
[0034] Figure 3 This is a schematic structural diagram of the high-dimensional mapping module - spatial feature extraction network - spatial attention mechanism module provided by the embodiments of the present application.
[0035] Figure 4 This is a schematic structural diagram of the information fusion module provided by the embodiments of the present application.
[0036] Figure 5 This is a schematic structural diagram of the spatial-spectral joint model provided by the embodiments of the present application.
[0037] Figure 6 This is an internal structure diagram of the computer system provided by the embodiments of the present application. Specific embodiments
[0038] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0039] The purpose of the present application is to provide an image recognition method, system and medium based on a spatial-spectral joint model, which improves the accuracy of recognizing multiple types of target objects in an image.
[0040] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0041] As Figure 1 shown, this embodiment provides an image recognition method based on a spatial-spectral joint model, which is specifically as follows.
[0042] Step S1: Obtain a plurality of different sample images; each sample image contains multiple types of target objects; the target objects of each type have similar colors or are colorless and transparent.
[0043] In this embodiment, the acquired sample image is an image of foreign fibers in unginned cotton, and the image of foreign fibers in unginned cotton contains various types of target objects such as unginned cotton fibers, catkins, white polypropylene filaments, and transparent plastic films. Specifically, the foreign fibers in unginned cotton to be identified and classified are placed on the plane to be measured in the dark box, the GaiaField Pro-V10E hyperspectral camera is fixed on the top of the dark box, the distance between the lens of the hyperspectral camera and the plane to be measured is 500 mm, 4 halogen light sources with a power of 50 watts are used to supplement light for the shooting area, and the computer is used to control the hyperspectral camera to perform scanning and shooting, and a standard whiteboard image (i.e., the image obtained by shooting a pure white nylon cloth), a standard blackboard image (i.e., the image obtained by covering the lid of the hyperspectral camera) and multiple images of foreign fibers in unginned cotton are obtained respectively.
[0044] Step S2: Preprocess all the sample images to obtain multiple sample slice images; the preprocessing includes at least: black and white correction, three-dimensional cropping, standardization, multiplicative scatter correction, and slicing.
[0045] In this embodiment, the process of preprocessing the image of foreign fibers in unginned cotton is specifically as follows.
[0046] In the first step, the black and white correction calculation formula is used to perform black and white correction processing on the image of foreign fibers in unginned cotton to obtain a black and white corrected image of foreign fibers in unginned cotton.
[0047] 。
[0048] In the formula, is the black and white corrected image of foreign fibers in unginned cotton, is the image of foreign fibers in unginned cotton, is the standard blackboard image, is the standard whiteboard image.
[0049] In the second step, three-dimensional cropping is performed on the black and white corrected image of foreign fibers in unginned cotton to remove part of the useless black background and balance the number of positive and negative samples (i.e., balance the number of pixels contained in unginned cotton fibers, various impurity fibers, background, etc.) to obtain a hyperspectral image of foreign fibers in unginned cotton.
[0050] In the third step, standardization processing and multiplicative scatter correction processing are sequentially performed on the hyperspectral image of foreign fibers in unginned cotton. The standardization processing is based on the hyperspectral image of foreign fibers in unginned cotton, and uses the relationship between its mean value and standard deviation to obtain a standardized image of foreign fibers in unginned cotton. Among them, the calculation formula for the standardization processing is as follows.
[0051] 。
[0052] In the formula, is the standardized image of foreign fibers in unginned cotton, is the hyperspectral image of foreign fibers in unginned cotton, is the mean value of the reflectance of the hyperspectral image of foreign fibers in unginned cotton, is the standard deviation of the reflectance of the hyperspectral raw cotton image with foreign fibers.
[0053] The multiplicative scatter correction (MSC) process is based on the standardized raw cotton image with foreign fibers. By calculating its average spectrum and performing a simple linear regression, the regression constant and regression coefficient are obtained. Then, the relative baseline shift and offset between the near-infrared spectra of each pixel are corrected to obtain the MSC raw cotton image with foreign fibers. The calculation formula for MSC is as follows.
[0054] 。
[0055] 。
[0056] 。
[0057] In the formula, is the average spectrum vector of the spectral data of all pixel points in the standardized raw cotton image with foreign fibers, is the spectral data of the i th pixel point in the standardized raw cotton image with foreign fibers, is the spectral band in the standardized raw cotton image with foreign fibers, is the regression coefficient under the spectral data of the i th pixel point, is the regression constant under the spectral data of the i th pixel point, is the MSC raw cotton image with foreign fibers under the spectral data of the i th pixel point.
[0058] In the fourth step, principal component analysis (PCA) is used to reduce the dimension of the MSC raw cotton image with foreign fibers. Through linear transformation, the high-dimensional space features are mapped to a low-dimensional space. While retaining as much of the original data as possible, the image dimension is reduced, enabling the image to be compressed to a large extent. The specific process is as follows: First, a sample set is established for the MSC raw cotton image with foreign fibers and centralized processing is performed; then, the covariance matrix of the samples is calculated and eigenvalue decomposition is carried out. The eigenvectors corresponding to the largest eigenvalues are taken out. After standardizing all the eigenvectors, an eigenvector matrix is formed; finally, each sample is transformed into a new sample, and the new sample set is output.
[0059] In the fifth step, the MSC raw cotton image with foreign fibers after dimension reduction by PCA is sliced according to a set size to generate multiple three-dimensional image blocks of the same size (i.e., the raw cotton image slices with foreign fibers).
[0060] Step S3: Divide all the sample slice images into a training set, a validation set, and a test set according to a certain proportion.
[0061] In this embodiment, 5000 lint cotton foreign fiber slice images obtained through preprocessing are divided into a training set, a validation set, and a test set according to a ratio of 6:2:2.
[0062] Step S4: Construct a basic image recognition model; the basic image recognition model at least includes: a spectral feature extraction network and a spatial feature extraction network; the spectral feature extraction network includes: a first hierarchical residual structure and a second hierarchical residual structure; the spatial feature extraction network includes: a third hierarchical residual structure and a fourth hierarchical residual structure; the first hierarchical residual structure and the second hierarchical residual structure are connected in a cascaded manner; the third hierarchical residual structure and the fourth hierarchical residual structure are connected in a cascaded manner; both the first hierarchical residual structure and the third hierarchical residual structure contain n layer structures, n and the layer structures are connected in a cascaded manner; both the first hierarchical residual structure and the third hierarchical residual structure have n -1 layer structures containing convolutional blocks; both the second hierarchical residual structure and the fourth hierarchical residual structure contain m layer structures, m and the layer structures are connected in a cascaded manner; both the second hierarchical residual structure and the fourth hierarchical residual structure have m -1 layer structures containing convolutional blocks; n , m are integers greater than 2 and satisfy n > m .
[0063] Furthermore, the basic image recognition model further includes: a high-dimensional mapping module, a channel attention mechanism module, a spatial attention mechanism module, and an information fusion module; the high-dimensional mapping module is respectively connected to the spectral feature extraction network and the spatial feature extraction network; the spectral feature extraction network is connected to the channel attention mechanism module; the spatial feature extraction network is connected to the spatial attention mechanism module; the channel attention mechanism module and the spatial attention mechanism module are respectively connected to the information fusion module.
[0064] The high-dimensional mapping module is used to perform dimensionality increase on the sample slice image to obtain a hyperspectral sample slice image; the spectral feature extraction network is used to extract the spectral features of the hyperspectral sample slice image to obtain an initial spectral feature map; the spatial feature extraction network is used to extract the spatial features of the hyperspectral sample slice image to obtain an initial spatial feature map; the channel attention mechanism module is used to change the channel attention weights of the initial spectral feature map to obtain an optimized spectral feature map; the spatial attention mechanism module is used to change the spatial attention weights of the initial spatial feature map to obtain an optimized spatial feature map; the information fusion module is used to perform information fusion on the optimized spectral feature map and the optimized spatial feature map to obtain feature fusion information; the feature fusion information is used to identify multiple types of target bodies in the sample slice image.
[0065] Among them, the first hierarchical residual structure satisfies the following operation relationships inside.
[0066] .
[0067] In the formula, is the output of the i -th layer structure in the first hierarchical residual structure, is the input of the n -th layer structure in the first hierarchical residual structure, is the output of the n -1-th layer structure in the first hierarchical residual structure, is the first convolution operator.
[0068] The second hierarchical residual structure satisfies the following operation relationships inside.
[0069] .
[0070] In the formula, is the output of the i -th layer structure in the second hierarchical residual structure, is the input of the m -th layer structure in the second hierarchical residual structure, is the output of the m -1-th layer structure in the second hierarchical residual structure, is the first convolution operator.
[0071] The third hierarchical residual structure satisfies the following operation relationships inside.
[0072] .
[0073] In the formula, is the output of the i -th layer structure in the third hierarchical residual structure, is the input of the n -th layer structure in the third hierarchical residual structure, is the output of the n -1-th layer structure in the third hierarchical residual structure, is the second convolution operator.
[0074] The fourth hierarchical residual structure satisfies the following operation relationships inside.
[0075] .
[0076] In the formula, is the output of the i -th layer structure in the fourth hierarchical residual structure, is the input of the m -th layer structure in the fourth hierarchical residual structure, is the output of the m -1 layer structure in the fourth hierarchical residual structure, and is the second convolution operator.
[0077] In this embodiment, as Figure 2 shown, the high-dimensional mapping module is composed of a 1×1 convolutional layer, a batch normalization layer, and an activation function layer. The high-dimensional mapping module can perform dimensionality increase on the sliced image of seed cotton foreign fiber in the training set, map the spectral features in the sliced image of seed cotton foreign fiber
[0078] to a high-dimensional space, facilitating the spectral feature extraction network and the spatial feature extraction network to learn spectral information. When the hyperspectral sliced image of seed cotton foreign fiber obtained through the high-dimensional mapping module is input into the spectral feature extraction network, the input of each layer structure in the first hierarchical residual structure is of the number of channels of the hyperspectral sliced image of seed cotton foreign fiber, and the values are respectively , , . After the following formula operations, the output of each layer structure in the first hierarchical residual structure is respectively
[0079] .
[0080] The overall output of the first hierarchical residual structure satisfies: , and after is input into the second hierarchical residual structure, the input of each layer structure in the second hierarchical residual structure is of the number of channels of , and the values are respectively , , and after the following formula operations, the output of each layer structure in the second hierarchical residual structure is respectively .
[0081] .
[0082] The overall output of the second hierarchical residual structure satisfies: , That is the initial spectral feature map obtained through the spectral feature extraction network.
[0083] Input it into the channel attention mechanism module. After global max pooling and global average pooling respectively, it is successively fed into a 3×3 convolutional layer, an activation function layer, and a 3×3 convolutional layer to obtain the first weight and the second weight respectively. After the following formula operation, it is input Sigmoid into the activation function layer to obtain the optimized spectral feature map .
[0084] .
[0085] .
[0086] .
[0087] .
[0088] In the formula, is the output weight of the channel attention mechanism module, represents element-wise multiplication, is Sigmoid the activation function, and represent the values obtained by global max pooling and global average pooling of the channel attention mechanism module respectively, represents global max pooling, represents global average pooling.
[0089] As Figure 3 shown, the third hierarchical residual structure in the spatial feature extraction network contains 5 layer structures, and 4 layer structures contain convolutional blocks with 6 convolutional blocks in each layer structure; the fourth hierarchical residual structure contains 3 layer structures, and 2 layer structures contain convolutional blocks with 6 convolutional blocks in each layer structure. Among them, the convolutional blocks in the third hierarchical residual structure and the fourth hierarchical residual structure are composed of a 3×3 convolutional layer, a batch normalization layer, and an activation function layer. When the hyperspectral cottonseed with foreign fiber slice image obtained through the high-dimensional mapping module is input into the spatial feature extraction network, the input of each layer structure in the third hierarchical residual structure is of the number of channels of the hyperspectral cottonseed with foreign fiber slice image, and the values are respectively , , , . After the following formula operation, the output of each layer structure in the third hierarchical residual structure is respectively .
[0090] 。
[0091] The overall output of the third hierarchical residual structure Satisfies: , After inputting into the fourth hierarchical residual structure, the inputs of each layer structure in the fourth hierarchical residual structure are of the number of channels , and the values taken are respectively , , After performing the following formula operations, the outputs of each layer structure in the fourth hierarchical residual structure are respectively 。
[0092] 。
[0093] The overall output of the fourth hierarchical residual structure Satisfies: , which is the initial spatial feature map obtained through the spatial feature extraction network.
[0094] Input into the spatial attention mechanism module. After performing global max pooling and global average pooling respectively, the feature maps are concatenated, and then passed into a 3×3 filter for convolution operation, and using Sigmoid the activation function and the following formula operations, the optimized spatial feature map is obtained.
[0095] 。
[0096] 。
[0097] 。
[0098] 。
[0099] In the formula, is the output weight of the spatial attention mechanism module, represents the 3×3 convolution operation, represents the feature map concatenation, and represent the values obtained by global max pooling and global average pooling of the spatial attention mechanism module respectively.
[0100] As Figure 4 shown, the optimized spectral feature map and the optimized spatial feature map are respectively fused through two layer structures (composed of a batch normalization layer, an activation function layer, and a dropout layer) in the information fusion module, and then input into the fully connected layer to complete the image of the foreign fiber slices in the unginned cotton Recognition and classification of various target objects in
[0101] Step S5: Based on the training set and the label set, train the basic image recognition model until the average accuracy and overall accuracy of the recognition of multiple types of target objects in the sample slice images are both greater than the set threshold; the label set contains multiple sample slice images annotated with various target objects.
[0102] In this embodiment, the label set is determined after preprocessing the raw cotton foreign fiber images, that is, various target objects to be recognized are manually calibrated in the preprocessed raw cotton foreign fiber images, and finally the label set is determined based on multiple raw cotton foreign fiber images with various target objects calibrated.
[0103] As a preferred implementation manner, this embodiment determines the average accuracy and overall accuracy from two aspects. On the first aspect, continuously change the convolution kernel size in the spatial feature extraction network, then use the basic image recognition model to conduct recognition and classification experiments on multiple types of target objects, and calculate the average accuracy and overall accuracy of the recognition and classification results of multiple types of target objects based on the experimental results and the following formula until both the average accuracy and overall accuracy are greater than 98%, and record the optimal convolution kernel size in the spatial feature extraction network at this time.
[0104] .
[0105] .
[0106] In the formula, is the average accuracy, is the number of samples of the j th type of target object correctly recognized, is the j th type of target object's total number of samples, is the total number of categories of target objects, is the overall accuracy, is the total number of samples of target objects of all categories, is the number of samples correctly recognized among target objects of all categories.
[0107] On the second aspect, keep the optimal convolution kernel size in the above spatial feature extraction network unchanged, continuously change the slice size of the sample images during the preprocessing, then use the basic image recognition model to conduct recognition and classification experiments on multiple types of target objects, and calculate the average accuracy and overall accuracy of the recognition and classification results of multiple types of target objects again based on the experimental results and the above formula until both the average accuracy and overall accuracy are greater than 98%, and record the optimal slice size of the sample images during the preprocessing at this time.
[0108] Step S6: Use the test set and the validation set to test and validate the trained basic image recognition model, and determine the trained basic image recognition model after training, testing, and validation as the hyperspectral joint model.
[0109] In this embodiment, based on the optimal convolutional kernel size and the optimal slicing size determined in the above step S5, use the test set and the validation set to test and validate the trained basic image recognition model, and continuously adjust the parameters in the trained basic image recognition model until the optimal is reached, and finally obtain the Figure 5 hyperspectral joint model as shown.
[0110] Step S7: Obtain the image to be recognized, and input the image to be recognized into the hyperspectral joint model to recognize multiple types of target objects in the image to be recognized.
[0111] In this embodiment, multiple types of recognized target objects will be marked in different colors in the image to be recognized, which is convenient for operators to observe and classify.
[0112] Based on the above analysis, this application mainly has the following advantages.
[0113] 1) In the spectral feature extraction network and the spatial feature extraction network, each hierarchical residual structure transmits information to adjacent layer structures through different branches, and finally cascades the information in each layer structure. At the same time, the cascading method of the front and rear hierarchical residual structures and the fusion design of the channel attention mechanism fully learn the spectral feature information.
[0114] 2) The hyperspectral joint model can obtain a rich receptive field, and can also fully learn and extract the spectral information and spatial position information of the hyperspectral image. Compared with traditional recognition methods, the feature description and classification recognition results of the target object to be recognized in this application are more accurate, reasonable, and have better robustness.
[0115] In another exemplary embodiment, this application also provides a computer system, which can be a server or a terminal, and its internal structure diagram can be as Figure 6As shown. The computer system includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer system is used to provide computing and control capabilities. The memory of the computer system includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer system is used to store video tag processing data. The input / output interface of the computer system is used to exchange information between the processor and external devices. The communication interface of the computer system is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements an image recognition method based on a spatial-spectral joint model.
[0116] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer system to which the solution of this application is applied. The specific computer system may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0117] In another exemplary embodiment, the present application also provides a computer-readable storage medium storing a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0118] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0119] The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0120] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.
[0121] Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An image recognition method based on an air-spectrum joint model, characterized in that The image recognition method based on the spatial-spectral joint model includes: Obtaining a plurality of different sample images; each sample image contains multiple types of target objects; the target objects of each type have similar colors or are colorless and transparent; Preprocessing all the sample images to obtain a plurality of sample slice images; the preprocessing at least includes: black and white correction, three-dimensional cropping, normalization, multiplicative scatter correction, and slicing; Dividing all the sample slice images into a training set, a validation set, and a test set according to a ratio; Constructing an image recognition basic model; the image recognition basic model at least includes: a spectral feature extraction network and a spatial feature extraction network; the spectral feature extraction network includes: a first hierarchical residual structure and a second hierarchical residual structure; the spatial feature extraction network includes: a third hierarchical residual structure and a fourth hierarchical residual structure; the first hierarchical residual structure and the second hierarchical residual structure are connected in a cascaded manner; the third hierarchical residual structure and the fourth hierarchical residual structure are connected in a cascaded manner; both the first hierarchical residual structure and the third hierarchical residual structure contain n layer structures, and the n layer structures are connected in a cascaded manner; n-1 layer structures in both the first hierarchical residual structure and the third hierarchical residual structure contain convolutional blocks; both the second hierarchical residual structure and the fourth hierarchical residual structure contain m layer structures, and the m layer structures are connected in a cascaded manner; m-1 layer structures in both the second hierarchical residual structure and the fourth hierarchical residual structure contain convolutional blocks; n and m are integers greater than 2 and satisfy n>m; the internal operation relationship of the third hierarchical residual structure satisfies: In the formula, is the output of the i-th layer structure in the third hierarchical residual structure, is the input of the n-th layer structure in the third hierarchical residual structure, is the output of the (n - 1)-th layer structure in the third hierarchical residual structure, and K(·) is the second convolution operator; the operation relationship inside the fourth hierarchical residual structure satisfies: In the formula, is the output of the i-th layer structure in the fourth hierarchical residual structure, is the input of the m-th layer structure in the fourth hierarchical residual structure, is the output of the (m - 1)-th layer structure in the fourth hierarchical residual structure, and K(·) is the second convolution operator; Based on the training set and the label set, training the image recognition basic model until the average accuracy and the overall accuracy of the recognition of multiple types of target objects in the sample slice images are both greater than a set threshold; the label set contains a plurality of sample slice images annotating various target objects; Using the test set and the validation set to test and validate the trained image recognition basic model, and determining the trained, tested, and validated image recognition basic model as the spatial-spectral joint model; Obtaining an image to be recognized, and inputting the image to be recognized into the spatial-spectral joint model to recognize multiple types of target objects in the image to be recognized.
2. The image recognition method based on the spatial-spectral joint model according to claim 1, wherein The image recognition basic model further includes: a high-dimensional mapping module, a channel attention mechanism module, a spatial attention mechanism module, and an information fusion module; The high-dimensional mapping module is respectively connected to the spectral feature extraction network and the spatial feature extraction network; the spectral feature extraction network is connected to the channel attention mechanism module; the spatial feature extraction network is connected to the spatial attention mechanism module; the channel attention mechanism module and the spatial attention mechanism module are respectively connected to the information fusion module; The high-dimensional mapping module is used to increase the dimension of the sample slice image to obtain a hyperspectral sample slice image; the spectral feature extraction network is used to extract the spectral features of the hyperspectral sample slice image to obtain an initial spectral feature map; the spatial feature extraction network is used to extract the spatial features of the hyperspectral sample slice image to obtain an initial spatial feature map; the channel attention mechanism module is used to change the channel attention weights of the initial spectral feature map to obtain an optimized spectral feature map; the spatial attention mechanism module is used to change the spatial attention weights of the initial spatial feature map to obtain an optimized spatial feature map; the information fusion module is used to perform information fusion on the optimized spectral feature map and the optimized spatial feature map to obtain feature fusion information; and the feature fusion information is used to identify multiple types of target objects in the sample slice image.
3. The image recognition method based on the spatial-spectral joint model according to claim 1, wherein The operation relationship inside the first hierarchical residual structure satisfies: In the formula, is the output of the i-th layer structure in the first hierarchical residual structure, is the input of the n-th layer structure in the first hierarchical residual structure, is the output of the (n - 1)-th layer structure in the first hierarchical residual structure, and C(·) is the first convolution operator.
4. The image recognition method based on the spatial-spectral joint model according to claim 1, wherein The operation relationship inside the second hierarchical residual structure satisfies: In the formula, is the output of the i-th layer structure in the second hierarchical residual structure, is the input of the m-th layer structure in the second hierarchical residual structure, is the output of the (m - 1)-th layer structure in the second hierarchical residual structure, and C(·) is the first convolution operator.
5. The image recognition method based on the spatial-spectral joint model according to claim 1, wherein The calculation formula for the average accuracy rate is: where average_a is the average accuracy rate, A j is the number of samples of the j-th type of target object correctly recognized, B j is the total number of samples of the j-th type of target object, and C is the total number of categories of target objects; The calculation formula for the overall accuracy rate is: In the formula, overall_a is the overall accuracy rate, E is the total number of samples of target objects of all categories, and D is the number of samples correctly identified among the target objects of all categories.
6. A computer system, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the image recognition method based on the spatial-spectral joint model according to any one of claims 1-5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image recognition method based on the spatial-spectral joint model according to any one of claims 1-5.
Citation Information
Patent Citations
Early warning method and system for vehicle lane departure in night scene
CN115880658A