A method, device, equipment and medium for identifying the age of tangerine peel
Through artificial intelligence technology, feature extraction and global context feature processing of dried tangerine peel images are solved, and the problem of inaccurate identification of the year in complex environments is achieved, achieving higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202310861712.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-07-13
AI Technical Summary
When identifying the year of tangerine peel, existing network models are easily disturbed by factors such as light, temperature, and humidity, resulting in inaccurate identification of features such as color, size and shape, making it difficult to accurately distinguish the year of tangerine peel.
Using artificial intelligence technology, feature extraction is performed by acquiring dried tangerine peel images, including the formation of initial feature information, global context feature extraction, dimensionality reduction processing and year probability calculation, and the convolutional neural network and attention mechanism are used to strengthen feature extraction, form target feature images and perform year recognition.
It improves the accuracy and robustness of the identification of tangerine peel year, and can effectively distinguish the years of tangerine peel in complex environments, prevent fraudulent behavior, and ensure product quality and safety.
Smart Images

Figure CN116958796B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment, and medium for identifying the age of tangerine peel. Background Art
[0002] Tangerine peel is an agricultural product obtained by drying and sunning citrus fruits of the Rutaceae family. It is rich in flavonoid substances, which have obvious effects in promoting qi flow, promoting digestion, stopping vomiting, and resolving phlegm. The total flavonoid content of tangerine peel varies in different years. The higher the storage years, the greater the flavonoid content. The quality and efficacy of tangerine peel are closely related to its age.
[0003] There are some unscrupulous merchants in society who, in order to obtain higher profits, will mix low-age tangerine peel into high-age tangerine peel, or use industrial dyeing technology to pass off low-age tangerine peel as high-age tangerine peel. In order to ensure product quality, prevent fraud, ensure product safety, and improve market competitiveness, it is extremely important to correctly identify the age of tangerine peel.
[0004] Currently, with the development of technology, network models have been introduced for identifying the age of tangerine peel. These network models identify the age based on the color, size, and shape of tangerine peel in images. However, in actual application scenarios, tangerine peel is affected by factors such as light, temperature, and humidity, making it easy for existing network models to be unable to accurately identify the age of tangerine peel due to changes in features such as color, size, and shape. Summary of the Invention
[0005] The main purpose of the embodiments of this application is to propose a method, device, equipment, and medium for identifying the age of tangerine peel, aiming to improve the accuracy of the identification result of the age of tangerine peel.
[0006] To achieve the above object, in the first aspect of the embodiments of this application, a method for identifying the age of tangerine peel is proposed. The method includes the following steps:
[0007] Obtain a tangerine peel image;
[0008] Extract features from the tangerine peel image to obtain initial feature information, and form an initial feature image based on the initial feature information;
[0009] Extract global context features from the initial feature image to obtain global context feature information, and form a target feature image based on the global context feature information;
[0010] Perform dimensionality reduction processing on the target feature image to obtain a feature map array, where the feature map array includes multiple feature map representative values;
[0011] Calculate the year probabilities based on multiple feature map representative values to obtain multiple year probability results;
[0012] Determine the Chenpi year recognition result according to the multiple year probability results, and output the Chenpi year recognition result.
[0013] In some possible embodiments of the present application, the global context feature extraction of the initial feature image to obtain global context feature information, and forming a target feature image according to the global context feature information includes:
[0014] Perform first-type feature extraction on the initial feature image to obtain a first-type feature map;
[0015] Perform second-type feature extraction on the first-type feature map to obtain a second-type feature map;
[0016] Perform feature map fusion according to the first-type feature map and the second-type feature map to obtain an output feature image.
[0017] In some possible embodiments of the present application, after performing feature map fusion according to the first-type feature map and the second-type feature map to obtain an output feature image, the method further includes:
[0018] Perform iterative feature extraction on the output feature image through third-type feature extraction processing and fourth-type feature extraction processing to obtain the output feature image for the next iterative feature extraction;
[0019] When the iteration number corresponding to the iterative feature extraction is equal to the preset iteration number threshold, determine the output feature image obtained by the last iterative feature extraction as the target feature image.
[0020] In some possible embodiments of the present application, the output feature image includes a plurality of first-channel feature information;
[0021] When the iteration number corresponding to the iterative feature extraction meets the preset downsampling condition, the multiple iterative feature extractions of the output feature image through third-type feature extraction and fourth-type feature extraction to obtain the output feature image for the next iterative feature extraction include:
[0022] Perform normalization processing on each of the first-channel feature information to obtain a normalized feature image;
[0023] Perform downsampling processing on the normalized feature image to obtain the output feature image for the current iterative feature extraction.
[0024] In some possible embodiments of the present application, the dimensionality reduction processing of the target feature image to obtain a feature array includes:
[0025] Perform a pooling operation on the target feature image to obtain multiple feature map values;
[0026] Normalize each of the feature map values to obtain a representative value of the feature map corresponding to each of the feature map values, and obtain the feature map array based on the multiple representative values of the feature maps.
[0027] In some possible embodiments of the present application, before performing the first type of feature extraction on the initial feature image to obtain the first type of feature map, the method includes:
[0028] Normalize the initial feature image to obtain the initial feature image for the first iteration of feature extraction.
[0029] In some possible embodiments of the present application, the iterative feature extraction of the tangerine peel in the output feature image through the third type of feature extraction and the fourth type of feature extraction to obtain the output feature image for the next iteration of feature extraction includes:
[0030] Perform the first type of feature extraction on the output feature image for the current iteration of feature extraction to obtain the corresponding first type of feature map;
[0031] Perform the second type of feature extraction on the first type of feature map to obtain the corresponding second type of feature map;
[0032] Perform feature map fusion based on the corresponding first type of feature map and the second type of feature map to obtain the output feature image for the next iteration of feature extraction.
[0033] To achieve the above object, a second aspect of the embodiments of the present application proposes an identification device for the year of tangerine peel, and the identification device includes:
[0034] An image acquisition module, configured to acquire a tangerine peel image;
[0035] An initial feature image acquisition module, configured to perform feature extraction on the tangerine peel image to obtain initial feature information, and form an initial feature image according to the initial feature information;
[0036] A target feature image acquisition module, configured to perform global context feature extraction on the initial feature image to obtain global context feature information, and form a target feature image according to the global context feature information;
[0037] A feature map array acquisition module, configured to perform dimensionality reduction processing on the target feature image to obtain a feature map array, where the feature map array includes multiple representative values of feature maps;
[0038] A year probability result acquisition module, configured to calculate year probabilities based on multiple representative values of the feature maps, and obtain multiple year probability results;
[0039] An aged tangerine peel year prediction module, configured to determine an aged tangerine peel year recognition result based on multiple year probability results, and output the aged tangerine peel year recognition result.
[0040] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0041] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0042] An aged tangerine peel year recognition method, device, equipment and medium provided by the present application, the method includes: acquiring an aged tangerine peel image; performing feature extraction on the aged tangerine peel image to obtain initial feature information, and forming an initial feature image according to the initial feature information; performing global context feature extraction on the initial feature image to obtain global context feature information, and forming a target feature image according to the global context feature information; performing dimensionality reduction processing on the target feature image to obtain a feature map array, where the feature map array includes multiple representative values of the feature maps; calculating year probabilities based on multiple representative values of the feature maps to obtain multiple year probability results; determining an aged tangerine peel year recognition result based on multiple year probability results, and outputting the aged tangerine peel year recognition result. By recognizing global context features, the effect of aged tangerine peel feature extraction is enhanced, and the accuracy of the aged tangerine peel year recognition result is improved. Description of the Drawings
[0043] Figure 1 is a step schematic diagram of the recognition method provided by the embodiments of the present application;
[0044] Figure 2 is Figure 1 a sub-step schematic diagram of step S103 in
[0045] Figure 3 is Figure 1 a sub-step schematic diagram of another embodiment of step S103 in
[0046] Figure 4 is Figure 3 a sub-step schematic diagram of step S301 in
[0047] Figure 5 is Figure 1Schematic diagram of sub-steps in step S104;
[0048] Figure 6 is Figure 5 Flowchart of step S502 in;
[0049] Figure 7 Schematic structural diagram of the recognition device provided by an embodiment of the present application;
[0050] Figure 8 Schematic hardware structure diagram of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0051] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, but not to limit the present application.
[0052] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the description, claims and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence.
[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0054] Tangerine peel is an agricultural product obtained by drying and sunning citrus fruits of the Rutaceae family. It is rich in flavonoids, which have obvious effects in promoting qi flow, promoting digestion, stopping vomiting, and resolving phlegm. The total flavonoid content of tangerine peel in different years will vary. The higher the storage years, the greater the flavonoid content. The quality and efficacy of tangerine peel are closely related to its years.
[0055] There are some unscrupulous merchants in society who sell tangerine peel. In order to obtain higher profits, they will mix low-year tangerine peel into high-year tangerine peel, or use industrial dyeing technology to pass off low-year tangerine peel as high-year tangerine peel. In order to ensure product quality, prevent fraud, ensure product safety, and improve market competitiveness, it is extremely important to correctly identify the years of tangerine peel.
[0056] Currently, with the development of technology, network models are introduced for the identification of the age of tangerine peel. These network models identify the age based on the color, size, and shape of tangerine peel in images. However, in actual application scenarios, tangerine peel is interfered by factors such as light, temperature, and humidity, making it difficult for existing network models to accurately identify the age of tangerine peel due to the easy change of features such as color, size, and shape.
[0057] Based on this, the embodiments of the present application provide a recognition method, device, equipment, and medium, aiming to improve the accuracy of the recognition result of the age of tangerine peel.
[0058] The recognition method, device, equipment, and medium provided by the embodiments of the present application are specifically described through the following embodiments. First, the recognition method in the embodiments of the present application is described.
[0059] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0060] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0061] The recognition method provided by the embodiments of the present application relates to the field of artificial intelligence technology. The recognition method provided by the embodiments of the present application can be applied to terminals, or to server sides, or can also be software running on terminals or server sides. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the recognition method, etc., but is not limited to the above forms.
[0062] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0063] See also Figure 1 , Figure 1 It is a schematic diagram of the steps of the identification method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S106.
[0064] Step S101, obtaining a tangerine peel image.
[0065] It should be understood that the specific methods for obtaining the tangerine peel image here are various, and can be the examples given below or other feasible examples. Those skilled in the art can determine the specific method for obtaining the tangerine peel image according to actual needs, and this application does not limit this.
[0066] For example, when the present application is applied to a terminal, an image of tangerine peel is photographed by a camera built into the terminal, and the year is identified by the photographed image of tangerine peel.
[0067] For example, when the present application is applied to the cloud or server, the tangerine peel image can be photographed by a terminal device, and the tangerine peel image can be uploaded to the cloud or server for online year identification.
[0068] Step S102, extracting features from the tangerine peel image to obtain initial feature information, and forming an initial feature image based on the initial feature information.
[0069] It should be understood that the tangerine peel image used for feature extraction here can be pre-processed to facilitate the extraction of initial feature information. Exemplarily, the tangerine peel image of step S101 is cropped to obtain a tangerine peel image of a specified size.
[0070] It should be understood that the initial feature information here is used to represent the characteristics of tangerine peel. These features are obtained by extracting features from multiple original channels in the tangerine peel image through a set convolution kernel. Some are low-level features of tangerine peel, and some are features irrelevant to the identification of the tangerine peel year. The number of original channels is expanded by means of feature extraction to extract useful information in the follow-up.
[0071] It should be understood that the features of tangerine peel here include low-level features and high-level features. Low-level features refer to the concave oil chamber points on the surface of tangerine peel, the bristle-like lines around the concave oil chamber points, and the texture features such as typhoon scars and color features; high-level features refer to the surface shape features of tangerine peel and the features related to the size of tangerine peel.
[0072] Taking a common RGB three-primary-color picture as an example, the tangerine peel image of the three primary colors is formed by the fusion of three channels of different colors. Feature extraction is performed on each channel through a preset convolution kernel to obtain multiple extended channels. The number of extended channels obtained is related to the size of the convolution kernel. The multiple extended channels are fused to form an initial feature image. Since the convolution kernel has performed convolution operations with the pixel points of each channel, for the initial feature image formed by the fusion of multiple extended channels, it will deepen the places with deep pixel points and weaken the places with shallow pixel points. Therefore, the contrast of the pixel points on the initial feature image will be amplified by feature extraction, making the contrast between features and non-features more obvious and reducing the loss of effective information.
[0073] However, the specific manifestations of each feature on the tangerine peel image are related not only to the tangerine peel itself of the year to be identified but also to the photographing factors. The ambient light when photographing the tangerine peel will affect the color performance of the tangerine peel in the tangerine peel image, and the shooting angle of the tangerine peel will affect the performance of the high-level features of the tangerine peel in the tangerine peel image. If only high-level features and color features are used for year identification, the above-mentioned photographing factors will reduce the accuracy of identifying the tangerine peel year. Therefore, the utilization of image information can be improved by extracting the texture features of tangerine peel, and the accuracy of year identification can be improved by the low-level features of tangerine peel composed of color features and texture features.
[0074] Step S103: Extract global context features from the initial feature image to obtain global context feature information, and form a target feature image according to the global context feature information.
[0075] Specifically, each channel of the initial feature image contains high-level features and low-level features. Through the method of global context feature extraction, the low-level features of tangerine peel are strengthened by using the high-level features of tangerine peel.
[0076] It should be understood that the specific methods for global context feature extraction here are diverse. Exemplarily, such as the global context extraction through the linear multi-head self-attention mechanism (LMS) in a neural network model established based on the convolutional neural networks (CNN) architecture and the Transformer model. Those skilled in the art can determine the technical methods for global context feature extraction according to the actual situation.
[0077] In the actual application embodiments of this application, in order to strengthen the extraction of low-level features of tangerine peel to enhance the accuracy of identifying the age of tangerine peel and improve the recognition speed, the global context features are extracted through the channel attention mechanism and the spatial attention mechanism. Exemplarily, the global context feature extraction here is divided into the first type of feature extraction processing and the second type of feature extraction processing. The first type of feature extraction processing is used to extract hierarchical features, and the second type of processing is used to extract high-level context features. Through the fusion of feature maps with different types of context features, the extraction of low-level context features is guided by high-level context features, and the finally formed target feature image carries global context features.
[0078] It should be understood that the first type of feature extraction processing is based on the method of grouped convolution. By setting convolutional kernels with the same number of channels as the initial feature image, each convolutional kernel processes a corresponding channel to extract hierarchical feature information, and all the convolved channels will form a feature map carrying hierarchical features.
[0079] It should be understood that the context features represent the relationships between pixel points composed of each channel. Other feature information of tangerine peel, such as high-level features of tangerine peel, can be extracted through context feature extraction. The low-level features of tangerine peel are strengthened through other feature information of tangerine peel, thereby improving the accuracy of tangerine peel age recognition. However, simply directly performing context feature extraction cannot effectively utilize the information carried on the image and may interfere with the extraction of low-level features of tangerine peel. For example, on the tangerine peel image, due to image shooting reasons, there are some traces of non-tangerine peel patterns outside the tangerine peel area, and these traces will interfere with the tangerine peel age recognition. Therefore, it is necessary to suppress these interferences through the attention mechanism.
[0080] For the feature map obtained after the first type of feature extraction processing, there are several features in each extended channel it contains. Some channels only contain useless information, some channels only contain useful information, and some channels contain a mixture of useful and useless information. Useless information has no effect on tangerine peel year recognition and may even reduce the effect of tangerine peel year recognition. Therefore, by introducing a channel attention mechanism to identify the useful information in the initial feature image, the channel weight corresponding to the useful information is larger than that corresponding to the useless information. The importance of each extended channel is determined by the weight, and the different importance levels result in different feature extraction convolution kernels being used, thereby enhancing the channels corresponding to the pixel points of the relevant information and suppressing the channels corresponding to the pixel points of the useless features, thus emphasizing the useful information on the image and making the feature extraction more biased towards the target features.
[0081] For the spatial attention mechanism, based on the channel attention mechanism, the image can pay specific attention to certain positions in the image channels. For tangerine peel images, the features for tangerine peel year recognition should be located in the area where the tangerine peel is in the image. The channel attention mechanism cannot only emphasize the useful information in the area where the tangerine peel is in the image. By using the spatial attention mechanism to identify the tangerine peel area, that is, to obtain the high-level features of the tangerine peel and suppress the information outside the tangerine peel area, the position of the useful information can be emphasized and the high-level features of the tangerine peel can be obtained. For this application, each channel in the initial feature image carries various features and is mixed. After the first type of feature extraction processing, multiple first type feature channels are obtained. Each first type feature channel carries hierarchical features after the first type of feature extraction processing. For example, a certain channel carries bristle texture features, a certain channel carries a mixture of bristle texture features and oil chamber dot features, and a certain channel carries color features, etc.; through the second type of feature extraction processing for focused extraction, finer-grained features of this channel are extracted, emphasizing the useful information on the image and its position, as well as obtaining the high-level features of the tangerine peel. For example, for a channel carrying bristle texture features, the bristle texture features belong to useful information, emphasizing the bristle texture features on this channel, and by identifying the tangerine peel area, sharing the weight of the tangerine peel area into this channel to emphasize the position of the bristle texture features on this channel. The same applies to other channels, so that all image channels emphasize the positions of the useful information they carry, making the useful information strengthened again, and then generating high-level context features.
[0082] Both the first type of feature extraction process and the advanced context feature extraction process will form a feature map. After obtaining the feature map carrying the advanced context features, based on the advanced context features, low-level context feature extraction is performed on the feature map obtained after the first type of feature extraction process. The global context features are obtained by guiding the extraction of low-level context features with the advanced context features, and then a feature map carrying the global context features will be formed. After multiple iterations of the global context features through the first type of feature extraction process and the second type of feature extraction process, the finally obtained feature map is the target feature image. Compared with the initial feature image, the relevant information in the tangerine peel features is extracted and enhanced, and the useless features are suppressed, improving the contrast of the two types of features, thereby ensuring the smooth and accurate progress of the subsequent steps and enhancing the importance of tangerine peel year recognition.
[0083] Step S104: Perform dimensionality reduction processing on the target feature image to obtain a feature map array.
[0084] It should be understood that the feature map array here includes multiple feature map representative values, and each feature map representative value corresponds to each preset year, and there is a non-linear relationship between the two.
[0085] It should be understood that the specific content of the dimensionality reduction processing here is diverse. It can be the following examples or other implementation manners. Those skilled in the art can determine the specific content of the dimensionality reduction processing according to the actual situation, and the present application does not limit this.
[0086] Exemplarily, if the dimensionality reduction processing includes global pooling and equalizing each channel. Assuming the target feature image includes X feature channels, X feature map representative values are obtained, and the target feature image is represented by the X feature map representative values. Each value represents the global feature formed by carrying the low-level features and the high-level features.
[0087] Exemplarily, if the dimensionality reduction processing includes global pooling and weighted calculation. Equalize each channel. Assuming the target feature image includes X feature channels, X values are obtained. Parameter weighting is performed on the values corresponding to some channels, and these channels are represented by the finally weighted values, while other channels are represented by the equalized values of each channel, so that the subsequent discrimination results are more biased.
[0088] Step S105: Calculate the year probabilities based on the multiple feature map representative values to obtain multiple year probability results.
[0089] It should be understood that the year probability calculation here is realized by non-linearly mapping the multiple feature map representative values through an activation function. The specific activation function here is diverse. It can be the softmax function or the sigmoid function, etc. Those skilled in the art can determine the specific activation function according to the actual situation.
[0090] Step S106: Determine the Chenpi year recognition result based on multiple year probability results, and output the Chenpi year recognition result.
[0091] It should be understood that the determination of the Chenpi year recognition result based on multiple year probability results here is specifically diverse. It can be the following examples or other implementation manners. Those skilled in the art can determine the specific manner of determining the Chenpi year recognition result according to the actual situation.
[0092] Exemplarily, if the target feature image includes X feature channels, and each year probability result corresponds to a representative value of a feature map, the year corresponding to the maximum probability value among all year probability results is used as the Chenpi year recognition result.
[0093] Exemplarily, if the target feature image includes X feature channels, and each year probability result corresponds to a representative value of a feature map, in order to compensate for the error in the recognition process, some year probability results are weighted, and the year corresponding to the maximum probability value among all year probability results including the weighted year probability results is used as the Chenpi year recognition result.
[0094] It should be understood that the output structure of the Chenpi year here is diverse and specifically depends on the application device of this application.
[0095] Exemplarily, when this application is applied to a terminal, the Chenpi year recognition result can be displayed on the terminal screen.
[0096] Exemplarily, when this application is applied to the cloud or the server, the Chenpi year recognition result can be sent to the terminal that requests the Chenpi year recognition service.
[0097] The following is an embodiment of this application applied to an actual Chenpi recognition scenario.
[0098] Exemplarily, steps S102 to S106 are executed based on a trained Chenpi year recognition model. The Chenpi year recognition model includes a backbone convolutional layer, a channel normalization (Layer Normalization, LN) layer, multiple csCAM modules, a downsampling layer, a global pooling layer, and a linear layer.
[0099] In the above model, the csCAM module is used to execute steps S102 to S103. Each time the image passes through the csCAM module, one iteration of feature extraction is completed. After one iteration of feature extraction of the image, the number of channels will increase. In order to prevent the loss of effective information caused by the increase in channels, downsampling needs to be performed through the downsampling layer.
[0100] Take a picture of a piece of dried tangerine peel of unknown year to obtain the dried tangerine peel image. For step S102, input the dried tangerine peel image into the dried tangerine peel year recognition model. In the dried tangerine peel year recognition model, the dried tangerine peel image will undergo feature extraction through the backbone convolutional layer, mapping each pixel point of the model's dried tangerine peel image into a specific multi-dimensional mapping space. Feature extraction will obtain an initial feature information, and an initial feature image with a channel number more than that of the dried tangerine peel image will be formed according to the initial feature information. To prevent gradient explosion, the initial feature image is sent to the LN layer for processing and then undergoes iterative feature extraction through the csCAM module.
[0101] For step 103, it is carried out through the csCAM module. The csCAM module includes a ConvNeXt block, a Channel Squeeze and Spatial Excitation Block (cSE), and a Spatial Squeeze and Channel Excitation Block (sSE) block. The above functional blocks can complete one iteration of feature extraction. Specifically, input the initial feature image into 1 csCAM module. First, it will pass through the ConvNeXt block, which can extract hierarchical features. At this time, each channel carries the hierarchical features and together forms a feature map L1. After that, the cSE block is used to extract the first high-level context features. The first high-level context features are used to emphasize the useful information in each channel. For this purpose, the channel weights corresponding to the useful information pixel points are strengthened, and the importance of the channels is determined according to the weights. Different granularity feature extraction methods are adopted for channels with different weights to obtain the first high-level context features. All the channels after feature extraction form a feature map L2 to strengthen the information related to the required dried tangerine peel existing in the feature map L1. After inputting the feature map L2 into the sSE block, the sSE block is used to extract the second high-level context features. The second context features are used to emphasize the positions where the useful information in each channel is located, so as to strengthen again the useful information emphasized in the feature map L1. The second high-level context features will identify the dried tangerine peel area and share the area weights to each channel. All channels emphasize the positions of the useful information they carry, and strengthen again the useful information emphasized in the feature map L1. The dried tangerine peel high-level features generated by the useful information strengthened again and the identified dried tangerine peel area form a feature map L3. At this time, L3 contains the first high-level context features and the second high-level context features, and the two form the overall high-level context features. Feature fusion is carried out between L3 and L1, and the low-level context features are extracted and guided by the overall high-level context features to form the global context features. At this time, each channel will form a feature map L4, completing one iteration of feature extraction.
[0102] After completing one iteration of feature extraction, the feature map L4 generated in that iteration serves as the object for the next iteration of extraction. The feature map L4 is input into the next csCAM module, and the above steps are repeated to complete multiple iterations of extraction until passing through the last csCAM module. The feature map L4 output by the last csCAM module is the target feature image mentioned in step S103.
[0103] It should be noted that since the feature extraction operation will cause changes in the feature dimension, in order to prevent the loss of effective information, downsampling is required to retain the effective information. After completing several iterations of feature extraction, the feature map L4 needs to be sent to the downsampling layer for downsampling to expand the number of channels of the feature map L4 to retain the effective information.
[0104] For step S104, the target feature image is dimension-reduced through the global pooling layer and the LN layer to reduce it to a one-dimensional feature map array. Each value in the array represents each channel in the target feature map that carries global context features.
[0105] For steps S105 to S106, the actual year determination through the generated feature map array is essentially a multi-classification problem. The feature map array is input into the linear decision layer, and the year probabilities of the feature map array are calculated through relevant activation functions to obtain multiple year probability results. The year corresponding to the maximum year probability result is determined as the Chenpi year recognition result and then output.
[0106] Please refer to Figure 2 , Figure 2 is Figure 1 a schematic diagram of the sub-steps of step S103 in. In some possible embodiments of the present application, step S103 includes but is not limited to the following sub-steps.
[0107] Step S201, perform the first type of feature extraction on the initial feature image to obtain the first type of feature map.
[0108] For the initial feature image, the initial feature image expands several image channels after the feature extraction in step S102. Each image channel carries some image information of Chenpi features, but the degree of carrying is different for each. Taking the image formed by one image channel as an example, the entire Chenpi may not be fully displayed on the image, and some features on the Chenpi may be shown, such as some scars on the Chenpi. These scars may be important or unimportant. It is necessary to strengthen the useful information through the second type of feature extraction processing, change the weight of each channel through the context features, and each channel will be extracted by the set convolution kernel according to the weight, weakening some channels carrying useless information and strengthening some channels carrying useful information, and changing the contrast between the two. All the processed channels are synthesized to obtain the first type of feature map.
[0109] It should be understood that the specific implementation methods of the first type of feature extraction processing here are diverse. Exemplarily, hierarchical feature extraction can be performed based on the Transformer convolutional architecture; or hierarchical feature extraction can be performed based on the ConvNet architecture, etc. Those skilled in the art can determine the specific implementation method of the first type of feature extraction processing according to the actual situation, and the present application does not make any limitations in this regard.
[0110] Step S202: Perform second type of feature extraction on the first type of feature map to obtain a second type of feature map.
[0111] For the first type of feature map, the first type of feature map contains hierarchical features, but the hierarchical features are mixed in each channel of the first type of feature map. For one channel, there are some useful information and useless information in the carried hierarchy, and the proportions of the two are different. Moreover, there are multiple types of low-level tangerine peel features. The hierarchical features in some channels may only include one of them. For example, only oil chamber dots are included, then the feature map formed by this one channel only contains some small dots. Another example is that the low-level tangerine peel features are not included, etc. It is necessary to change the weights and pay attention to relevant positions based on the channel attention mechanism and the spatial attention mechanism.
[0112] First, perform feature extraction on the first type of feature map based on the channel attention mechanism. Based on different weights for each channel, for important channels, the receptive field needs to be increased to more accurately obtain the useful information under the current channel.
[0113] It should be understood that the specific forms of performing feature extraction based on the channel attention mechanism here are diverse. Exemplarily, when the computing performance is sufficient, set convolution kernels of different sizes. For channels with large weights, use a 7×7 convolution kernel to extract features. On the contrary, for channels with small weights, use a smaller convolution kernel such as 1×1; or when computing performance is prioritized, perform feature extraction on channels with larger weights by allocating multiple convolution kernels of specified sizes. For example, allocate 3 3×3 convolution kernels to replace 1 7×7 convolution kernel, etc. Those skilled in the art can determine the specific form of performing feature extraction based on the channel attention mechanism according to the actual situation, and the present application does not make any limitations in this regard.
[0114] After performing feature extraction based on the channel attention mechanism, all channels will form a feature map. Perform low-level tangerine peel feature extraction based on the spatial attention mechanism, share the weight of the tangerine peel area to each channel, enable each channel to pay attention to the location of the useful information, emphasize the location of the useful information, and enhance the useful information in each channel again, thereby obtaining a second type of feature map carrying high-level context features.
[0115] Step S203: Perform feature map fusion based on the first type of feature map and the second type of feature map to obtain an output feature image.
[0116] Specifically, the first type of feature map in step S201 carries hierarchical features, and the second type of feature map in step S202 carries high-level context features. By fusing the first type of feature map and the second type of feature map, the high-level context features act on the first type of feature map to extract low-level context features from the first type of feature map. The two are combined to form global context features, strengthening the distinction between the low-level features of tangerine peel and the non-tangerine peel features on the image, making the low-level features on the tangerine peel image more obvious, enhancing the extraction effect of tangerine peel features, and improving the accuracy of tangerine peel year recognition.
[0117] In the embodiment of the present application, by performing the first type of feature extraction and the second type of feature extraction, the context feature extraction of the tangerine peel image is completed, strengthening the distinction effect between the low-level features of the tangerine peel and the surrounding environment, making the low-level features on the tangerine peel image more obvious, enhancing the extraction effect of the tangerine peel features, and improving the accuracy of tangerine peel year recognition.
[0118] Please refer to Figure 3 , Figure 3 For Figure 1 a schematic diagram of sub-steps of another embodiment of step S103 in
[0119] Step S301: Perform iterative feature extraction on the output feature image through the third type of feature extraction process and the fourth type of feature extraction process to obtain an output feature image for the next iterative feature extraction.
[0120] It should be understood that a new feature map will be formed after each feature extraction process. Here, the multiple iterative feature extractions refer to alternately performing the third type of feature extraction and the fourth type of feature extraction on the generated feature maps.
[0121] It should be understood that the third type of feature extraction and the fourth type of feature extraction here are essentially used to extract the global context features carried in the output feature image to achieve iterative feature extraction, so as to more deeply extract the essence of tangerine peel features.
[0122] It should be understood that the third type of feature extraction processing and the fourth type of feature extraction processing here are diverse. Exemplarily, the third type of feature processing can be downsampling to retain the information missing due to feature extraction in steps S201 to S203. The fourth type of feature can be the feature extraction method in step S102, such as increasing the number of channels of the feature map, etc. Those skilled in the art can determine the specific forms of the third type of feature extraction and the fourth type of feature extraction according to the actual situation.
[0123] In the actual application embodiment of the present application, the third type of feature extraction is the same as the first type of feature extraction, and the fourth type of feature extraction is the same as the second type of feature extraction. Step S301 is implemented through but not limited to the following sub-steps:
[0124] Step S301 a, perform the first type of feature extraction on the output feature image for the current iteration of feature extraction to obtain the corresponding first type of feature map.
[0125] Step S301 b, perform the second type of feature extraction on the first type of feature map to obtain the corresponding second type of feature map.
[0126] Step S301 c, perform feature map fusion according to the corresponding first type of feature map and the second type of feature map to obtain the output feature image for the next iteration of feature extraction.
[0127] Assume that the initial feature image in step S103 is A. Perform the first type of feature extraction processing on A to obtain the feature map B (corresponding to the first type of feature map). Perform the second type of feature extraction processing on the feature map B to obtain the feature map C (corresponding to the second type of feature map). Perform feature map fusion on the feature map B and the feature map C to obtain the feature map D (corresponding to the output feature image). At this time, one iteration is completed. Perform the first type of feature extraction processing on the feature map D to obtain the feature map E (corresponding to the first type of feature map). Perform the second type of feature extraction processing on the feature map E to obtain the feature map F (corresponding to the second type of feature map). Perform feature map fusion on the feature map E and the feature map F to obtain the feature map G (corresponding to the output feature image) to complete another iteration, and so on. When the iteration stop condition is satisfied, stop the iteration and perform other processing.
[0128] Exemplarily, for another example, the third type of feature processing includes the first type of feature extraction and other processing, and the fourth type of feature processing includes the second type of feature extraction and other processing.
[0129] Step S302, when the iteration times corresponding to the iterative feature extraction is equal to the preset iteration times threshold, determine the output feature image obtained by the last iterative feature extraction as the target feature image.
[0130] It should be understood that the number of iterations corresponding to the iterative feature extraction here is used to represent that the current iterative feature extraction belongs to the nth iterative feature extraction, where n is a positive integer. The specific value of the preset iteration number threshold here is diverse, and those skilled in the art can determine the specific value of the preset iteration number threshold according to the actual performance of the computing device. This application does not make any limitations in this regard.
[0131] After multiple iterative feature extractions, the number of channels that make up the target feature image is much larger than that of the initial feature image. The large number of channels maps the features of tangerine peel to a sufficiently high feature space dimension, enabling the features of tangerine peel to be linearly classified, so that the features of tangerine peel can be classified into one of the preset years, thereby realizing year recognition in subsequent steps.
[0132] In the embodiments of this application, through multiple iterative feature extractions, the features of tangerine peel are mapped into a feature space with a higher dimension, enabling the effective utilization of the high-level features of tangerine peel and reducing the influence of the low-level features of tangerine peel on year recognition.
[0133] Please refer to Figure 4 , Figure 4 is Figure 3 a schematic diagram of the sub-steps of step S301 in. In some possible embodiments of this application, when the number of iterations corresponding to the iterative feature extraction meets the preset downsampling condition, step S301 includes but is not limited to the following sub-steps.
[0134] Step S401, perform normalization processing on each first-channel feature information to obtain a normalized feature image.
[0135] It should be understood that the specific form of the normalization processing here is diverse. Exemplarily, such as Batch Normalization (BN), local response normalization, etc. Those skilled in the art can determine the specific form of the normalization processing according to the actual situation. This application does not make any limitations in this regard.
[0136] In the embodiments of the actual application of this application, the layer normalization (Layer Normalization, LN) is adopted for the normalization processing here. By performing normalization processing on each channel through layer normalization, the feature information on each channel can be retained, improving the accuracy of the next iterative feature extraction and possibly preventing the problem of gradient explosion caused by iterative feature extraction.
[0137] Step S402, perform downsampling processing on the normalized feature image to obtain an output feature image for the current iterative feature extraction.
[0138] In an embodiment of the actual application of the present application, downsampling processing is performed on the normalized feature image after LN processing. In addition to preventing the problem of gradient explosion in subsequent iterative feature extraction, the normalized feature image after LN processing also reduces the numerical values of each channel representing the normalized feature image, reducing the computational resources consumed by the downsampling processing, and effectively retaining the advanced tangerine peel features extracted at the same time.
[0139] It should be understood that the preset downsampling conditions in this embodiment are diverse. Exemplarily, when the number of iterative feature extraction times is not 0, downsampling processing is performed on the normalized feature image after each normalization, etc.
[0140] In a specific application embodiment of the present application, the downsampling condition is that the number of iterative feature extraction times reaches a predetermined number. For example, if three iterative feature extractions are required for one downsampling, then downsampling is required after the 3rd, 6th, 9th... iterative feature extractions. The iterative feature extraction is based on the output feature image. Since the output feature image has already undergone feature extraction, some information will be missing in the output feature image entering an iterative feature. In order to prevent too much useful information missing in the subsequent iterative feature extraction, which affects the extraction accuracy and quality of the features, the channels included in the output feature image are expanded and the size of the output feature image is reduced by means of downsampling, so that the output feature image carries more useful information and prevents the loss of useful information.
[0141] Please refer to Figure 5 , Figure 5 For Figure 1 the schematic diagram of the sub-steps of step S104 in
[0142] Step S501, perform a pooling operation on the target feature image to obtain multiple feature map values.
[0143] Specifically, the target feature image contains multiple channels, and each channel carries different tangerine peel feature information. However, too high a feature dimension will cause overfitting during the final decision classification, losing generality and being unable to accurately identify the years of other tangerine peels. The dimension in each channel is reduced through the pooling operation. Finally, each channel is represented by a feature map value, and each feature map value represents the tangerine peel features carried by the corresponding channel in a low-dimensional manner. For the feature map values, non-linear mapping can be performed through an activation function so that each feature map value represents a certain tangerine peel year relationship.
[0144] It should be understood that the specific form of the pooling operation here is diverse. Exemplarily, such as max pooling, overlapping pooling, etc. Those skilled in the art can determine the specific form of the pooling operation according to actual needs, and the present application does not limit this.
[0145] Step S502: Normalize each feature map value to obtain a representative value of the feature map corresponding to each feature map value, and obtain a feature map array based on multiple representative feature map values.
[0146] It should be understood that the normalization process here is diverse. Exemplarily, such as BN processing, local response normalization, etc. Those skilled in the art can determine the specific form of the normalization process according to the actual situation, and the present application does not limit this.
[0147] In the embodiments of the actual application of the present application, the normalization process here adopts LN processing. By performing layer normalization on each channel, the feature information on each channel can be retained. For the feature map array, the feature map array represents the target feature image. Performing LN processing on each feature map value is to perform normalization processing on each channel, which not only retains the feature information of the high-level features of tangerine peel, but also speeds up the speed of identifying the year based on the high-level features of tangerine peel and prevents overfitting of the tangerine peel year recognition result.
[0148] Please refer to Figure 6 , Figure 6 which is a schematic diagram of the steps of another embodiment of the present application. In some possible embodiments of the present application, before step S201, the recognition method includes but is not limited to the following steps.
[0149] Step S601: Normalize the initial feature image to obtain an initial feature image for the first iterative feature extraction.
[0150] Specifically, when the number of iterative feature extraction times is 0, that is, the global context feature extraction of the initial feature image is completed. In order to prevent gradient explosion from occurring in the subsequent iterative feature extraction process and affecting the contrast between relevant information and irrelevant information, resulting in a reduction in subsequent feature extraction, it is necessary to normalize the output feature image that is about to enter the iterative feature extraction process.
[0151] It should be understood that the normalization process here is diverse. Exemplarily, such as BN processing, local response normalization, etc. Those skilled in the art can determine the specific form of the normalization process according to the actual situation, and the present application does not limit this.
[0152] In the embodiments of the actual application of the present application, the normalization process here adopts LN processing.
[0153] Please refer to Figure 7 , the embodiments of the present application also provide a device for identifying the year of tangerine peel, which can implement the above recognition method. The device 700 includes:
[0154] An image acquisition module 701 for acquiring tangerine peel images.
[0155] An initial feature image acquisition module 702 for extracting features from the tangerine peel image to obtain initial feature information and forming an initial feature image based on the initial feature information.
[0156] A target feature image acquisition module 703 for extracting global context features from the initial feature image to obtain global context feature information and forming a target feature image based on the global context feature information.
[0157] A feature map array acquisition module 704 for performing dimensionality reduction processing on the target feature image to obtain a feature map array.
[0158] It should be understood that the feature map array here includes multiple feature map representative values.
[0159] A year probability result acquisition module 705 for calculating year probabilities based on multiple feature map representative values to obtain multiple year probability results.
[0160] A tangerine peel year prediction module 706 for determining a tangerine peel year recognition result based on multiple year probability results and outputting the tangerine peel year recognition result.
[0161] The specific implementation manner of this recognition device is basically the same as the specific embodiments of the above recognition method and will not be elaborated here.
[0162] This application embodiment also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above recognition method is implemented. This electronic device can be any intelligent terminal including a tablet computer, etc.
[0163] Please refer to Figure 8 , Figure 8 which illustrates the hardware structure of an electronic device in another embodiment. The electronic device 900 includes:
[0164] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by this application embodiment;
[0165] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the recognition method of the embodiments of this application;
[0166] The input / output interface 903 is used to implement information input and output;
[0167] The communication interface 904 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0168] The bus 905 transmits information between the various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);
[0169] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 achieve communication connections with each other inside the device through the bus 905.
[0170] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above recognition method is implemented.
[0171] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0172] The embodiments described in the embodiments of this application are for more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0173] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0174] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0175] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0176] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0177] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0178] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0179] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0180] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0181] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM for short), random access memory (RAM for short), magnetic disks, or optical discs.
[0182] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A method for identifying the age of tangerine peel, characterized in that, The method includes the following steps: Obtain a tangerine peel image; Extract features from the tangerine peel image to obtain initial feature information, and form an initial feature image according to the initial feature information; Extract global context features from the initial feature image to obtain global context feature information, and form a target feature image according to the global context feature information; Perform dimensionality reduction processing on the target feature image to obtain a feature map array, where the feature map array includes multiple feature map representative values; Calculate the year probabilities according to the multiple feature map representative values to obtain multiple year probability results; Determine the tangerine peel year recognition result according to the multiple year probability results, and output the tangerine peel year recognition result; Wherein, the extracting global context features from the initial feature image to obtain global context feature information and forming a target feature image according to the global context feature information includes: Extract first-type features from the initial feature image to obtain first-type feature maps; the first-type feature extraction includes hierarchical feature extraction through a preset convolution kernel with the same number of channels as the initial feature image; Extract second-type features from the first-type feature maps to obtain second-type feature maps; the second-type feature extraction is used to extract high-level context feature information from the first-type feature maps; the high-level context feature information is used to perform low-level context feature extraction in the first-type feature maps to extract the global context feature information; Perform feature map fusion according to the first-type feature maps and the second-type feature maps to obtain an output feature image.
2. The recognition method according to claim 1, wherein After performing feature map fusion according to the first-type feature maps and the second-type feature maps to obtain an output feature image, the method further includes: Perform iterative feature extraction on the output feature image through third-type feature extraction processing and fourth-type feature extraction processing to obtain the output feature image for the next iterative feature extraction; When the number of iterations corresponding to the iterative feature extraction is equal to a preset iteration number threshold, determine the output feature image obtained by the last iterative feature extraction as the target feature image.
3. The recognition method according to claim 2, wherein The output feature image includes multiple first-channel feature information; When the number of iterations corresponding to the iterative feature extraction meets a preset downsampling condition, the performing iterative feature extraction on the output feature image through third-type feature extraction processing and fourth-type feature extraction processing to obtain the output feature image for the next iterative feature extraction includes: Perform normalization processing on each of the first-channel feature information to obtain a normalized feature image; Perform downsampling processing on the normalized feature image to obtain the output feature image for the current iterative feature extraction.
4. The recognition method according to claim 1, wherein The performing dimensionality reduction processing on the target feature image to obtain a feature map array includes: Perform a pooling operation on the target feature image to obtain multiple feature map values; Normalize each of the feature map values to obtain a representative value of the feature map corresponding to each feature map value, and obtain the feature map array based on multiple representative values of the feature maps.
5. The recognition method according to claim 1, wherein Before performing the first type of feature extraction on the initial feature image to obtain the first type of feature map, the method includes: Normalize the initial feature image to obtain the initial feature image for the first iteration of feature extraction.
6. The recognition method according to claim 2, wherein Performing iterative feature extraction on the output feature image through the third type of feature extraction process and the fourth type of feature extraction process to obtain the output feature image for the next iteration of feature extraction includes: Perform the first type of feature extraction on the output feature image for the current iteration of feature extraction to obtain the corresponding first type of feature map; Perform the second type of feature extraction on the first type of feature map to obtain the corresponding second type of feature map; Perform feature map fusion based on the corresponding first type of feature map and the second type of feature map to obtain the output feature image for the next iteration of feature extraction.
7. An identification device for the age of tangerine peel, characterized in that, The recognition device includes: An image acquisition module for acquiring tangerine peel images; An initial feature image acquisition module for performing feature extraction on the tangerine peel image to obtain initial feature information, and forming an initial feature image based on the initial feature information; A target feature image acquisition module for performing global context feature extraction on the initial feature image to obtain global context feature information, and forming a target feature image based on the global context feature information; A feature map array acquisition module for performing dimensionality reduction processing on the target feature image to obtain a feature map array, where the feature map array includes multiple representative values of the feature maps; A year probability result acquisition module for calculating year probabilities based on multiple representative values of the feature maps to obtain multiple year probability results; A tangerine peel year prediction module for determining a tangerine peel year recognition result based on multiple year probability results and outputting the tangerine peel year recognition result; Among them, performing global context feature extraction on the initial feature image to obtain global context feature information, and forming a target feature image based on the global context feature information includes: Perform the first type of feature extraction on the initial feature image to obtain the first type of feature map; the first type of feature extraction includes hierarchical feature extraction through a preset convolution kernel having the same number of channels as the initial feature image; Perform the second type of feature extraction on the first type of feature map to obtain the second type of feature map; the second type of feature extraction is used to extract high-level context feature information from the first type of feature map; the high-level context feature information is used to perform low-level context feature extraction in the first type of feature map to extract the global context feature information; Perform feature map fusion based on the first type of feature map and the second type of feature map to obtain an output feature image.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Method for identifying production place and aging year of pericarpium citri reticulatae
CN121540743A