Target object classification method and device based on image processing, medium and equipment
By performing block processing on water chestnut images and extracting texture features using the LBP algorithm, combined with the target object classification model, the problems of environmental interference and local feature recognition in water chestnut quality detection are solved, and the sorting accuracy is improved.
Patent Information
- Application Number
- CN202510931227.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-07
AI Technical Summary
In the existing technology of water chestnut quality inspection, the image color analysis method is easily affected by external environmental factors and cannot accurately identify local subtle features, resulting in low sorting accuracy.
The LBP (local binary pattern) algorithm is used to divide the water chestnut image into blocks, extract the texture feature vector, and combine it with the target object classification model to distinguish good and bad fruits and reduce the impact of environmental interference.
The accuracy of water chestnut quality detection has been improved, local subtle features can be accurately identified, and the impact of external environmental factors on the judgment results can be reduced.
Smart Images

Figure CN120599374A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to a method, device, medium and equipment for classifying target objects based on image processing. Background Art
[0002] In the field of water chestnut quality inspection and sorting, quickly and accurately distinguishing good from bad fruit is crucial for improving product quality and reducing labor costs. Currently, mainstream methods for distinguishing good from bad water chestnuts rely on image color analysis. These methods capture water chestnut images and use image processing algorithms to extract color features, thereby determining the quality of the water chestnuts. However, this method has significant technical defects. First, the image acquisition process is easily affected by external environmental factors, such as light intensity, shooting angle, environmental reflection, etc. These factors will cause the color of the collected water chestnut images to deviate, resulting in large errors in the judgment results of good and bad fruits based on color features. Secondly, the quality defects of water chestnuts often present localized and subtle features, such as local mildew, insect damage, etc. The existing global analysis method based on color can only obtain the overall color distribution information of water chestnuts, and cannot accurately capture these local subtle features, resulting in some water chestnuts with local defects being misjudged as good fruits, seriously affecting the sorting accuracy and product quality. Therefore, there is an urgent need for a good and bad fruit differentiation technology that can overcome external environmental interference and accurately identify the local subtle features of water chestnuts to meet the actual needs of water chestnut quality detection. Summary of the Invention
[0003] In response to the above technical problems, the present application provides a method, device, medium and equipment for object classification based on image processing, which at least partially solve the problems existing in the prior art.
[0004] In a first aspect of the present application, a method for classifying an object based on image processing is provided, the method comprising: S100 , obtaining an image of a target object to be processed; wherein the size of the image of the target object to be processed is a pixels×a pixels.
[0005] S200 , performing horizontal and vertical sliding block division on the target object image to be processed according to a sliding window of b pixels×b pixels with a step size of c pixels to obtain a plurality of pixel blocks to be processed; wherein b<a.
[0006] S300, performing LBP processing on each pixel block to be processed to obtain the target object texture feature vector T = (T1, T2, ..., T i ,…,T n ); i = 1, 2, ..., n; where n is the number of pixel blocks to be processed; n = a / b; T iis the texture feature corresponding to the i-th pixel block to be processed; each pixel point in each pixel block to be processed has a corresponding LBP value; T i =(T i,1 , T i,2 ,…,T i,j ,…,T i,m ) j = 1, 2, ..., m; m is the number of possible LBP values for each pixel in the pixel block to be processed; T i,j is the number of pixels corresponding to the possible values of the j-th LBP value in the i-th pixel block to be processed.
[0007] S400, input T into the object classification model to obtain a classification result; wherein the classification result is that the object in the to-be-processed object image is a good fruit or the object in the to-be-processed object image is a bad fruit.
[0008] In a second aspect of the present application, a device for classifying an object based on image processing is provided, the device comprising: The image acquisition unit is used to acquire an image of the target object to be processed; wherein the size of the image of the target object to be processed is a pixels×a pixels.
[0009] The blocking unit is used to perform horizontal and vertical sliding blocking on the target image to be processed according to a sliding window of b pixels × b pixels with a step size of c pixels to obtain a plurality of pixel blocks to be processed; wherein b < a.
[0010] The processing unit is used to perform LBP processing on each pixel block to be processed to obtain the target object texture feature vector T=(T1, T2, ..., T i ,…,T n ); i = 1, 2, ..., n; where n is the number of pixel blocks to be processed; n = a / b; T i is the texture feature corresponding to the i-th pixel block to be processed; each pixel point in each pixel block to be processed has a corresponding LBP value; T i =(T i,1 , T i,2 ,…,T i,j ,…,T i,m ) j = 1, 2, ..., m; m is the number of possible LBP values for each pixel in the pixel block to be processed; T i,j is the number of pixels corresponding to the possible values of the j-th LBP value in the i-th pixel block to be processed.
[0011] The classification unit is used to input T into the target object classification model to obtain a classification result; wherein the classification result is that the target object in the target object image to be processed is a good fruit or the target object in the target object image to be processed is a bad fruit.
[0012] In a third aspect of the present application, a non-transitory computer-readable storage medium is provided, in which at least one instruction or at least one program is stored, and the at least one instruction or at least one program is loaded and executed by a processor to implement the aforementioned image processing-based target object classification method.
[0013] In a fourth aspect of the present application, an electronic device is provided, comprising a processor and the above-mentioned non-transitory computer-readable storage medium.
[0014] This application has at least the following beneficial effects: The object classification method based on image processing provided in this application first acquires a fixed-size (a pixels × a pixels) image of the target object to be processed, providing a standardized data foundation for subsequent processing. Then, a sliding window of b pixels × b pixels is used to partition the target object image with a step size of c pixels, resulting in several pixel blocks to be processed. This method can subdivide the water chestnut image into multiple local regions. Since water chestnut quality defects are often localized, small features, the block-based processing allows these local defect features to be individually extracted and analyzed, effectively avoiding the problem of missing local defects due to global analysis. LBP (local binary pattern) processing is then performed on each pixel block to obtain the target object's texture feature vector T. The LBP algorithm is highly robust to lighting variations and can, to a certain extent, eliminate the influence of external environmental factors such as light intensity and ambient reflections on image feature extraction. Stable texture features can be extracted, reducing the impact of image color deviation caused by environmental changes on the judgment results, thereby improving the accuracy and reliability of feature extraction. Finally, the extracted texture feature vector T is input into the object classification model to obtain the classification result. Since stable texture information containing local subtle features has been effectively extracted, compared with the misjudgment problem caused by global color analysis in the background technology, this classification model can make judgments based on more accurate feature information, thereby greatly improving the accuracy of distinguishing good and bad water chestnuts and meeting the actual needs of water chestnut quality detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0016] Figure 1 A flow chart of the target object classification method based on image processing provided in an embodiment of the present application; Figure 2 This is a structural block diagram of the target object classification device based on image processing provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0018] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or devices.
[0019] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on this application, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement an apparatus and / or practice a method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this apparatus and / or practice this method.
[0020] Please refer to Figure 1 As shown, an embodiment of the present application provides a method for classifying an object based on image processing, the method comprising: S100 , obtaining an image of a target object to be processed; wherein the size of the image of the target object to be processed is a pixels×a pixels.
[0021] Specifically, the image of the target object to be processed is an image containing the target object to be processed. In one embodiment, the target object to be processed is a water chestnut, and the image of the target object to be processed is an image containing the water chestnut. The image of the target object to be processed can be an image containing a water chestnut to be processed that is captured from an image containing several water chestnuts; as an example: a pixel × a pixel can be 256 pixels × 256 pixels.
[0022] S200 , performing horizontal and vertical sliding block division on the target object image to be processed according to a sliding window of b pixels×b pixels with a step size of c pixels to obtain a plurality of pixel blocks to be processed; wherein b<a.
[0023] Specifically, as an example, b pixels × b pixels can be 16 pixels × 16 pixels, and the step size is 8 pixels. That is, in this embodiment, a sliding window of 16 pixels × 16 pixels with a step size of 8 pixels is used to perform horizontal and vertical sliding partitioning on the 256 pixel × 256 pixel water chestnut image to obtain multiple pixel blocks to be processed.
[0024] S300, perform LBP processing on each pixel block to be processed to obtain a texture feature vector of the target object T=(T1, T2, …, Ti, …, Tn); i=1, 2, …, n; wherein n is the number of pixel blocks to be processed; n=a / b; Ti is the texture feature corresponding to the i-th pixel block to be processed; each pixel point in each pixel block to be processed has a corresponding LBP value; Ti=(Ti, 1, Ti, 2, …, Ti, j, …, Ti, m) j=1, 2, …, m; m is the number of possible LBP values of each pixel point in the pixel block to be processed; Ti, j is the number of pixels corresponding to the possible j-th LBP value in the i-th pixel block to be processed.
[0025] Specifically, LBP (Local Binary Pattern) is an operator used to describe local texture features in an image. Its basic principle is to select a fixed-size neighborhood (such as a 3×3 neighborhood) around each image pixel, use the grayscale value of the neighborhood's center pixel as a threshold, and compare it with the grayscale values of other pixels in the neighborhood. If the grayscale value of a neighboring pixel is greater than or equal to the grayscale value of the center pixel, it is marked as 1; if it is less than the grayscale value of the center pixel, it is marked as 0. In this way, the grayscale relationship of pixels in the neighborhood is converted into a binary code, which is the LBP value corresponding to the center pixel. For example, within a 3×3 neighborhood, the grayscale values of neighboring pixels are compared clockwise with the center pixel to obtain an 8-bit binary number. Converting this to decimal is the LBP value for that pixel. Thus, the LBP value of each pixel point in each pixel block to be processed is obtained. At this time, the LBP value has 2α modes; α=8(9-1); that is, the neighborhood is 3×3 and the number of sampling points is 8, so the LBP value has 256(28) modes (expressed in 0-255). At this time, in each pixel block to be processed, the frequency of occurrence of the LBP values of all pixels is counted, that is, the number of pixels corresponding to each mode is obtained, and the texture feature vector corresponding to the pixel block to be processed is obtained, that is, m=256 at this time. In another embodiment, if the neighborhood is 5×5 and the number of sampling points is 16, then the LBP value has 216 modes, that is, m=216 at this time.
[0026] This embodiment converts image texture features into quantized codes through local pixel grayscale comparison and block statistics, which can not only accurately describe local texture details but also adapt to multi-scale texture analysis; it also has low computational complexity and can flexibly control feature dimensions by adjusting neighborhood and block parameters; it also has grayscale invariance and noise resistance, and is highly robust to illumination changes and image noise.
[0027] S400, input T into the object classification model to obtain a classification result; wherein the classification result is that the object in the to-be-processed object image is a good fruit or the object in the to-be-processed object image is a bad fruit.
[0028] Specifically, the extracted texture feature vector T is input into the object classification model to obtain the classification result. Because stable texture information containing local subtle features has been effectively extracted, compared to the misjudgment caused by global color analysis in the background art, this classification model can make judgments based on more precise feature information, significantly improving the accuracy of distinguishing good and bad water chestnuts, and meeting the actual needs of water chestnut quality inspection.
[0029] This embodiment first obtains an image of the target object to be processed of a fixed size (a pixels × a pixels), providing a standardized data foundation for subsequent processing. Then, a sliding window of b pixels × b pixels is used with a step size of c pixels to block the target object image, resulting in several pixel blocks to be processed. This allows the water chestnut image to be subdivided into multiple local regions. Since water chestnut quality defects are often localized, small features, block processing allows these local defect features to be individually extracted and analyzed, effectively avoiding the problem of missing local defects during global analysis. LBP (Local Binary Pattern) processing is then performed on each pixel block to obtain the target object's texture feature vector T. The LBP algorithm is highly robust to lighting variations and can, to a certain extent, eliminate the influence of external environmental factors such as light intensity and ambient reflections on image feature extraction. This extracts stable texture features, reduces the impact of image color deviation caused by environmental changes on the judgment results, and improves the accuracy and reliability of feature extraction. Finally, the extracted texture feature vector T is input into the object classification model to obtain the classification result. Since stable texture information containing local subtle features has been effectively extracted, compared with the misjudgment problem caused by global color analysis in the background technology, this classification model can make judgments based on more accurate feature information, thereby greatly improving the accuracy of distinguishing good and bad water chestnuts and meeting the actual needs of water chestnut quality detection.
[0030] In an exemplary embodiment of the present application, step S300 includes: S310, performing a Fourier transform on each pixel block to be processed to obtain a high-frequency component ratio list G = (G1, G2, ..., Gi, ..., Gn); wherein Gi is the high-frequency component ratio of the image corresponding to the i-th pixel block to be processed after Fourier transform; the high-frequency component ratio is the ratio of the energy in the high-frequency region to the total energy.
[0031] Specifically, a Fourier transform is performed on each pixel block to be processed. The Fourier transform decomposes the image into sine and cosine components of different frequencies, where low-frequency components correspond to smooth areas in the image (such as large color blocks and backgrounds). High-frequency components correspond to rapidly changing areas in the image, such as edges, texture details, and noise. The proportion of high-frequency components is the ratio of the energy in the high-frequency area to the total energy. The total energy is the sum of the energy of all frequency components in the frequency domain, and the energy of the high-frequency area is the sum of the energy of all frequency components in the high-frequency area (corresponding to rapidly changing parts of the image, such as edges, texture details, and noise). The larger the proportion of high-frequency components, the more complex the texture of the area (such as edges and fine textures). Conversely, the smaller the proportion of high-frequency components, the smoother the area (such as background and textureless areas).
[0032] S320: If Gi is greater than a preset high-frequency component ratio threshold, the LBP value of each pixel in Gi is obtained according to the first neighborhood radius.
[0033] Specifically, if Gi is greater than a preset high-frequency component ratio threshold, indicating a large high-frequency component ratio, this indicates a larger area of more complex textures (such as edges or fine textures). In this case, the LBP value is obtained based on the first neighborhood radius. For example, if the first neighborhood radius is 1, the LBP value of each pixel is obtained based on a 3×3 neighborhood.
[0034] S330 , obtaining the number of pixels corresponding to each possible value of the LBP value obtained according to the first neighborhood radius, to obtain T=(T1, T2, …, Ti, …, Tn).
[0035] Specifically, as an example: for each pixel point in each pixel block to be processed, in a 3×3 neighborhood, the grayscale values of the neighboring pixels and the central pixel are compared in a clockwise direction to obtain an 8-bit binary number, which is converted into a decimal number, which is the LBP value of the pixel point. At this time, the LBP value has 2α modes; α = 8 (9-1); that is, the neighborhood is 3×3 and the number of sampling points is 8, so the LBP has 256 (28) modes (expressed in 0-255). At this time, in each pixel block to be processed, the frequency of occurrence of the LBP values of all pixels is counted, that is, the number of pixels corresponding to each mode is obtained, and the texture feature vector T corresponding to the pixel block to be processed is obtained.
[0036] In this embodiment, when Gi is greater than a preset high-frequency component ratio threshold, meaning the high-frequency component accounts for a significant proportion, a smaller domain radius is selected to accurately capture microtexture features, as regions with more complex textures (such as edges and fine textures) have larger areas. Using a large neighborhood (e.g., 5×5 or 7×7) may contain multiple different high-frequency features, resulting in averaging of the grayscale comparison results and an inability to accurately reflect local details. A small neighborhood (e.g., 3×3) covers only the local area of a single high-frequency feature and can more accurately capture its grayscale variation pattern. LBP generates a binary code by comparing the grayscale of pixels within the neighborhood with the central pixel. The eight sampling points in a small neighborhood (e.g., 3×3) are closer to the central pixel and are more sensitive to local microstructures (e.g., individual edges and spots). Furthermore, when the high-frequency component ratio is high, the region may contain dense high-frequency features (e.g., fine textures). A large neighborhood will superimpose the grayscale variations of multiple features, making it difficult for LBP encoding to distinguish between different texture primitives, ultimately reducing the discriminability of the feature vector.
[0037] In an exemplary embodiment of the present application, after step S310, the method further includes: S340: If Gi is equal to or less than a preset high-frequency component ratio threshold, the LBP value of each pixel in Gi is obtained according to the second neighborhood radius; wherein the number of sampling points corresponding to the second neighborhood radius is greater than the number of sampling points corresponding to the first neighborhood radius.
[0038] S350 , obtaining the number of pixels corresponding to each possible value of the LBP value obtained according to the second neighborhood radius, to obtain T=(T1, T2, …, Ti, …, Tn).
[0039] In this embodiment, as an example, the grayscale values of neighboring pixels and the central pixel are compared in a clockwise direction to obtain a 16-bit binary number, which is converted to a decimal number, namely the LBP value of the pixel. In this case, the LBP value has 2α modes; α = 16; that is, the neighborhood is 5×5 and the number of sampling points is 16, so there are 216 LBP modes. At this time, within each pixel block to be processed, the frequency of occurrence of the LBP values of all pixels is counted, that is, the number of pixels corresponding to each mode is obtained, thereby obtaining the texture feature vector T corresponding to the pixel block to be processed. Here, the sampling points are 16 points distributed in a circular pattern.
[0040] In this embodiment, when Gi is equal to or less than a preset high-frequency component ratio threshold, that is, when the high-frequency component ratio is relatively low, it indicates that the texture changes in the local area of the image are relatively smooth (such as large smooth areas and backgrounds). In this case, due to the slow grayscale changes in areas dominated by low frequencies (such as fruit surfaces and uniform backgrounds), there is a lack of clear edges or details. If a small neighborhood (such as 3×3) is used, the LBP code may be concentrated in a few patterns (such as all 0s or all 1s) due to the high similarity of pixel values within the neighborhood, resulting in extremely weak feature expression. Increasing the neighborhood radius can introduce more pixels for comparison, artificially creating grayscale differences and enriching the LBP code. For example, in a 5×5 neighborhood, even if the central area is smooth, subtle changes in edge pixels can be captured, generating more distinctive LBP patterns.
[0041] In an exemplary embodiment of the present application, the classification result obtained according to the target object classification model has a corresponding first confidence level YZ; YZ>50%; after step S400, the method further includes: S500, obtaining the shadow ratio P of the target object in the image of the target object to be processed; wherein P meets the following condition: P=Syin / Szheng; wherein Syin is the number of pixels contained in the shadow of the target object in the image of the target object to be processed; Szheng is the number of pixels contained in the target object to be processed in the image of the target object to be processed.
[0042] Specifically, shadows can cause local grayscale anomalies in an image, potentially misclassifying them as texture defects (such as rotten areas on rotten fruit). By calculating the shadow ratio, we can assess the impact of lighting interference and camera angle interference on classification results. We then establish a mapping between shadow ratio and classification confidence, quantifying the impact of shadows on the credibility of classification results.
[0043] Syin can be determined using a method based on color space conversion and threshold segmentation. Specifically, the target image is converted from the common RGB color space to a color space more suitable for shadow analysis, such as HSI (hue, saturation, intensity). In HSI space, shadow regions generally have lower intensity values, while hue and saturation are relatively stable. A shadow index is calculated based on the various components in HSI space. The image index is then thresholded using adaptive threshold calculation methods such as the maximum between-class variance method (OTSU), or manually set based on experience. Finally, a logical judgment is performed on the corresponding image index based on the threshold, and the shadow region is extracted through masking. Finally, the number of pixels within the shadow region is counted.
[0044] S600, obtaining a second confidence level EZ corresponding to P based on P and a preset shadow ratio mapping table; wherein the preset shadow ratio mapping table includes several shadow ratios and a second confidence level corresponding to each shadow ratio; EZ∈[0,1]; P is any one of the several shadow ratios; and P is inversely proportional to EZ.
[0045] S700, obtaining an updated classification result based on YZ and EZ; wherein, if the confidence SZ of the updated classification result is greater than a preset confidence threshold, the updated classification result is the same as the classification result obtained by the target object classification model; SZ=YZ×EZ.
[0046] Specifically, P is inversely proportional to EZ. The higher the shadow ratio of the target object, the greater the probability that the classification model will misjudge due to lighting interference. Therefore, a lower EZ is used to mitigate the misleading effects of shadows on the classification results. By quantifying the shadow ratio and adjusting the confidence level, the system can identify false defects caused by uneven lighting (such as shadows being misclassified as bad fruit), preventing good fruit from being misclassified as bad.
[0047] This embodiment no longer relies on the output of a single classification model, but instead dynamically adjusts the confidence level by incorporating shadow interference factors, making the results more consistent with actual scenarios (such as natural light shadows during outdoor picking and strong light reflections on production lines), thereby improving the accuracy of the judgment results.
[0048] In another embodiment, if the result output by the model is a good fruit (i.e., the first confidence level of the result output by the model is a good fruit is greater than 50%), and the confidence level SZ of the last updated classification result is equal to or less than the preset confidence threshold, it means that the shadow ratio of the target object is relatively high, resulting in a high probability that the classification model will misjudge due to light interference, i.e., misjudge it as a good fruit. At this time, the updated classification result may be a pending fruit, which can be manually processed to reduce the misjudgment rate. As an example: if the result output by the model is a good fruit, the corresponding first confidence level YZ is 80%, and the second confidence level EZ is 0.9, then SZ=72%, and the preset confidence threshold is 75%, then the last updated classification result is a pending fruit.
[0049] Please refer to Figure 2 As shown, an embodiment of the present application provides an object classification device 100 based on image processing, the device comprising: The image acquisition unit 110 is used to acquire an image of the target object to be processed; wherein the size of the image of the target object to be processed is a pixels×a pixels.
[0050] The blocking unit 120 is used to perform horizontal and vertical sliding blocking on the target image to be processed according to a sliding window of b pixels×b pixels with a step size of c pixels to obtain a plurality of pixel blocks to be processed; wherein b<a.
[0051] The processing unit 130 is used to perform LBP processing on each pixel block to be processed to obtain a texture feature vector of the target object T=(T1, T2, …, Ti, …, Tn); i=1, 2, …, n; wherein n is the number of pixel blocks to be processed; n=a / b; Ti is the texture feature corresponding to the i-th pixel block to be processed; each pixel point in each pixel block to be processed has a corresponding LBP value; Ti=(Ti, 1, Ti, 2, …, Ti, j, …, Ti, m) j=1, 2, …, m; m is the number of possible LBP values of each pixel point in the pixel block to be processed; Ti, j is the number of pixels corresponding to the possible j-th LBP value in the i-th pixel block to be processed.
[0052] The classification unit 140 is used to input T into the object classification model to obtain a classification result; wherein the classification result is that the object in the target object image to be processed is a good fruit or the object in the target object image to be processed is a bad fruit.
[0053] An embodiment of the present application further provides a computer program product, which includes program code. When the program product is run on an electronic device, the program code is used to enable the electronic device to execute the steps of the method according to various exemplary embodiments of the present application described above in this specification.
[0054] Furthermore, although the steps of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0055] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0056] In an exemplary embodiment of the present application, an electronic device capable of implementing the above method is also provided.
[0057] Those skilled in the art will appreciate that various aspects of the present application can be implemented as systems, methods, or program products. Therefore, various aspects of the present application can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."
[0058] The electronic device according to this embodiment of the present application is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0059] The electronic device is implemented as a general-purpose computing device. Components of the electronic device may include, but are not limited to, the at least one processor, the at least one memory, and a bus connecting different system components (including the memory and the processor).
[0060] The storage stores program codes, which can be executed by the processor, so that the processor executes the steps described in the above “Exemplary Method” section of this specification according to various exemplary embodiments of the present application.
[0061] The memory may include readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory, and may further include read only memory (ROM).
[0062] The storage may also include a program / utility having a set (at least one) of program modules, such program modules including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0063] The bus may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures.
[0064] The electronic device may also communicate with one or more external devices (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device, and / or any device that enables the electronic device to communicate with one or more other computing devices (e.g., a router, modem, etc.). This communication may occur via an input / output (I / O) interface. Furthermore, the electronic device may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter. As shown in the figure, the network adapter communicates with other modules of the electronic device via a bus. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0065] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0066] In exemplary embodiments of the present application, a computer-readable storage medium is also provided, on which is stored a program product capable of implementing the aforementioned methods of this specification. In some possible implementations, various aspects of the present application may also be implemented in the form of a program product, which includes program code. When the program product is executed on a terminal device, the program code is used to cause the terminal device to execute the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of the present application.
[0067] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0068] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0069] The program code embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0070] The program code used to perform the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0071] Furthermore, the above-mentioned figures are merely illustrative of the processes included in the methods according to exemplary embodiments of the present application and are not intended to be limiting. It is readily understood that the processes illustrated in the above-mentioned figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0072] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiment of the application, the features and functions of two or more modules or units described above can be concretized in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0073] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for object classification based on image processing, characterized in that: The method comprises: S100, obtaining an image of a target object to be processed; wherein the size of the image of the target object to be processed is a pixels × a pixels; S200 , performing horizontal and vertical sliding block division on the target object image to be processed according to a sliding window of b pixels×b pixels with a step size of c pixels to obtain a plurality of pixel blocks to be processed; wherein b<a; S300, performing LBP processing on each pixel block to be processed to obtain the target object texture feature vector T = (T1, T2, ..., T i ,…,T n ); i = 1, 2, ..., n; where n is the number of pixel blocks to be processed; n = a / b; T i is the texture feature corresponding to the i-th pixel block to be processed; each pixel point in each pixel block to be processed has a corresponding LBP value; T i =(T i,1 , T i,2 ,…,T i,j ,…,T i,m ) j = 1, 2, ..., m; m is the number of possible LBP values for each pixel in the pixel block to be processed; T i,j is the number of pixels corresponding to the possible values of the j-th LBP value in the i-th pixel block to be processed; S400, input T into the object classification model to obtain a classification result; wherein the classification result is that the object in the to-be-processed object image is a good fruit or the object in the to-be-processed object image is a bad fruit.
2. The method for classifying objects based on image processing according to claim 1, characterized in that: Step S300 includes: S310, performing Fourier transform on each pixel block to be processed to obtain a high frequency component proportion list G = (G1, G2, ..., G i ,…,G n ); where G i The proportion of high-frequency components after Fourier transformation of the image corresponding to the i-th pixel block to be processed; the proportion of high-frequency components is the ratio of the energy in the high-frequency area to the total energy; S320, if G i If it is greater than the preset high-frequency component ratio threshold, G is obtained according to the first neighborhood radius. i The LBP value of each pixel in; S330, obtaining the number of pixels corresponding to each possible value of the LBP value obtained according to the first neighborhood radius, to obtain T=(T1, T2, ..., T i ,…,T n ).
3. The method for object classification based on image processing according to claim 2, characterized in that: After step S310, the method further includes: S340, if G i If it is equal to or less than the preset high-frequency component ratio threshold, G is obtained according to the second neighborhood radius. i The LBP value of each pixel in the , where the number of sampling points corresponding to the second neighborhood radius is greater than the number of sampling points corresponding to the first neighborhood radius: S350, the number of pixels corresponding to each possible value of the LBP value obtained according to the second neighborhood radius is used to obtain T=(T1, T2, ..., T i ,…,T n ).
4. The method for object classification based on image processing according to claim 3, characterized in that: The target object classification model includes a first target object classification model and a second target object classification model, wherein the first target object classification model is applicable to the target object texture feature vector obtained according to the first neighborhood, and the second target object classification model is applicable to the target object texture feature vector obtained according to the second neighborhood. Step S400 includes: S410, input T into the first target classification model or the second target classification model to obtain a classification result; wherein, if T is obtained based on the first neighborhood, then T is input into the first target classification model; if T is obtained based on the second neighborhood, then T is input into the second target classification model.
5. The method for classifying objects based on image processing according to claim 4, characterized in that: The classification result obtained according to the target object classification model has a corresponding first confidence level YZ; YZ>50%; after step S400, the method further includes: S500, obtaining the shadow ratio P of the target object in the target object image to be processed; wherein P meets the following conditions: P=S 阴 / S 整 ; Among them, S 阴 is the number of pixels contained in the shadow of the target object in the target image to be processed; S 整 is the number of pixels of the target object to be processed in the image of the target object to be processed; S600, obtaining a second confidence level EZ corresponding to P based on P and a preset shadow ratio mapping table; wherein the preset shadow ratio mapping table includes a plurality of shadow ratios and a second confidence level corresponding to each shadow ratio; EZ∈[0,1]; and P is any one of the plurality of shadow ratios; S700, obtaining an updated classification result based on YZ and EZ; wherein, if the confidence SZ of the updated classification result is greater than a preset confidence threshold, the updated classification result is the same as the classification result obtained by the target object classification model; SZ=YZ×EZ.
6. An object classification device based on image processing, characterized in that: The device comprises: An image acquisition unit is used to acquire an image of the target object to be processed; wherein the size of the image of the target object to be processed is a pixels × a pixels; A blocking unit is used to perform horizontal and vertical sliding blocking of the target image to be processed according to a sliding window of b pixels × b pixels with a step size of c pixels to obtain a plurality of pixel blocks to be processed; wherein b < a; The processing unit is used to perform LBP processing on each pixel block to be processed to obtain the target object texture feature vector T=(T1, T2, ..., T i ,…,T n ); i = 1, 2, ..., n; where n is the number of pixel blocks to be processed; n = a / b; T i is the texture feature corresponding to the i-th pixel block to be processed; each pixel point in each pixel block to be processed has a corresponding LBP value; T i =(T i,1 , T i,2 ,…,T i,j ,…,T i,m ) j = 1, 2, ..., m; m is the number of possible LBP values for each pixel in the pixel block to be processed; T i,j is the number of pixels corresponding to the possible values of the j-th LBP value in the i-th pixel block to be processed; The classification unit is used to input T into the target object classification model to obtain a classification result; wherein the classification result is that the target object in the target object image to be processed is a good fruit or the target object in the target object image to be processed is a bad fruit.
7. A non-transitory computer-readable storage medium, characterized in that The storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement the target object classification method based on image processing as described in any one of claims 1-5.
8. An electronic device, characterized in that: The device comprises a processor and the non-transitory computer-readable storage medium as claimed in claim 7.
Citation Information
Patent Citations
Fruit tree disease diagnosis method and system
CN111539293A
Image multi-classification method, device and equipment based on prior attribute atlas and medium
CN113449784A
Steel wire rope surface defect identification method based on feature fusion
CN115565011A
Plate visual positioning, identifying and grabbing method, control system and plate production line
CN117689716A
Automobile part intelligent identification method based on picture classification model
CN120148014A
Cited By
Image optimization transmission method, device, field programmable gate array and system
CN121217929A