Fruit damage image data enhancement method and system

By constructing a joint optimization framework for generation and segmentation, high-fidelity apple image-mask pairs are generated, solving the problems of sample scarcity and imbalance in structured light apple image detection, and achieving efficient data augmentation and improved segmentation performance.

CN122049567APending Publication Date: 2026-05-15ANHUI AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI AGRICULTURAL UNIVERSITY
Filing Date
2026-02-04
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies lack a generative data framework for structured light apple image detection, making it difficult to generate high-quality, semantically consistent image-mask pairs. This fails to effectively address the issues of sample scarcity and imbalance, and existing generative models do not directly serve the optimization of segmentation performance.

Method used

A joint optimization framework for generation and segmentation is constructed. Images are acquired and preprocessed through a structured light vision system to generate dual-channel control masks. Conditional generative adversarial networks are trained, and the segmentation network is optimized by combining cross-entropy and Jaccard loss to achieve synchronous generation and consistency between the generated images and the masks.

Benefits of technology

Generating high-fidelity, high-quality apple image-mask data with limited labeled samples improves segmentation performance and dataset applicability, and significantly enhances the accuracy and robustness of damage detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049567A_ABST
    Figure CN122049567A_ABST
Patent Text Reader

Abstract

The invention discloses a fruit damage image data enhancement method and system, and the method comprises the steps: collecting a reflection image of the surface of an apple, and carrying out the normalization and size adjustment preprocessing of the image and a marked mask; executing geometric enhancement operation on the real mask, and constructing a dual-channel control mask; taking the dual-channel control mask as condition input, and generating a corresponding structured light apple image; jointly inputting the real image and the generated image into a segmentation network for training; in the training process, the generative network and the segmentation network are incorporated into a multi-stage optimization framework for joint iterative optimization, and an optimal model is dynamically selected according to segmentation indexes of a verification set. By utilizing the embodiment of the invention, the effective data amplification can be realized under the limited labeled sample, and the texture authenticity, the region consistency and the downstream applicability of the generated data are improved through the joint optimization of the generation and segmentation models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically a method and system for enhancing fruit damage image data. Background Technology

[0002] Apple appearance quality inspection is a crucial step in fruit grading and sales, with surface damage playing a decisive role in quality assessment. Traditional manual sorting is inefficient and highly subjective, failing to meet the real-time inspection needs of large-scale production lines. Therefore, automated apple defect segmentation technology based on machine vision and deep learning has attracted significant attention. Among these technologies, structured light illumination can highlight differences in fruit surface reflection, improving defect visibility. Existing research has achieved a certain degree of automated detection using semantic segmentation networks.

[0003] However, effective training of existing segmentation models relies on a large number of real images with masked annotations. Apple damage varies in shape, size, and location, and collecting high-quality datasets is costly. Traditional data augmentation methods can only perform local transformations and cannot generate new samples to compensate for the contradiction of "healthy samples being dominant and damaged samples being scarce and diverse." Manual image synthesis, on the other hand, cannot guarantee lighting consistency and realism, and cannot simultaneously generate matching semantic masks, thus failing to support reliable training.

[0004] In the field of generative learning, image-mask joint enhancement based on generative models has become an important means to improve segmentation performance. It has been applied in medical, industrial vision, and agricultural fruit and vegetable detection to alleviate the problems of sample scarcity and imbalance. However, these methods are mostly aimed at visible light or uniformly illuminated images, optimizing the overall classification / detection accuracy of target focusing, and lack specialized mask design and pairwise synthesis mechanisms for "health-damage" dual-class structured light images.

[0005] While structured light imaging has proven effective in characterizing optical differences in fruit surface and near-surface tissues, enabling enhancement and segmentation of dark-damaged areas, current work remains within the paradigm of "enhancing and segmenting a given image." A generative data framework specifically for structured light apple images is lacking: a unified model for simultaneously generating "structured light image - healthy / damaged mask" is absent, as is a joint optimization strategy that integrates the physical characteristics of structured light imaging into the generation process. This makes it difficult to guarantee the authenticity of the synthesized data and its applicability to downstream tasks. Therefore, an integrated generation-segmentation framework combining generative models, domain optical characteristics, and mask design is urgently needed. Summary of the Invention

[0006] The purpose of this invention is to provide a method and system for enhancing fruit damage image data, which addresses the shortcomings of existing technologies. It can achieve effective data augmentation with limited labeled samples and improve the texture realism, regional consistency, and downstream applicability of the generated data through joint optimization of generation and segmentation models.

[0007] One embodiment of this application provides a method for enhancing fruit damage image data, the method comprising:

[0008] The reflection image of the apple surface is acquired by a structured light vision system, and the AC component image with enhanced visibility of the damaged area is obtained by demodulation based on the three-phase phase shift method. At the same time, the image and its labeled mask are normalized and resized.

[0009] Perform a geometric enhancement operation that preserves category information on the preprocessed real mask to generate diverse semantic masks, and construct a dual-channel control mask containing an apple channel and a damage channel based on the semantic mask;

[0010] Using the dual-channel control mask as a conditional input, the generator in the conditional generative adversarial network is trained to generate the corresponding structured light apple image, and the adversarial loss and reconstruction loss are calculated by the discriminator to constrain the realism of the generated image.

[0011] The real image and the generated image are input into the segmentation network for training. The real mask and the generated mask are used as supervision information. The segmentation performance of the damaged region is jointly optimized by cross-entropy loss and Jaccard loss.

[0012] During training, the generator network and the segmentation network are incorporated into a multi-level optimization framework for joint iterative optimization, so that the quality of the generated data directly serves to improve the segmentation performance, and the optimal model is dynamically selected based on the segmentation index of the validation set.

[0013] Optionally, the step of acquiring a reflection image of the apple surface using a structured light vision system and demodulating it using a three-phase phase-shifting method to obtain an AC component image that enhances the visibility of the damaged area, while simultaneously performing normalization and size adjustment preprocessing on the image and its labeled mask, includes:

[0014] Deploy structured light vision system hardware, including a projector, an area array industrial camera, a lens, and a polarizer. Project structured light patterns onto the surface of an apple using the projector, and simultaneously acquire reflected images using the camera.

[0015] Three grayscale reflection images with fixed phase intervals were acquired, each corresponding to a different phase shift angle, for subsequent demodulation processing;

[0016] The three acquired images are demodulated using the three-phase phase shift method to calculate the AC component image that enhances the visibility of the damaged area, and the DC component image is also calculated.

[0017] The demodulated AC component images and their corresponding manually labeled masks are preprocessed, including pixel value normalization and affine transformation-based size adjustment, and the dataset is divided into training and test sets.

[0018] Optionally, the step of performing a geometric enhancement operation on the preprocessed real mask while preserving category information to generate diverse semantic masks, and constructing a dual-channel control mask containing an apple channel and a damage channel based on the semantic masks, includes:

[0019] Geometric enhancement operations, including flipping, rotating, scaling, translating, and shearing, are performed on the preprocessed real mask to generate enhanced semantic masks with diverse shapes.

[0020] During the geometric enhancement operation, the category information of the mask remains unchanged to ensure that the label values ​​of the background, healthy apple, and damaged area are correctly preserved after enhancement;

[0021] Based on the enhanced semantic mask, an apple channel mask is constructed, where the pixel values ​​corresponding to healthy apples and damaged areas are set to 1, and the background is set to 0.

[0022] Based on the enhanced semantic mask, a damage channel mask is constructed, where the pixel value corresponding to the damaged area is set to 1, and the healthy apple and background are set to 0. The apple channel and the damage channel are then concatenated to form a dual-channel control mask.

[0023] Optionally, the step of using the dual-channel control mask as a conditional input to train the generator in the conditional generative adversarial network to generate the corresponding structured light apple image, and calculating the adversarial loss and reconstruction loss through a discriminator to constrain the realism of the generated image, includes:

[0024] A conditional generative adversarial network is constructed, consisting of a generator and a discriminator, with a dual-channel control mask used as the conditional input to the generator.

[0025] The generator generates a corresponding three-channel structured light apple image based on the input dual-channel control mask;

[0026] The discriminator simultaneously receives real images, generated images, and conditional masks, and calculates adversarial loss to evaluate the realism of the generated images;

[0027] A reconstruction loss is introduced into the generator's loss function to constrain the consistency between the generated image and the real image at the pixel level, and together with the adversarial loss, the generator is optimized.

[0028] Optionally, the step of inputting both the real image and the generated image into the segmentation network for training, and using the real mask and the generated mask as supervision information, and jointly optimizing the segmentation performance of the damaged region through cross-entropy loss and Jaccard loss, includes:

[0029] Construct a three-class semantic segmentation network based on the U-Net architecture, which simultaneously receives real images and generated images as input;

[0030] The segmentation network extracts and decodes features from the input image and outputs the predicted probability that each pixel belongs to the background, a healthy apple, or a damaged area.

[0031] The difference between the predicted mask and the true mask is calculated using the cross-entropy loss function to optimize the classification performance of the segmentation network;

[0032] We introduce the Jaccard loss function as a supplement to optimize the performance of the predicted mask and the real mask in terms of region overlap, and train the segmentation network jointly with the cross-entropy loss.

[0033] Optionally, the step of incorporating the generator network and the segmentation network into a multi-level optimization framework for joint iterative optimization during training, so that the quality of the generated data directly contributes to the improvement of segmentation performance, and dynamically selecting the optimal model based on the segmentation metrics of the validation set, includes:

[0034] A multi-level optimization framework is established to integrate the generator, discriminator, and segmentation network of the generative network into a unified training process.

[0035] The generative network is trained in the inner layer optimization to improve the quality and realism of the generated images and provide more training data for the segmentation network;

[0036] The segmentation network is trained in the outer layer optimization, and the segmentation performance is improved by using real and generated images. The segmentation performance is then fed back to guide the optimization of the generation network.

[0037] After each training round, a segmentation metric is calculated using the validation set. The learning rate is dynamically adjusted based on the metric changes, and the best-performing generative and segmentation models are saved.

[0038] Another embodiment of this application provides a fruit damage image data enhancement system, the system comprising:

[0039] The acquisition module is used to acquire the reflection image of the apple surface through the structured light vision system, and demodulate the AC component image to enhance the visibility of the damaged area based on the three-phase phase shift method. At the same time, the image and its labeled mask are normalized and resized for preprocessing.

[0040] The enhancement module is used to perform geometric enhancement operations on the preprocessed real mask while preserving category information, generate diverse semantic masks, and construct a dual-channel control mask containing an apple channel and a damage channel based on the semantic mask.

[0041] The generation module is used to train the generator in the conditional generative adversarial network with the dual-channel control mask as a conditional input, generate the corresponding structured light apple image, and calculate the adversarial loss and reconstruction loss through the discriminator to constrain the realism of the generated image.

[0042] The training module is used to input real and generated images into the segmentation network for training, and uses real and generated masks as supervision information to jointly optimize the segmentation performance of damaged regions through cross-entropy loss and Jaccard loss.

[0043] The optimization module is used to incorporate the generator network and the segmentation network into a multi-level optimization framework for joint iterative optimization during the training process. This ensures that the quality of the generated data directly contributes to the improvement of segmentation performance, and dynamically selects the optimal model based on the segmentation metrics of the validation set.

[0044] Another embodiment of this application provides a storage medium storing a computer program, wherein the computer program is configured to execute the method described in any of the preceding claims when running.

[0045] Another embodiment of this application provides an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the method described in any of the preceding claims.

[0046] Compared with existing technologies, this invention provides a method for enhancing fruit damage image data. It acquires a reflection image of the apple surface and simultaneously performs normalization and size adjustment preprocessing on the image and its labeled mask. Geometric enhancement is performed on the real mask to construct a dual-channel control mask. Using the dual-channel control mask as conditional input, a corresponding structured light apple image is generated. The real image and the generated image are jointly input into a segmentation network for training. During training, the generation and segmentation networks are incorporated into a multi-level optimization framework for joint iterative optimization. The optimal model is dynamically selected based on the segmentation metrics of the validation set, thereby achieving effective data augmentation with limited labeled samples. Through joint optimization of the generation and segmentation models, the texture realism, regional consistency, and downstream applicability of the generated data are improved. Attached Figure Description

[0047] Figure 1 Hardware structure block diagram of a computer terminal for a fruit damage image data enhancement method provided in an embodiment of the present invention;

[0048] Figure 2 A flowchart illustrating a method for enhancing fruit damage image data according to an embodiment of the present invention;

[0049] Figure 3 A flowchart illustrating another method for enhancing fruit damage image data provided in an embodiment of the present invention;

[0050] Figure 4 A schematic diagram of a training mask-image generation model provided in an embodiment of the present invention;

[0051] Figure 5 This is a schematic diagram of a training semantic segmentation model provided in an embodiment of the present invention;

[0052] Figure 6 A schematic diagram of a synthesized image provided in an embodiment of the present invention;

[0053] Figure 7 A schematic diagram of a downstream application provided by an embodiment of the present invention;

[0054] Figure 8 This is a schematic diagram of a fruit damage image data enhancement system provided in an embodiment of the present invention. Detailed Implementation

[0055] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0056] In structured light apple appearance detection tasks, researchers typically only have access to a limited amount of labeled data with an imbalanced class distribution: healthy apple samples constitute the vast majority, while damaged samples are not only scarce but also exhibit high variability in morphology, extent, and reflectance characteristics. This makes segmentation networks prone to bias towards healthy regions during training, resulting in insufficient learning of damaged regions and difficulty in achieving robust defect segmentation performance. Simultaneously, structured light imaging exhibits specific reflectance and brightness distribution patterns, making it difficult for existing image enhancement methods or general generative models to accurately simulate fruit surface texture, differences in illumination reflection, and damaged structures based on these imaging characteristics. This leads to generated images that lack sufficient realism and structural consistency, and cannot guarantee a strict correspondence with the semantic mask. Especially in the scenario of multi-class mask augmentation for healthy and damaged apples, there is a lack of effective generation methods that can simultaneously guarantee image quality, semantic consistency, and fit of reflectance characteristics, and the problem of insufficient available training samples remains unresolved.

[0057] While existing generative models possess the ability to synthesize images from masks or conditional inputs, their generation process is typically independent of the segmentation task itself. These models cannot adjust their generation strategies based on segmentation performance feedback, resulting in generated data that fails to fully meet segmentation training requirements in terms of visual quality, class boundary sharpness, and the realism of damaged textures. Furthermore, general generative-segmentation frameworks are not specifically designed for the unique reflection patterns and semantic structures of structured light apple images. The generated results often lack realism, have insufficient high-frequency texture details, or fail to guarantee semantic and visual consistency between healthy and damaged regions, thus limiting their ability to improve segmentation performance.

[0058] Therefore, the core problem this invention aims to solve is: how to automatically generate high-fidelity, high-quality, and semantically consistent apple image-mask pair data in structured light apple scenarios with limited labeled samples, and ensure that the generation process directly serves to optimize damage segmentation performance. To this end, it is necessary to construct a domain-specific method that can jointly optimize the generation model and the segmentation model, so that the generated image significantly outperforms existing technologies in terms of visual realism, reflectance feature fitting, category structure consistency, and its contribution to the segmentation task, thus achieving a high-quality data augmentation mechanism.

[0059] This invention first provides a method for enhancing fruit damage image data. This method can be applied to electronic devices, such as computer terminals, specifically ordinary computers.

[0060] The following detailed explanation uses a computer terminal as an example. Figure 1 This is a hardware structure block diagram of a computer terminal for a fruit damage image data enhancement method provided in an embodiment of the present invention. Figure 1 As shown, the computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0061] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions that, when executed, cause the processor to perform any method for enhancing fruit damage image data.

[0062] The processor provides computing and control capabilities, supporting the operation of the entire computer device.

[0063] Internal memory provides an environment for the execution of computer programs in non-volatile storage media, which, when executed by a processor, enable the processor to perform any method of fruit damage image data enhancement.

[0064] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0065] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0066] See Figures 2-7 The present invention provides a method for enhancing fruit damage image data, which may include the following steps:

[0067] S201 acquires the reflection image of the apple surface through a structured light vision system, and obtains the AC component image that enhances the visibility of the damaged area based on the three-phase phase shift method. At the same time, the image and its labeled mask are normalized and resized.

[0068] Specifically, structured light vision system hardware can be deployed, including a projector, an area array industrial camera, a lens, and a polarizer. The projector projects a structured light pattern onto the surface of the apple, and the camera simultaneously captures the reflected images.

[0069] Three grayscale reflection images with fixed phase intervals were acquired, each corresponding to a different phase shift angle, for subsequent demodulation processing;

[0070] The three acquired images are demodulated using the three-phase phase shift method to calculate the AC component image that enhances the visibility of the damaged area, and the DC component image is also calculated.

[0071] The demodulated AC component images and their corresponding manually labeled masks are preprocessed, including pixel value normalization and affine transformation-based size adjustment, and the dataset is divided into training and test sets.

[0072] S202, Perform a geometric enhancement operation that preserves category information on the preprocessed real mask to generate a variety of semantic masks, and construct a dual-channel control mask containing an apple channel and a damage channel based on the semantic mask;

[0073] Specifically, geometric enhancement operations can be performed on the preprocessed real mask, including flipping, rotating, scaling, translating and shearing, to generate enhanced semantic masks with diverse shapes;

[0074] During the geometric enhancement operation, the category information of the mask remains unchanged to ensure that the label values ​​of the background, healthy apple, and damaged area are correctly preserved after enhancement;

[0075] Based on the enhanced semantic mask, an apple channel mask is constructed, where the pixel values ​​corresponding to healthy apples and damaged areas are set to 1, and the background is set to 0.

[0076] Based on the enhanced semantic mask, a damage channel mask is constructed, where the pixel value corresponding to the damaged area is set to 1, and the healthy apple and background are set to 0. The apple channel and the damage channel are then concatenated to form a dual-channel control mask.

[0077] S203, using the dual-channel control mask as a conditional input, train the generator in the conditional generative adversarial network to generate the corresponding structured light apple image, and calculate the adversarial loss and reconstruction loss through the discriminator to constrain the authenticity of the generated image;

[0078] Specifically, a conditional generative adversarial network can be constructed, including a generator and a discriminator, with a dual-channel control mask used as the conditional input to the generator;

[0079] The generator generates a corresponding three-channel structured light apple image based on the input dual-channel control mask;

[0080] The discriminator simultaneously receives real images, generated images, and conditional masks, and calculates adversarial loss to evaluate the realism of the generated images;

[0081] A reconstruction loss is introduced into the generator's loss function to constrain the consistency between the generated image and the real image at the pixel level, and together with the adversarial loss, the generator is optimized.

[0082] S204 inputs both real and generated images into the segmentation network for training, and uses real and generated masks as supervision information to jointly optimize the segmentation performance of damaged regions through cross-entropy loss and Jaccard loss.

[0083] Specifically, a three-class semantic segmentation network based on the U-Net architecture can be constructed, which can simultaneously receive real images and generated images as input;

[0084] The segmentation network extracts and decodes features from the input image and outputs the predicted probability that each pixel belongs to the background, a healthy apple, or a damaged area.

[0085] The difference between the predicted mask and the true mask is calculated using the cross-entropy loss function to optimize the classification performance of the segmentation network;

[0086] We introduce the Jaccard loss function as a supplement to optimize the performance of the predicted mask and the real mask in terms of region overlap, and train the segmentation network jointly with the cross-entropy loss.

[0087] S205 incorporates the generator network and the segmentation network into a multi-level optimization framework for joint iterative optimization during training. This ensures that the quality of the generated data directly contributes to the improvement of segmentation performance, and the optimal model is dynamically selected based on the segmentation metrics of the validation set.

[0088] Specifically, a multi-level optimization framework can be established to integrate the generator, discriminator, and segmentation network of the generative network into a unified training process;

[0089] The generative network is trained in the inner layer optimization to improve the quality and realism of the generated images and provide more training data for the segmentation network;

[0090] The segmentation network is trained in the outer layer optimization, and the segmentation performance is improved by using real and generated images. The segmentation performance is then fed back to guide the optimization of the generation network.

[0091] After each training round, a segmentation metric is calculated using the validation set. The learning rate is dynamically adjusted based on the metric changes, and the best-performing generative and segmentation models are saved.

[0092] The purpose of this invention is to provide an automatic generation method for paired structured light apple images based on generative learning and deep segmentation networks. By utilizing a small number of real structured light apple images and masks, a joint generation-segmentation optimization framework is constructed to achieve simultaneous generation and semantically consistent modeling of healthy and damaged regions of the apple. This invention can automatically generate high-fidelity apple images and their corresponding masks to augment the training dataset, and improve the effectiveness of the generated data through a segmentation performance feedback mechanism. This method has advantages such as high data construction efficiency, stable generation quality, and adaptability to multiple damage morphologies, and can significantly improve the accuracy and robustness of structured light apple damage segmentation tasks.

[0093] Step S1: This invention employs a structured light vision system for acquiring and demodulating apple images. The system hardware includes a projector, an area array industrial camera, a lens, and a polarizer. The projector projects a structured light pattern (such as sinusoidal or binary fringes) onto the apple surface, while the camera simultaneously captures the reflected image. To reduce specular reflection, a polarizer is installed in front of the camera and projector. The demodulation of the AC (alternating current) image is based on a three-phase phase-shifting method to enhance the visibility of damaged areas on the apple.

[0094] Step S2: Acquire images of healthy and damaged apples from the structured light imaging system, and manually label the corresponding masks to form three categories of labels: background (0), healthy apple (1), and damaged area (2). Read the images and masks through the data loading module, divide the data into training and test sets, and perform preprocessing operations such as normalization and size adjustment.

[0095] Step S3: Based on the real mask, geometric enhancement operations such as flipping, rotating, scaling, translating, and shearing are performed using the imgaug library to perform category-preserving transformations on the mask, thereby synthesizing diverse healthy / damaged region structures. The enhanced mask is used as input to the generative model to generate new apple image-mask pairs.

[0096] Step S4: Initialize the Pix2Pix generative model, using the enhanced dual-channel mask (apple channel and damage channel) as conditional input, and generate the corresponding three-channel structured light apple image through the generator network. The adversarial loss and L1 reconstruction loss are calculated through the discriminator to achieve the realism constraint of the generated image.

[0097] Step S5: A three-class segmentation U-Net model is adopted, using real and generated images as inputs, and corresponding masks as supervision information. Cross-entropy loss and Jaccard loss are jointly used to optimize segmentation performance. The output of the segmentation network is used to evaluate the quality of the generated images and to guide the generation model in reverse, forming a generation-segmentation co-optimization mechanism.

[0098] Step S6: In the generation stage, a "healthy apple-specific mask" is constructed, where all apple channels are set to 1 and damaged channels are set to 0. A healthy apple image is then generated by the generator, further alleviating the problem of scarce damaged samples. The generated healthy image is then predicted again by U-Net to construct additional optimization terms, improving the balanced segmentation performance between healthy and damaged classes.

[0099] Step S7: Employing a multi-level optimization engine, the generator (G), discriminator (D), segmentation network (U-Net), and structural parameter optimization module (Arch) are integrated into a unified optimization graph. Inner-layer optimization improves the quality of generated data, while outer-layer optimization enhances the final segmentation performance, achieving a closed-loop "generation-segmentation" training process. This mechanism ensures that the generation process consistently serves to maximize segmentation performance.

[0100] Step S8: After each training round, calculate the Jaccard metric using the validation set, dynamically adjust the learning rate based on the validation results, and save the best model to achieve adaptive optimization. After training, output the final segmentation model and generation model, which can be used to automatically construct high-quality structured light apple image-mask pair datasets.

[0101] Further, in step S1, the structured light vision system includes a projector, an area array industrial camera, a lens, and a polarizer. The projector projects a sinusoidal or binary fringe pattern onto the apple surface, and the camera simultaneously acquires the reflected image. To suppress specular reflection interference, a polarizer is installed in front of the camera and the projector. Three images are acquired with a phase shift interval of... grayscale reflection image , , Its mathematical model is:

[0102]

[0103]

[0104]

[0105] Where (x,y) are pixel coordinates. DC component (background lighting) For AC component (enhancing damage contrast). denoted as the spatial frequency along the x-axis.

[0106] The AC image was obtained by demodulation using the three-phase phase-shift method:

[0107]

[0108] DC shunt image Calculated by the following formula:

[0109]

[0110] The demodulated AC image is used for subsequent damage area extraction and annotation.

[0111] Further, in step S2, a dataset of structured light apple images is established. This dataset, containing apple images and corresponding mask images, will be used for model training and testing, facilitating better pairwise image and mask generation to support data augmentation. The labeled dataset is divided into training and testing sets in an 8:2 ratio. Data preprocessing includes image normalization and resizing. Image normalization maps the original 8-bit image to the normalization range required by the deep learning model by dividing the pixels by 255, using the following formula:

[0112]

[0113] Where x and y refer to the two-dimensional spatial coordinates of a pixel. Resizing includes scaling, translation, rotation, and shearing, and is a standard affine transformation, with the formula:

[0114]

[0115] Where x, y are the pixel coordinates of the original image or original mask, and x′, y′ are the pixel coordinates of the transformed image or mask. This refers to the scaling ratio in the x and y directions, where θ is the image rotation angle in radians. This refers to the translation offset. This refers to the shearing transformation parameter, which characterizes the degree of shearing (skewing) in the affine transformation. In this embodiment, it is introduced in the form of an angle and acts together with the rotation angle θ on the trigonometric function terms of the affine transformation matrix to achieve a joint geometric transformation of rotation and shearing of the image / mask.

[0116] Furthermore, in step S3, diverse semantic masks are constructed by performing serialized geometric enhancement on the real mask to control subsequent image generation processes. The mask category mapping uses numerical quantization for enhancement processing, and the mapping relationship is as follows: ,in, These correspond to the background, healthy apple, and damaged area, respectively. The enhanced mask recovers the semantic category using the following formula:

[0117]

[0118] Here, `clip` refers to the truncation / limiting function, used to restrict the input value to a preset range [0,2]. When the input value is less than 0, the output is 0; when the input value is greater than 2, the output is 2; otherwise, the output is the input value itself, ensuring that the restored mask category index falls within the valid category range. `round` refers to the rounding function, used to map continuous input values ​​to the nearest integer value, thus restoring the enhanced grayscale mask value (normalized by 127.5) to a discrete semantic category label. Subsequently, a dual-channel control mask for conditional generation is constructed, defined as follows:

[0119]

[0120]

[0121] in, This represents the binary value of the Apple region control channel at pixel (x, y); when When the value is 1, it indicates that the pixel belongs to the apple (including healthy and damaged areas); when... When the value is 0, it indicates that the pixel belongs to the background area; This represents the semantic mask value (i.e., category label / category index) at the pixel with coordinates (x, y), used to characterize the semantic category to which the pixel belongs. This indicates the binary value of the control channel at pixel (x, y) in the damaged area; when A value of 1 indicates that the pixel belongs to the damaged area; otherwise, a value of 0 indicates that the pixel does not belong to the damaged area. The control mask input is:

[0122]

[0123] Furthermore, in step S4, a conditional generative network is used to map the semantic mask to the structured light apple image. The generator loss function includes adversarial loss and L1 reconstruction loss, defined as follows:

[0124]

[0125] Where C is the two-channel control mask and I is the real image. =20 is the reconstruction loss weight, and G is... Let represent the total loss function of the generator, used to constrain the generator output to both satisfy the adversarial discrimination requirements and approximate the real image as closely as possible. To represent the adversarial loss term, which characterizes the consistency between the generated results and the real data distribution (generated by the adversarial learning mechanism), the discriminator uses the following loss:

[0126]

[0127] Furthermore, in step S5, the U-Net network is used to perform three types of semantic segmentation. Its loss function is composed of cross-entropy and Jaccard loss, as shown in the formula:

[0128]

[0129] in, This represents the total loss function value for the semantic segmentation task. The cross-entropy loss function is used to measure the difference between the predicted mask and the true mask in terms of class distribution. For the real mask, s is the prediction mask, and s is the smoothing coefficient.

[0130] Furthermore, in step S6, to address the sample class imbalance problem, the damaged area is forcibly mapped to a healthy area to generate additional healthy apple images. The mapping relationship is as follows:

[0131]

[0132] This allows for the generation of "purely healthy" apple samples, improving the model's robustness to the health category.

[0133] Furthermore, in step S7, the generator G, discriminator D, segmentation network S, and structural parameter optimization module are jointly trained through a multi-level optimization framework. The upper-level optimization objective is to improve the performance of the segmentation network, and the lower-level optimization objective is to improve the image quality generated by the generator. Their coupling relationship can be expressed as: , , where θ represents the network parameters. This framework ensures that the generated data always serves to maximize segmentation performance.

[0134] Furthermore, in step S8, the optimal model is selected using the Jaccard index of the validation set, defined as follows:

[0135]

[0136] The learning rate is adaptively adjusted based on the validation results, and the optimal generation model and segmentation model are finally output.

[0137] This invention achieves efficient generation of structured light apple images and their masks by jointly optimizing a generative model and a segmentation network. Compared with traditional image enhancement methods, this invention can not only generate more diverse images, but also ensure semantic consistency between the generated images and the real masks, thereby effectively expanding the dataset and solving the problems of insufficient sample quantity and class imbalance.

[0138] This invention employs a specially designed mask enhancement mechanism, enabling the generated images to not only distinguish between healthy and damaged apples but also accurately identify different types of damage regions. This multi-class mask generation method significantly improves the model's performance in detecting diverse apple damage, avoiding the performance bottleneck of traditional methods on imbalanced datasets.

[0139] This invention employs coupled training of a generative model and a segmentation model, enabling the generated data to directly contribute to improved segmentation performance. Compared to traditional deep learning methods, this invention achieves higher-quality segmentation results with only a small number of real-labeled samples through a joint optimization strategy of generation and segmentation, significantly reducing the cost of data annotation.

[0140] This invention, through joint optimization of the generator and segmentation network, not only augments the dataset but also effectively ensures the quality of the generated samples. Traditional data augmentation methods often lead to inconsistencies between the image and the mask, while the generator-segmentation optimization framework of this invention guarantees semantic and structural consistency between the generated image and the mask, thereby improving the quality of the dataset and the segmentation accuracy.

[0141] As can be seen, the process involves acquiring a reflection image of the apple surface, and simultaneously performing normalization and size adjustment preprocessing on the image and its labeled mask; performing geometric enhancement on the real mask to construct a dual-channel control mask; using the dual-channel control mask as conditional input to generate the corresponding structured light apple image; inputting both the real image and the generated image into the segmentation network for training; during training, incorporating the generation network and the segmentation network into a multi-level optimization framework for joint iterative optimization, and dynamically selecting the optimal model based on the segmentation metrics of the validation set, thereby achieving effective data augmentation with limited labeled samples. Through joint optimization of the generation and segmentation models, the texture realism, regional consistency, and downstream applicability of the generated data are improved.

[0142] Another embodiment of the present invention provides a fruit damage image data enhancement system, see [link to documentation]. Figure 8 The system may include:

[0143] The acquisition module 801 is used to acquire the reflection image of the apple surface through the structured light vision system, and demodulate the AC component image to enhance the visibility of the damaged area based on the three-phase phase shift method. At the same time, the image and its labeled mask are normalized and size adjusted for preprocessing.

[0144] The enhancement module 802 is used to perform a geometric enhancement operation that preserves category information on the preprocessed real mask, generate a variety of semantic masks, and construct a dual-channel control mask containing an apple channel and a damage channel based on the semantic mask.

[0145] The generation module 803 is used to train the generator in the conditional generative adversarial network with the dual-channel control mask as a conditional input, generate the corresponding structured light apple image, and calculate the adversarial loss and reconstruction loss through the discriminator to constrain the realism of the generated image.

[0146] Training module 804 is used to input real images and generated images into the segmentation network for training, and uses real masks and generated masks as supervision information to jointly optimize the segmentation performance of damaged regions through cross-entropy loss and Jaccard loss.

[0147] The optimization module 805 is used to incorporate the generator network and the segmentation network into a multi-level optimization framework for joint iterative optimization during the training process, so that the quality of the generated data directly serves to improve the segmentation performance, and dynamically selects the optimal model based on the segmentation index of the validation set.

[0148] This invention also provides a storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.

[0149] This invention also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.

[0150] Specifically, the aforementioned electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the aforementioned processor, and the input / output device is connected to the aforementioned processor.

[0151] The above description, based on the embodiments shown in the figures, details the structure, features, and effects of the present invention. The above description is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or equivalent embodiments modified to have equivalent changes, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.

Claims

1. A method for enhancing fruit damage image data, characterized in that, The method includes: The reflection image of the apple surface is acquired by a structured light vision system, and the AC component image with enhanced visibility of the damaged area is obtained by demodulation based on the three-phase phase shift method. At the same time, the image and its labeled mask are normalized and resized. Perform a geometric enhancement operation that preserves category information on the preprocessed real mask to generate diverse semantic masks, and construct a dual-channel control mask containing an apple channel and a damage channel based on the semantic mask; Using the dual-channel control mask as a conditional input, the generator in the conditional generative adversarial network is trained to generate the corresponding structured light apple image, and the adversarial loss and reconstruction loss are calculated by the discriminator to constrain the realism of the generated image. The real image and the generated image are input into the segmentation network for training. The real mask and the generated mask are used as supervision information. The segmentation performance of the damaged region is jointly optimized by cross-entropy loss and Jaccard loss. During training, the generator network and the segmentation network are incorporated into a multi-level optimization framework for joint iterative optimization, so that the quality of the generated data directly serves to improve the segmentation performance, and the optimal model is dynamically selected based on the segmentation index of the validation set.

2. The method according to claim 1, characterized in that, The process involves acquiring a reflection image of the apple surface using a structured light vision system, demodulating it using a three-phase phase-shifting method to obtain an AC component image that enhances the visibility of the damaged area, and simultaneously performing normalization and size adjustment preprocessing on the image and its labeled mask, including: Deploy structured light vision system hardware, including a projector, an area array industrial camera, a lens, and a polarizer. Project structured light patterns onto the surface of an apple using the projector, and simultaneously acquire reflected images using the camera. Three grayscale reflection images with fixed phase intervals were acquired, each corresponding to a different phase shift angle, for subsequent demodulation processing; The three acquired images are demodulated using the three-phase phase shift method to calculate the AC component image that enhances the visibility of the damaged area, and the DC component image is also calculated. The demodulated AC component images and their corresponding manually labeled masks are preprocessed, including pixel value normalization and affine transformation-based size adjustment, and the dataset is divided into training and test sets.

3. The method according to claim 2, characterized in that, The process involves performing a geometric enhancement operation on the preprocessed real mask while preserving category information, generating diverse semantic masks, and constructing a dual-channel control mask containing an apple channel and a damage channel based on the semantic masks, including: Geometric enhancement operations, including flipping, rotating, scaling, translating, and shearing, are performed on the preprocessed real mask to generate enhanced semantic masks with diverse shapes. During the geometric enhancement operation, the category information of the mask remains unchanged to ensure that the label values ​​of the background, healthy apple, and damaged area are correctly preserved after enhancement; Based on the enhanced semantic mask, an apple channel mask is constructed, where the pixel values ​​corresponding to healthy apples and damaged areas are set to 1, and the background is set to 0. Based on the enhanced semantic mask, a damage channel mask is constructed, where the pixel value corresponding to the damaged area is set to 1, and the healthy apple and background are set to 0. The apple channel and the damage channel are then concatenated to form a dual-channel control mask.

4. The method according to claim 3, characterized in that, The process of training the generator in the conditional generative adversarial network using the dual-channel control mask as a conditional input to generate a corresponding structured light apple image, and calculating the adversarial loss and reconstruction loss through a discriminator to constrain the realism of the generated image, includes: A conditional generative adversarial network is constructed, consisting of a generator and a discriminator, with a dual-channel control mask used as the conditional input to the generator. The generator generates a corresponding three-channel structured light apple image based on the input dual-channel control mask; The discriminator simultaneously receives real images, generated images, and conditional masks, and calculates adversarial loss to evaluate the realism of the generated images; A reconstruction loss is introduced into the generator's loss function to constrain the consistency between the generated image and the real image at the pixel level, and together with the adversarial loss, the generator is optimized.

5. The method according to claim 4, characterized in that, The process of inputting both real and generated images into a segmentation network for training, and using both real and generated masks as supervision information, optimizes the segmentation performance of damaged regions through a combination of cross-entropy loss and Jaccard loss, including: Construct a three-class semantic segmentation network based on the U-Net architecture, which simultaneously receives real images and generated images as input; The segmentation network extracts and decodes features from the input image and outputs the predicted probability that each pixel belongs to the background, a healthy apple, or a damaged area. The difference between the predicted mask and the true mask is calculated using the cross-entropy loss function to optimize the classification performance of the segmentation network; We introduce the Jaccard loss function as a supplement to optimize the performance of the predicted mask and the real mask in terms of region overlap, and train the segmentation network jointly with the cross-entropy loss.

6. The method according to claim 5, characterized in that, The process of integrating the generator network and the segmentation network into a multi-level optimization framework for joint iterative optimization during training ensures that the quality of the generated data directly contributes to improving segmentation performance. Furthermore, the optimal model is dynamically selected based on the segmentation metrics of the validation set. This includes: A multi-level optimization framework is established to integrate the generator, discriminator, and segmentation network of the generative network into a unified training process. The generative network is trained in the inner layer optimization to improve the quality and realism of the generated images and provide more training data for the segmentation network; The segmentation network is trained in the outer layer optimization, and the segmentation performance is improved by using real and generated images. The segmentation performance is then fed back to guide the optimization of the generation network. After each training round, a segmentation metric is calculated using the validation set. The learning rate is dynamically adjusted based on the metric changes, and the best-performing generative and segmentation models are saved.

7. A system for enhancing fruit damage image data, characterized in that, The system includes: The acquisition module is used to acquire the reflection image of the apple surface through the structured light vision system, and demodulate the AC component image to enhance the visibility of the damaged area based on the three-phase phase shift method. At the same time, the image and its labeled mask are normalized and resized for preprocessing. The enhancement module is used to perform geometric enhancement operations on the preprocessed real mask while preserving category information, generate diverse semantic masks, and construct a dual-channel control mask containing an apple channel and a damage channel based on the semantic mask. The generation module is used to train the generator in the conditional generative adversarial network with the dual-channel control mask as a conditional input, generate the corresponding structured light apple image, and calculate the adversarial loss and reconstruction loss through the discriminator to constrain the realism of the generated image. The training module is used to input real and generated images into the segmentation network for training, and uses real and generated masks as supervision information to jointly optimize the segmentation performance of damaged regions through cross-entropy loss and Jaccard loss. The optimization module is used to incorporate the generator network and the segmentation network into a multi-level optimization framework for joint iterative optimization during the training process. This ensures that the quality of the generated data directly contributes to the improvement of segmentation performance, and dynamically selects the optimal model based on the segmentation metrics of the validation set.

8. The system according to claim 7, characterized in that, The acquisition module is specifically used for: Deploy structured light vision system hardware, including a projector, an area array industrial camera, a lens, and a polarizer. Project structured light patterns onto the surface of an apple using the projector, and simultaneously acquire reflected images using the camera. Three grayscale reflection images with fixed phase intervals were acquired, each corresponding to a different phase shift angle, for subsequent demodulation processing; The three acquired images are demodulated using the three-phase phase shift method to calculate the AC component image that enhances the visibility of the damaged area, and the DC component image is also calculated. The demodulated AC component images and their corresponding manually labeled masks are preprocessed, including pixel value normalization and affine transformation-based size adjustment, and the dataset is divided into training and test sets.

9. A storage medium, characterized in that, The storage medium stores a computer program, wherein the computer program is configured to execute the method of any one of claims 1-6 when it is run.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method of any one of claims 1-6.