A target labeling method and system based on multi-feature loss function fusion

By using a target labeling method based on multi-feature loss function fusion and employing the multi-dimensional loss function of the entropy weight method to constrain the target transformation model, the problems of low fruit detection efficiency and poor generalization in existing technologies are solved, and efficient and accurate automatic fruit labeling is achieved.

CN116681921BActive Publication Date: 2025-12-05BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310504776.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-06
Publication Date
2025-12-05
Estimated Expiration
2043-05-06

AI Technical Summary

Technical Problem

The difficulty of existing technologies lies in how to detect fruit. Existing target detection technologies require a large amount of manually labeled data, resulting in low efficiency. Furthermore, deep learning models have poor generalization performance and are difficult to adapt to complex and diverse orchard environments and different types of fruits, especially in terms of shape and texture feature description.

Method used

A target annotation method based on multi-feature loss function fusion is adopted. The multi-dimensional loss function of entropy weight method is used to constrain the target transformation model. The latent spatial feature map is extracted through a pre-trained feature extraction network. Color, shape and texture feature loss functions are fused to achieve efficient automatic annotation of fruit images.

Benefits of technology

It improves the generalization and domain adaptability of the target detection model, reduces the cost and time of manual annotation, and achieves accurate description of shape and texture features, making it suitable for automatic annotation of multiple types of fruits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116681921B_ABST
    Figure CN116681921B_ABST
Patent Text Reader

Abstract

The application discloses a target labeling method based on multi-feature loss function fusion, wherein the multi-feature loss function is a multi-dimensional loss function based on an entropy weight method and is used for respectively restricting the generation direction of the color, shape and texture of multiple category targets in a target conversion model training process.The method comprises the following steps: acquiring a single-category best source domain background-free target image; performing feature map visualization on the single-category best source domain background-free target image, so as to extract a feature map based on latent space; inputting the feature map based on latent space into a target conversion model supervised by a multi-dimensional loss function based on an entropy weight method, so as to obtain a subset of multi-category target domain background-free target images; fusing the single-category best source domain background-free target image and the feature map based on latent space to form a multi-modal input signal; inputting the multi-modal input signal into a target conversion network; and performing target labeling based on the target conversion network.The application further discloses a system, an electronic device and a computer readable storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing and intelligent information extraction, and particularly relates to a target labeling method and system based on multi-feature loss function fusion. BACKGROUND

[0002] With the combination of traditional agriculture and artificial intelligence technology, the construction of smart orchard has received more extensive attention in the development of fruit industry. High-precision fruit detection technology is an important basic technology in the practical application of modern smart orchard, and has wide application value in fruit positioning, fruit sorting, fruit yield prediction, fruit automatic picking and other intelligent work of smart orchard. The general method of target labeling and its application in smart orchard are becoming more and more important.

[0003] On the one hand, the target detection technology at present mostly adopts the method of deep learning, which needs to rely on a large number of labeled data sets to support the training and learning of deep learning model. Therefore, a large number of sample images need to be labeled manually to train the image labeling model, which consumes manpower and time, resulting in low image labeling efficiency and low training efficiency of image detection model. Therefore, although the target detection technology based on deep learning at present has been widely applied, it needs to rely on a large number of labeled data sets to support the training and learning of detection model, resulting in high cost of manual labeling.

[0004] Secondly, the fruit trees in real scene are densely distributed, the growth of fruit is irregular, the size is small and the occlusion is serious, resulting in strong diversity of scene environment. Due to the poor generalization performance of the deep learning model at present, researchers need to make new fruit data sets for different scene environments and different types of fruits, which greatly increases the difficulty of data set labeling work and is more time-consuming and laborious.

[0005] Thirdly, when selecting the most suitable source domain data, sometimes it is difficult to select the most suitable source domain because there is only one target in some clusters. Since the original CycleGAN network can only train the generator to achieve the effect of re-painting, it is difficult to accurately describe the shape and texture features, so there is a lack of shape and texture feature information of real target image for network fitting training.

[0006] The current technical direction includes: (1) introducing instance-level loss constraints to better regulate the generation direction of the foreground target in the image, but such a method introduces an additional manual labeling process and is not suitable for automatic fruit labeling tasks based on unsupervised learning; (2) using an Across-CycleGAN fruit conversion model that compares paths across cycles, which realizes the conversion of circular fruits to elliptical fruits by introducing a structural similarity loss function; however, the automatic target labeling method has low generalization, and cannot realize the automatic labeling task of the target domain target with large feature differences, especially shape differences.

[0007] Therefore, there is an urgent need for how to establish an automatic labeling method for a target dataset with higher generalization and stronger domain adaptability, while optimizing the generation model, so as to realize realistic conversion and reduce domain differences when the shape, color and texture change greatly. SUMMARY

[0008] In order to solve the problems in the prior art, the present application provides a target labeling method and system based on multi-feature loss function fusion, further improves the performance of the unsupervised fruit conversion model, enhances the description ability of the algorithm for fruit phenotype characteristics, and thus controls the generation direction of the fruit in the cross-fruit image conversion task with large phenotype feature differences.

[0009] The first aspect of the present application provides a target labeling method based on multi-feature loss function fusion, wherein the method is used for target labeling tasks of multiple categories, and the multi-feature loss function is a multi-dimensional loss function based on an entropy weight method, which is used to constrain the generation direction of the color, shape and texture of the target of multiple categories in the training process of the target conversion model, and includes:

[0010] S1, obtaining a single-class best source domain background-free target image; the single-class best source domain background-free target image is represented by an original RGB image;

[0011] S2, visualizing the feature map of the single-class best source domain background-free target image, so as to extract a feature map based on latent space;

[0012] S3, fusing the original RGB image and the feature map based on latent space to form a multi-modal input signal, and inputting the multi-modal input signal into a target conversion model supervised by a multi-dimensional loss function based on an entropy weight method to obtain a subset of multi-class target domain background-free target images;

[0013] S4, inputting the subset of multi-class target domain background-free target images into a target detection model, and performing target labeling based on the target detection model.

[0014] Preferably, the S2 comprises:

[0015] S21, using a pre-trained feature extraction network or a pre-trained feature encoding network as an encoder to mine the latent space of the target image;

[0016] S22, using a reverse-oriented feature visualization mapping as a decoder to highlight the latent space representation of the target feature in the target image, thereby discovering the latent feature in the target image in an unsupervised manner;

[0017] S23, extracting a latent space-based feature map based on the latent feature.

[0018] Preferably, the encoder is a serialized network VGG16, and the S21 comprises: extracting high-level semantic information of the vectorized representation of the image from the deep convolutional layer of the last layer of VGG16, the vectorized representation being a vector value y; and decoupling the features using the latent code z;

[0019] The S22 comprises: performing feature map mapping through the decoder to obtain gradient information y' of each feature in the deep convolutional layer, the gradient information y' being represented as the contribution of each channel in the convolutional layer to y, and the greater the contribution, the more important the channel, and the weight proportion of the c channels in the feature layer Conv being denoted as weight c ;weight c is represented as:

[0020]

[0021] The S23 comprises: performing back propagation, calculating the activation gradient of the image through the ReLU activation function and weighted summation, normalizing the mean value of y' in the width and height of the feature map to obtain the importance of each channel, maximizing the high-level semantic feature image in the activation target, and obtaining the shape and texture feature map FeatureMap of each type of target image after spatial decoupling, the calculation process being:

[0022]

[0023] wherein weight c represents the weight proportion of the c channels in the feature layer Conv, y represents the vector value obtained after the original image is forward propagated through the serialized network VGG16 encoder, w and h represent the width and height of the high-level semantic feature image, represents the data at the coordinate position (i, j) in the channel c of the feature layer.

[0024] Preferably, the S3 comprises:

[0025] S31, supervising the generator of the target conversion model by a multi-dimensional loss function comprising three types of loss functions, i.e., a color feature loss function L Color (), a shape feature loss function L Shape () and a texture feature loss function L Texture ();

[0026] S32, balancing the weights of the multi-dimensional loss function based on a dynamic adaptive weight method of quantifiable target phenotype features to obtain an entropy weight method-based multi-dimensional loss function;

[0027] S33, fusing the original RGB image and the latent space-based feature map to form a multi-modal input signal, inputting the multi-modal input signal into the target conversion model supervised by the entropy weight method-based multi-dimensional loss function with balanced weights to obtain a subset of multi-class target domain background-free target images.

[0028] Preferably, in S31, the color feature loss function is a cycle consistency loss function and a self-mapping loss function in a CycleGAN network; the color feature loss function is represented as:

[0029] L Color (G ST +G TS )=L Cycle (G ST +G TS )+L Identity (G ST +G TS ) (4)

[0030] The cycle consistency loss is represented as:

[0031] I Cycle (G ST +G TS )=E s~pdata(s) ||G TS (G ST (s))-s||1+E t~pdata(t) ||G ST (G TS (t))-t||1(5)

[0032] The self-mapping loss function is represented as:

[0033] L Identity (G ST +G TS )=E s~pdata(t) || s -G ST (s)||1+E s~pdata(t) ||t-G TS (t)||1 (6)

[0034] wherein G ST represents source domain feature, G TS represents target domain feature, E s~pdata(s) and E t~pdata(t) represent data distribution in source domain and target domain respectively, t and s represent image information of target domain and source domain respectively;

[0035] The shape feature loss function is based on multi-scale structural similarity index MS-SSIM, and the shape feature loss function is represented as:

[0036] L Shape (G ST +G TS )=(1-MS_SSIM(G ST (s),t))+(1-MS_SSIM(G TS (t),s)) (7)

[0037] wherein MS_SSIM represents multi-scale structural similarity index loss calculation;

[0038] The texture feature loss function is based on local binary pattern (LBP) descriptor texture feature loss function, and the texture feature loss function is represented as:

[0039] L Texture (G ST +G TS )=Pearson(LBP(G ST (s),t)+Pearson(G TS (t),s)) (8)

[0040] LBP(X,Y)=N(LBP(x C ,y C )) (9)

[0041]

[0042]

[0043] wherein Pearson represents difference size between target texture features calculated by Pearson correlation coefficient, N represents all pixel values in the whole image, x C ,y C represent center pixels, i p and i c represent two different gray values in binary pattern, s is a sign function, and P represents P neighborhood selected from the center pixel.

[0044] Preferably, the S32 comprises:

[0045] (1) Calculate the quantifiable descriptor values of the shape, color and texture features of the i-th target in the source domain and the target domain in turn, and normalize them. The normalized shape, color and texture features of the i-th target are denoted as S i ,C i ,T i , respectively.

[0046] (2) Calculate the proportion P ij of each target under different feature values, which is used to describe the difference between different feature descriptor values, as shown in formula (12):

[0047]

[0048] where P ij represents the proportion of each target under different feature values; Y ij represents different feature descriptor values, i is the target number, and j takes shape, color and texture features as three different indexes in turn.

[0049] (3) Calculate the information entropy of a group of data as shown in formula (13):

[0050]

[0051] (4) Obtain the weight of each index according to the calculation formula of information entropy as shown in formula (14):

[0052]

[0053] (5) The overall loss function L Guided-GAN of the multi-dimensional loss function based on the entropy weight method is shown as formula (3):

[0054] L Guided-GAN = W s *L Shape (G ST +G TS )+W c ·L Color (G ST +G TS )

[0055] +W t ·L Texture (G ST +G TS ) (3)

[0056] where G TS represents the generator for mapping the source domain to the target domain, G TS represents the generator for mapping the target domain to the source domain, and W sW c and W t respectively represent the weight proportion of shape, color and texture loss function assigned to the shape, color and texture by entropy weight method in the model training process.

[0057] Preferably, the method further comprises: obtaining the optimal source domain in the single-class optimal source domain background-free target image, wherein the method for obtaining the optimal source domain comprises:

[0058] extracting the appearance features of each class of target from the multi-class target foreground image;

[0059] abstracting the appearance features into specific shapes, colors and textures, and calculating the relative distances of specific shapes, colors and textures as the analysis description set of the appearance features of different targets based on multi-dimensional feature quantitative analysis method;

[0060] constructing different class description models based on multi-dimensional feature space reconstruction and feature difference division of the analysis description set, and selecting a single-class optimal source domain target image therefrom;

[0061] obtaining the optimal source domain of the target based on the single-class description model, comprising: classifying different targets according to the appearance features based on the single-class description model; and selecting the optimal source domain target image from the classification for the target domain category according to actual needs.

[0062] The second aspect of the application provides a target labeling system based on a multi-dimensional space feature model optimal source domain, comprising:

[0063] a first image acquisition module for acquiring a single-class optimal source domain background-free target image; the single-class optimal source domain background-free target image is represented by an original RGB image;

[0064] a feature map extraction module for visualizing the single-class optimal source domain background-free target image to extract a feature map based on latent space;

[0065] a second image acquisition module for fusing the original RGB image and the feature map based on latent space to form a multi-modal input signal, inputting the multi-modal input signal into a target conversion model supervised by a multi-dimensional loss function based on entropy weight method to obtain a subset of multi-class target domain background-free target images;

[0066] a target labeling module for inputting the subset of multi-class target domain background-free target images into a target detection model, and performing target labeling based on the target detection model.

[0067] The third aspect of the present application provides an electronic device comprising a processor and a memory, the memory storing a plurality of instructions, and the processor being configured to read the instructions and perform the method according to the first aspect.

[0068] The fourth aspect of the present application provides a computer-readable storage medium storing a plurality of instructions, the plurality of instructions being readable by a processor and executable to perform the method according to the first aspect.

[0069] The present application provides a target labeling method, system, electronic device and computer-readable storage medium based on a multi-dimensional space feature model optimal source domain, which has the following beneficial technical effects:

[0070] An automatic labeling method with higher generalization and stronger domain adaptability and capable of meeting different types of fruit data sets is established, the labels of the target domain target can be automatically obtained, thereby being applied to downstream intelligent agricultural projects, and the monetary cost and time cost generated when manually labeling target frames are greatly reduced (compared with the average 0.2 yuan per labeled frame in the market in the single-scene data set labeling in the prior art, 30 fruits per image, 3 minutes of labeling time per image, and at least 10,000 images per data set). BRIEF DESCRIPTION OF DRAWINGS

[0071] Figure 1 The Guided-GAN overall network architecture diagram of the present application.

[0072] Figure 2 The multi-dimensional phenotype feature extraction method flowchart based on the latent space of the present application.

[0073] Figure 3 The multi-dimensional loss function schematic diagram in the Guided-GAN model of the present application.

[0074] Figure 4 The target labeling method flowchart based on the multi-feature loss function fusion of the present application.

[0075] Figure 5 The target labeling system architecture diagram based on the multi-feature loss function fusion of the present application.

[0076] Figure 6 The electronic device structure schematic diagram of the present application. DETAILED DESCRIPTION

[0077] In order to better understand the above technical solutions, the above technical solutions will be described in detail below in combination with the drawings in the specification and specific embodiments.

[0078] Embodiment One

[0079] Referring to Figure 4 The embodiment provides a target labeling method based on multi-feature loss function fusion, wherein the target has multiple categories, the multi-feature loss function is a multi-dimensional loss function based on an entropy weight method, and the multi-dimensional loss function based on the entropy weight method is respectively used to constrain the generation direction of the color, shape and texture of the target of multiple categories in a target conversion model training process. The method comprises the following steps: S1, acquiring a single-category best source domain background-free target image; in the embodiment, the single-category best source domain background-free target image is represented by an original RGB image; S2, performing feature map visualization on the single-category best source domain background-free target image, so as to extract a feature map based on latent space; in order to effectively extract multi-dimensional phenotype features of different categories of fruits, improve the feature learning ability of an unsupervised network model, and make the converted target domain fruit image more realistic, the embodiment provides a multi-dimensional phenotype feature extraction method based on latent space, referring to Figure 1 part ② in the embodiment. The required target features are separated from an original image in a latent space decoupling manner and input into a network model for training. The implementation process of the method is as shown in Figure 2 .

[0080] At present, unsupervised learning uses unlabeled data for training and learning, so that the network is difficult to extract important semantic features, resulting in poor target feature representation ability of the unsupervised learning method. With the strong potential of latent space technology in more and more fields, applying it to the generation network to extract important features of targets in different domains can further improve the network performance, so as to realize some more complex tasks. At present, the method is widely used in the conversion task of face images. Shen et al. proposed an InterFaceGAN framework to explain the disentangled face representation information learned by the existing GAN model, and studied the properties of the encoded face semantics in the latent space, so as to realize realistic conversion of face images in different poses. Sainburg et al. proposed an automatic encoder (AE) and GAN generation network structure, which promotes the convex latent distribution by performing adversarial training on the latent space interpolation, so as to control different attributes in the target to achieve more detailed changes in the face image. However, in most automatic target labeling fields, more attention needs to be paid to the multi-dimensional phenotype features of the target. The features of the target are decomposed into multiple interpretable attributes through the latent space, so as to better extract shape and texture features.

[0081] As a preferred implementation, the S2 comprises: S21, using a pre-trained feature extraction network or a pre-trained feature encoding network as an encoder to mine the latent space of the target image; S22, using a reverse-oriented feature visualization mapping as a decoder to highlight the latent space representation of the target feature in the target image, so as to discover the latent feature in the target image in an unsupervised manner; and S23, extracting a latent space-based feature map based on the latent feature.

[0082] Since the original CycleGAN network can only train the generator to achieve the effect of re-coloring, it is difficult to accurately describe the shape and texture features, and the shape and texture feature information of the real target (fruit in this embodiment) image is also lacking for network fitting training, therefore, the latent space-based feature map in this embodiment is preferably a shape and texture feature map, of course, a person skilled in the art can also select the latent space-based feature map to be a full feature map containing color, shape and texture.

[0083] In this embodiment, in terms of the selection of the backbone network, considering that the encoder needs to correspond to the decoder structure, the serialized network VGG16 is used as the encoder. In order to better decouple the image shape and texture semantic features, the high-level semantic information of the image is extracted from the vectorized representation of the deep convolutional layer of the last layer of VGG16, and the vectorized representation is a vector value y; and the vector value y is decoupled by using the latent code z, and the feature map is mapped by the decoder to obtain the gradient information y' of each feature in the deep convolutional layer, the gradient information y' is represented as the contribution of each channel in the convolutional layer to y, and the greater the contribution, the more important the channel, and the contribution value of each channel is denoted as weight c ; then the activation gradient of the image is calculated by ReLU activation function and weighted summation, which has the advantage that the input image does not need to be adjusted, and the deep and complex feature information can also be effectively learned, and the back propagation process is guided, which limits the back propagation of the gradient less than 0, and the importance of each channel can be obtained by normalizing the average value of y' in the width and height of the feature map, which can maximize the high-level semantic feature image in the activation target, and finally the shape and texture feature map of each type of target image after spatial decoupling is obtained. The calculation process of the shape and texture feature map FeatureMap can be represented as:

[0084]

[0085]

[0086] where weight crepresents the proportion of weights for c channels in the feature layer Conv, y represents the vector value obtained after the original image is encoded by the serialized network VGG16, and w and h represent the width and height of the high-level semantic feature image, respectively, represents the data of the feature layer at the coordinate position (i, j) in the channel c.

[0087] S3, the original RGB image is fused with the feature map based on the latent space to form a multi-modal input signal, which is input into a target conversion model supervised by a multi-dimensional loss function based on an entropy weight method to obtain a subset of multi-class target domain background-free target images. In this embodiment, the implementation of S3 is to more accurately describe the phenotype of the target with large feature difference (fruit in this embodiment), and to solve the problem of single function of the loss function. The present application proposes a multi-dimensional loss function based on an entropy weight method, as shown in Figure 1 in part ③. In the fruit image conversion model, the generation direction of multi-dimensional features is better controlled, and ultimately better results are achieved in the cross-fruit conversion task with large feature differences.

[0088] As described above, since the original CycleGAN network can only train the generator to achieve the effect of re-painting, it is difficult to accurately describe the shape and texture features, and there is a lack of shape and texture feature information of real fruit images for network fitting training. The prior art may introduce an instance-level loss constraint to better regulate the generation direction of the foreground target in the image, but such an approach is not suitable for automatic fruit labeling tasks based on unsupervised learning due to the introduction of an additional manual labeling process. There is also a fruit conversion model Across-CycleGAN that compares the paths of the cycles, which realizes the conversion of a circular target to an elliptical target by introducing a structural similarity loss function, and is applied to the scene of fruit labeling. In order to better improve the generalization of the automatic fruit labeling method and realize the automatic labeling task of more types of target domain fruits, it is necessary to further improve the performance of the unsupervised fruit conversion model and enhance the description ability of the algorithm for fruit phenotype features, so as to accurately control the generation direction of the fruit in the cross-fruit image conversion task with large differences in phenotype features.

[0089] Based on this, the embodiment of the present application uses a multi-dimensional loss function to constrain the generation direction of the color, shape and texture of the fruit in the fruit conversion model training process. The design diagram of the multi-dimensional loss function in the generator of the model is as shown in Figure 3 The embodiment of the present application uses two generators and two discriminators to construct two A and B cycle training structures, and combines intra-cycle training (such as Figure 3 the direction of the intra-cycle arrow) and cross-cycle training (such as Figure 3The two loss function comparison schemes respectively accurately describe the color, shape and texture features.

[0090] As a preferred embodiment, the S3 comprises: S31, supervising the generator of the target conversion model by a multi-dimensional loss function, wherein the multi-dimensional loss function comprises three types of loss functions, respectively L Color (), L Shape () and L Texture ().

[0091] As shown in Figure 3 , the regions Domain Cycle A and Domain Cycle B are two domain cycle directions of the source domain to the target domain and the target domain to the source domain, which are used to control the generation of the color feature of the target (fruit in this embodiment); the region Across Cycle represents a cross-cycle loss function comparison path, in which the image feature information of the real target (fruit in this embodiment) is used to train the fitting network to generate simulated fruit image data, helping the model to better learn and constrain the generation of shape and texture features.

[0092] In this embodiment:

[0093] (1) For the color feature loss function: the cycle consistency loss function and the self-mapping loss function in the CycleGAN network are used in this embodiment, and the coloring effect can help the target conversion model to better control the generation of the color feature, wherein the color feature loss function is represented as:

[0094] L Color (G ST +G TS )=L Cycle (G ST +G TS )+L Identtity (G ST +G TS ) (4);

[0095] The cycle consistency loss is represented as:

[0096] L Cycle (G ST +G TS )=E s~pdata(s) ||G TS (G ST (s))-s||1+E t~psata(t) ||G ST (G TS (t))-t||1(5)

[0097] The self-mapping loss function is represented as:

[0098] L Identity (G ST +G TS )=E s~psata(t) ||s-G ST (s)||1+E s~pdata(t) ||t-G TS (t)||1 (6)

[0099] Wherein, s, G ST indicates the source domain feature, G TS indicates the target domain feature, E s~pdata(s and E t~pdata(t) respectively indicate the data distribution in the source domain and the target domain, and t and s respectively indicate the image information of the target domain and the source domain.

[0100] (2) For shape feature loss function: this embodiment adopts multi-scale structural similarity index MS-SSIM based on different size convolution kernel to adjust the image receptive field size and count the shape structure feature information of the corresponding region of the image under different scale conditions, so as to effectively distinguish the geometric difference of different category fruit images, and train the model to better adapt to the difference change of shape features between different categories of targets (fruit in this embodiment). This embodiment uses a cross-cycle comparison method to compare the original image with the converted image in another cycle, so as to better constrain the generation process of the shape feature of the target (fruit in this embodiment), and the shape feature loss function is represented as:

[0101] L Shape (G ST +G TS )=(1-MS_SSIM(G ST (s),t))+(1-MS_SSIM(G TS (t),s)) (7)

[0102] Wherein MS_SSIM represents multi-scale structural similarity index loss calculation.

[0103] (3) For texture feature loss function: in the scene of target labeling taking fruit as target, the texture feature in the fruit image is too detailed, and if only the original RGB image is used for loss function comparison, the texture feature cannot be fully expressed. Moreover, the resolution of the fruit in the data set is smaller, and the texture feature cannot be well expressed, which increases the difficulty of the image conversion model. Therefore, this embodiment designs a texture feature loss function based on local binary pattern (LBP) descriptor, which can highlight the texture loss calculation method of the regular arrangement of the target texture and its regularity, accurately describe the texture feature, and better play the performance of the image conversion model. The texture feature loss function is represented as:

[0104] LTexture (G ST +G TS )=Pearson(LBP(G ST (s),t)+Pearson(G TS (t),s)) (8)

[0105] LBP(X,Y)=N(LBP(x C ,y C )) (9)

[0106]

[0107]

[0108] Wherein Pearson represents the difference size between the texture features of the target (fruit in the embodiment) by using the Pearson correlation coefficient, N represents traversing all pixel values in the whole image, x C ,y C represents the center pixel, i p and i c respectively represent two different gray values in the binary mode, s is a sign function, P represents a P neighborhood selected from the center pixel, and it is verified by experiments that the effect is best when P is 16.

[0109] Without the constraint of paired supervision information, the distribution of two image domains is highly discrete and irregular, and in the present application, the generation direction of the visual attributes such as color, shape and texture of the fruit in the training process of the fruit conversion model is constrained by designing and using a multi-dimensional loss function, so that the multi-dimensional phenotype characteristics in the fruit conversion process can be described more accurately.

[0110] S32, the weight of the multi-dimensional loss function is balanced by the dynamic adaptive weight method based on the quantifiable target phenotype characteristics, and the multi-dimensional loss function based on the entropy weight method is obtained.

[0111] In step S31, the multi-dimensional feature loss function is added to accurately describe the characteristics of the target (fruit in the embodiment) in the training process. However, in the training process of the generative adversarial network, the total loss value is obtained by adding the loss values of each dimension of the loss function, so the weight of the loss value of each loss function when added affects the network model effect, and if the weight is not set reasonably, the model cannot be normally fitted in the training stage, so that the generation direction of describing the target characteristics is lost. Therefore, in order to balance the multi-dimensional loss function added in the embodiment of the present application and make it converge stably and accurately describe the multi-dimensional fruit phenotype characteristics, the dynamic adaptive weight method based on the quantifiable target (fruit in the embodiment) phenotype characteristics is introduced in the embodiment of the present application, which is used to balance the weight of the multi-dimensional loss function. The specific process of S32 is as follows:

[0112] (1) Calculate the quantifiable descriptor values of the shape, color and texture features of the i-th target (fruit in this embodiment) in the source domain and the target domain in turn, and normalize them. The normalized shape, color and texture features of the i-th target are denoted as S i ,C i ,T i ;

[0113] (2) Calculate the proportion P ij of each target (fruit in this embodiment) sample under different feature value labels, which is used to describe the difference between different feature descriptor values, as shown in formula (12):

[0114]

[0115] where P ij represents the proportion of each target under different feature values; Y ij represents different feature descriptor values, i is the target number, and j takes shape, color and texture features (S, C and T) as three different indexes in turn;

[0116] (3) According to the definition of information entropy in information theory, the greater the descriptor difference value of different target (fruit in this embodiment) samples, the more information they can provide in the training of the GAN model, so they need to be allocated more weights in the model training process. The information entropy of a group of data is calculated as shown in formula (13):

[0117]

[0118] (4) The weight of each index obtained according to the calculation formula of information entropy is shown in formula (14):

[0119]

[0120] The overall loss function L Guided-GAN of the multi-dimensional loss function based on the entropy weight method generated by the model generator can be represented as formula (3):

[0121] L Guided-GAN =W s ·L Shape (G ST +G TS )+W c ·L Color (G ST +G TS )

[0122] +W t ·L Texture (G ST +GTS ) (3)

[0123] wherein G TS represents a generator mapping source domain to target domain, G TS represents a generator mapping target domain to source domain, W s ,W c and W t respectively represent the weight proportion of shape, color and texture loss function assigned to the shape, color and texture loss function by entropy weight method in the model training process.

[0124] In the fruit labeling application scenario, when converting between two types of fruits, the differences between the shape, color and texture descriptors of all samples of the two types of fruits are directly compared, the specific numerical values of the differences between the fruits are automatically calculated, and the weight proportion W s ,W c ,W t of the multi-dimensional loss function is dynamically adjusted each time the training is performed, so as to better assist the network model in fitting, accelerate the convergence process, and make the generated target domain fruit image better in quality.

[0125] S33, the original RGB image is fused with the feature map based on the latent space to form a multi-modal input signal, which is input into a target conversion model supervised by a multi-dimensional loss function based on entropy weight method after weighting to obtain a subset of multi-class target domain background-free target images.

[0126] S4, the subset of multi-class target domain background-free target images is input into a target detection model, and target labeling is performed based on the target detection model.

[0127] In this embodiment, the single-class best source domain background-free target image is a single-class best source domain background-free fruit image.

[0128] As a preferred embodiment, the single-class best source domain background-free target image can be an image pre-stored by the computer device, or an image downloaded by the computer device from other devices, or an image uploaded to the computer device by other devices, or an image currently collected by the computer device.

[0129] As a preferred implementation, the method further comprises: obtaining the optimal source domain in the single-category optimal source domain background-free target image, wherein the method for obtaining the optimal source domain comprises: extracting the appearance features of each category of target from the multi-category target foreground image; abstracting the appearance features into specific shapes, colors and textures, and calculating the relative distances of specific shapes, colors and textures as the analysis description set of the appearance features of different targets based on a multi-dimensional feature quantitative analysis method; constructing different category description models based on multi-dimensional feature space reconstruction and feature difference division on the analysis description set, and selecting a single-category optimal source domain target image therefrom; and obtaining the optimal source domain of the target based on the single-category description model.

[0130] As a preferred implementation, the method for obtaining the optimal source domain of the target based on the single-category description model comprises: classifying different targets according to the appearance features based on the single-category description model; and selecting the optimal source domain target image from the classification for the target domain category of actual demand.

[0131] Embodiment two

[0132] Referring to Figure 5 , the embodiment provides a target labeling system based on a multi-dimensional space feature model optimal source domain, comprising: a first image acquisition module 101 configured to acquire a single-category optimal source domain background-free target image; in the embodiment, the single-category optimal source domain background-free target image is represented by an original RGB image; a feature map extraction module 102 configured to visualize the single-category optimal source domain background-free target image to extract a feature map based on a latent space; a second image acquisition module 103 configured to fuse the original RGB image and the feature map based on the latent space to form a multi-modal input signal, input the multi-modal input signal into a target conversion model supervised by a multi-dimensional loss function based on an entropy weight method to obtain a subset of multi-category target domain background-free target images; and a target labeling module 104 configured to input the subset of multi-category target domain background-free target images into a target detection model and perform target labeling based on the target detection model.

[0133] The embodiment further provides a memory storing a plurality of instructions for implementing the method.

[0134] As Figure 6 shown, the embodiment further provides an electronic device comprising a processor 301 and a memory 302 connected to the processor 301, wherein the memory 302 stores a plurality of instructions, the instructions can be loaded and executed by the processor to enable the processor to perform the method.

[0135] While the preferred embodiments of the application have been described, additional variations and modifications can be made to these embodiments by those skilled in the art once they have the benefit of the foregoing description without departing from the spirit and scope of the application. Accordingly, it is intended that the appended claims be interpreted as including all such variations and modifications as fall within the spirit and scope of the application. It is further intended that the disclosure of all such modifications and variations be included within the scope of the application, the terms used herein being defined solely for purposes of the description being applied thereto unless otherwise indicated.

Claims

1. A target annotation method based on multi-feature loss function fusion, characterized in that, The method is used for target annotation tasks of multiple categories. The multi-feature loss function is a multi-dimensional loss function based on entropy weighting. The multi-dimensional loss function based on entropy weighting is used to constrain the generation direction of color, shape, and texture of targets of multiple categories during the training of the target transformation model, including: S1, Obtain the best source domain background-free target image for a single category; the best source domain background-free target image for a single category is represented by the original RGB image; S2, visualize the feature map of the single-category best source domain background-free target image, thereby extracting the feature map based on the latent space; S3, the original RGB image is fused with the feature map based on the latent space to form a multimodal input signal, which is input to a subset of background-free target images in the multi-class target domain obtained by the target transformation model supervised by the multidimensional loss function based on the entropy weight method; S4, input a subset of background-free target images from the multi-category target domain into the target detection model, and perform target labeling based on the target detection model.

2. The target annotation method based on multi-feature loss function fusion according to claim 1, characterized in that, S2 includes: S21, using a pre-trained feature extraction network or a pre-trained feature encoding network as an encoder to mine the latent space of the target image; S22, using the reverse guided feature visualization map as the solution space representation of the target features in the target image to highlight the decoder, thereby discovering the latent features in the target image in an unsupervised manner; S23, extract a feature map based on the latent space based on the latent features.

3. The target annotation method based on multi-feature loss function fusion according to claim 2, characterized in that, The encoder is a serialization network VGG16, and S21 includes: extracting high-level semantic information from the vectorized representation of the image output from the last deep convolutional layer of VGG16, wherein the vectorized representation is a vector value y; and decoupling the vector value y using a latent code z. S22 includes: mapping feature maps using a decoder to obtain gradient information y' of each feature in the deep convolutional layer. The gradient information y' represents the contribution of each channel in the convolutional layer to y. The larger the contribution, the more important the channel. The weight ratio of c channels in the feature layer Conv is denoted as weight. c weight c Represented as: S23 includes: performing backpropagation, calculating the activation gradient of the image using the ReLU activation function and weighted summation, normalizing the mean of the width and height of y' in the feature map to obtain the importance of each channel, maximizing the activation of high-level semantic feature images in the target, and obtaining the shape and texture feature maps of various target images after spatial decoupling. The calculation process is as follows: Where weight c Let represent the weight percentage of each of the c channels in the feature layer Conv, y represent the vector value obtained after forward propagation of the original image through the VGG16 encoder of the serialization network, and w and h represent the width and height of the high-level semantic feature image, respectively. This represents the data of the feature layer at coordinate position (i,j) in channel c.

4. The target annotation method based on multi-feature loss function fusion according to claim 3, characterized in that, S3 includes: S31, the generator of the target transformation model is supervised by a multidimensional loss function, which includes three types of loss functions, namely the color feature loss function L. Color (), Shape feature loss function L Shape () and texture feature loss function L Texture (); S32, after balancing the weights of the multidimensional loss function based on the dynamic adaptive weighting method of quantifiable target phenotypic features, a multidimensional loss function based on the entropy weighting method is obtained. S33, the original RGB image is fused with the feature map based on the latent space to form a multimodal input signal, which is then input into the target transformation model supervised by the multidimensional loss function based on the entropy weight method after the weighting is balanced, to obtain a subset of background-free target images in the multi-class target domain.

5. The target annotation method based on multi-feature loss function fusion according to claim 4, characterized in that, In step S31, the color feature loss function is the cycle-consistent loss function and the self-mapping loss function in the CycleGAN network; the color feature loss function is expressed as: L Color (G ST +G TS )=L Cycle (G ST +G TS )+L Identity (G ST +G TS ) (4) The cycle-consistent loss is expressed as: L Cycle (G ST +G TS )=E s~pdata(s) ||G TS (G ST (s))-s||1+E t~pdata(t) ||G ST (G TS (t))-t||1 (5) The self-mapping loss function is expressed as: L Identity (G ST +G TS )=E s~pdata(t) ||s-G ST (s)||1+E s~pdata(t) ||t-G TS (t)||1 (6) Among them, G ST G represents the source domain features. TS E represents the target domain features. s~pdata(s) and E t~pdata(t) Let represent the data distribution in the source domain and the target domain, respectively, and t and s represent the image information in the target domain and the source domain, respectively; The shape feature loss function is based on the multi-scale structural similarity index MS-SSIM, and the shape feature loss function is expressed as follows: L Shape (G ST +G TS )=(1-MS_SSIM(G ST (s),t))+(1-MS_SSIM(G TS (t),s)) (7) MS_SSIM represents the loss calculation based on the multi-scale structural similarity index; The texture feature loss function is a texture feature loss function based on the Local Binary Pattern (LBP) descriptor, and the texture feature loss function is expressed as follows: L Texture (G ST +G TS )=Pearson(LBP(G ST (s),t)+Pearson(G TS (t),s)) (8) LBP(X,Y)=N(LBP(x C ,y C )) (9) Where Pearson represents the magnitude of the difference between target texture features calculated using the Pearson correlation coefficient, N represents the number of pixel values ​​traversed throughout the entire image, and x C ,y C Indicates the center pixel, i p and i c These represent two different grayscale values ​​in binary mode, where s is the sign function and P represents the P-neighborhood selected from the center pixel.

6. The target annotation method based on multi-feature loss function fusion according to claim 5, characterized in that, S32 includes: (1) Calculate the quantifiable descriptor values ​​of the shape, color, and texture features of the i-th target in the source and target domains respectively, and normalize them. The normalized shape, color, and texture features of the i-th target are denoted as S. i C i ,T i ; (2) Calculate the proportion P of each target under different eigenvalues. ij , is used to describe the magnitude of the differences in numerical values ​​of different feature descriptors, as shown in formula (12): P ij And ij / Σ j AND ij (12) Among them, P ij Y represents the proportion of each target under different eigenvalues; ij The values ​​represent different feature descriptors, where i is the target number and j takes shape, color and texture features as three different indicators in turn. (3) Calculate the information entropy of a set of data as shown in formula (13): (4) The weights of each indicator are obtained according to the formula for calculating information entropy, as shown in formula (14): (5) The overall loss function L of the multidimensional loss function based on the entropy weighting method Guided-GAN Represented as formula (3): L Guided-GAN =W s ·L Shape (G ST +G TS )+W c ·L Color (G ST +G TS )+W t ·L Texture (G ST +G ST ) (3) Among them G ST G represents a generator that maps the source domain to the target domain. ST W represents a generator that maps the target domain to the source domain. s W c and W t These represent the weight proportions assigned to the shape, color, and texture loss functions using the entropy weighting method during model training.

7. The target annotation method based on multi-feature loss function fusion according to claim 1, characterized in that, The method further includes: The optimal source region in the background-free target image of the single category is obtained, wherein the optimal source region is obtained by means of: Extract the appearance features of each target category from the multi-category target foreground images; The appearance features are abstracted into specific shapes, colors, and textures. Based on the multidimensional feature quantitative analysis method, the relative distances of specific shapes, colors, and textures are calculated for different target features as an analysis description set of appearance features for different targets. Based on the multi-dimensional feature space reconstruction and feature difference classification of the analysis description set, different category description models are constructed, and the best source domain target image of a single category is selected from them. Obtaining the optimal source domain of a target based on the single-category description model includes: classifying different targets according to their appearance features based on the single-category description model; and selecting the optimal source domain target image from the classifications based on the target domain types required in practice.

8. A target labeling system based on an optimal source domain using a multidimensional spatial feature model, used to implement the method described in any one of claims 1-7, characterized in that, include: The first image acquisition module is used to acquire the best source domain background-free target image for a single category. The single-category optimal source domain background-free target image is represented by the original RGB image; The feature map extraction module is used to visualize the feature map of the single-category best source domain background-free target image, thereby extracting a feature map based on the latent space. The second image acquisition module is used to fuse the original RGB image with a feature map based on latent space to form a multimodal input signal, which is input to a subset of background-free target images in the multi-class target domain obtained by a target conversion model supervised by a multidimensional loss function based on entropy weighting. The target annotation module is used to input a subset of background-free target images from the multi-category target domain into the target detection model and to perform target annotation based on the target detection model.

9. An electronic device comprising a processor and a memory, the memory storing a plurality of instructions, the processor being configured to read the instructions and execute the method as claimed in claims 1-7.

10. A computer-readable storage medium storing a plurality of instructions that can be read by a processor and executed according to the method of claims 1-7.

Citation Information

Patent Citations

  • Unsupervised unpaired image translation method based on attention generator network

    CN113837290A

  • High-resolution SAR (Synthetic Aperture Radar) image ground feature element extraction method based on depth unsupervised multi-step adversarial domain self-adaption

    CN115049841A