A multi-category target domain dataset annotation method and system based on zero annotation

Through multi-dimensional spatial feature analysis and cross-category common description models, a target domain synthetic dataset is constructed and the generation model is optimized, which solves the problem of deep learning models' dependence on manual labeling and realizes efficient and low-cost multi-category target domain dataset labeling.

CN116778223BActive Publication Date: 2025-09-30BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310505349.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-06
Publication Date
2025-09-30
Estimated Expiration
2043-05-06

AI Technical Summary

Technical Problem

In existing technologies, deep learning-based target detection models require a large amount of manually labeled data sets, resulting in high costs and poor generalization performance, making it difficult to achieve efficient automatic labeling in different scenarios and target types.

Method used

A multi-category target domain dataset annotation method based on zero annotation is adopted. Through multi-dimensional spatial feature analysis and cross-category common description model, a target domain synthetic dataset is constructed and automatically annotated. The generation model is optimized using a multi-dimensional loss function to achieve unsupervised target conversion.

Benefits of technology

It enables the training of deep learning models without manual labeling, improves the generalization and domain adaptability of the model, reduces the cost and time of manual labeling, and is suitable for labeling datasets in multiple farms, multiple varieties, and multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116778223B_ABST
    Figure CN116778223B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for labeling a multi-category target domain dataset based on zero labeling, comprising: obtaining target domain foreground images of different categories; performing quantitative analysis of multidimensional spatial features based on the target domain foreground images of different categories and constructing a cross-category common description model based on the analyzed multidimensional spatial features; obtaining the optimal source domain of the target based on the cross-category common description model; converting the image of the obtained optimal source domain; constructing a target domain synthetic dataset based on the converted image; generating target domain labels based on the target domain synthetic dataset, and training a detection model based on the target domain synthetic dataset and the target domain labels; and labeling the target based on the target domain label-trained detection model to obtain a labeled target domain dataset. The present invention also discloses a system, an electronic device, and a computer-readable storage medium, which can realize deep learning model training without manual labeling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing and intelligent information extraction, and in particular to a zero-labeling-based multi-category target domain data set labeling method and system. Background Art

[0002] With the combination of traditional agriculture and artificial intelligence technology, the construction of smart orchards has received more widespread attention in the development of the fruit industry. Among them, high-precision fruit detection technology is an important basic technology in the practical application of modern smart orchards. It has wide application value in many smart orchard intelligent tasks such as fruit positioning, fruit sorting, fruit yield prediction, and automatic fruit picking.

[0003] While deep learning-based object detection technology is currently widely used, it relies on large, labeled datasets to support the training of detection models, resulting in high manual labeling costs. Furthermore, due to the poor generalization performance of current deep learning models, applying the models to different scenarios, environments, shooting methods, and target types requires independently creating new target datasets and training new detection models, which is time-consuming and labor-intensive.

[0004] Current technical directions include: (1) introducing instance-level loss constraints to better regulate the generation direction of foreground targets in images. However, this approach is not suitable for the automatic fruit labeling task based on unsupervised learning due to the introduction of a manual labeling process; (2) using a fruit conversion model Across-CycleGAN with a cross-cycle comparison path, and realizing the conversion from round fruits to elliptical fruits by introducing a structural similarity loss function; however, the generalization of the target automatic labeling method is not high, and it is impossible to achieve the automatic labeling task of more types of target domain targets.

[0005] Therefore, there is an urgent need to establish a zero-cost data automatic labeling method with higher generalization and stronger domain adaptability, while optimizing the generative model so as to achieve realistic conversion in multi-category cases (manifested in large variations in shape, color, and texture) and reduce domain differences. Summary of the Invention

[0006] In order to solve the problems existing in the prior art, the present invention provides a multi-category target domain dataset labeling method and system based on zero labeling, which can realize the training of deep learning models without manual labeling costs, further improve the performance of unsupervised target conversion models, and enhance the algorithm's ability to describe target phenotypic characteristics, so as to control the model to accurately control the target generation direction in leapfrog target image conversion tasks with large differences in phenotypic characteristics. It can be applied to the rapid labeling of datasets in multiple farms, multiple varieties, and multiple scenarios.

[0007] A first aspect of the present invention provides a method for labeling a multi-category target domain dataset based on zero labeling, comprising:

[0008] S1, obtain target domain foreground images of different categories;

[0009] S2, performing a quantitative analysis of multidimensional spatial features based on the target domain foreground images of different categories and constructing a cross-category commonality description model based on the multidimensional spatial features after the quantitative analysis; obtaining an optimal source domain of the target based on the cross-category commonality description model;

[0010] S3, transforms the best source domain image based on the multi-category target generation model;

[0011] S4, constructs a target domain synthetic dataset based on the converted images;

[0012] S5, detecting the target based on the target domain synthetic data set, obtaining bounding box information of the target, and obtaining a target domain label based on the target domain synthetic data set and the bounding box information of the target to train a detection model;

[0013] S6: Perform automatic target labeling on a detection model trained based on the target domain label to obtain a labeled target domain dataset.

[0014] Preferably, the S2 includes:

[0015] S21, extracting appearance features of the target from foreground images of target domains of different categories, wherein the appearance features include edge contours, global colors, and local details;

[0016] S22, abstracting the appearance features into specific shapes, colors, and textures, and calculating relative distances of the specific shapes, colors, and textures for the features of different targets based on a multi-dimensional feature quantitative analysis method as an analysis description set of the appearance features of different target individuals;

[0017] S23, constructing a cross-category commonality description model based on multi-dimensional feature space reconstruction and feature difference division of the analysis description set;

[0018] S24, obtaining an optimal source domain of the target based on the cross-category commonality description model;

[0019] Preferably, the S22 includes:

[0020] S221, extracting the target shape based on the Fourier descriptor and discretizing the Fourier descriptor;

[0021] S222, extracting the spatial distribution and proportion of Lab colors in the target foreground, and drawing a CIELab space color distribution histogram;

[0022] S223, extracting pixel value gradients and directional derivative information of the target foreground to obtain texture information description based on the LBP algorithm;

[0023] S224, performing a relative distance calculation of a single appearance feature based on correlation and spatial distribution based on the discretization of the Fourier descriptor, the drawn CIELab space color distribution histogram, and the texture information description based on the LBP algorithm;

[0024] S225, constructing a relative distance matrix based on the calculated relative distance values ​​of the single appearance features;

[0025] The S23 includes:

[0026] S231, Multidimensional Feature Space Reconstruction: Construct a multidimensional feature space by using the relative distances between each pair of target features, thereby converting the relative distances between different target features into absolute distances in the same feature space, so as to concisely and accurately describe the phenotypic characteristics of each target image through a pair of two-dimensional plane position coordinates;

[0027] S232, feature difference division: divide the feature points in the space based on the clustering algorithm, and cluster them according to the similarity of the target features, so as to obtain the best source domain targets in different clusters;

[0028] The S24 includes:

[0029] S241, classifying different objects according to the appearance features based on the cross-category common description model;

[0030] S242, based on the target domain type actually required, select the best source domain type from the classification, input it into the generation model for target conversion, and optimize the generation model; the optimization of the generation model includes obtaining multi-category target domain background-free multimodal images through feature map extraction based on latent space and visual feature map extraction based on guided gradient information.

[0031] Preferably, the S3 includes performing image preprocessing and image conversion on the target sample data, obtaining the target domain simulation target, background and other components to form a target domain image synthesis component, including:

[0032] S31, the model generator generates a multi-dimensional loss function, which includes three types of loss functions, namely L Color (),L Shape () and L Texture ();

[0033] S32, balancing the weights of the multidimensional loss function using a dynamic adaptive weighting method based on quantifiable target phenotypic characteristics to obtain a multidimensional loss function based on an entropy weighting method;

[0034] S33, inputting the feature map based on the latent space into the multi-dimensional loss function based on the entropy weight method after weight balancing to obtain a subset of the background-free target images in the multi-category target domain.

[0035] Preferably, the S4 includes:

[0036] Based on the target domain image synthesis component, a knowledge graph system based on hierarchical component synthesis rules is established;

[0037] Constructing synthetic images based on a knowledge graph system of hierarchical component synthesis rules;

[0038] Record the target's location information, size information, and category information to form label data information;

[0039] A target domain synthetic dataset is formed based on synthetic images and label data information.

[0040] Preferably, the S5 includes:

[0041] Detect the target based on the target domain synthetic dataset to obtain the pre-trained model of the detection algorithm and the bounding box information of the target;

[0042] Based on the bounding box information of the detected target, pseudo-label self-learning is performed to generate target domain labels and obtain a labeled target domain dataset.

[0043] Preferably, the target domain label training detection model in S6 is built based on a multi-category target automatic labeling method, including:

[0044] The target domain label training detection model is used to automatically label the target to obtain a labeled target domain dataset.

[0045] A second aspect of the present invention is to provide an object labeling system based on an optimal source domain of a multidimensional spatial feature model, comprising:

[0046] The first image acquisition module is used to acquire target domain foreground images of different categories;

[0047] An optimal source domain selection module is used to perform a quantitative analysis of multidimensional spatial features based on the target domain foreground images of different categories and to construct a cross-category commonality description model based on the multidimensional spatial features after the quantitative analysis; and to obtain the optimal source domain of the target based on the cross-category commonality description model;

[0048] Image conversion module, used to convert the image of the optimal source domain based on the multi-category target generation model;

[0049] The target domain synthetic dataset construction module is used to construct the target domain synthetic dataset based on the converted images;

[0050] A target domain label generation module is used to detect the target based on the target domain synthetic data set, obtain the target's bounding box information, and obtain the target domain label training detection model based on the target domain synthetic data set and the target's bounding box information;

[0051] The target labeling module is used to automatically label the target based on the target domain label training detection model to obtain a labeled target domain data set.

[0052] A third aspect of the present invention provides an electronic device, comprising a processor and a memory, wherein the memory stores a plurality of instructions, and the processor is configured to read the instructions and execute the method described in the first aspect.

[0053] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a plurality of instructions, and the plurality of instructions can be read by a processor to execute the method described in the first aspect.

[0054] The target labeling method, system, electronic device, and computer-readable storage medium provided by the present invention have the following beneficial technical effects:

[0055] Establish an automatic labeling method with higher generalization, stronger domain adaptability, and the ability to meet the needs of different categories of fruit datasets; be able to automatically obtain labels for target domain targets, so as to be applied to downstream smart agriculture projects; and greatly reduce the monetary and time costs incurred in manually labeling target boxes (compared to the existing technology for labeling single scene datasets, the market average cost is 0.2 yuan / labeled box, each image has an average of 30 fruits, each image takes an average of 3 minutes to label, and each dataset contains at least 10,000 images). BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 This is a flow chart of a method for labeling a multi-category target domain dataset based on zero labeling according to the present invention.

[0057] Figure 2 This is a data logic diagram of the zero-labeled multi-category target domain dataset labeling method described in the present invention.

[0058] Figure 3 This is the overall flow chart of the image generation model described in the present invention.

[0059] Figure 4 This is an architecture diagram of a multi-category target domain dataset annotation system based on zero annotation according to the present invention.

[0060] Figure 5 The figure is a schematic diagram of the structure of the electronic device according to the present invention. DETAILED DESCRIPTION

[0061] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.

[0062] The method provided by the present invention can be implemented in the following terminal environment, which may include one or more of the following components: a processor, a memory, and a display screen. The memory stores at least one instruction, which is loaded and executed by the processor to implement the method described in the following embodiments.

[0063] A processor can include one or more processing cores. It connects various components within the terminal using various interfaces and circuits. It executes instructions, programs, code sets, or instruction sets stored in memory, and accesses data stored in memory to perform various terminal functions and process data.

[0064] The memory may include random access memory (RAM) or read-only memory (ROM). The memory may be used to store instructions, programs, codes, code sets, or instructions.

[0065] The display is used to show the user interface of each application.

[0066] In addition, those skilled in the art will appreciate that the structure of the terminal described above does not limit the terminal. The terminal may include more or fewer components, or a combination of certain components, or a different arrangement of components. For example, the terminal may also include a radio frequency circuit, an input unit, a sensor, an audio circuit, a power supply, and other components, which will not be described in detail here.

[0067] Example 1

[0068] See also Figure 1 and Figure 2 This embodiment provides a method for labeling a multi-category target domain dataset based on zero labeling, including:

[0069] S1, obtain target domain foreground images of different categories;

[0070] In this embodiment, the target domain foreground images of different categories can be images pre-stored by the computer device, or images downloaded by the computer device from other devices, or images uploaded to the computer device by other devices, or the target domain foreground images of different categories can be images currently collected by the computer device, and the embodiments of the present application do not limit this. For example, in this embodiment, the fruit annotation in the orchard is used as a specific application scenario, and a high-definition camera device is used to assist drones and other high-altitude shooting methods to obtain a wide-area orchard image as the target image. In addition, the target image and the final annotation image have the same size, for example, both are 96px*96px.

[0071] S2, performing a quantitative analysis of multidimensional spatial features based on the target domain foreground images of the different categories and constructing a cross-category commonality description model based on the multidimensional spatial features after the quantitative analysis; and obtaining the optimal source domain of the target based on the cross-category commonality description model.

[0072] As a preferred embodiment, the S2 includes: S21, extracting the appearance features of the target from target domain foreground images of different categories, the appearance features including but not limited to edge contours, global colors and local details; S22, abstracting the appearance features into specific shapes, colors and textures, and calculating the relative distances of specific shapes, colors and textures for the features of different targets based on a multidimensional feature quantitative analysis method as an analysis description set of the appearance features of different target individuals; S23, constructing a cross-category common description model based on multidimensional feature space reconstruction and feature difference division of the analysis description set; S24, obtaining the optimal source domain of the target based on the cross-category common description model.

[0073] In this embodiment, the optimal source domain selection module is used to design, describe, and analyze the phenotypic characteristics of different fruit categories. This module primarily calculates the commonalities between these features as prior knowledge for deep learning algorithms, providing guidance for deep learning dataset selection and training parameter setting. This module primarily consists of two parts: first, a multidimensional feature quantitative analysis method is proposed to analyze and describe the appearance characteristics of different fruit individuals; second, a cross-category commonality description model is constructed to classify different fruits according to their phenotypic characteristics and select the optimal source domain fruit species from these categories.

[0074] As a preferred embodiment, the S22, abstracting the appearance features into specific shapes, colors and textures, and calculating the relative distances of specific shapes, colors and textures for the features of different targets based on a multidimensional feature quantitative analysis method as the analysis description set of the appearance features of different target individuals includes: S221, extracting the shape of the target (fruit in this embodiment) based on the Fourier descriptor, and discretizing the Fourier descriptor; S222, extracting the spatial distribution and proportion of the Lab color in the foreground of the target (fruit in this embodiment), and drawing a CIELab space color distribution histogram; S223, extracting the pixel value gradient and directional derivative information of the foreground of the target (fruit in this embodiment) to obtain a texture information description based on the LBP algorithm; S224, calculating the relative distance of a single appearance feature based on correlation and spatial distribution based on the discretization of the Fourier descriptor, the drawn CIELab space color distribution histogram and the texture information description based on the LBP algorithm; S225, constructing a relative distance matrix based on the calculated relative distance values ​​of the single appearance feature.

[0075] As a preferred embodiment, the S23, based on the multidimensional feature space reconstruction and feature difference division of the analysis description set to construct a cross-category common description model, includes: S231, multidimensional feature space reconstruction: constructing a multidimensional feature space through the relative distance between the features of each target (fruit in this embodiment), thereby converting the relative distance between the features of different targets (fruit in this embodiment) into an absolute distance in the same feature space, so as to facilitate the concise and accurate description of the phenotypic characteristics of each target (fruit in this embodiment) image through a pair of two-dimensional plane position coordinates.

[0076] In this embodiment, the multidimensional feature space reconstruction adopts the MDS algorithm, including: projecting points in high-dimensional coordinates into low-dimensional coordinates based on distance, maintaining the relative distance between the points in the high-dimensional coordinates and the points in the low-dimensional coordinates unchanged, and projecting the points in the low-dimensional coordinates into a two-dimensional plane space to convert the relative distance into an absolute distance. Of course, those skilled in the art may also use other algorithms, as long as the relative distance can be converted into an absolute distance through coordinate projection and relative distance relationship, it is within the scope of protection of this field.

[0077] S232, feature difference division: Divide the feature points in the space based on the clustering algorithm, and cluster them according to the similarity of the target (fruit in this embodiment) features, so as to obtain the best source domain target (fruit in this embodiment) in different clusters.

[0078] In this embodiment, the clustering algorithm used for the feature difference division is the DBSCAN algorithm, including: clustering according to the closeness of samples in the multidimensional feature space, automatically dividing and selecting the category of the source domain target (fruit in this embodiment) and the number of source domains; automatically determining the number of clusters according to the distribution difference of the target (fruit in this embodiment) characteristics, and the target (fruit in this embodiment) type at the geometric center of each cluster is used as the optimal source domain target (fruit in this embodiment) type.

[0079] As a preferred implementation, the S24, obtaining the optimal source domain of the target based on the cross-category common description model includes: S241, classifying different targets according to the appearance features based on the cross-category common description model; S242, selecting the best source domain type from the classification according to the target domain type actually required, inputting it into the generation model for target conversion, and optimizing the generation model.

[0080] As a preferred implementation, for step S242, since it may sometimes be impossible to select the most suitable source domain when selecting the most suitable source domain data (some clusters only have one target or fruit), the generation model needs to be optimized so that realistic conversion can be achieved even when the shape, color, and texture vary greatly, thereby reducing domain differences.

[0081] The optimization of the generative model includes obtaining multi-category target domain background-free target multimodal images by extracting feature maps based on latent space and visual feature maps based on guided gradient information, thereby solving the problem of single-category optimal source domain background-free target images.

[0082] S3, converting the best source domain image based on the multi-category target generation model; including: performing image preprocessing and image conversion on the target sample data, obtaining the target domain simulated target (fruit), background and other (leaves) components to form the target domain image synthesis component;

[0083] S31, the model generator generates a multi-dimensional loss function, which includes three types of loss functions, namely L Color (),L Shape () and L Texture ().

[0084] In this embodiment:

[0085] (1) Regarding the color feature loss function: This embodiment uses the cycle-consistent loss function and the self-mapping loss function (not shown in the figure) in the CycleGAN network. The coloring effect can help the target conversion model better control the generation of color features, where the cycle-consistent loss is expressed as:

[0086] L Color (GST +G TS )=L Cycle (G ST +G TS )+L Identity (G ST +G TS ) (1)

[0087] L Cycle (G ST +G TS =E s~pdata(s) ||G TS (G ST (s))-s||1+E t~pdata(t) ||G ST (G TS (t))-t||1(2)

[0088] The self-mapping loss function is expressed as:

[0089] L Identity (G ST +G TS )=E s~pdata(t) ||s-G ST (s)||1+E s~pdata(t) ||t-G TS (t)||1 (3)

[0090] Where s~pdata(s) and t~pdata(t) represent the data distribution in the source domain and the target domain respectively, and t and s represent the image information in the target domain and the source domain respectively.

[0091] (2) Regarding the shape feature loss function: This embodiment adopts the multi-scale structural similarity index MS-SSIM, uses convolution kernels of different sizes to adjust the image receptive field size and counts the shape structural feature information of the corresponding area of ​​the image under different scale conditions, thereby effectively distinguishing the geometric differences of different categories of fruit images, and training the model to better adapt to the differences in shape features between different categories of targets (fruits in this embodiment). This embodiment uses a cross-cycle comparison method to compare the original image with the image converted in another cycle, so as to better constrain the generation process of the target (fruit in this embodiment). The shape feature loss function is expressed as:

[0092] L Shape (G ST +G TS )=(1-MS_SSIM(G ST (s),t))+(1-MS_SSIM(G TS (t),s)) (4)

[0093] MS_SSIM represents the calculation based on the multi-scale structural similarity index loss.

[0094] (3) Regarding the texture feature loss function: In the scenario where fruit is used as the target for target labeling, since the texture features in the fruit image are too detailed, the texture features cannot be fully expressed if the loss function is only compared from the original RGB image; and the resolution of the fruit in the dataset is smaller, the texture features cannot be well expressed, which adds a certain degree of difficulty to the image conversion model. Therefore, this embodiment designs a texture feature loss function based on the local binary pattern (LBP) descriptor, so that it can better highlight the texture of the target and its regular arrangement of the texture loss calculation method, accurately describe its texture features, and better play the performance of the image conversion model. The texture feature loss function is expressed as:

[0095] L Texture (G ST +G TS )=Pearson(LBP(G ST (s),t)+Pearson(G TS (t),s)) (5)

[0096] LBP(X,Y)=N(LBP(x C ,y C )) (6)

[0097]

[0098]

[0099] Pearson means using the Pearson correlation coefficient to calculate the difference between the texture features of the fruit, N means traversing all the pixel values ​​in the entire image, x C ,y C Represents the center pixel, g represents the grayscale value, s is the sign function, and P represents the P neighborhood selected from the center pixel. Experimental verification shows that the best effect is achieved when P is 16.

[0100] In the absence of paired supervision information constraints, the distribution of the two image domains is highly discrete and irregular. The present invention designs and uses a multidimensional loss function to constrain the generation direction of visual attributes such as color, shape, and texture of the fruit during the training process of the fruit transformation model, which can more accurately describe the multidimensional phenotypic characteristics of the fruit transformation process.

[0101] S32, after balancing the weights of the multidimensional loss function based on the dynamic adaptive weight method of quantifiable target phenotypic characteristics, a multidimensional loss function based on the entropy weight method is obtained; in step S32, a multidimensional feature loss function is added to accurately describe the characteristics of the target (the present embodiment is the fruit) during the training process. However, in the generative adversarial network training process, it is not the case that the more loss functions, the better the network model effect. If an excessive loss function is added, the model in the training phase cannot be properly fitted, thereby losing the generation direction of describing the target characteristics. Therefore, in order to balance the multidimensional loss function added in the embodiment of the present invention, so that it can converge stably and accurately describe the multidimensional fruit phenotypic characteristics, the embodiment of the present invention introduces a dynamic adaptive weight method based on the phenotypic characteristics of the quantifiable target (the present embodiment is the fruit) for balancing the weights of the multidimensional loss function. The specific process of the S32 is as follows:

[0102] (1) Calculate the quantifiable descriptor values ​​of the shape, color, and texture features of the i-th target (fruit in this example) in the source domain and the target domain in turn, and perform normalization processing on them, which are denoted as S i ,C i ,T i ;

[0103] (2) Calculate the proportion P of each target sample (fruit in this example) under different characteristic value standards ij , is used to describe the difference in the values ​​of different feature descriptors, as shown in formula (9):

[0104]

[0105] Among them, j takes shape, color and texture features (S, C, T) as three different indicators; Y represents the fruit sample, Y i Represents each different fruit sample, Y ij Indicates different phenotypic characteristics of different fruit samples;

[0106] (3) According to the definition of information entropy in information theory, the greater the difference in the descriptor values ​​of different target samples (in this example, fruit), the more information can be provided in training the GAN model. Therefore, it is necessary to assign more weight to them during the model training process. At this time, the information entropy of a set of data is calculated as shown in formula (10):

[0107]

[0108] (4) According to the calculation formula of information entropy, the weight of each indicator is obtained as shown in formula (11):

[0109]

[0110] The overall loss function L of the multidimensional loss function based on the entropy weight method generated by the model generator Guided-GAN It can be expressed as formula (12):

[0111] L Guided-GAN =W s ·L Shape (G ST +G TS )+W c ·L Color (G ST +G TS )W t ·L Texture (G ST +G TS ) (12)

[0112] Among them G ST Denotes the generator that maps the source domain to the target domain, G TS represents the generator that maps the target domain to the source domain, W s ,W c and W t They respectively represent the weight ratios assigned to the shape, color, and texture loss functions using the entropy weight method during model training.

[0113] In the fruit labeling application scenario, when converting between two types of fruits, we directly compare the differences in shape, color, and texture descriptors of all samples of the two types of fruits, automatically calculate the specific values ​​of the differences between the fruits, and dynamically adjust the weight ratio W of the multidimensional loss function during each training. s ,W c ,W t , thereby better assisting the network model in fitting, accelerating the convergence process, and making the generated target domain fruit image quality better.

[0114] S33, inputting the feature map based on the latent space into the multi-dimensional loss function based on the entropy weight method after weight balancing to obtain a subset of the background-free target images in the multi-category target domain.

[0115] The multi-category target generation model in S3 is built using a method based on the fusion of multiple feature loss functions. The targets have multiple categories, and the multi-feature loss function is a multidimensional loss function based on the entropy weight method. The multidimensional loss function based on the entropy weight method is used to constrain the generation direction of the color, shape, and texture of the targets of multiple categories during the training process of the target conversion model, thereby more accurately describing the phenotypic characteristics of targets with large feature differences (fruits in this embodiment) and solving the problem of the single function of the loss function. In the fruit image conversion model, the generation direction of multidimensional features is better controlled, and finally, good results can be achieved in the leapfrog fruit conversion task with large feature differences.

[0116] Since the original CycleGAN network can only train the generator to achieve the effect of recoloring, it is difficult to accurately describe features such as shape and texture, and there is a lack of shape and texture feature information of real fruit images for network fitting training; existing technologies may introduce instance-level loss constraints to better regulate the generation direction of foreground targets in images, but such practices are not suitable for unsupervised learning-based automatic fruit labeling tasks due to the introduction of a manual labeling process; there is also a fruit conversion model Across-CycleGAN with a cross-cycle comparison path, which realizes the conversion from circular targets to elliptical targets by introducing a structural similarity loss function, and is applied to scenarios such as fruit labeling. In order to better improve the generalization of the fruit automatic labeling method and thus realize the automatic labeling task of more types of target domain fruits, it is necessary to further improve the performance of the unsupervised fruit conversion model and enhance the algorithm's ability to describe the phenotypic characteristics of the fruit, so that the control model can accurately control the fruit generation direction in the leap-forward fruit image conversion task with large differences in phenotypic characteristics.

[0117] S4, constructing a target domain synthetic dataset based on the converted image, including: establishing a knowledge graph system based on hierarchical component synthesis rules based on the target domain image synthesis component; in this embodiment, the knowledge graph refers to a knowledge graph system based on hierarchical component synthesis rules constructed by setting growth rules for each component according to the natural semantic structure, growth semantic structure and target domain background characteristics; constructing a synthetic image based on the knowledge graph system based on the hierarchical component synthesis rules; recording the target's position information, size information and category information, and forming it into label data information; forming a target domain synthetic dataset based on the synthetic image and label data information.

[0118] As a preferred embodiment, the target domain image synthesis component is based on the establishment of a knowledge graph system based on hierarchical component synthesis rules so that the constructed target domain synthesis dataset follows certain rules, including: composition rules based on natural semantics, construction rules based on growth semantics, and domain adaptation rules based on scene environment to form a construction process from components to scenes.

[0119] In this embodiment, due to the complexity of the orchard scene and the changeable environment, it is very difficult to achieve automatic data set synthesis by relying entirely on random placement methods. Therefore, this method classifies each component of the orchard scene more finely according to different situations based on the structural and regular relationships between components, forming a knowledge graph based on the hierarchical structure of the orchard scene, so as to reasonably divide the synthesis weights between different components.

[0120] The domain adaptation rules based on scene environment form the basic components of orchard scene distribution, including land, sky, skeleton, leaves and fruits; the component rules based on growth semantics form the basic building components of fruit tree growth status (including trees and occluded fruits) and the combined components of fruit tree growth status (including trees with fruits), among which the occluded fruits are formed by the domain adaptation sub-rules of fruits based on scene environment, the trees are formed by the component sub-rules of skeleton and leaves based on growth semantics, and the trees with fruits are formed by the component sub-rules of trees and occluded fruits based on growth semantics; the composition rules based on natural semantics form the orchard scene with natural semantic structure, among which the domain adaptation rules based on scene environment and the composition rules based on natural semantics of the trees with fruits, sky and land finally form the target domain synthetic image.

[0121] S5, detecting the target based on the target domain synthetic data set, obtaining the target's bounding box information, and obtaining the target domain label training detection model based on the target domain synthetic data set and the target's bounding box information; including: detecting the target based on the target domain synthetic data set, obtaining a pre-trained model of the detection algorithm and the target's bounding box information; performing pseudo-label self-learning based on the detected target's bounding box information to generate a target domain label, and obtaining a labeled target domain data set.

[0122] S6: Automatically labeling targets based on the target domain label training detection model to obtain a labeled target domain dataset. The target domain label training detection model in S6 is built based on a multi-category target automatic labeling method, including: automatically labeling targets based on the target domain label training detection model to obtain a labeled target domain dataset.

[0123] Based on this, the embodiment of the present invention uses a multidimensional loss function to constrain the generation direction of the color, shape and texture of the fruit during the training process of the fruit conversion model. The schematic diagram of the multidimensional loss function design in the generator of the model is as follows Figure 3 shown.

[0124] S6 includes: automatically labeling the target domain label training detection model to obtain a labeled target domain data set.

[0125] In this embodiment, the single-category optimal source domain background-free target image is a single-category optimal source domain background-free fruit image. The loss function is then calculated after feature map visualization, which includes feature map extraction based on the latent space and visual feature map acquisition based on guided gradient information.

[0126] As a preferred embodiment, the single-category optimal source domain background-free target images can all be images pre-stored in the computer device, or images downloaded by the computer device from other devices, or images uploaded to the computer device by other devices, or images currently captured by the computer device.

[0127] As a preferred embodiment, the method for obtaining the optimal source domain in the single-category optimal source domain background-free target image includes: extracting the appearance features of the target from the single-category target domain foreground image; abstracting the appearance features into specific shapes, colors and textures, and calculating the relative distances of specific shapes, colors and textures for the features of different targets based on a multidimensional feature quantitative analysis method as an analysis description set of the appearance features of different target individuals; constructing a single-category description model based on multidimensional feature space reconstruction and feature difference division of the analysis description set; and obtaining the optimal source domain of the target based on the single-category description model.

[0128] As a preferred embodiment, the method of obtaining the optimal source domain of the target based on the single category description model includes: classifying different targets according to the appearance features based on the single category description model; selecting the optimal source domain type from the classification according to the target domain type actually required, and inputting it into the single category description model for target conversion to obtain the optimal source domain of the target.

[0129] Example 2

[0130] See also Figure 4 The present embodiment provides a target labeling system for an optimal source domain based on a multidimensional spatial feature model, comprising: a first image acquisition module 101, configured to acquire target domain foreground images of different categories; an optimal source domain selection module 102, configured to perform a quantitative analysis of multidimensional spatial features based on the target domain foreground images of different categories and to construct a cross-category commonality description model based on the multidimensional spatial features after the quantitative analysis; and to obtain an optimal source domain for the target based on the cross-category commonality description model; an image conversion module 103, configured to convert the image of the optimal source domain based on a multi-category target generation model; a target domain synthetic dataset construction module 104, configured to construct a target domain synthetic dataset based on the converted image; a target domain label generation module 105, configured to detect the target based on the target domain synthetic dataset, obtain the target's bounding box information, and obtain a target domain label training detection model based on the target domain synthetic dataset and the target's bounding box information; and a target labeling module 106, configured to automatically label the target based on the target domain label training detection model to obtain a labeled target domain dataset.

[0131] The present invention also provides a memory storing a plurality of instructions, wherein the instructions are used to implement the method described in the first embodiment.

[0132] like Figure 5 As shown, the present invention also provides an electronic device, including a processor 301 and a memory 302 connected to the processor 301, wherein the memory 302 stores multiple instructions, which can be loaded and executed by the processor to enable the processor to execute the method described in Example 1.

[0133] Although preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic inventive concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the invention. Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the invention. Thus, the present invention is intended to include such changes and modifications as fall within the scope of the claims and their equivalents.

Claims

1. A zero-labeled multi-category target domain dataset annotation method, characterized in that: include: S1, obtain target domain foreground images of different categories; S2, performing a quantitative analysis of multi-dimensional spatial features based on the target domain foreground images of different categories and constructing a cross-category commonality description model based on the multi-dimensional spatial features after the quantitative analysis; Obtain the optimal source domain of the target based on the cross-category commonality description model; S3, transforms the best source domain image based on the multi-category target generation model; S4, constructs a target domain synthetic dataset based on the converted images; S5, detecting the target based on the target domain synthetic data set, obtaining bounding box information of the target, and obtaining a target domain label based on the target domain synthetic data set and the bounding box information of the target to train a detection model; S6: Perform automatic target labeling on a detection model trained based on the target domain label to obtain a labeled target domain dataset.

2. The zero-labeled multi-category target domain dataset labeling method according to claim 1, characterized in that: The S2 includes: S21, extracting appearance features of the target from foreground images of target domains of different categories, wherein the appearance features include edge contours, global colors, and local details; S22, abstracting the appearance features into specific shapes, colors and textures, and calculating the relative distances of specific shapes, colors and textures for the features of different targets based on a multidimensional feature quantitative analysis method as an analysis description set of the appearance features of different target individuals; S22 specifically includes: S221, extracting the target shape based on the Fourier descriptor, and discretizing the Fourier descriptor; S222, extracting the spatial distribution and proportion of the Lab color in the target foreground, and drawing a CIELab space color distribution histogram; S223, extracting the pixel value gradient and directional derivative information of the target foreground to obtain a texture information description based on the LBP algorithm; S224, calculating the relative distance of a single appearance feature based on correlation and spatial distribution based on the discretization of the Fourier descriptor, the drawn CIELab space color distribution histogram and the texture information description based on the LBP algorithm; S225, constructing a relative distance matrix based on the calculated relative distance values ​​of the single appearance feature; S23, constructing a cross-category commonality description model based on multi-dimensional feature space reconstruction and feature difference division of the analysis description set; S24, obtaining an optimal source domain of the target based on the cross-category commonality description model.

3. The zero-labeled multi-category target domain dataset labeling method according to claim 2, characterized in that: The S23 includes: S231, Multidimensional Feature Space Reconstruction: Construct a multidimensional feature space by using the relative distances between each pair of target features, thereby converting the relative distances between different target features into absolute distances in the same feature space, so as to concisely and accurately describe the phenotypic characteristics of each target image through a pair of two-dimensional plane position coordinates; S232, feature difference division: divide the feature points in the space based on the clustering algorithm, and cluster them according to the similarity of the target features, so as to obtain the best source domain targets in different clusters; The S24 includes: S241, classifying different objects according to the appearance features based on the cross-category common description model; S242, based on the target domain type actually required, select the best source domain type from the classification, input it into the generation model for target conversion, and optimize the generation model; the optimization of the generation model includes obtaining multi-category target domain background-free multimodal images through feature map extraction based on latent space and visual feature map extraction based on guided gradient information.

4. The zero-labeled multi-category target domain dataset labeling method according to claim 3, characterized in that: S3 includes performing image preprocessing and image conversion on the target sample data, obtaining the target domain simulation target, background and other components to form a target domain image synthesis component, including: S31, the model generator generates a multi-dimensional loss function, which includes three types of loss functions: color feature loss function L Color (), shape feature loss function L Shape () and texture feature loss function L Texture (); S32, balancing the weights of the multidimensional loss function using a dynamic adaptive weighting method based on quantifiable target phenotypic characteristics to obtain a multidimensional loss function based on an entropy weighting method; S33, inputting the feature map based on the latent space into the multi-dimensional loss function based on the entropy weight method after weight balancing to obtain a subset of the background-free target images in the multi-category target domain.

5. The zero-labeled multi-category target domain dataset labeling method according to claim 1, characterized in that: The S4 includes: Based on the target domain image synthesis component, a knowledge graph system based on hierarchical component synthesis rules is established; Constructing synthetic images based on a knowledge graph system of hierarchical component synthesis rules; Record the target's location information, size information, and category information to form label data information; A target domain synthetic dataset is formed based on synthetic images and label data information.

6. The method for labeling a multi-category target domain dataset based on zero labeling according to claim 1, characterized in that: The S5 includes: Detect the target based on the target domain synthetic dataset to obtain the pre-trained model of the detection algorithm and the bounding box information of the target; Based on the bounding box information of the detected target, pseudo-label self-learning is performed to generate target domain labels and obtain a labeled target domain dataset.

7. The zero-labeled multi-category target domain dataset labeling method according to claim 1, characterized in that: The target domain label training detection model in S6 is built based on a multi-category target automatic labeling method, including: The target domain label training detection model is used to automatically label the target to obtain a labeled target domain dataset.

8. An object labeling system based on an optimal source domain of a multidimensional spatial feature model, used to implement the method according to any one of claims 1 to 7, characterized in that: include: The first image acquisition module is used to acquire target domain foreground images of different categories; An optimal source domain selection module is used to perform a quantitative analysis of multi-dimensional spatial features based on the target domain foreground images of different categories and to construct a cross-category commonality description model based on the multi-dimensional spatial features after the quantitative analysis; Obtain the optimal source domain of the target based on the cross-category commonality description model; Image conversion module, used to convert the image of the optimal source domain based on the multi-category target generation model; The target domain synthetic dataset construction module is used to construct the target domain synthetic dataset based on the converted images; A target domain label generation module is used to detect the target based on the target domain synthetic data set, obtain the target's bounding box information, and obtain the target domain label training detection model based on the target domain synthetic data set and the target's bounding box information; The target labeling module is used to automatically label the target based on the target domain label training detection model to obtain a labeled target domain data set.

9. An electronic device comprising a processor and a memory, wherein the memory stores a plurality of instructions, characterized in that: The processor is configured to read the instruction and execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a plurality of instructions, characterized in that: The plurality of instructions may be read by a processor and executed by the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image description model training and description method, system and device and storage medium

    CN115147644A

  • Multi-source domain adaptation with mutual learning

    US20220076074A1