Image Data Augmentation Method, Device, Electronic Device and Storage Medium

By generating adversarial networks, the edge image data is characterized and randomly encoded offset processing, and the expanded image data is generated, which solves the problems of low edge data expansion efficiency and poor model enhancement effect in the prior art, and improves the generalization ability of the model, especially in road recognition scenarios in vehicle autonomous driving, and improves recognition accuracy and robustness.

CN115346089BActive Publication Date: 2025-07-18BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211062121.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-31
Publication Date
2025-07-18
Estimated Expiration
2042-08-31

AI Technical Summary

Technical Problem

The method of edge data expansion based on data augmentation or feature mining in the prior art is inefficient and has poor model enhancement effect, making it difficult to effectively improve the model generalization ability in small sample image data learning scenarios.

Method used

The edge image data is characterized by using the generative adversarial network, and the expanded image data is generated through random encoding offset processing, and combined with the generation and decoding of the preset probability distribution function, the expansion efficiency of edge image data is improved.

Benefits of technology

It realizes efficient expansion of edge image data, improves the generalization ability of corresponding models, especially in road recognition scenarios in vehicle autonomous driving, and improves the recognition accuracy and robustness of models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115346089B_ABST
    Figure CN115346089B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image data augmentation method, apparatus, electronic device, and storage medium, which relate to the technical fields of artificial intelligence such as data mining, data synthesis, and data generation, and at least solve the technical problems of low data augmentation efficiency and poor model enhancement effect in the prior art for the method of augmenting edge data based on data enhancement or feature mining. The specific implementation solution is as follows: Obtain an original image data set, where the original image data set includes edge image data and non-edge image data, and the number of edge image data is less than the number of non-edge image data; Use a generative adversarial network to perform feature encoding on the edge image data to obtain a first encoding result, where the generative adversarial network is trained using the original image data set; Based on the first encoding result, perform an offset process on a random encoding to obtain a second encoding result; Use the generative adversarial network to decode the second encoding result to obtain augmented image data corresponding to the edge image data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technologies such as data mining, data synthesis, and data generation, and particularly relates to an image data augmentation method, apparatus, electronic device, and storage medium. Background Art

[0002] In deep learning tasks, it is difficult to collect marginal data in some small-sample image data learning scenarios (such as project cold start scenarios in the industrial field, machine learning model optimization and iteration scenarios, etc.). The data in the training set usually cannot cover the test scenarios, resulting in poor generalization ability of the trained model. In response to this, those skilled in the art have continuously tried various marginal data augmentation methods.

[0003] The marginal data augmentation methods provided by related technologies mainly include:

[0004] First, a marginal data augmentation method based on data augmentation: augment a small amount of marginal data by means of random pasting, random affine transformation, etc. However, the defect of this method is that the augmented data only contains the features of a small amount of marginal data, and the transformation methods are few, so the augmented data has a poor enhancement effect on the model;

[0005] Second, a marginal data augmentation method based on feature mining: use the existing data to train a feature encoding network in an unsupervised manner, then encode the marginal data, and re-label the data similar to the marginal data features in the test set, so as to obtain augmented data similar to the marginal data. However, the defect of this method is that a large amount of real unlabeled data needs to be introduced, the task iteration period is long, and the enhancement effect of the augmented data on the model is still limited by data collection.

[0006] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0007] The present disclosure provides an image data augmentation method, apparatus, electronic device, and storage medium to at least solve the technical problems of low data augmentation efficiency and poor model enhancement effect in the prior art for augmenting marginal data based on data augmentation or feature mining.

[0008] According to one aspect of the present disclosure, there is provided an image data augmentation method, including: obtaining an original image data set, where the original image data set includes edge image data and non-edge image data, and the number of edge image data is less than the number of non-edge image data; using a generative adversarial network to perform feature encoding on the edge image data to obtain a first encoding result, where the generative adversarial network is trained using the original image data set; based on the first encoding result, performing an offset process on a random encoding to obtain a second encoding result, where the random encoding is generated by the generative adversarial network and a preset probability distribution function; using the generative adversarial network to decode the second encoding result to obtain augmented image data corresponding to the edge image data.

[0009] According to another aspect of the present disclosure, there is also provided an apparatus for image data augmentation, including: an obtaining module for obtaining an original image data set, where the original image data set includes edge image data and non-edge image data, and the number of edge image data is less than the number of non-edge image data; an encoding module for using a generative adversarial network to perform feature encoding on the edge image data to obtain a first encoding result, where the generative adversarial network is trained using the original image data set; an offset module for performing an offset process on a random encoding based on the first encoding result to obtain a second encoding result, where the random encoding is generated by the generative adversarial network and a preset probability distribution function; a decoding module for using the generative adversarial network to decode the second encoding result to obtain augmented image data corresponding to the edge image data.

[0010] According to another aspect of the present disclosure, there is also provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image data augmentation method proposed by the present disclosure.

[0011] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the image data augmentation method proposed by the present disclosure.

[0012] According to another aspect of the present disclosure, there is also provided a computer program product, including a computer program, where the computer program, when executed by a processor, executes the image data augmentation method proposed by the present disclosure.

[0013] In the present disclosure, an original image dataset is obtained, where the original image dataset includes edge image data and non-edge image data, and the number of edge image data is less than that of non-edge image data. By using a generative adversarial network to perform feature encoding on the edge image data, a first encoding result is obtained, where the generative adversarial network is trained using the original image dataset. A method of offsetting a random encoding based on the first encoding result is adopted to obtain a second encoding result, where the random encoding is generated by the generative adversarial network and a preset probability distribution function. Then, the generative adversarial network is used to decode the second encoding result to obtain augmented image data corresponding to the edge image data, achieving the purpose of generating augmented image data corresponding to the edge image data in the original image dataset based on the generative adversarial network and the random encoding, realizing the technical effect of improving the augmentation efficiency of the edge image data and thus enhancing the generalization ability of the corresponding model, and solving the technical problems of low data augmentation efficiency and poor model enhancement effect in the related art methods for augmenting edge data based on data enhancement or feature mining.

[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0016] Figure 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing an image data augmentation method;

[0017] Figure 2 is a flowchart of an image data augmentation method according to an embodiment of the present disclosure;

[0018] Figure 3 is a schematic diagram of an optional generative adversarial network training process according to an embodiment of the present disclosure;

[0019] Figure 4 is a schematic diagram of an optional process of generating augmented image data according to an embodiment of the present disclosure;

[0020] Figure 5 is a structure block diagram of an image data augmentation device provided according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted below for clarity and conciseness.

[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0023] According to an embodiment of the present disclosure, an image data augmentation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0024] The method embodiments provided by the embodiments of the present disclosure can be executed in a mobile terminal, a computer terminal, or a similar electronic device. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing the image data augmentation method is shown.

[0025] As Figure 1As shown, the computer terminal 100 includes a computing unit 101, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 102 or the computer program loaded from the storage unit 108 into the random access memory (RAM) 103. In the RAM 103, various programs and data required for the operation of the computer terminal 100 can also be stored. The computing unit 101, the ROM 102, and the RAM 103 are connected to each other via a bus 104. The input / output (I / O) interface 105 is also connected to the bus 104.

[0026] Multiple components in the computer terminal 100 are connected to the I / O interface 105, including: an input unit 106, such as a keyboard, a mouse, etc.; an output unit 107, such as various types of displays, speakers, etc.; a storage unit 108, such as a magnetic disk, an optical disc, etc.; and a communication unit 109, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 109 allows the computer terminal 100 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0027] The computing unit 101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 101 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 101 executes the image data augmentation method described herein. For example, in some embodiments, the image data augmentation method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the computer terminal 100 via the ROM 102 and / or the communication unit 109. When the computer program is loaded into the RAM 103 and executed by the computing unit 101, one or more steps of the method for locating a faulty hard disk described herein can be executed. Alternatively, in other embodiments, the computing unit 101 can be configured to execute the method for locating a faulty hard disk in any other appropriate manner (e.g., by means of firmware).

[0028] The various embodiments of the systems and techniques described herein may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0029] It should be noted here that in some alternative embodiments, the above Figure 1 shown electronic device may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance and is intended to illustrate the types of components that may exist in the above electronic device.

[0030] Under the above operating environment, the present disclosure provides an Figure 2 shown image data augmentation method, which can be executed by an Figure 1 shown computer terminal or a similar electronic device. Figure 2 is a flowchart of an image data augmentation method provided according to an embodiment of the present disclosure. As Figure 2 shown, the method may include the following steps:

[0031] Step S20, obtaining an original image data set, where the original image data set includes edge image data and non-edge image data, and the number of edge image data is less than the number of non-edge image data;

[0032] The above original image data set may be a large amount of image data collected during actual application. The above edge image data may be image data of an edge scene in the original image data set. The above non-edge image data may be other image data in the original image data set except the edge image data.

[0033] For example, the original image dataset for training a pet recognition model can contain a large amount of image data. Among them, non-edge image data are image data showing common pets (such as cats, dogs, rabbits, ducks, birds, etc.), and edge image data are image data showing uncommon pets (such as snakes, spiders, snails, etc.). Usually, the number of non-edge image data is large, and the number of edge image data is small.

[0034] It is easy to note that since the edge image data corresponding to the edge scenario in the original image dataset is less, the generalization ability of the trained model for this edge scenario is usually poor. By augmenting the edge image data and using the augmented image data to train the model, the generalization ability of the model for this edge scenario can be further improved.

[0035] Step S22: Use a generative adversarial network to perform feature encoding on the edge image data to obtain a first encoding result, where the generative adversarial network is trained using the original image dataset;

[0036] The above-mentioned generative adversarial network (Generative Adversarial Network, abbreviated as GAN) can be a neural network trained through machine learning using the above-mentioned original image data. In an actual application scenario, the above-mentioned generative adversarial network can also be a style-GAN.

[0037] Using the above-mentioned generative adversarial network to encode the features of the above-mentioned edge image data can obtain the above-mentioned first encoding result. This first encoding result can represent the features of the edge image data in the feature space.

[0038] Specifically, the specific implementation process of using a generative adversarial network to perform feature encoding on the edge image data to obtain a first encoding result can refer to the further introduction of the embodiments of the present disclosure and will not be elaborated here.

[0039] Step S24: Based on the first encoding result, perform an offset process on the random encoding to obtain a second encoding result, where the random encoding is generated by a generative adversarial network and a preset probability distribution function;

[0040] The above-mentioned random encoding can be generated by a generative adversarial network and the above-mentioned preset probability distribution function. This preset probability distribution function can be pre-specified by those skilled in the art according to the actual application scenario (for example: it can be a Gaussian distribution function).

[0041] Based on the above-mentioned first encoding result corresponding to the edge image data, perform an offset process on the above-mentioned random encoding in the feature space to obtain the above-mentioned second encoding result. This second encoding result is the encoding corresponding to the augmented image data generated based on the edge image data.

[0042] The above offset processing may be to offset the random code in the direction closer to the first coding result, and this offset processing can improve the data correlation between the generated augmented image data and the edge image data.

[0043] It should be noted that based on the offset processing of the random code in the feature space, the editing of the image data corresponding to the code can be realized by mapping the offset of the code to the image space.

[0044] Specifically, based on the first coding result, the specific implementation process of offset processing the random code to obtain the second coding result can refer to the further introduction of the embodiments of the present disclosure and will not be elaborated here.

[0045] Step S26: Use the generative adversarial network to decode the second coding result to obtain the augmented image data corresponding to the edge image data.

[0046] Using the generative adversarial network to decode the above second coding result can obtain the above augmented image data corresponding to the above edge image data. Since the augmented image data is obtained by offsetting the random code based on the features of the edge image data (equivalent to the above first coding result), the augmented image data has both randomness and good data correlation with the edge image data.

[0047] Specifically, the specific implementation process of using the generative adversarial network to decode the second coding result to obtain the augmented image data corresponding to the edge image data can refer to the further introduction of the embodiments of the present disclosure and will not be elaborated here.

[0048] According to the above steps S20 to S26 of the present disclosure, an original image data set is obtained, wherein the original image data set includes edge image data and non-edge image data, and the number of edge image data is less than the number of non-edge image data. By using the generative adversarial network to perform feature coding on the edge image data, a first coding result is obtained, wherein the generative adversarial network is trained using the original image data set. A method of offsetting the random code based on the first coding result is adopted to obtain a second coding result, wherein the random code is generated by the generative adversarial network and a preset probability distribution function. Then, the generative adversarial network is used to decode the second coding result to obtain the augmented image data corresponding to the edge image data, achieving the purpose of generating the augmented image data corresponding to the edge image data in the original image data set based on the generative adversarial network and the random code, realizing the technical effect of improving the augmentation efficiency of the edge image data and further enhancing the generalization ability of the corresponding model, and solving the technical problems of low data augmentation efficiency and poor model enhancement effect in the related art method of performing edge data augmentation based on data enhancement or feature mining.

[0049] The above method of this embodiment will be further introduced below.

[0050] As an alternative implementation, a generative adversarial network is used to perform feature encoding on the edge image data to obtain a first encoding result, and the following method steps are further included:

[0051] Step S221: Use the original image dataset to train a generative adversarial network through self-supervised machine learning, where the generative adversarial network includes an encoder and a decoder;

[0052] Step S222: Use the encoder to encode the features of the edge image data into the feature space to obtain a first encoding result.

[0053] Using the above original image dataset, a generative adversarial network is trained through self-supervised machine learning. The generative adversarial network includes a generator and a discriminator. During the above self-supervised machine learning training process, the discriminator is used to perform self-supervision on the training process, so that the trained generative adversarial network learns good feature encoding ability and image generation ability.

[0054] The generative adversarial network obtained by the above self-supervised machine learning training includes an encoder and a decoder. In an actual application scenario, the encoder can be used to encode image data into the feature space, and the decoder can decode the encoding in the feature space to obtain corresponding image data.

[0055] Using the encoder of the above generative adversarial network, the encoded features of the edge image data in the above original image dataset are encoded into the feature space, and then the above first encoding result is obtained. The first encoding result can represent the features of the edge image data in the feature space.

[0056] Figure 3 It is a schematic diagram of an alternative generative adversarial network training process according to an embodiment of the present disclosure. As Figure 3 shown, based on the collected image data (equivalent to the above original image dataset), self-supervised machine learning training can be performed to obtain a generative adversarial network. The generative adversarial network includes an encoder and a decoder. The encoder can encode the result of Gaussian sampling to obtain a hidden encoding. The decoder can decode the hidden encoding to obtain a corresponding augmented image.

[0057] As an alternative implementation, based on the first encoding result, an offset process is performed on the random encoding to obtain a second encoding result, and the following method steps are further included:

[0058] Step S241: Perform clustering processing on the first encoding result to obtain the clustering center corresponding to the edge image data;

[0059] Step S242: Based on a preset probability distribution function, use a generative adversarial network to generate a random code;

[0060] Step S243: Perform an offset process on the random code according to an offset strategy to obtain a second coding result, where the offset strategy is used to determine to move the random code in a direction closer to the clustering center.

[0061] Performing a clustering process on the above first coding result can obtain the clustering center corresponding to the above edge image data. This clustering center can be used to represent the characteristic position of the edge image data in the feature space. That is to say, in the feature space, the closer the code corresponding to the clustering center is, the higher the similarity between the corresponding image data and the edge image data.

[0062] It should be noted that the clustering method for performing the clustering process on the first coding result can be a simple averaging clustering method, or other unsupervised clustering methods (such as the k-means clustering algorithm, the LOF algorithm based on local density, the DBSCAN algorithm, etc.).

[0063] Figure 4 is a schematic diagram of an optional process for generating augmented image data according to an embodiment of the present disclosure. As Figure 4 shown, the process of generating augmented image data based on the collected edge image data includes: inputting the edge image data in the collected image data into the encoder of the generative adversarial network to obtain the corresponding edge data code (equivalent to the above first coding result); performing clustering on the edge data code in the feature space of the generative adversarial network to obtain the clustering center.

[0064] Based on the above preset probability distribution function, the above random code can be generated using the above generative adversarial network. Generating the augmented image data corresponding to the edge image data based on the random code can ensure the randomness of the augmented image data.

[0065] Performing an offset process on the above random code according to the above offset strategy can obtain the above second coding result. The above offset strategy can be used to determine to move the above random code in a direction closer to the above clustering center. Through the above offset process, the random code with randomness can be closer to the clustering center in the feature middle, and thus the augmented image data corresponding to the random code has a high data correlation with the edge image data.

[0066] As an optional implementation manner, the preset probability distribution function is a Gaussian distribution function. Based on the preset probability distribution function, using the generative adversarial network to generate a random code further includes the following method steps:

[0067] Step S2421: Perform random sampling on the Gaussian distribution function to obtain a sampling result;

[0068] Step S2422, using an encoder to encode the sampling result to obtain a random code.

[0069] The generation process of the above-mentioned random code can be: designating a preset probability distribution function as a Gaussian distribution function; randomly sampling the Gaussian distribution function to obtain a sampling result; inputting the sampling result into the encoder of the above-mentioned generative adversarial network for encoding, and then obtaining the above-mentioned random code.

[0070] Still like Figure 4 As shown, the process of generating the extended image data based on the collected edge image data also includes: randomly sampling the result of the Gaussian distribution function (such as Figure 4 The Gaussian sampling shown in ) is input into the encoder of the generative adversarial network to obtain random codes.

[0071] It is easy to notice that since the above sampling results are obtained by random sampling, the random codes corresponding to the above sampling results are random, which can ensure that the expanded image data corresponding to the random codes are random, thereby avoiding the situation where the expanded image data is too similar to the edge image data, resulting in poor enhancement effect of the expanded image data on the trained model.

[0072] As an optional implementation, the random code is offset processed according to the offset strategy to obtain a second encoding result, and the method further includes the following steps:

[0073] Step S2431, determining an offset weight according to an offset strategy;

[0074] Step S2432: Based on the offset weight, perform weighted calculation on the cluster center and the random code to obtain a second coding result.

[0075] The above offset strategy can be pre-specified by a technician according to the actual application scenario requirements. According to the offset strategy, the offset weight for offsetting the above random coding can be determined. The offset weight is used to control the distance between the second coding result after offset and the above cluster center.

[0076] Based on the offset weights, the cluster centers corresponding to the edge image data and the random codes can be weighted to obtain the second coding result. That is, the second coding result is determined by calculating the cluster centers, the random codes and the offset weights.

[0077] Still like Figure 4 As shown, the process of generating the extended image data based on the collected edge image data also includes: performing weighted offset calculation on the cluster center and the random code to obtain the offset code (equivalent to the above second coding result).

[0078] Specifically, the above weighted offset calculation can be shown as in the following formula (1):

[0079] w′ = a × w + (1 - a) × c Formula (1)

[0080] In the above formula (1), a represents the offset weight, c represents the clustering center of the edge image data in the feature space, w represents the random encoding in the feature space, and w′ represents the offset encoding obtained after offset (equivalent to the above second encoding result).

[0081] It is easy to note that the above offset weight can be a constant value determined by the offset strategy, or a variable value that can be adjusted by a technician according to the visual result of the augmented image data after being determined by the offset strategy. As shown in formula (1), the smaller the value of the offset weight, the more similar the augmented image data is to the edge image data.

[0082] It is easy to note that by moving the above random encoding towards the clustering center in the feature space, the data correlation between the augmented image data corresponding to the random encoding and the edge image data can be improved, thereby avoiding the situation where the augmented image data has a poor enhancement effect on the trained model due to being too different from the edge image data.

[0083] As an alternative implementation, using a generative adversarial network to decode the second encoding result to obtain the augmented data corresponding to the edge image data further includes the following method steps:

[0084] Step S261, using the decoder of the generative adversarial network to obtain the augmented image data corresponding to the edge image data from the second encoding result.

[0085] Still as Figure 4 shown, the process of generating augmented image data based on the collected edge image data further includes: inputting the offset encoding (equivalent to the above second encoding result) obtained by weighted offset calculation of the clustering center and the random encoding into the decoder of the generative adversarial network, and the augmented image data corresponding to the edge image data can be obtained.

[0086] It is easy to note that through the above embodiments of the present disclosure, based on the generative adversarial network and the random encoding, generating augmented image data corresponding to the edge image data in the original image dataset can flexibly control the randomness of the augmented image data and the data correlation between the augmented image data and the edge image data. Therefore, the beneficial effects of the present disclosure are at least: improving the augmentation efficiency of the edge image data and enhancing the generalization ability of the corresponding model, thereby solving the technical problems of low data augmentation efficiency and poor model enhancement effect in the related art methods for augmenting edge data based on data enhancement or feature mining.

[0087] In summary, the present disclosure realizes the effects of improving the augmentation efficiency of edge image data and enhancing the generalization ability of the corresponding model through a technical solution that combines the use of a generative adversarial network and random coding for augmenting edge image data. The technical solution provided by the present disclosure can be applied to the scenario of road recognition in vehicle autonomous driving. The following takes this scenario as an example to detail the key technologies of the embodiments of the present disclosure.

[0088] First, obtain the original image dataset in the scenario of road recognition in vehicle autonomous driving. The original image dataset can be road images captured in real time by the vehicle's camera or road images stored in the vehicle's memory. The original image dataset can be used to recognize the current road where the vehicle is located to make autonomous driving-related decisions.

[0089] Exemplarily, the above original image dataset contains N road images, where e road images are uncommon road images (equivalent to the above-mentioned edge image data, such as images showing swamps, railways, woods, etc.), and f road images are common images (equivalent to the above-mentioned edge image data, such as images showing asphalt roads, gravel roads, wooden plank roads, bridges, green spaces, etc.). Here, e < f, and e + f = N.

[0090] Furthermore, through a training image set (equivalent to the above-mentioned original image dataset) pre-collected for training a road recognition model, self-supervised machine learning training is performed on the initial neural network to obtain a generative adversarial network (denoted as styleGAN01), and this styleGAN01 has an encoder structure and a decoder structure.

[0091] Using the encoder structure of styleGAN01, feature encoding is performed on the above-mentioned e uncommon road images, that is, the image features of the e uncommon road images are encoded into the hidden space of styleGAN01 to obtain edge feature encoding (equivalent to the above-mentioned first encoding result) w01.

[0092] Furthermore, in the feature space of the above-mentioned styleGAN01, feature clustering is performed on the edge feature encoding w01, and the feature clustering center c01 of the above-mentioned e uncommon road images can be obtained.

[0093] Further, in the above-mentioned StyleGAN01, a fully connected mapping layer is added to map the traditional Gaussian distribution noise to the w-space with stronger semantic continuity, so as to achieve semantic decoupling of image features. Thus, by randomly sampling in the above-mentioned w-space, the w-space coordinates po01 of the sampling points can be obtained (equivalent to the above-mentioned sampling result); using the encoder of StyleGAN01 to encode the w-space coordinates ra01 of the sampling points, the sampling code w01 can be obtained (equivalent to the above-mentioned random code).

[0094] Further, for the feature space of StyleGAN01, the offset strategy is set as weighted offset, as shown in the following formula (2):

[0095] w′01 = a × w01+(1 - a) × c01 Formula (2)

[0096] In the above formula (2), a represents the offset weight, and w′01 represents the offset code obtained after offset (equivalent to the above-mentioned second coding result).

[0097] Further, according to the above offset strategy, in the feature space of the above-mentioned StyleGAN01, weighted offset calculation is performed on the feature clustering center c01 and the sampling code w01 of the above-mentioned e uncommon road images, and the offset code w′01 obtained after offset can be obtained.

[0098] Further, using the decoder of the above-mentioned StyleGAN01 to decode the above offset code w′01, e2 uncommon road images corresponding to the above-mentioned e uncommon road images can be obtained (equivalent to the above-mentioned expanded image data).

[0099] It should be noted that since the above-mentioned e2 uncommon road images are generated based on the feature clustering centers and randomly sampled points of the above-mentioned e uncommon road images, these e2 uncommon road images have both a certain degree of randomness and a correlation with existing uncommon road images. Therefore, these e2 uncommon road images can be used to further train and enhance the effect of the road recognition model, which can improve the recognition accuracy of the road recognition model for uncommon road images, enhance the robustness of the road recognition model, and further improve the safety of vehicle autonomous driving, which is beneficial to practical scenario applications. Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present disclosure.

[0100] In the present disclosure, an image data augmentation device is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0101] Figure 5 is a structural block diagram of an image data augmentation device provided according to an embodiment of the present disclosure. As Figure 5 shown, the image data augmentation device 500 includes: an acquisition module 501, an encoding module 502, an offset module 503, and a decoding module 504.

[0102] The acquisition module 501 is used to acquire an original image data set, where the original image data set includes edge image data and non-edge image data, and the number of edge image data is less than the number of non-edge image data;

[0103] The encoding module 502 is used to perform feature encoding on the edge image data by using a generative adversarial network to obtain a first encoding result, where the generative adversarial network is trained by using the original image data set;

[0104] The offset module 503 is used to perform offset processing on a random code based on the first encoding result to obtain a second encoding result, where the random code is generated by the generative adversarial network and a preset probability distribution function;

[0105] The decoding module 504 is configured to use a generative adversarial network to decode the second encoded result to obtain augmented image data corresponding to the edge image data.

[0106] Optionally, the encoding module 502 is further configured to: train a generative adversarial network by self-supervised machine learning using the original image dataset, where the generative adversarial network includes an encoder and a decoder; use the encoder to encode the features of the edge image data into the feature space to obtain a first encoded result.

[0107] Optionally, the offset module 503 is further configured to: perform clustering processing on the first encoded result to obtain a clustering center corresponding to the edge image data; generate a random code using the generative adversarial network based on a preset probability distribution function; perform offset processing on the random code according to an offset strategy, where the offset strategy is used to determine to move the random code in a direction closer to the clustering center to obtain a second encoded result.

[0108] Optionally, the offset module 503 is further configured to: perform random sampling on the Gaussian distribution function to obtain a sampling result; use the encoder to encode the sampling result to obtain a random code.

[0109] Optionally, the offset module 503 is further configured to: determine an offset weight according to the offset strategy; perform weighted calculation on the clustering center and the random code based on the offset weight to obtain a second encoded result.

[0110] Optionally, the decoding module 504 is further configured to: use the decoder of the generative adversarial network to process the second encoded result to obtain augmented image data corresponding to the edge image data.

[0111] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: all the above-mentioned modules are located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.

[0112] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, including a memory and at least one processor. Computer instructions are stored in the memory, and the processor is configured to run the computer instructions to execute the steps in any one of the above method embodiments.

[0113] Optionally, the above-mentioned electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.

[0114] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:

[0115] Step S1, obtain an original image dataset, where the original image dataset includes edge image data and non-edge image data, and the number of edge image data is less than that of non-edge image data;

[0116] Step S2, perform feature encoding on the edge image data by using a generative adversarial network to obtain a first encoding result, where the generative adversarial network is trained by using the original image dataset;

[0117] Step S3, based on the first encoding result, perform an offset process on a random encoding to obtain a second encoding result, where the random encoding is generated by the generative adversarial network and a preset probability distribution function;

[0118] Step S4, decode the second encoding result by using the generative adversarial network to obtain augmented image data corresponding to the edge image data.

[0119] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and details are not described herein again.

[0120] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are stored in the non-transitory computer-readable storage medium, and the computer instructions are configured to execute the steps in any one of the above method embodiments when running.

[0121] Optionally, in this embodiment, the above non-transitory computer-readable storage medium may be configured to store a computer program for executing the following steps:

[0122] Step S1, obtain an original image dataset, where the original image dataset includes edge image data and non-edge image data, and the number of edge image data is less than that of non-edge image data;

[0123] Step S2, perform feature encoding on the edge image data by using a generative adversarial network to obtain a first encoding result, where the generative adversarial network is trained by using the original image dataset;

[0124] Step S3, based on the first encoding result, perform an offset process on a random encoding to obtain a second encoding result, where the random encoding is generated by the generative adversarial network and a preset probability distribution function;

[0125] Step S4, decode the second encoding result by using the generative adversarial network to obtain augmented image data corresponding to the edge image data.

[0126] Optionally, in this embodiment, the non-transitory computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing. More specific examples of the readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0127] According to an embodiment of the present disclosure, the present disclosure also provides a computer program product. The program code for implementing the image data augmentation method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatuses, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0128] The serial numbers of the above embodiments of the present disclosure are merely for description and do not represent the superiority or inferiority of the embodiments.

[0129] In the above embodiments of the present disclosure, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0130] In several embodiments provided by the present disclosure, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0131] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0132] In addition, in each embodiment of the present disclosure, each functional unit may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0133] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present disclosure. The foregoing storage medium includes: various media such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store program codes.

[0134] The above are only the preferred embodiments of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present disclosure, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present disclosure.

Claims

1. An image data augmentation method, characterized in that, Including: Obtain an original image dataset, where the original image dataset includes edge image data and non-edge image data, and the number of the edge image data is less than the number of the non-edge image data; Use a generative adversarial network to perform feature encoding on the edge image data to obtain a first encoding result, where the generative adversarial network is trained using the original image dataset; Based on the first encoding result, perform an offset process on a random encoding to obtain a second encoding result, where the random encoding is generated by the generative adversarial network and a preset probability distribution function; Use the generative adversarial network to decode the second encoding result to obtain augmented image data corresponding to the edge image data.

2. The method according to claim 1, wherein The using a generative adversarial network to perform feature encoding on the edge image data to obtain a first encoding result includes: Use the original image dataset to train the generative adversarial network through self-supervised machine learning, where the generative adversarial network includes an encoder and a decoder; Use the encoder to encode the features of the edge image data into a feature space to obtain the first encoding result.

3. The method according to claim 2, characterized in that, The based on the first encoding result, performing an offset process on a random encoding to obtain a second encoding result includes: Perform a clustering process on the first encoding result to obtain a clustering center corresponding to the edge image data; Based on the preset probability distribution function, use the generative adversarial network to generate the random encoding; Perform an offset process on the random encoding according to an offset strategy to obtain the second encoding result, where the offset strategy is used to determine to move the random encoding in a direction close to the clustering center.

4. The method according to claim 3, wherein The preset probability distribution function is a Gaussian distribution function, and the based on the preset probability distribution function, using the generative adversarial network to generate the random encoding includes: Perform random sampling on the Gaussian distribution function to obtain a sampling result; Use the encoder to encode the sampling result to obtain the random encoding.

5. The method according to claim 3, wherein The performing an offset process on the random encoding according to an offset strategy to obtain the second encoding result includes: Determine an offset weight according to the offset strategy; Based on the offset weight, perform a weighted calculation on the clustering center and the random encoding to obtain the second encoding result.

6. The method according to claim 2, wherein The using the generative adversarial network to decode the second encoding result to obtain augmented data corresponding to the edge image data includes: Use the decoder of the generative adversarial network to the second encoding result to obtain the augmented image data corresponding to the edge image data.

7. An image data augmentation device, characterized in that, Including: An acquisition module, configured to acquire an original image dataset, where the original image dataset includes edge image data and non-edge image data, and the number of the edge image data is less than the number of the non-edge image data; An encoding module, configured to use a generative adversarial network to perform feature encoding on the edge image data to obtain a first encoding result, where the generative adversarial network is trained using the original image dataset; An offset module, configured to perform an offset process on a random code based on the first coding result to obtain a second coding result, where the random code is generated by the generative adversarial network and a preset probability distribution function; A decoding module, configured to decode the second coding result by using the generative adversarial network to obtain the augmented image data corresponding to the edge image data.

8. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

10. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • A training sample data expansion method and device based on a variational auto-encoder

    CN109886388A

  • Recognizer processing method based on picture data expansion

    CN112200307A