Image sample augmentation method and device
Through the style transfer algorithm, the content image and style image are integrated in the VGG neural network model, the problem of sea image scene replacement is solved, the image samples are augmented, and the real and diverse sample data is generated.
Patent Information
- Application Number
- CN202411884752.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is difficult to effectively replace the scene after taking pictures of sea images, resulting in a lack of diversity in image samples.
The style transfer algorithm is used to input the content image and style image into the pre-trained VGG neural network model, and the target content image is generated by the fusion of the style feature map and the content feature map to augment the image sample.
More realistic and diverse data on sea scarce image samples are generated, improving the diversity and authenticity of image samples.
Smart Images

Figure CN120236161A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and particularly relates to a method and device for augmenting image samples. Background Art
[0002] Currently, for maritime images, an image of a certain scene on the sea surface may be captured. However, how to replace the scene of the captured image is an urgent problem to be solved at present. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and device for augmenting image samples.
[0004] According to a first aspect of the present invention, there is provided an augmentation of an image sample, including:
[0005] Obtaining a content image and a style image;
[0006] Using a style transfer algorithm, inputting the content image and the style image into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image;
[0007] Augmenting the image sample according to the target content image.
[0008] Optionally, the step of using a style transfer algorithm, inputting the content image and the style image into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image includes:
[0009] Inputting the style image and the content image into the VGG neural network model respectively, and outputting a style feature map of the style image and a content feature map of the content image at different layers;
[0010] Copying the content image, and fusing the copied content image and the style image to obtain the target content image.
[0011] Optionally, the method further includes:
[0012] Calculating a style loss and a content loss of the style feature map of the style image;
[0013] Calculating a total loss function according to the style loss and the content loss;
[0014] Updating parameters of the pre-trained VGG neural network model according to the total loss function.
[0015] Optionally, the step of calculating the style loss of the style feature map of the style image includes:
[0016] After using the Gram matrix to represent the style features and then calculating the mean square error, the style loss of the style feature map of the style image is obtained.
[0017] Optionally, inputting the style image and the content image into the VGG neural network model respectively, and outputting the style feature map of the style image and the content feature map of the content image at different layers includes:
[0018] Select the conv_2, conv_4, conv_8, conv_12, and conv_16 layers of VGG-19 to extract the style in the style picture and the target picture.
[0019] According to the second aspect of the present invention, an image sample augmentation device is provided, including:
[0020] An acquisition module for acquiring a content image and a style image;
[0021] A processing module for using a style transfer algorithm to input the content image and the style image into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image;
[0022] An augmentation module for augmenting the image sample according to the target content image.
[0023] Optionally, the processing module is used for:
[0024] Input the style image and the content image into the VGG neural network model respectively, and output the style feature map of the style image and the content feature map of the content image at different layers;
[0025] Copy the content image, and fuse the copied content image and the style image to obtain the target content image.
[0026] Optionally, the processing module is used for:
[0027] Calculate the style loss and content loss of the style feature map of the style image;
[0028] Calculate the total loss function according to the style loss and the content loss;
[0029] Update the parameters of the pre-trained VGG neural network model according to the total loss function.
[0030] Optionally, the processing module is used for:
[0031] After using the Gram matrix to represent the style features and then calculating the mean square error, the style loss of the style feature map of the style image is obtained.
[0032] Optionally, the processing module is configured to:
[0033] Select the conv_2, conv_4, conv_8, conv_12, and conv_16 layers of VGG-19 to extract the style in the style image and the target image.
[0034] In a third aspect, the present application discloses an electronic device, which includes: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to execute the method described in any of the above aspects.
[0035] In a fourth aspect, the present application discloses a non-transitory computer-readable storage medium, which when the instructions in the storage medium are executed by the processor of the electronic device, enables the electronic device to execute the method described in any of the above aspects.
[0036] In a fifth aspect, the present application discloses a computer program product, which when the instructions in the computer program product are executed by the processor of the electronic device, enables the electronic device to execute the method described in any of the above aspects.
[0037] The beneficial effects brought by the present invention are as follows:
[0038] As can be seen from the above solution, the embodiments of the present invention provide a method and device for augmenting image samples. By obtaining a content image and a style image; using a style transfer algorithm, inputting the content image and the style image into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image; augmenting the image samples according to the target content image, the augmentation of scarce marine image samples based on style transfer can generate more real and diverse sample data. Description of the Drawings
[0039] Figure 1 It is a schematic flowchart of a method for augmenting image samples provided according to an embodiment;
[0040] Figure 2 It is a schematic diagram of the effect of a method for augmenting image samples provided according to an embodiment;
[0041] Figure 3 It is a block diagram of a device for augmenting image samples of the present application.
[0042] Figure 4 It is a block diagram of an electronic device of the present application.
[0043] Figure 5 It is a block diagram of a computer-readable storage medium of the present application. Detailed Embodiments
[0044] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] Referring to Figure 1 , a flowchart of steps of an image sample augmentation method according to the present application is shown. This method can be applied to an electronic device. Specifically, the method may specifically include the following steps:
[0046] S101. Obtain a content image and a style image;
[0047] Specifically, style refers to the texture, color, and visual patterns at different spatial scales in an image, and content refers to the high-level macro structure of the image.
[0048] S102. Use a style transfer algorithm to input the content image and the style image into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image;
[0049] S103. Augment the image sample according to the target content image.
[0050] Specifically, the style transfer algorithm selects two pictures as the style picture and the content picture respectively. By using a pre-trained VGG neural network, the style information of the style picture and the content information of the content picture are extracted by passing the style picture and the content picture through the network and output by different layers. Different from other deep learning algorithms, this algorithm does not update the parameters of the VGG network used for feature extraction, but selects to copy the content picture and use it as the target picture, calculates the style loss and content loss of the target picture, and directly updates the parameters of the target picture, so as to obtain a picture with the target style.
[0051] Another embodiment of the present application further supplements and explains the image sample augmentation method provided in the above embodiment.
[0052] Optionally, using a style transfer algorithm to input the content image and the style image into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image includes:
[0053] Input the style image and the content image into the VGG neural network model respectively, and output the style feature map of the style image and the content feature map of the content image at different layers;
[0054] Copy the content image, and fuse the copied content image with the style image to obtain the target content image.
[0055] Optionally, the method further includes:
[0056] Calculate the style loss and content loss of the style feature map of the style image;
[0057] Calculate the total loss function according to the style loss and content loss;
[0058] Update the parameters of the pre-trained VGG neural network model according to the total loss function.
[0059] Specifically, in the lower layers of the VGG_19 network, the feature map can relatively completely preserve information such as the color, corners, and lines of the original image. In the higher layers of the VGG_19 network, a large amount of low-level features will be lost, but the position and general outline of the ship in the original image can still be preserved. Therefore, it is possible to choose to pass the target image and the content image through the pre-trained network respectively, output by the higher layer of the network, and calculate the loss of the output feature map to guide the generation of the target image, so as to ensure that the content part of the target image is consistent with the content image. The loss function of the content loss part:
[0060]
[0061] The content of an image is represented by its local features. Therefore, the mean squared error loss can be directly used to restore the feature map of the middle layer. However, the style of an image is represented by the global features of the entire image, and the convolutional neural network is based on local connections. Although the receptive field becomes larger as the depth of the neural network increases, it is still a local feature. Therefore, the style loss cannot be calculated in the same way as the content loss by directly calculating the mean squared error of the feature map. Instead, the mean squared error is calculated after representing the style features by the gram matrix.
[0062]
[0063] It can be seen from this matrix that each element in a feature vector is multiplied by each element in other feature vectors to establish a connection, thereby extracting the global features, so that the style features of the picture can be highlighted.
[0064] Optionally, calculating the style loss of the style feature map of the style image includes:
[0065] Use the gram matrix to represent the style features and then calculate the mean squared error to obtain the style loss of the style feature map of the style image.
[0066] Specifically, the noisy image and the style image are simultaneously passed through the pre-trained VGG_19 network. After obtaining the feature maps output by the intermediate layers, the corresponding Gram matrices are calculated for the feature maps. Then, the mean squared error loss is calculated for the Gram matrices of the two to obtain the style loss between them, and this style loss is used to guide the generation of the target image to restore the style of the style image. A large amount of information such as color and texture is lost in the high-level layers of the VGG_19 network, while more color and texture information is retained in its low-level layers, which can better restore the style of the original style image. However, the texture of the style image restored only using the low-level feature maps often has a small scale, only retaining the color and local texture of the original style image, while the style restored by combining the low-level and high-level feature maps.
[0067] At this time, the style loss can be expressed as:
[0068]
[0069] Among them, \(L_{style}\) represents the style loss, which is the sum of the Gram matrices of the feature maps of the target image on 5 convolutional layers. \(\sum_{l=1}^{5}G_{s}^{l}\) is the sum of the Gram matrices of the feature maps of the style image on 5 convolutional layers. \(G_{s}^{l}\) represents the feature map of the style image in the \(l\)-th convolutional layer. In the denominator, \(N\) and \(M\) are the two dimensions of the Gram matrix \(G\), \(4N\) 2 M 2 is the normalization term, and its purpose is to prevent the order of magnitude of the style loss from differing greatly from that of the content loss.
[0070] Optionally, the style image and the content image are respectively input into the VGG neural network model, and the style feature maps of the style image and the content feature maps of the content image are output at different layers, including:
[0071] Select the conv_2, conv_4, conv_8, conv_12, and conv_16 layers of VGG-19 to extract the style in the style image and the target image.
[0072] Finally, the network structure is as Figure 2 shown. After the content image and the target image pass through the VGG-19 network respectively, they are output by the conv_12 layer, and the content information of both is extracted simultaneously, and the content loss is calculated for them; for the style loss, it has been proven by experiments that the "style image" restored by combining the deep layer and the shallow layer is more real and closer to the original image. Therefore, the conv_2, conv_4, conv_8, conv_12, and conv_16 layers of VGG-19 are selected to extract the style in the style image and the target image.
[0073] After obtaining the content loss and the style loss, the two are weighted and summed to obtain the complete loss:
[0074] Ltotle (t, c, s) = αL content (t, c, l) + βL style (t, s, l)
[0075] Wherein, L totle is the total loss function, and L content and L style are the weights of the content loss and the style loss respectively. The parameters of the target image are adjusted according to the total loss function, and after multiple cycles, the target image with the style of the style image and the content of the content image can be obtained.
[0076] It should be noted that each feasible method in this embodiment can be implemented alone or in any combination without conflict. The present application does not make any limitations.
[0077] The embodiment of the present invention provides a method for augmenting image samples. By obtaining a content image and a style image; using a style transfer algorithm, the content image and the style image are input into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image; the image samples are augmented according to the target content image. The augmentation of scarce marine image samples based on style transfer can generate more real and diverse sample data.
[0078] Another embodiment of the present application provides an apparatus for augmenting image samples, which is used to execute the method for augmenting image samples provided in the above embodiment.
[0079] As Figure 3 shown, it is a schematic structural diagram of the apparatus for augmenting image samples provided in the embodiment of the present application. The xx apparatus includes an acquisition module 301, a processing module 302, and an augmentation module 303, wherein:
[0080] The acquisition module 301 is used to acquire a content image and a style image;
[0081] The processing module 302 is used to use a style transfer algorithm to input the content image and the style image into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image;
[0082] The augmentation module 303 is used to augment the image samples according to the target content image.
[0083] Regarding the apparatus in this embodiment, the specific manners in which each module performs operations have been described in detail in the embodiment related to the method, and will not be elaborated here.
[0084] Another embodiment of the present application further supplements the apparatus for augmenting image samples provided in the above embodiment.
[0085] Optionally, the processing module is configured to:
[0086] Input the style image and the content image into the VGG neural network model respectively, and output the style feature map of the style image and the content feature map of the content image at different layers;
[0087] Copy the content image, and fuse the copied content image and the style image to obtain a target content image.
[0088] Optionally, the processing module is configured to:
[0089] Calculate the style loss and the content loss of the style feature map of the style image;
[0090] Calculate the total loss function according to the style loss and the content loss;
[0091] Update the parameters of the pre-trained VGG neural network model according to the total loss function.
[0092] Optionally, the processing module is configured to:
[0093] Adopt the mean square error after representing the style features by the gram matrix to obtain the style loss of the style feature map of the style image.
[0094] Optionally, the processing module is configured to:
[0095] Select the conv_2, conv_4, conv_8, conv_12, and conv_16 layers of VGG-19 to extract the style in the style picture and the target picture.
[0096] The embodiment of the present invention provides a method for augmenting image samples. By obtaining a content image and a style image; adopting a style transfer algorithm, inputting the content image and the style image into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image; augmenting the image samples according to the target content image, the augmentation of scarce marine image samples based on style transfer can generate more real and diverse sample data.
[0097] For the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the related parts, please refer to the partial description of the method embodiment.
[0098] Optionally, an embodiment of the present application further provides an electronic device, including: a processor, a memory, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0099] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above method embodiments and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0100] The block diagram of an electronic device 800 shown in the present application. For example, the electronic device 800 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0101] The electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0102] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0103] The memory 804 is configured to store various types of data to support the operation of the device 800. Examples of these data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, images, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0104] The power supply component 806 provides power for various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0105] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0106] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0107] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0108] The sensor component 814 includes one or more sensors for providing status assessments of various aspects of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the device 800, the relative positioning of components, such as the display and the keypad of the electronic device 800. The sensor component 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0109] The communication component 816 is configured to facilitate communication, either wired or wirelessly, between the electronic device 800 and other devices. The electronic device 800 may access a wireless network based on communication standards, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast operation information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0110] In an exemplary embodiment, the electronic device 800 may be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described methods.
[0111] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions that can be executed by a processor 820 of the electronic device 800 to complete the above-described methods. For example, the non-transitory computer-readable storage medium may be a ROM, a Random Access Memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, among others.
[0112] A block diagram of a computer-readable storage medium 1900 illustrated in the present application. For example, the computer-readable storage medium 1900 may be provided as a server.
[0113] The computer-readable storage medium 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described methods.
[0114] The computer-readable storage medium 1900 may further include a power supply component 1926 configured to perform power management of the computer-readable storage medium 1900, a wired or wireless network interface 1950 configured to connect the computer-readable storage medium 1900 to a network, and an input / output (I / O) interface 1958. The computer-readable storage medium 1900 may operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, or the like.
[0115] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article, or apparatus including such element.
[0116] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0117] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.
[0118] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0119] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0120] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical, or other forms.
[0121] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0122] In addition, the functional units in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0123] If the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0124] As described above, the above are only the specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0125] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A method for augmenting an image sample, characterized in that: include: Get content image and style image; Using a style transfer algorithm, the content image and the style image are input into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image; Image samples are augmented according to the target content image.
2. The method for augmenting an image sample according to claim 1, characterized in that: The adopting of a style transfer algorithm to input the content image and the style image into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image includes: Inputting the style image and the content image into the VGG neural network model respectively, and outputting a style feature map of the style image and a content feature map of the content image at different layers; The content image is copied, and the copied content image is merged with the style image to obtain the target content image.
3. The method for augmenting an image sample according to claim 2, characterized in that: The method further comprises: Calculating the style loss and content loss of the style feature map of the style image; Calculating a total loss function according to the style loss and the content loss; According to the total loss function, the parameters of the pre-trained VGG neural network model are updated.
4. The method for augmenting an image sample according to claim 3, characterized in that: The calculating the style loss of the style feature map of the style image includes: The style features are represented by a gram matrix and then the mean square error is calculated to obtain the style loss of the style feature graph of the style image.
5. The method for augmenting an image sample according to claim 2, characterized in that: The step of inputting the style image and the content image into the VGG neural network model respectively, and outputting a style feature map of the style image and a content feature map of the content image at different layers, comprises: Select conv_2, conv_4, conv_8, conv_12, and conv_16 layers of VGG-19 to extract the styles in the style images and target images.
6. An image sample augmentation device, characterized in that: include: An acquisition module, used to acquire content images and style images; A processing module, configured to use a style transfer algorithm to input the content image and the style image into a pre-trained VGG neural network model to obtain a target content image corresponding to the style image; An augmentation module is used to augment image samples according to the target content image.
7. The image sample augmentation device according to claim 6, characterized in that: The processing module is used for: Inputting the style image and the content image into the VGG neural network model respectively, and outputting a style feature map of the style image and a content feature map of the content image at different layers; The content image is copied, and the copied content image is merged with the style image to obtain the target content image.
8. The method for augmenting an image sample according to claim 7, characterized in that: The processing module is used for: Calculating the style loss and content loss of the style feature map of the style image; Calculating a total loss function according to the style loss and the content loss; According to the total loss function, the parameters of the pre-trained VGG neural network model are updated.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the method according to any one of claims 1 to 5 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.