Animal image generation method, device, equipment and storage medium
By using animal image generation models in video interactive applications, integrating and encoding animal image feature information, and generating target animal image images, the problem of limited animal image transformation types in the prior art is solved, personalized image transformation is realized, and user experience is improved.
Patent Information
- Application Number
- CN202111152039.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-09-29
AI Technical Summary
The types of transformation of animal images in existing video interaction applications are limited and cannot meet users' personalized image transformation needs.
By generating the animal image based on the animal image, at least two animal image images and corresponding image feature information are obtained, fused to generate mixed image feature information, and attribute encoding is generated through an encoder based on the preset attribute information, and finally input this information into the animal image generation model to generate the target animal image image.
It realizes the animal image that generates users' personalized needs and improves user experience and satisfaction.
Smart Images

Figure CN113850890B_ABST
Abstract
Description
Technical Field
[0001] The disclosed embodiments relate to the field of image processing technology, and in particular to a method, device, equipment and storage medium for generating an animal image. Background Art
[0002] With the development of science and technology, more and more application software have entered the lives of users and gradually enriched their spare time, such as short video applications (Application, APP), photo editing APP Qingyan, Xingtu, etc.
[0003] Currently, some users like to upload photos of small animals (such as cats and dogs) or use photos as avatars. By transforming the animal image, users can obtain their favorite animal image. However, the types of transformations of animal images in existing video interaction applications are still limited and cannot meet the personalized image transformation needs of users. Summary of the invention
[0004] The embodiments of the present disclosure provide a method, device, equipment and storage medium for generating an animal image, which can generate an animal image that meets the personalized needs of a user and improve the user experience.
[0005] In a first aspect, an embodiment of the present disclosure provides a method for generating an animal image, comprising:
[0006] Based on the animal image generation model, obtaining at least two animal image images and at least two sets of image feature information corresponding to the animal image images;
[0007] fusing the at least two sets of image feature information to obtain mixed image feature information;
[0008] Inputting the preset attribute information into a preset encoder to obtain an attribute code;
[0009] The mixed image feature information and the attribute code are input into the animal image generation model to obtain a target animal image and target image feature information.
[0010] In a second aspect, the embodiment of the present disclosure further provides a device for generating an animal image, comprising:
[0011] An animal image acquisition module, used to obtain at least two animal image images and at least two sets of image feature information corresponding to the animal image images based on an animal image generation model;
[0012] A mixed image feature information obtaining module, used for fusing the at least two sets of image feature information to obtain mixed image feature information;
[0013] An attribute encoding module, used for inputting preset attribute information into a preset encoder to obtain an attribute code;
[0014] The target animal image acquisition module is used to input the mixed image feature information and the attribute code into the animal image generation model to obtain the target animal image and target image feature information.
[0015] In a third aspect, an embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0016] one or more processing devices;
[0017] A storage device for storing one or more programs;
[0018] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the method for generating an animal image as described in the embodiment of the present disclosure.
[0019] In a fourth aspect, an embodiment of the present disclosure discloses a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the method for generating an animal image as described in the embodiment of the present disclosure.
[0020] The disclosed embodiment discloses a method, device, equipment and storage medium for generating an animal image. The method comprises: based on an animal image generation model, obtaining at least two animal image images and at least two sets of image feature information corresponding to the animal image images; fusing at least two sets of image feature information to obtain mixed image feature information; inputting preset attribute information into a preset encoder to obtain attribute coding; inputting mixed image feature information and attribute coding into an animal image generation model to obtain a target animal image image and target image feature information. The method for generating an animal image provided by the disclosed embodiment inputs mixed image feature information and attribute coding into an animal image generation model to obtain a target animal image image and target image feature information, and can generate an animal image that meets the personalized needs of a user, thereby improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of a method for generating an animal image in an embodiment of the present disclosure;
[0022] Figure 2 is an example diagram of generating an animal image in an embodiment of the present disclosure;
[0023] Figure 3 is a schematic structural diagram of a device for generating an animal image in an embodiment of the present disclosure;
[0024] Figure 4 It is a structural schematic diagram of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0026] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0027] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0028] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0029] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0030] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0031] Figure 1 is a flow chart of a method for generating an animal image provided by an embodiment of the present disclosure. This embodiment is applicable to the case where an animal image is transformed according to the personalized needs of a user. The method can be executed by an animal image generation device, which can be composed of hardware and / or software and can generally be integrated in a device with a function of transforming an animal image, which can be an electronic device such as a server, a mobile terminal or a server cluster. Figure 1 As shown, the method specifically comprises the following steps:
[0032] Step 110: obtaining at least two animal image images and at least two sets of image feature information corresponding to the animal image images based on the animal image generation model.
[0033] Among them, the animal image generation model has a neural network model that can be understood as a function of generating animal image images.
[0034] The image feature information can be specifically understood as the encoding of the features of the animal image. In this embodiment, for various animal image features, corresponding encoding can be performed, that is, a quantization process. For example, the image feature information can be represented in the form of a matrix or a vector.
[0035] Specifically, at least two groups of image feature information are input into the animal image generation model, and the animal image generation model generates at least two animal image images and at least two groups of image feature information corresponding to the animal image images according to the image feature information.
[0036] The image feature information input into the animal image generation model may be a random feature code or image feature information output by the animal image generation model.
[0037] The random feature code can be understood as a feature code generated by a computer according to a set random algorithm. The image feature information output by the animal image generation model is input into the animal image generation model again, so as to realize multi-generation fusion of animal images.
[0038] In this embodiment, the animal image generation model can be obtained based on the training of the generative adversarial model. The animal image generation model corresponds to the generation model in the generative adversarial model.
[0039] Specifically, the training method of the animal image generation model can be: cross-iterative training of the generation model and the discrimination model until the accuracy of the discrimination result output by the discrimination model meets the set conditions, and then the trained generation model is determined as the animal image generation model.
[0040] In this embodiment, the cross-iterative training of the generation model and the discrimination model can be performed by inputting random noise to obtain the discrimination result, keeping the parameters of the discrimination model unchanged, training the generation model according to the discrimination result, and adjusting the parameters of the generation model. Then, the discrimination result is obtained by inputting random noise again; keeping the parameters of the generation model unchanged, training the discrimination model according to the discrimination result, and adjusting the parameters of the generation model. This is repeated until the accuracy of the discrimination result output by the discrimination model meets the set conditions, and the trained generation model is determined as the animal image generation model.
[0041] The process of cross-iteration training is:
[0042] a1) inputting the first random noise data into the generation model to obtain the first animal image data; inputting the first animal image data and the first animal image sample data into the discrimination model to obtain the first discrimination result; adjusting the parameters in the generation model based on the first discrimination result.
[0043] The animal image sample data is specifically understood as an animal image that displays the characteristics of a real animal image, which can be obtained by collecting images of animals taken on the Internet. The animal image sample data involved in the animal image generation model can be for different animal types or for different animal species under the same animal type, and multiple animal image generation models can be trained separately. The animal image output data can include the animal image and the image feature information corresponding to the animal image.
[0044] Specifically, the first random noise data is input into the generation model to obtain the first animal image data, and the first animal image data and the first animal image sample data are input into the discrimination model. According to the discrimination result, the parameters in the initial generation model are adjusted so that the animal image output data can better restore the input animal image sample data, thereby obtaining a more accurate animal image generation model.
[0045] Exemplarily, the discrimination result may be expressed in terms of simulation degree. The higher the simulation degree, the more accurate the generated model is; the lower the simulation degree, the less accurate the generated model is.
[0046] b1) Inputting the second random noise data into the adjusted generation model to obtain second animal image data; inputting the second animal image data and the second animal image sample into the discrimination model to obtain a second discrimination result, and determining the true discrimination result between the second animal image data and the second animal image sample; adjusting the parameters in the discrimination model according to the loss function of the second discrimination result and the true discrimination result.
[0047] Among them, the discriminant model can be understood as the discriminant model in the generative adversarial network, which is trained adversarially with the generative model.
[0048] In this step, the loss function is obtained by comparing the second discrimination result obtained by the discriminant model with the true discrimination result, and the parameters of the discriminant model are adjusted according to the loss function to make the discriminant model more accurate.
[0049] For example, the smaller the loss function is, the more accurate the discriminant model is.
[0050] Step 120, fusing at least two sets of image feature information to obtain mixed image feature information.
[0051] Specifically, weighted sum calculation may be performed on at least two groups of image feature information according to preset weights to obtain mixed image feature information.
[0052] The preset weights can be set arbitrarily by the user. For example, assuming that there are currently three groups of image feature information, namely e1, e2 and e3, and the weights set by the user are 0.5, 0.2 and 0.3 respectively, then the calculation formula for the mixed image feature information is e=0.5*e1+0.2*e2+0.3*e3. In this embodiment, the weight can represent the proportion of the group of image features in the mixed image features.
[0053] Step 130: input the preset attribute information into a preset encoder to obtain the attribute code.
[0054] The attribute information can be specifically understood as information characterizing the characteristics of the animal's image, and the attribute information includes at least one of the following: age, hair color, image angle, and breed. The preset attribute information can be set according to user needs. The encoder has the function of editing the attribute information into a digital code, that is, it has the function of quantifying the attribute information. In this embodiment, the encoder can be a neural network with encoding function.
[0055] Specifically, according to user needs, the preset attribute information is input into the preset encoder, and the encoder compiles and converts the preset attribute information to obtain the attribute code. The attribute code can be represented in the form of a matrix. For example, assuming that the preset attribute information is age 10 years old, the age 10 years old is input into the encoder, and the encoder outputs the encoding information corresponding to the age 10 years old.
[0056] In this embodiment, the encoder may be trained in the following manner:
[0057] a2) Input the real attribute information into the initial encoder to obtain the initial attribute code.
[0058] Specifically, the real attribute information of the animal is input into the initial encoder, and the initial encoder encodes the input real attribute according to the stored rules to obtain the initial attribute code.
[0059] b2) Inputting the initial attribute code and the preset animal image feature information into the trained animal image generation model to obtain the training animal image and the training image feature information.
[0060] Specifically, the initial attribute code represents the attribute information of the animal, and the preset animal image feature information represents the image features of the animal image. The initial attribute code and the preset animal image feature information are input into the trained animal image generation model to obtain the training animal image and the training image feature information. The animal image in the training animal image carries the attribute features in the real initial attribute code.
[0061] c2) Determine the encoding attribute information based on the training animal image.
[0062] After obtaining the training animal image, the training animal image is recognized to obtain attribute information of the training animal image, that is, coded attribute information.
[0063] Specifically, the method of determining the coded attribute information according to the training animal image may be: inputting the training animal image into a preset attribute recognition model to obtain the coded attribute information.
[0064] Among them, the attribute recognition model has the function of identifying the encoded attribute information.
[0065] d2) Train the initial encoder according to the loss function of the real attribute information and the encoded attribute information to obtain a trained encoder.
[0066] Among them, the loss function can also be called a cost function, which can be specifically understood as a function that characterizes the difference between real attribute information and encoded attribute information.
[0067] Specifically, the loss function of the real attribute information and the encoded attribute information is calculated, and the parameters of the initial encoder are adjusted according to the loss function until the loss function meets the set conditions, and the encoder training is completed.
[0068] Step 140, inputting the mixed image feature information and attribute code into the animal image generation model to obtain the target animal image and target image feature information.
[0069] The target animal image refers to an animal image obtained by mixing and transforming at least two animal image images, and correspondingly, the target image feature information refers to image feature information corresponding to the obtained animal image.
[0070] Specifically, the mixed image feature information represents the features of the animal image, and the attribute coding represents the attribute information of the animal image, which is input into the animal image generation model to obtain the target animal image and the target image feature information.
[0071] In order to more clearly describe the embodiments of the present disclosure, Figure 2 is an example diagram of generating an animal image in an embodiment of the present disclosure, for example, Figure 2As shown, the animal image generation model is represented by G1, and the encoder is represented by E. The specific process of generating an animal image can be described as follows: based on the animal image generation model G1, obtain an animal image x1, an animal image x2, image feature information e1 corresponding to the animal image x1, and image feature information e2 corresponding to the animal image x2; fuse the image feature information e1 and e2 to obtain mixed image feature information; edit the age attribute and input it into the preset encoder E to obtain the attribute code; input the mixed image feature information and the attribute code into the animal image generation model G1 to obtain the target animal image x3 and the target image feature information e3.
[0072] The disclosed embodiment discloses a method, device, equipment and storage medium for generating an animal image. The method comprises: based on an animal image generation model, obtaining at least two animal image images and at least two sets of image feature information corresponding to the animal image images; fusing at least two sets of image feature information to obtain mixed image feature information; inputting preset attribute information into a preset encoder to obtain attribute coding; inputting mixed image feature information and attribute coding into an animal image generation model to obtain a target animal image image and target image feature information. The method for generating an animal image provided by the disclosed embodiment inputs mixed image feature information and attribute coding into an animal image generation model to obtain a target animal image image and target image feature information, and can generate an animal image that meets the personalized needs of a user, thereby improving the user experience.
[0073] Figure 3 Schematic diagram of a device for generating an animal image disclosed in an embodiment of the present disclosure. Figure 3 As shown, the device comprises:
[0074] The animal image acquisition module 210 is used to obtain at least two animal image images and at least two sets of image feature information corresponding to the animal image images based on the animal image generation model;
[0075] A mixed image feature information obtaining module 220 is used to fuse at least two sets of image feature information to obtain mixed image feature information;
[0076] The attribute encoding module 230 is used to input the preset attribute information into a preset encoder to obtain the attribute code;
[0077] The target animal image acquisition module 240 is used to input the mixed image feature information and attribute code into the animal image generation model to obtain the target animal image and target image feature information.
[0078] Optionally, the animal image acquisition module 210 is further used for:
[0079] The random feature code or the image feature information output by the animal image generation model is input into the animal image generation model to obtain at least two animal image images and at least two groups of image feature information corresponding to the animal image images.
[0080] Optionally, the mixed image feature information obtaining module 220 is further used for:
[0081] A weighted sum calculation is performed on at least two groups of image feature information according to preset weights to obtain mixed image feature information.
[0082] Optionally, the device further comprises:
[0083] A training module for animal image generation models, used to:
[0084] The generative model and the discriminative model are cross-trained iteratively until the accuracy of the discrimination result output by the discriminative model meets the set conditions, and then the trained generative model is determined as the animal image generation model;
[0085] The process of cross-iteration training is:
[0086] Inputting first random noise data into a generation model to obtain first animal image data;
[0087] Inputting the first animal image data and the first animal image sample data into a discrimination model to obtain a first discrimination result;
[0088] Adjusting parameters in the generation model based on the first discrimination result;
[0089] inputting the second random noise data into the adjusted generation model to obtain second animal image data;
[0090] Inputting the second animal image data and the second animal image sample into the discrimination model to obtain a second discrimination result, and determining a true discrimination result between the second animal image data and the second animal image sample;
[0091] The parameters in the discrimination model are adjusted according to the loss function of the second discrimination result and the true discrimination result.
[0092] Optionally, the device further comprises:
[0093] The encoder training module includes:
[0094] An initial attribute code obtaining unit, used for inputting the real attribute information into an initial encoder to obtain an initial attribute code;
[0095] A training animal image acquisition unit is used to input the initial attribute code and the preset animal image feature information into the trained animal image generation model to obtain the training animal image and the training image feature information;
[0096] A coding attribute information determination unit, used to determine coding attribute information based on the training animal image;
[0097] The encoder acquisition unit is used to train the initial encoder according to the loss function of the real attribute information and the encoded attribute information to obtain the trained encoder.
[0098] Optionally, the encoding attribute information determining unit is further used to:
[0099] The training animal image is input into a preset attribute recognition model to obtain the encoded attribute information.
[0100] Optionally, the attribute information includes at least one of the following: age, hair color, image angle and breed.
[0101] The above device can execute the methods provided by all the above embodiments of the present disclosure, and has the corresponding functional modules and beneficial effects of executing the above methods. For technical details not fully described in this embodiment, please refer to the methods provided by all the above embodiments of the present disclosure.
[0102] Reference below Figure 4 , which shows a schematic diagram of the structure of an electronic device 300 suitable for implementing the embodiment of the present disclosure. The electronic device in the embodiment of the present disclosure may include but is not limited to mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc., or various forms of servers, such as independent servers or server clusters. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0103] like Figure 4 As shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only storage device (ROM) 302 or a program loaded from a storage device 308 to a random access storage device (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0104] Typically, the following devices may be connected to the I / O interface 305: input devices 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 308 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 309. The communication devices 309 may allow the electronic device 300 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 300 is shown with various devices, but it should be understood that it is not required to implement or possess all the devices shown. More or fewer devices may be implemented or possessed instead.
[0105] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing a method for recommending words. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device 309, or installed from a storage device 308, or installed from a ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0106] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0107] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0108] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0109] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains at least two animal image images and at least two sets of image feature information corresponding to the animal image images based on the animal image generation model; fuses the at least two sets of image feature information to obtain mixed image feature information; inputs preset attribute information into a preset encoder to obtain attribute code; inputs the mixed image feature information and the attribute code into the animal image generation model to obtain a target animal image and target image feature information.
[0110] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including, but not limited to, object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0111] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0112] The units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a unit does not, in some cases, limit the unit itself.
[0113] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0114] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0115] According to one or more embodiments of the present disclosure, the present disclosure discloses a method for generating an animal image, including:
[0116] Based on the animal image generation model, obtaining at least two animal image images and at least two sets of image feature information corresponding to the animal image images;
[0117] fusing the at least two sets of image feature information to obtain mixed image feature information;
[0118] Inputting the preset attribute information into a preset encoder to obtain an attribute code;
[0119] The mixed image feature information and the attribute code are input into the animal image generation model to obtain a target animal image and target image feature information.
[0120] Furthermore, based on the animal image generation model, at least two animal image images and at least two sets of image feature information corresponding to the animal image images are obtained, including:
[0121] The random feature code or the image feature information output by the animal image generation model is input into the animal image generation model to obtain at least two animal image images and at least two groups of image feature information corresponding to the animal image images.
[0122] Further, the at least two sets of image feature information are fused to obtain mixed image feature information, including:
[0123] The at least two groups of image feature information are weighted and summed according to preset weights to obtain mixed image feature information.
[0124] Furthermore, the training method of the animal image generation model is:
[0125] The generative model and the discriminative model are cross-trained iteratively until the accuracy of the discrimination result output by the discriminative model meets the set conditions, and then the trained generative model is determined as the animal image generation model;
[0126] The process of cross-iteration training is:
[0127] Inputting first random noise data into a generation model to obtain first animal image data;
[0128] Inputting the first animal image data and the first animal image sample data into a discrimination model to obtain a first discrimination result;
[0129] Adjusting parameters in the generation model based on the first discrimination result;
[0130] inputting the second random noise data into the adjusted generation model to obtain second animal image data;
[0131] Inputting the second animal image data and the second animal image sample into the discrimination model to obtain a second discrimination result, and determining a true discrimination result between the second animal image data and the second animal image sample;
[0132] The parameters in the discrimination model are adjusted according to the loss function of the second discrimination result and the true discrimination result.
[0133] Furthermore, the encoder is trained in the following manner:
[0134] Input the real attribute information into the initial encoder to obtain the initial attribute code;
[0135] Inputting the initial attribute code and the preset animal image feature information into a trained animal image generation model to obtain a training animal image and training image feature information;
[0136] Determining encoding attribute information according to the training animal image;
[0137] The initial encoder is trained according to the loss function of the real attribute information and the encoded attribute information to obtain a trained encoder.
[0138] Further, determining the encoding attribute information according to the training animal image includes:
[0139] The training animal image is input into a preset attribute recognition model to obtain coded attribute information.
[0140] Furthermore, the attribute information includes at least one of the following: age, hair color, image angle and breed.
[0141] Note that the above are only preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art will understand that the present disclosure is not limited to the specific embodiments herein, and that various obvious changes, readjustments and substitutions can be made by those skilled in the art without departing from the scope of protection of the present disclosure. Therefore, although the present disclosure is described in more detail through the above embodiments, the present disclosure is not limited to the above embodiments, and may include more other equivalent embodiments without departing from the concept of the present disclosure, and the scope of the present disclosure is determined by the scope of the appended claims.
Claims
1. A method for generating an animal image, characterized in that: include: Based on the animal image generation model, at least two animal image images and at least two sets of image feature information corresponding to the animal image images are obtained, wherein the animal image is a non-human animal image, and the image feature information is a code representing the features of the animal image images; The input of the animal image generation model is a random feature code or image feature information output by the animal image generation model; fusing the at least two sets of image feature information to obtain mixed image feature information; Inputting the preset attribute information into a preset encoder to obtain an attribute code; Inputting the mixed image feature information and the attribute code into the animal image generation model to obtain a target animal image and target image feature information, wherein the mixed image feature information represents the features of the animal image, and the attribute code represents the animal attribute information, wherein the attribute information includes: age, hair color, image angle and breed, and the target image feature information is the image feature information corresponding to the target animal image; The encoder is trained in the following way: Input the real attribute information into the initial encoder to obtain the initial attribute code; Inputting the initial attribute code and the preset animal image feature information into an animal image generation model to obtain a training animal image and training image feature information; Determining encoding attribute information according to the training animal image; The initial encoder is trained according to the loss function of the real attribute information and the encoded attribute information to obtain a trained encoder.
2. The method according to claim 1, characterized in that: Based on the animal image generation model, at least two animal image images and at least two sets of image feature information corresponding to the animal image images are obtained, including: The random feature code or the image feature information output by the animal image generation model is input into the animal image generation model to obtain at least two animal image images and at least two groups of image feature information corresponding to the animal image images.
3. The method according to claim 1, characterized in that The at least two sets of image feature information are fused to obtain mixed image feature information, including: The at least two groups of image feature information are weighted and summed according to preset weights to obtain mixed image feature information.
4. The method according to claim 1, characterized in that The training method of the animal image generation model is: The generative model and the discriminative model are cross-trained iteratively until the accuracy of the discrimination result output by the discriminative model meets the set conditions, and then the trained generative model is determined as the animal image generation model; The process of cross-iteration training is: Inputting first random noise data into a generation model to obtain first animal image data; Inputting the first animal image data and the first animal image sample data into a discrimination model to obtain a first discrimination result; Adjusting parameters in the generation model based on the first discrimination result; inputting the second random noise data into the adjusted generation model to obtain second animal image data; Inputting the second animal image data and the second animal image sample into the discrimination model to obtain a second discrimination result, and determining a true discrimination result between the second animal image data and the second animal image sample; The parameters in the discrimination model are adjusted according to the loss function of the second discrimination result and the true discrimination result.
5. The method according to claim 1, characterized in that Determining encoding attribute information according to the training animal image includes: The training animal image is input into a preset attribute recognition model to obtain coded attribute information.
6. A device for generating an animal image, characterized in that: include: An animal image acquisition module is used to obtain at least two animal image images and at least two sets of image feature information corresponding to the animal image images based on an animal image generation model, wherein the animal image is a non-human animal image, and the image feature information is a code representing the features of the animal image images; the input of the animal image generation model is a random feature code or the image feature information output by the animal image generation model; A mixed image feature information obtaining module, used for fusing the at least two sets of image feature information to obtain mixed image feature information; An attribute encoding module, used for inputting preset attribute information into a preset encoder to obtain an attribute code; A target animal image acquisition module is used to input the mixed image feature information and the attribute code into the animal image generation model to obtain a target animal image and target image feature information, wherein the mixed image feature information represents the features of the animal image, and the attribute code represents the animal attribute information, and the attribute information includes: age, hair color, image angle and breed, and the target image feature information is the image feature information corresponding to the target animal image; The encoder is trained in the following way: Input the real attribute information into the initial encoder to obtain the initial attribute code; Inputting the initial attribute code and the preset animal image feature information into an animal image generation model to obtain a training animal image and training image feature information; Determining encoding attribute information according to the training animal image; The initial encoder is trained according to the loss function of the real attribute information and the encoded attribute information to obtain a trained encoder.
7. An electronic device, characterized in that: The electronic device comprises: one or more processing devices; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the method for generating an animal image as described in any one of claims 1-5.
8. A computer readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processing device, the method for generating an animal image as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Image generation model training method and device, electronic equipment and storage medium
CN113096055A
Face image synthesis method and device
CN113327191A