Controllable pedestrian generation method and device based on generative adversarial neural network, and medium

Through the adversarial joint training method of generating an adversarial neural network, the problem of semantic segmentation labeling dependence in the existing technology is solved, high-quality controllable pedestrian images and mask generation is achieved, and the performance of pedestrian detection tasks is improved.

CN120047454AActive Publication Date: 2025-05-27SOUTH CHINA UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510105687.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing controllable image generation neural network is limited by semantic segmentation annotation during training, making it difficult to achieve high-quality controllable pedestrian generation at a lower cost.

Method used

Using a method based on the generation of adversarial neural network, high-quality generation of pedestrian images and masks are achieved through the adversarial joint training of the mapping network, the mask encoder, the three-branch generation network and the discriminator.

Benefits of technology

It realizes the generation of high-quality controllable pedestrian images and masks without relying on pixel-by-pixel semantic annotation, improving the performance of downstream semi-supervised pedestrian detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047454A_ABST
    Figure CN120047454A_ABST
Patent Text Reader

Abstract

The invention discloses a controllable pedestrian generation method and device based on a generative adversarial neural network and a medium. The synthesis effect is improved in a mode that a discriminator and a generator confront each other. The whole generator is of a coding and decoding structure, and a three-branch structure is introduced into a decoder: a foreground branch head synthesizes a pedestrian feature map, a background branch head synthesizes a street view feature map, and a mask branch head synthesizes a pedestrian segmentation mask and guides feature map fusion of foreground branches and background branches. And the discriminator comprises a pedestrian discriminator and a mask discriminator, and is used for adversarial training to improve the trueness of pedestrians and masks synthesized by the generator, and measuring whether the masks can accurately segment pedestrians in the synthesized image or not. In addition, the model constrains an input mask to be consistent with a synthesized mask through a loss function, so that synthesis of a pedestrian image is controlled according to a specified semantic mask. The method can be widely applied to the technical field of pedestrian generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian generation, and in particular to a controllable pedestrian generation method, device and medium based on a generative adversarial neural network. Background Art

[0002] In recent years, generative adversarial neural networks have shown excellent performance in generating diverse and high-fidelity images. To enhance the constraint on the structure of the generated images, many works have attempted to use the structure mask of the input model as control information to train the model to synthesize images that match the structure, so as to realize the use of semantic masks to segment the specified content in the images and serve downstream tasks.

[0003] When the amount of data is sufficient, the effect of controllable image generation is relatively ideal. However, in practical applications, the training of controllable image generation neural networks is subject to some limitations. For the street scene pedestrian dataset, images can be collected by on-vehicle cameras, and the common annotation is low-cost detection box annotation, but generally, a large amount of manpower, material resources and financial resources are not consumed for semantic segmentation annotation. Semantic segmentation annotation is a difficult task. It is easy to obtain a large amount of pedestrian data with rich scenes by collecting street scene images with on-vehicle cameras and annotating pedestrian detection boxes. However, most existing controllable image generation neural networks rely on datasets where the mask matches the image. Even if some of the pedestrians in the images are semantically segmented and annotated, it is difficult to provide a large amount of data for the training of the generation model, and it is impossible to achieve high-quality controllable pedestrian generation at a low cost. Summary of the Invention

[0004] To at least to some extent solve one of the technical problems existing in the prior art, an object of the present invention is to provide a controllable pedestrian generation method, device and medium based on a generative adversarial neural network.

[0005] The first technical solution adopted by the present invention is:

[0006] A controllable pedestrian generation method based on a generative adversarial neural network, comprising the following steps:

[0007] Obtain a pedestrian instance dataset and a pedestrian pose mask dataset, and denote the pedestrian instance dataset as Denote the pedestrian pose mask dataset as

[0008] Construct a mapping network M for mapping an input control vector z that follows a Gaussian distribution into an intermediate vector w;

[0009] Construct a mask encoder E m for extracting features e m from the input pedestrian pose mask x i ;

[0010] Construct a three-branch generation network G for generating pedestrian images and corresponding pedestrian segmentation masks according to the intermediate vector w and the features of the pedestrian pose mask.

[0011] Use the generated pedestrian images and pedestrian segmentation masks as fake data, and use the data in the pedestrian instance dataset X p and the pedestrian pose mask dataset X m as real data to train the image discriminator D img and the mask discriminator D mask ; where the image discriminator D img is used to distinguish the generated pedestrian images and the images x from the pedestrian instance dataset p , and the mask discriminator D mask is used to distinguish the generated pedestrian segmentation masks and the masks x from the pedestrian pose mask dataset m ;

[0012] Through the adversarial process between the image discriminator D img and the mask discriminator D mask with the mapping network M, mask encoder E m , and three-branch generation network G for synthesizing pedestrian images and masks, the learning of the neural network is constrained. When the adversarial learning reaches equilibrium, the generator can synthesize pedestrian images of different styles and their corresponding masks according to the input style vector and control mask.

[0013] Furthermore, the pedestrian instance dataset X p is obtained from the pedestrian detection dataset: according to the pedestrian detection box annotations in the pedestrian detection dataset, extract the regions at the corresponding positions in the dataset images to construct the pedestrian instance dataset X p ;

[0014] The data in the pedestrian pose mask dataset X m comes from multiple different pedestrian segmentation datasets.

[0015] Furthermore, use a multi-layer perceptron MLP as the network structure of the mapping network M;

[0016] The intermediate vector w output by the mapping network M is used as the control condition of the style after affine transformation and input into each synthesis block of the three-branch generation network G to affect the clothing of the generated pedestrian images.

[0017] ​​​​Further, a multi-layer convolutional neural network is used as the mask encoder E m 's network structure;

[0018] The input pedestrian pose mask x m extracts features through convolution will be introduced into the synthesis blocks on the backbone of the three-branch generation network G and the synthesis blocks of the mask branch to affect the generated pedestrian image and the generated pedestrian segmentation mask 's structure, that is, the pedestrian pose; where L represents the mask encoder E m 's number of convolutional layers.

[0019] Further, the three-branch generation network G consists of four parts: a backbone, a foreground branch, a background branch, and a mask branch;

[0020] Each branch is implemented by multiple synthesis blocks, and the synthesis blocks adopt the structure of the synthesis blocks in StyleGAN2; the synthesis blocks of the backbone synthesize a low-resolution common feature map f according to the style obtained by the affine transformation of the intermediate vector w i out , and introduce the extraction features e of the corresponding resolution of the pedestrian pose mask x under the regulation of the control factor α m ; after reaching a certain resolution, the three-branch generation network G differentiates into a foreground branch, a background branch, and a mask branch; the synthesis blocks in the mask branch continue to receive the style and the features e of the pedestrian pose mask x j to synthesize the pedestrian segmentation mask m ; then, the pedestrian segmentation mask i will control the fusion of the feature maps of the foreground branch synthesis block and the background branch synthesis block of the corresponding resolution; after reaching a satisfactory resolution, the fused feature map c will be converted into the generated pedestrian image fused

[0021]

[0022] Further, the expression of the fused feature map c fused is:

[0022]

[0023] where σ is a downsampling function, c f , c b respectively represent the output feature maps of the foreground synthesis block and the background synthesis block of the corresponding resolution; c fused and c b will be used as the inputs of the next foreground synthesis block and background synthesis block respectively.

[0024] Further, the image discriminator D imgand the mask discriminator D mask adopts the same structure as the discriminator of StyleGAN2;

[0025] The image discriminator D img scores the generated pedestrian images and expects the score to be close to 0, and scores the image x from the pedestrian instance dataset p and expects the score to be close to 1;

[0026] The mask discriminator D mask scores the generated pedestrian segmentation mask and expects the score to be close to 0, and scores the mask x from the pedestrian pose mask dataset m and expects the score to be close to 1.

[0027] Furthermore, the mapping network M, mask encoder E of the mask m , the three-branch generation network G and the image discriminator D img and the mask discriminator D mask carry out confrontation to improve the authenticity of the synthetic data, and the expression of the loss function is:

[0028]

[0029] where S represents the softplus function; represents the expected value; Z represents the set of control vectors subject to Gaussian distribution input to the mapping network M;

[0030] In addition, the reconstruction loss of the mask is introduced to strengthen the ability to synthesize pedestrians with corresponding poses according to the input mask x m

[0031] The second technical solution adopted by the present invention is:

[0032] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a controllable pedestrian generation method based on a generative adversarial neural network as described above.

[0033] The third technical solution adopted by the present invention is:

[0034] ​A computer-readable storage medium stores at least one instruction, at least one program, a code set or an instruction set. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a controllable pedestrian generation method based on a generative adversarial neural network as described above.

[0035] The fourth technical solution adopted by the present invention is:

[0036] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned controllable pedestrian generation method based on a generative adversarial neural network.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] (1) The present invention combines a mapping network M, a mask encoder E m , a three-branch generation network G, an image discriminator D img and a mask discriminator D mask Five neural networks. Through the adversarial joint training among the three, the controllable high-definition pedestrian image synthesis is finally realized, so that the generated pedestrian images can be segmented according to the mask and better serve the downstream semi-supervised pedestrian detection task.

[0039] (2) Most of the existing methods for realizing pedestrian synthesis control through the input mask rely on per-pixel semantic annotation of training data, and cannot well synthesize pedestrians with specified poses through unsegmented and unannotated images, lacking the promotion ability on pedestrian datasets only containing detection box annotations. The present invention enables the mask-controlled pedestrian synthesis to get rid of the dependence on per-pixel semantic annotation training data, and the requirement for the training set is reduced from a semantic segmentation dataset to an object detection dataset, having good promotion ability on different pedestrian datasets.

[0040] (3) The present invention controls the style information and structural information of the synthesized pedestrian images through the mapping network M and the mask encoder E m respectively, and realizes the decoupled control of the style and structure of the generated images. This is of great significance for improving the downstream semi-supervised pedestrian detection task, and can improve the ability of the detection model to detect pedestrians with different clothes in different scenarios. Description of the Drawings

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0042] Figure 1 It is the model structure diagram of a controllable pedestrian generation method based on a generative adversarial neural network in an embodiment of the present invention. Detailed implementation manners

[0043] The following details the embodiments of the present invention. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation to the present invention. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0044] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention.

[0045] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is two or more, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence of the indicated technical features.

[0046] In the description of the present invention, unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense. Those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.

[0047] In view of the existing technical problems, the present invention provides a controllable pedestrian generation method, device and medium based on a generative adversarial neural network. In the present invention, through appropriate structural constraints, the model can realize the synthesis of pedestrian images according to a specified semantic mask without performing pixel-by-pixel semantic annotation on the pedestrian image dataset, so that a large amount of street scene pedestrian data can be fully utilized to synthesize real and diverse pedestrian images, and the generated data can be applied to data augmentation of pedestrian detection models.

[0048] Specifically, the present invention improves the synthesis effect in a way that the discriminator and the generator compete with each other. The generator has an overall encoder-decoder structure, and a three-branch structure is introduced in the decoder: the foreground branch head synthesizes the pedestrian feature map, the background branch head synthesizes the street scene feature map, and the mask branch head synthesizes the segmentation mask of the pedestrian and uses it to guide the fusion of the feature maps of the foreground branch and the background branch. The discriminator includes a pedestrian discriminator and a mask discriminator, which are used for adversarial training to improve the authenticity of the pedestrians and masks synthesized by the generator, and to measure whether the mask can accurately segment the pedestrians in the synthesized image. In addition, the model constrains the input mask to be consistent with the synthesized mask through a loss function, so as to realize the synthesis of pedestrian images according to the specified semantic mask.

[0049] Embodiment 1

[0050] As Figure 1 shown, this embodiment provides a controllable pedestrian generation method based on a generative adversarial neural network. By using the pedestrian data in the pedestrian detection dataset and the data in an independent pedestrian pose mask dataset, the model is trained to synthesize real and diverse pedestrian images that meet the corresponding pose constraints according to the specified pedestrian pose mask through the competition between the generator and the discriminator, so as to serve downstream tasks. The method specifically includes the following steps:

[0051] S1. Prepare a pedestrian instance dataset and a pedestrian pose mask dataset, and there is no matching relationship between the two datasets. Denote the pedestrian instance dataset as and the pedestrian pose mask dataset as

[0052] In some embodiments, the pedestrian instance dataset X p can be obtained from a pedestrian detection dataset. Specifically, according to the pedestrian detection box annotations in the pedestrian detection dataset, the regions at the corresponding positions in the dataset images are extracted to construct the pedestrian instance dataset X p . The data in the pedestrian pose mask dataset X m can come from multiple different pedestrian segmentation datasets.

[0053] Exemplarily, the pedestrian instance data is from the CityPersons dataset, and there are 2,975 RGB street view images with pedestrian detection box annotations in its training set. The pedestrian pose mask data comes from multiple datasets such as CityScapes and Penn-FUDAN. According to the position of the pedestrian, the pedestrian image and the pedestrian pose mask can be extracted from the dataset, and their resolutions are uniformly changed to 128×128 to obtain the pedestrian instance dataset and the pedestrian pose mask dataset In this embodiment, the number of pedestrian instances N p = 12,329, and the number of pedestrian pose masks N m = 2,645.

[0054] S2. Train a mapping network M: z→w implemented by a neural network to map the input control vector z~N(0, I) that follows a Gaussian distribution to an intermediate vector w.

[0055] As an alternative implementation, an MLP is selected as the network structure of the mapping network M. The intermediate vector w output by the mapping network M, after an affine transformation, is used as the control condition of the "style" and input into each synthesis block of the three-branch generation network G, which can affect the clothing of the generated pedestrian image of the clothing.

[0056] S3. Train a mask encoder implemented by a neural network to extract the feature e m ∈X m from the input pedestrian pose mask x i of the extraction.

[0057] As an alternative implementation, a multi-layer convolutional neural network is selected as the network structure of the mask encoder E m The input pedestrian pose mask x m extracts the feature through convolution will be introduced into the synthesis blocks on the backbone of the three-branch generation network G and the synthesis blocks of the mask branch, which can affect the generated pedestrian image and the generated pedestrian segmentation mask of the structure, that is, the pedestrian pose. In this embodiment, the number of convolutional layers of the mask encoder E m i.e., the number of output features L = 5, and e 1 , e 2 , e 3 , e 4 , e 5 represent feature maps with resolutions of 64×64, 32×32, 16×16, 8×8, and 4×4 respectively.

[0058] S4. Train a three-branch generation network implemented by a neural network The intermediate vector w obtained according to step S2 and the features of the pedestrian pose mask obtained in step S3 are used to generate a pedestrian image and the corresponding pedestrian segmentation mask

[0059] As a specific implementation, the three-branch generation network G consists of four parts: a backbone, a foreground branch, a background branch, and a mask branch. Each branch is implemented by multiple synthesis blocks, and the synthesis blocks adopt the structure of the synthesis blocks in StyleGAN2. The synthesis blocks of the backbone synthesize the common feature maps f with resolutions of 4×4, 8×8, 16×16, and 32×32 at a low resolution according to the "style" obtained by the affine transformation of the intermediate vector w i out and introduce the extraction features e with the corresponding resolutions of the pedestrian pose mask x m under the regulation of the control factor α. The formula is j :

[0060]

[0061] where represents the input feature map of the next synthesis block. In this embodiment, the control factor α = 0.9

[0062] After reaching the 32×32 resolution, the three-branch generation network G branches out into a foreground branch, a background branch, and a mask branch. The synthesis blocks in the mask branch continue to receive the "style" and the features e m of the pedestrian pose mask x j synthesize the feature maps with resolutions of 64×64 and 128×128 and convert the feature map with a resolution of 128×128 into a pedestrian segmentation mask with a resolution of 128×128 Then, the pedestrian segmentation mask controls the feature map fusion of the foreground branch synthesis block and the background branch synthesis block with the corresponding resolution at the 64×64 resolution. The formula is

[0063]

[0064] where σ is a downsampling function respectively represent the feature maps with a resolution of 64×64 output by the foreground synthesis block and the background synthesis block with the corresponding resolution represents the fused feature map, and and will be used as the inputs of the next foreground synthesis block and background synthesis block respectively. Furthermore, the pedestrian segmentation mask controls the feature map fusion of the foreground branch synthesis block and the background branch synthesis block with the corresponding resolution at the 128×128 resolution. The formula is

[0065]

[0066] Fused feature map will be transformed into an RGB pedestrian image with a resolution of 128×128

[0067] S5. Use the pedestrian image and the pedestrian segmentation mask generated in step S4 as fake data, and the data in the pedestrian instance dataset X p and the pedestrian pose mask dataset X m as real data to train the image discriminator D img implemented by a neural network: x→[0,1] and the mask discriminator D mask : x→[0,1], where the image discriminator D img is used to distinguish the generated pedestrian image and the image x p ∈X p from the pedestrian instance dataset, and the mask discriminator D mask is used to distinguish the generated pedestrian segmentation mask and the mask x m ∈X m .

[0068] As an alternative implementation, the image discriminator D img and the mask discriminator D mask adopt the same structure as the discriminator of StyleGAN2. The image discriminator D img scores the generated pedestrian image and expects the score to be close to 0, and scores the image x p ∈X p from the pedestrian instance dataset and expects the score to be close to 1; the mask discriminator D mask scores the generated pedestrian segmentation mask and expects the score to be close to 0, and scores the mask x m ∈X m from the pedestrian pose mask dataset and expects the score to be close to 1.

[0069] S6. Through the confrontation between the image discriminator D img and the mask discriminator D mask and the mapping network M, mask encoder E m used to synthesize pedestrian images and masks, and the three-branch generator G. To constrain the learning of the neural network, when the adversarial learning of the three reaches equilibrium, the generator can generate according to the input style vector z and the control mask x mSynthesize a large number of controllable pedestrian images with different styles and their corresponding masks

[0070] Specifically, the mapping network M and encoder E of the mask m , the three-branch generation network G, image discriminator D img and mask discriminator D mask carry out adversarial training to improve the authenticity of the synthesized data. The loss function formulas are as follows:

[0071]

[0072] where S represents the softplus function S(t) = log(1 + exp(t)). In addition, the reconstruction loss of the mask is introduced to strengthen the ability to synthesize pedestrians with corresponding poses according to the input mask x m .

[0073] After training, the performance of this method is quantitatively evaluated on the pedestrian instance dataset X constructed from the CityPersons dataset p . The evaluation metrics include the Fréchet Inception Distance (FID) and Inception Score (IS). FID represents the similarity of the generated images and real images in terms of feature distribution. The lower the value, the more real the generated images are. IS represents the overall distribution of the generated images. The higher the value, the more real and diverse the generated images are. After evaluation, the effects of the present invention on the two evaluation criteria are significantly higher than those of the baseline method and are worthy of promotion.

[0074] In summary, this embodiment provides a controllable pedestrian generation method based on a generative adversarial neural network. Without performing pixel-by-pixel semantic annotation on the pedestrian image dataset, this method realizes the synthesis of pedestrian images according to a specified semantic mask. Considering that there is a large amount of semantic information in the intermediate feature maps of the generative model, we introduce a mask encoder and a mapping network to edit the structural semantics and style semantics in the intermediate feature maps respectively. Specifically, the present invention adopts an encoder-decoder three-branch structure in the generative model: the mask encoder encodes the mask image that controls the structural semantics, the mapping network encodes the vector that controls the style semantics, and the backbone structure of the decoder synthesizes a low-resolution basic feature map. The three branch heads respectively synthesize a high-resolution foreground image, a background image, and a mask image. The mapping network passes the style semantics into the feature maps synthesized by the decoder, and the mask encoder passes the structural semantics into the backbone structure of the decoder and the mask branch head. The mask branch head affects the feature map fusion of the foreground branch head and the background branch head, thereby realizing the control of the structure and style of the generated image. Without the need for semantic segmentation annotation of the pedestrian dataset, the present invention realizes the synthesis of a large number of pedestrian images with different styles but conforming to the corresponding poses by specifying any mask. The synthesized pedestrians can be segmented by the specified mask and can be applied to downstream tasks such as the training of semi-supervised pedestrian detection models in traffic scenarios.

[0075] Embodiment 2

[0076] The embodiment of the present invention also provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement a Figure 1 controllable pedestrian generation method based on a generative adversarial neural network as shown.

[0077] It can be understood that the memory may include a random access memory (RAM) and may also include a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function, instructions for implementing the above method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.

[0078] The processor may include one or more processing cores. The processor connects various parts within the entire server using various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by invoking data stored in the memory, the processor performs various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate a combination of one or several of a central processing unit (CPU) and a modem, etc. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor and may be implemented separately by a single chip.

[0079] Since this electronic device is an electronic device corresponding to a controllable pedestrian generation method based on a generative adversarial neural network in an embodiment of the present invention, and the principle by which this electronic device solves problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and repeated parts will not be elaborated.

[0080] Embodiment 3

[0081] An embodiment of the present invention further provides a computer-readable storage medium. At least one instruction, at least one segment of program, code set, or instruction set is stored in the storage medium. The at least one instruction, the at least one segment of program, the code set, or the instruction set is loaded and executed by a processor to implement Figure 1 a controllable pedestrian generation method based on a generative adversarial neural network as shown.

[0082] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other computer-readable medium capable of carrying or storing data.

[0083] Since this storage medium is the storage medium corresponding to a controllable human motion generation method based on a generative adversarial neural network in an embodiment of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be elaborated.

[0084] Embodiment 4

[0085] In some possible implementation manners, various aspects of the method in an embodiment of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a controllable human motion generation method according to various exemplary implementation manners of the present application described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or written in various other programming languages.

[0086] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0087] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0088] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and their purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.

Claims

1. A controllable pedestrian generation method based on generative adversarial neural network, characterized in that: The following steps are involved: Obtain pedestrian instance dataset and pedestrian posture mask dataset, and record the pedestrian instance dataset as The pedestrian pose mask dataset is denoted as Construct a mapping network M to map the input control vector z that follows Gaussian distribution to an intermediate vector w; Constructing the mask encoder E m , used for the input pedestrian posture mask x m Perform feature i Extraction of Construct a three-branch generative network G to generate the features of the intermediate vector w and the pedestrian posture mask To generate pedestrian images And the corresponding pedestrian segmentation mask ; The generated pedestrian image and pedestrian segmentation mask As false data, pedestrian instance dataset X p and pedestrian pose mask dataset X m The data in is used as the real data to train the image discriminator D img and mask discriminator D mask ; The image discriminator D img Used to identify the generated pedestrian images and an image x from a pedestrian instance dataset p , mask discriminator D mask Used to identify the generated pedestrian segmentation mask and the mask x from the Pedestrian Pose Mask Dataset m ; Through the image discriminator D img and mask discriminator D mask And the mapping network M and mask encoder E for synthesizing pedestrian images and masks m , the three-branch generative network G, the adversarial network is used to constrain the learning of the neural network. When the adversarial learning reaches a balance, the generator can synthesize pedestrian images of different styles according to the input style vector and the control mask. and its corresponding mask 2. The controllable pedestrian generation method based on generative adversarial neural network according to claim 1 is characterized in that: The pedestrian instance dataset X p Obtained from the pedestrian detection dataset: According to the pedestrian detection box annotation in the pedestrian detection dataset, the area at the corresponding position in the dataset image is extracted to construct the pedestrian instance dataset X p ; The pedestrian pose mask dataset X m The data in comes from multiple different pedestrian segmentation datasets.

3. The controllable pedestrian generation method based on generative adversarial neural network according to claim 1 is characterized in that: A multi-layer perceptron MLP is used as the network structure of the mapping network M; The intermediate vector w output by the mapping network M is used as the style control condition after affine transformation and input into each synthesis block of the three-branch generation network G to affect the generated pedestrian image. clothing.

4. The controllable pedestrian generation method based on generative adversarial neural network according to claim 1 is characterized in that: A multi-layer convolutional neural network is used as the mask encoder E m The network structure; Input pedestrian pose mask x m Extract features through convolution It will be introduced into the synthesis block on the trunk of the three-branch generation network G and the synthesis block of the mask branch to affect the generated pedestrian image And the generated pedestrian segmentation mask The structure of pedestrian pose, where L represents the mask encoder E m The number of convolutional layers.

5. The controllable pedestrian generation method based on generative adversarial neural network according to claim 1 is characterized in that: The three-branch generation network G consists of four parts: a trunk, a foreground branch, a background branch, and a mask branch; Each branch is implemented by multiple synthesis blocks, and the synthesis block adopts the structure of the synthesis block in StyleGAN2; the synthesis block of the trunk synthesizes the low-resolution common feature map f according to the style obtained by affine transformation of the intermediate vector w i out , and introduce the pedestrian posture mask x under the control of the control factor α m The extracted features of the corresponding resolution e j ; After reaching a certain resolution, the three-branch generative network G differentiates into a foreground branch, a background branch, and a mask branch; the synthetic block in the mask branch continues to receive the style and pedestrian posture mask x m Features i , synthesize pedestrian segmentation mask Next, the pedestrian segmentation mask It will control the fusion of the feature maps of the foreground branch synthesis block and the background branch synthesis block of the corresponding resolution; after reaching a satisfactory resolution, the fusion feature map c fused Will be converted into the generated pedestrian image 6. The controllable pedestrian generation method based on generative adversarial neural network according to claim 5 is characterized in that: The fused feature map c fused The expression is: Where σ is the downsampling function, c f 、c b Respectively represent the output feature maps of the foreground synthesis block and the background synthesis block of the corresponding resolution; c fused and c b They will be used as input for the next foreground synthesis block and background synthesis block respectively.

7. The controllable pedestrian generation method based on generative adversarial neural network according to claim 1 is characterized in that: The image discriminator D img and mask discriminator D mask Uses the same structure as the discriminator of StyleGAN2; The image discriminator D img For the generated pedestrian image Scoring and expecting the score to be close to 0, for the image x from the pedestrian instance dataset p Score and expect the score to be close to 1; The mask discriminator D mask Generate pedestrian segmentation mask Score and expect the score to be close to 0, for the mask x from the pedestrian pose mask dataset m Score it and expect the score to be close to 1.

8. The controllable pedestrian generation method based on generative adversarial neural network according to claim 1 is characterized in that: The mask mapping network M, the mask encoder E m , three-branch generation network G and image discriminator D img and mask discriminator D mask To improve the authenticity of synthetic data, the loss function is expressed as: Where S represents the softplus function; represents the expected value; Z represents the set of control vectors of the input mapping network M that obey the Gaussian distribution; In addition, the mask reconstruction loss is introduced To strengthen the mask x according to the input m The ability to synthesize pedestrians in corresponding poses.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Pedestrian detection data expansion method based on generative adversarial network

    CN111950346A

  • Image style migration method and device based on generative adversarial network and related equipment

    CN113538224A

  • Face attribute editing method based on mask denoising and feature selection

    CN115546461A

  • Methods, systems, and computer readable media for mask embedding for realistic high-resolution image synthesis

    US11580673B1