Unmanned aerial vehicle performance pattern generation method and device, computer device, and storage medium

By constructing a training dataset and utilizing a text-based graph model and a control model, drone performance patterns are automatically generated, solving the problems of low generation efficiency and poor style matching in existing technologies, and achieving efficient and accurate pattern generation.

CN120747288BActive Publication Date: 2025-11-07SHENZHEN DAMO DAZHI CONTROL TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511255210.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-11-07
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

Existing technologies for generating drone performance patterns suffer from problems such as heavy manual workload, difficulty in matching style requirements, and mismatch between line density and performance frequency. Furthermore, existing models have low generation efficiency and cannot quickly generate patterns that meet the requirements.

Method used

Historical drone performance patterns are collected, text-labeled, and a training dataset is constructed. The model is then trained using a text-to-image model and a control model to generate a pattern generation model. Based on the set pattern generation requirements, drone performance patterns are automatically generated.

Benefits of technology

Automated generation significantly shortens the generation cycle of drone performance patterns, improves generation efficiency, and generates patterns that better match style requirements. The density of lines is precisely matched with the number of performance flights, reducing the workload of subsequent modifications and improving generation quality and applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747288B_ABST
    Figure CN120747288B_ABST
Patent Text Reader

Abstract

The application discloses a UAV performance pattern generation method and device, computer equipment and a storage medium, and the method comprises the following steps: collecting historical UAV performance patterns, and performing text marking on the historical UAV performance patterns to construct a training data set; training a text-to-image model and a corresponding control model by using the training data set to construct a pattern generation model; and generating corresponding UAV performance patterns by using the pattern generation model based on a set pattern generation requirement. The application collects historical UAV performance patterns and performs text marking to construct a training data set, trains a text-to-image model and a corresponding control model by using the data set, constructs a pattern generation model, and finally generates patterns by using the model pattern generation model based on a set pattern generation requirement. In this way, the process of manual direct design or repeated adjustment can be reduced, the generation cycle of the UAV performance pattern is significantly shortened, and the generation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer software, in particular to a UAV performance pattern generation method and device, a computer device and a storage medium. BACKGROUND

[0002] In the process of continuous development of UAV light show technology, the performance times of UAVs are increasing, and now it has reached the scale of tens of thousands of times. The performance pattern plays an increasingly important role in the whole UAV light show. At present, there are mainly two ways to obtain the performance pattern of the UAV light show: one is to directly design by the designer, and the other is to use the design software to outline the performance pattern according to the specified image of the customer, and then determine the position of the UAV through discrete points to form a complete UAV performance picture.

[0003] However, with the increase of complexity and quantity of performance pictures, the workload of manually obtaining performance patterns is increasing. At the same time, the existing model has some shortcomings when generating UAV performance patterns: some models are difficult to meet the style requirements of UAV performance patterns, and some models cannot generate patterns with line density and performance times that match. Even if the text prompt words are carefully designed, it usually needs to be generated several times to get a pattern close to the requirements in terms of style and line density, and often needs to be modified a lot. In addition, the related prior art, such as CN118967855B, mainly solves the problems of position alignment and automatic discrete point distribution of existing performance patterns, and involves related devices, equipment and media, but does not mention the generation method of performance patterns; CN112596536A solves the problem of discrete point distribution of existing performance patterns according to the performance times of UAVs to form a performance picture, and like CN118967855B, it also does not involve the generation method of performance patterns. SUMMARY

[0004] The embodiments of the present application provide a UAV performance pattern generation method, device, computer device and storage medium, which aims to improve the generation efficiency and effect of UAV performance patterns.

[0005] In a first aspect, the embodiments of the present application provide a UAV performance pattern generation method, comprising:

[0006] Collecting historical UAV performance patterns and text tagging the historical UAV performance patterns to construct a training data set;

[0007] Training a text generation model and its corresponding control model using the training data set to construct a pattern generation model;

[0008] Generating a corresponding UAV performance pattern based on the set pattern generation requirements using the pattern generation model.

[0009] In a second aspect, an embodiment of the present application provides a device for generating a performance pattern of a UAV, comprising:

[0010] a data collection unit configured to collect historical performance patterns of a UAV and to mark the historical performance patterns of the UAV with texts to construct a training data set;

[0011] a model construction unit configured to train a text-to-image model and a corresponding control model by using the training data set to construct a pattern generation model;

[0012] a pattern generation unit configured to generate a corresponding performance pattern of a UAV by using the pattern generation model based on a set pattern generation requirement.

[0013] In a third aspect, an embodiment of the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for generating a performance pattern of a UAV as described in the first aspect when executing the computer program.

[0014] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executable on a processor to implement the method for generating a performance pattern of a UAV as described in the first aspect.

[0015] The embodiments of the present application provide a method, device, computer device, and storage medium for generating a performance pattern of a UAV, which comprises the following steps: collecting historical performance patterns of a UAV and marking the historical performance patterns of the UAV with texts to construct a training data set; training a text-to-image model and a corresponding control model by using the training data set to construct a pattern generation model; and generating a corresponding performance pattern of a UAV by using the pattern generation model based on a set pattern generation requirement. The embodiments of the present application collect historical performance patterns of a UAV and mark the historical performance patterns of the UAV with texts to construct a training data set, then train a text-to-image model and a corresponding control model by using the data set to construct a pattern generation model, and finally generate a corresponding performance pattern of a UAV by using the model based on a set pattern generation requirement, so as to reduce the process of manual direct design or repeated adjustment, to reduce the intensity of manual intervention by automatic generation through the model, to significantly shorten the generation cycle of the performance pattern of the UAV, and to improve the generation efficiency. Meanwhile, since the training data set is constructed based on historical performance patterns of a UAV and is trained by a control model, the generated pattern can better meet the style requirement of the performance of the UAV, and the line density and performance times can be more accurately matched, so as to reduce the workload of subsequent modification and to improve the generation quality and applicability of the performance pattern of the UAV. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0017] Figure 1 A flowchart of a UAV performance pattern generation method provided by an embodiment of the present application is shown in the figure.

[0018] Figure 2 A schematic block diagram of a UAV performance pattern generation device provided by an embodiment of the present application is shown in the figure.

[0019] Figure 3 A data set diagram in a UAV performance pattern generation method provided by an embodiment of the present application is shown in the figure.

[0020] Figure 4 Another data set diagram in a UAV performance pattern generation method provided by an embodiment of the present application is shown in the figure.

[0021] Figure 5 A network architecture diagram of a privacy publishing model in a UAV performance pattern generation method provided by an embodiment of the present application is shown in the figure.

[0022] Figure 6 A network architecture diagram of a control model in a UAV performance pattern generation method provided by an embodiment of the present application is shown in the figure.

[0023] Figure 7 A first generation effect diagram of a UAV performance pattern generation method provided by an embodiment of the present application is shown in the figure.

[0024] Figure 8 A second generation effect diagram of a UAV performance pattern generation method provided by an embodiment of the present application is shown in the figure.

[0025] Figure 9 A third generation effect diagram of a UAV performance pattern generation method provided by an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0026] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0027] It should be understood that the terms "comprises" and "comprising," when used in this specification and the following claims, indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0028] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used in this specification and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0029] It should also be further understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as to the lack of combinations when interpreted in the alternative.

[0030] Please see below Figure 1 The embodiment of the present application provides a kind of unmanned plane performance pattern generation method, specifically includes: step S101-S103.

[0031] Step S101, historical unmanned plane performance pattern is collected, and the historical unmanned plane performance pattern is text marked, to build training data set;

[0032] Step S102, training data set is used to train text generation graph model and its corresponding control model, to build pattern generation model;

[0033] Step S103, based on the pattern generation requirement set, the corresponding unmanned plane performance pattern is generated using the pattern generation model.

[0034] In the embodiment, first, historical unmanned plane performance pattern is collected and text marked to build training data set, then the data set is used to train text generation graph model and its control model, to build pattern generation model, finally, according to the pattern generation requirement set, the corresponding unmanned plane performance pattern is generated by the model. In this way, the process of manual direct design or repeated adjustment can be reduced, the intensity of manual intervention is reduced by automatic generation of the model, so that the generation cycle of unmanned plane performance pattern is significantly shortened, and the generation efficiency is improved. At the same time, since the training data set is built based on historical unmanned plane performance pattern, and is trained by the control model, the generated pattern can better meet the style requirements of unmanned plane performance, and the line density and performance times can be more accurately matched, so that the workload of subsequent modification is reduced, and the generation quality and applicability of unmanned plane performance pattern are improved.

[0035] In an embodiment, the step S101 includes:

[0036] The historical UAV performance pattern is divided into different levels based on the number of UAV performance frames; wherein the levels include minimal, simple, moderate, and intricate;

[0037] A text label is generated for the historical UAV performance pattern by a picture understanding model, and a first training data set is constructed in combination with the levels;

[0038] An original image corresponding to the historical UAV performance pattern is obtained, and a deep learning model is used to convert the original image into a depth map;

[0039] A second training data set is constructed in combination with the depth map, levels, and text label.

[0040] Since the UAV performance pattern is mainly a sketch and a line drawing, and is quite different from a sketch style, the UAV performance pattern is mainly clean black lines and a white background, which facilitates subsequent discrete point distribution of the lines to form a UAV performance picture, and the line density of the performance pattern is closely related to the number of performance frames. Therefore, when performing text labeling of the performance pattern, different numbers of performance frames should be distinguished in the text description to correspond to different line densities. Therefore, according to the number of performance frames, the data is divided into four different levels of minimal, simple, moderate, and intricate, and a picture understanding model and manual correction are used to generate a text label describing each performance pattern.

[0041] Based on the divided levels, a first training data set for training of a text-to-image model is first generated, which is mainly used to train the model, input text, and generate a UAV performance pattern. In actual application, the description words of the above four levels are placed before the text label to synthesize the text label of each performance pattern. The entire data set has 60,000 data of different aspect ratios, each of which is composed of a performance pattern and a text label. As the complexity of the performance picture increases, the number of UAV performance frames also increases, so the text label can be generated by a picture understanding model Qwen2.5-VL-7B-Instruct, and combined with manual correction and other methods to ensure that the text labeling is accurate and reliable. Among them, the four words minimal, simple, moderate, and intricate are used as prompt words to generate different line density performance patterns.

[0042] In combination with Figure 3 , the data in the first training data set includes performance frames, performance patterns, and text labels. Here, since the text label contains more content, a character is used instead in the figure (the same below), and since an English encoder is used, English description data is used (the same below). InFigure 3In the center, text label A represents "minimal, linedrawing style, white background, a blooming flower, with the petals occupyingthe main part, consists of five petals, each of which presents a soft curveand natural shape, with slightly curled edges and slight overlap between thepetals. In the center of the flower, there is a cluster of slender stamens,dotted with small dots at the top." which translates to "extremely simple, line drawing style, white background, a blooming flower, with the petals occupying the main part, consists of five petals, each of which presents a soft curve and natural shape, with slightly curled edges and slight overlap between the petals. In the center of the flower, there is a cluster of slender stamens, dotted with small dots at the top." and "minimal, linedrawing style, white background, a blooming flower, with the petals occupyingthe main part, consists of five petals, each of which presents a soft curveand natural shape, with slightly curled edges and slight overlap between thepetals. In the center of the flower, there is a cluster of slender stamens,dotted with small dots at the top." which translates to "extremely simple, line drawing style, white background, a blooming flower, with the petals occupying the main part, consists of five petals, each of which presents a soft curve and natural shape, with slightly curled edges and slight overlap between the petals. In the center of the flower, there is a cluster of slender stamens, dotted with small dots at the top." and "minimal, linedrawing style, white background, a blooming flower, with the petals occupyingthe main part, consists of five petals, each of which presents a soft curveand natural shape, with slightly curled edges and slight overlap between thepetals. In the center of the flower, there is a cluster of slender stamens,dotted with small dots at the top." which translates to "extremely simple, line drawing style, white background, a blooming flower, with the petals occupying the main part, consists of five petals, each of which presents a soft curve and natural shape, with slightly curled edges and slight overlap between the petals. In the center of the flower, there is a cluster of slender stamens, dotted with small dots at the top." and "minimal, linedrawing style, white background, a blooming flower, with the petals occupyingthe main part, consists of five petals, each of which presents a soft curveand natural shape, with slightly curled edges and slight overlap between thepetals. In the center of the flower, there is a cluster of slender stamens,dotted with small dots at the top." which translates to "extremely simple, line drawing style, white background, a blooming flower, with the petals occupying the main part, consists of five petals, each of which presents a soft curve and natural shape, with slightly curled edges and slight overlap between the petals. In the center of the flower, there is a cluster of slender stamens, dotted with small dots at the top." and "minimal, linedrawing style, white background, a blooming flower, with the petals occupyingthe main part, consists of five petals, each of which presents a soft curveand natural shape, with slightly curled edges and slight overlap between thepetals. In the center of the flower, there is a cluster of slender stamens,dotted with small dots at the top." which translates to "extremely simple, line drawing style, white background, a blooming flower,Different sizes of lotus leaves are distributed in a staggered manner throughout the picture, and each lotus leaf has radial veins and slightly curled edges. One large lotus leaf is nearly circular, and the other also shows delicate veins. A blooming lotus flower and a bud are connected by a slender stem. Text label C represents "moderate, line drawing style, white background, three blooming flowers and green leaves, with layers of petals and delicately carved stamens, show complexity and beauty. The leaves are of different shapes and have serrated edges, adding a sense of layering. The stems are interwoven to form a base." That is, "moderate size, line drawing style, white background, three blooming flowers and green leaves, with layers of petals and delicately carved stamens, show complexity and beauty. The leaves are of different shapes and have serrated edges, adding a sense of layering. The stems are interwoven to form a base." Text label D represents "intricate, line drawing style, white background, a steaming bowl of noodles, filled with various toppings, is lifted from the bowl with chopsticks, revealing a soft and springy texture. The bowl is filled with a variety of ingredients, including chunks of vegetables and slices of meat, and the surface is dotted with a few drops of broth." That is, "complex, line drawing, white background, a steaming bowl of noodles is lifted from the bowl with chopsticks, revealing a soft and springy texture. The bowl is filled with a variety of ingredients, including chunks of vegetables and slices of meat, and the surface is dotted with a few drops of broth."

[0043] Then a second training data set for controlling model training is generated, which is mainly used to train the control model, input control map and text, and generate UAV performance patterns with the same image structure as the existing image. Here, since the existing control model will convert the original image into a line drawing, a depth map, a normal map, a pose map, a semantic segmentation map, etc., and then use the converted image to control the generation of images, but the line drawing will introduce background noise and affect the generated image to be too complex, and the normal map, pose map and semantic segmentation map will cause too much loss of original image information, making it difficult to generate a performance pattern with the same composition as the original image. Only the depth map can provide accurate control information without losing much information and can adapt to all original image scenes, which is very suitable for the generation of performance patterns. Therefore, the embodiment uses a depth map as a control map, and in actual application, a training data set of 60,000 different aspect ratios is made, each data consisting of a depth map, a performance pattern and a text label. Through the depth model, the original image can be converted into a depth map, and through the picture understanding model, the text can be generated, and the text label is formed by combining the line density words.

[0044] In combination Figure 4 , the data in the second training data set includes performance times, a depth map, a performance pattern and a text label, Figure 4In this case, text label E represents "minimal, line drawing style, white background, a truck. The truck is composed of simple geometric shapes. The front section features a rectangular cab with the outlines of windows and doors. The body is a rectangular cargo box with a horizontal and vertical line on the side representing the doors. The wheels are represented by three circles, each with a small dot inside to represent the wheel hub." and text label F represents "simple, line drawing style, white background, three chrysanthemums of different shapes and a few leaves, each flower outlined with simple and smooth lines, the petals clearly layered, showing the states of bloom and bud. The flowers are connected by a slender stem, dotted with several leaves of various shapes, with naturally curved edges."The text tag G indicates "moderate, line drawing style, white background, a cluster of figs and their leaves, containing three ripe, oval figs with distinct longitudinal stripes, appearing plump and textured. Each fig is attached to a slender branch dotted with several broad leaves with wavy edges and clearly visible veins, showcasing a natural texture." The text tag H indicates "intricate, line drawing style, white background, a scene depicts an ancient sailing vessel sailing on thesea. Its streamlined hull, with distinct wavy lines on its underside, symbolizes its waterborne navigation. The vessel boasts three decks, each with neatly arranged windows and railings. The sails are hoisted high, thesails inflated, and a flag flutters from the mast." In other words, it describes a scene of an ancient sailing vessel sailing on thesea against a white background. The streamlined hull is decorated with striking wave patterns, symbolizing its waterborne navigation. The vessel has three decks, each with neatly arranged windows and railings. The sails are hoisted high, and a flag flutters from the mast.

[0045] In one embodiment, step S102 includes:

[0046] training the text-to-graph model using the first training dataset and training the control model using the second training dataset;

[0047] building the pattern generation model based on the trained text-to-graph model and control model.

[0048] The first training dataset and the second training dataset are used to train the text-to-graph model and the control model respectively, so that the text-to-graph model can generate corresponding drone performance patterns according to the input text label, and the control model can generate drone performance patterns with the same image structure as the existing image according to the input control graph and text. This not only improves the generation efficiency and accuracy of drone performance patterns, but also makes the generated patterns more in line with the expectations and needs of users. Through the trained text-to-graph model and control model, users only need to input the corresponding text label or control graph and text label to quickly generate the required drone performance pattern.

[0049] In an embodiment, the training the text-to-graph model using the first training dataset comprises:

[0050] inputting the first training data in the first training dataset into a stable diffusion model;

[0051] encoding the first training data into a latent space variable using a variational autoencoder in the stable diffusion model;

[0052] iteratively diffusing the latent space variable based on the latent space until the iterative diffusion requirement is met, to obtain an iteratively denoised latent space variable;

[0053] decoding the iteratively denoised latent space variable into a corresponding image using a decoder of the variational autoencoder.

[0054] Further, the training the text-to-graph model using the first training dataset further comprises:

[0055] constructing a first objective function according to the following formula, and updating parameters of the text-to-graph model based on the first objective function:

[0056] ;

[0057] wherein L LDM represents an objective function, represents an encoder of VAE, x represents the first training data, and y represents the text label, is noise added to the image, represents an encoding model of the text label, represents a U-net model structure, represents the latent variable added with t time step noise, t represents the time step.

[0058] This embodiment uses stable diffusion as a text-to-image model. Stable diffusion is a latent space text-to-image diffusion model. It encodes images into latent space variables through a variational autoencoder (VAE), and then performs an iterative diffusion process in the latent space. During inference, the latent space variable obtained by ending the iterative denoising process is decoded into an image through the VAE. At the same time, the image generation is constrained by text to generate specific picture content. The whole process is as shown in Figure 5 Figure 5 In the figure, the left red area represents the pixel space, the middle green area represents the latent space, and the right represents the constraint condition. x is the image, represents the encoder of the VAE, and z is the latent variable of the latent space, is the latent variable added with T time step noise, is the noise added to the image, and t represents the time step, represents the U-Net model structure, and the subscript θ represents the parameters of the neural network, is the decoder of the VAE, is the decoded generated image, is the encoding model of the constraint condition, which is used to map the constraint condition into an intermediate representation, and then map it to the intermediate layer of the U-Net through a cross attention layer.

[0059] To support training of different aspect ratios, improve text compliance, and balance image diversity and generation speed, an improved stable diffusion model, i.e., stable diffusion XL model, can be used. In practical applications, to adapt to different aspect ratios of images encountered in real application scenarios, different aspect ratio images can be used for bucket training. The image sizes include the following: [768*1408, 768*1344, 768*1280, 896*1152, 960*1024, 1024*1024, 1024*960, 1152*896, 1280*768, 1344*768, 1408*768]. At the same time, to achieve fast convergence, the existing model parameters can be used to initialize the model structure to be trained.

[0060] The specific training process of the stable diffusion model is as follows:

[0061] ​(1) Determine the model structure to be trained and the training hyperparameters, including batch size, maximum number of rounds, optimizer, learning rate, learning rate schedule, and other key hyperparameters;

[0062] (2) Bucket the training data by aspect ratio, and use VAE to encode the images into latent space representation in advance;

[0063] (3) Load the existing weights as the initialization weights of the training model;

[0064] (4) Loop to get the data of each batch, i.e. the previously encoded latent space representation and the text;

[0065] (5) The text label is encoded by the text encoder and input into the U-Net structure together, and the U-Net outputs the predicted noise residual , the obtained predicted noise residual and the added noise are substituted into the above objective function to calculate the loss value;

[0066] (6) Backpropagation to calculate the gradient and update the model parameters;

[0067] (7) Loop the above process until convergence or reach the maximum number of rounds.

[0068] In an embodiment, the training of the control model using the second training data set comprises:

[0069] Setting the model structure and training hyperparameters of the control model; wherein the model structure of the control model comprises the encoder block and the intermediate block in the stable diffusion model, and a convolutional layer is added before each block of the control model;

[0070] Bucketing the second training data in the second training data set and image encoding the second training data by a variational autoencoder;

[0071] Loading the weights of the stable diffusion model to initialize the control model, and initializing the convolutional layer to 0;

[0072] Using the control model to perform loop iteration training on the image encoded second training data until the loop iteration requirement is met;

[0073] Using the decoder of the variational autoencoder to decode the loop iteration result into the corresponding image.

[0074] Further, the training of the control model using the second training data set further comprises:

[0075] Constructing a second objective function according to the following formula, and updating the parameters of the control model using the second objective function:

[0076] ;

[0077] wherein, denotes a second objective function, denotes a latent variable without adding a time step, is a text constraint, is a control image constraint.

[0078] In order to extract and perform a pattern with an existing image, the control model needs to control the composition of the existing image. The control model (controlnet) can be an adaptive model of the SD model, which generates image content by inputting text and control images. In combination with Figure 6 , the control model copies the encoder block and the middle block of the U-Net structure of the existing SD model, and adds a 1×1 convolution layer with an initial weight of 0 before each block. The output of the control model is added to the middle block and the decoder block of the U-Net structure through skip-connections, thereby playing a control role. Only the 1×1 convolution layer, the copied encoder block and the middle block need to be trained, and other parts do not participate in training. During training, the control image constraint obtained by encoding the control image through the control model, the text constraint obtained by encoding the text through the text encoder, combined with each time step t and the corresponding latent variable are input into the U-Net structure of the SD model to predict noise residuals, and then the noise is substituted into the above-mentioned objective function (i.e. the second objective function) of the control model, the loss is calculated, then the gradient is calculated through back propagation, and the parameters of the control model are updated according to the gradient until the model converges. During inference, the control model encodes the constraint image, and inputs it together with the text constraint into the U-Net structure of the SD model, thereby controlling the generation of the image. After step-by-step denoising iteration, the latent variable with the same distribution as the training image is obtained, and then the decoder of the VAE is decoded to the pixel space, and the image is obtained.

[0079] In order to adapt to different aspect ratios of images encountered in real application scenarios, images with different aspect ratios can be used for bucket training. The image sizes include the following: [768*1408, 768*1344, 768*1280, 896*1152, 960*1024, 1024*1024, 1024*960, 1152*896, 1280*768, 1344*768, 1408*768]. The specific training process of the control model is as follows:

[0080] (1) Determine the control model structure to be trained and the training hyperparameters, including batch size, maximum number of rounds, optimizer, learning rate, learning rate schedule, and other key hyperparameters;

[0081] (2) Bucket the training data by aspect ratio, and use VAE to encode the images into latent space representation in advance;

[0082] (3) Load the weights of the SDXL model as the initialization weights of the training model, initialize the convolutional layer parameters to 0, and freeze the SDXL model parameters at the same time;

[0083] (4) Loop to get the data of each batch, i.e., the previously encoded latent space representation and the text;

[0084] (5) The text label is encoded by the text encoder and input into the U-Net structure together, and the U-Net outputs the predicted noise residual. The predicted noise parameters obtained and the added noise are substituted into the above objective function to calculate the loss value;

[0085] (6) Backpropagation to calculate the gradient and update the control model parameters;

[0086] (7) Loop the above process until convergence or reach the maximum number of rounds.

[0087] In addition, as can be known from the foregoing, the control model needs to match the text-to-image model structure, so when the present embodiment uses SDXL, the trained SDXL model is used as the control model SDXL model, but the model parameters thereof do not participate in updating, and only the encoder blocks and intermediate blocks of the U-Net structure thereof are copied as the control model, and the weights thereof are used for initialization.

[0088] Based on the unmanned aerial vehicle performance pattern generation method provided in the present embodiment, the user can quickly generate an unmanned aerial vehicle performance pattern that meets the user's expectations according to the input text label and control graph. In addition, the unmanned aerial vehicle performance pattern generation method provided in the present embodiment also has high flexibility and scalability. Users can freely adjust the content of the text label and the control graph according to their actual needs, thereby generating diversified unmanned aerial vehicle performance patterns. In combination Figures 7-9 , Figure 7 In combination with the foregoing, when the input texts are a1, a2, a3, and a4, the generated unmanned aerial vehicle performance patterns are as shown in FIG. 6. Figure 7As shown, a1 is "minimal, linedrawing style, white background, a classical car, round wheels, two windows, handlebars, side view.", i.e., "extremely simple, line drawing style, white background, a classic car, round wheels, two windows, handlebars, side view"; a2 is "simple, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.", i.e., "simple, line drawing style, white background, a classic car, round wheels, two windows, handlebars, side view"; a3 is "moderate, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.", i.e., "moderate, line drawing style, white background, a classic car, round wheels, two windows, handlebars, side view"; a4 is "intricate, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.", i.e., "intricate, line drawing style, white background, a classic car, round wheels, two windows, handlebars, side view"; Figure 8 As shown, b1 is "line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.", i.e., "line drawing style, white background, a classic car, round wheels, two windows, handlebars, side view"; b2 is "line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.", i.e., "line drawing style, white background, a classic car, round wheels, two windows, handlebars, side view"; Figure 9In the embodiment, the corresponding UAV performance pattern is generated based on the input text and the original picture, wherein the input text c1 is "line drawing style, white background, a cyberpunk panda, mechanical armor.", the input text c2 is "line drawing style, white background, a cyberpunk cat, mechanical armor.", the input text c3 is "line drawing style, white background, a blooming lotus.", and the input text c4 is "line drawing style, white background, an airplane in the sky".

[0089] Figure 2 A schematic block diagram of a UAV performance pattern generation device 200 provided by the embodiment of the present application is provided, and the device 200 comprises:

[0090] A data acquisition unit 201 is configured to acquire historical UAV performance patterns, and perform text labeling on the historical UAV performance patterns to construct a training data set.

[0091] A model construction unit 202 is configured to train a text-to-image model and a corresponding control model by using the training data set, so as to construct a pattern generation model.

[0092] A pattern generation unit 203 is configured to generate a corresponding UAV performance pattern by using the pattern generation model based on a set pattern generation requirement.

[0093] In an embodiment, the data acquisition unit 201 comprises:

[0094] A grade division unit is configured to divide the historical UAV performance patterns into different grades based on UAV performance times; wherein the grades include minimal, simple, medium and complex.

[0095] A first construction unit is configured to generate a text label for the historical UAV performance patterns by using a picture understanding model, and construct a first training data set in combination with the grades.

[0096] an image conversion unit, configured to acquire an original image corresponding to the historical unmanned aerial vehicle performance pattern, and convert the original image into a depth map by using a deep learning model;

[0097] a second construction unit, configured to construct a second training data set by combining the depth map, the level and the text label.

[0098] In an embodiment, the model construction unit 202 comprises:

[0099] a model training unit, configured to train the text-to-image model by using the first training data set, and train the control model by using the second training data set;

[0100] a model building unit, configured to build the pattern generation model based on the trained text-to-image model and the control model.

[0101] In an embodiment, the model training unit comprises:

[0102] a data input unit, configured to input first training data in the first training data set into a stable diffusion model;

[0103] a data encoding unit, configured to encode the first training data into latent space variables by using a variational autoencoder in the stable diffusion model;

[0104] an iterative diffusion unit, configured to iteratively diffuse the latent space variables based on the latent space until an iterative diffusion requirement is met, to obtain iteratively denoised latent space variables;

[0105] a first decoding unit, configured to decode the iteratively denoised latent space variables into corresponding images by using a decoder of the variational autoencoder.

[0106] In an embodiment, the model training unit further comprises:

[0107] a first updating unit, configured to construct a first objective function according to the following formula, and perform parameter updating on the text-to-image model based on the first objective function:

[0108] ;

[0109] wherein, L LDM represents an objective function, represents an encoder of the VAE, x represents the first training data, and y represents the text label, is noise added to the image, represents an encoding model of the text label, represents a U-net model structure, represents a latent variable to which t time step noises are added, and t represents a time step.

[0110] In an embodiment, the model training unit further comprises:

[0111] a model setting unit, configured to set a model structure and training hyperparameters of the control model; wherein the model structure of the control model comprises an encoder block and an intermediate block in the stable diffusion model, and the control model is added with a convolution layer before each block;

[0112] a data bucketing unit, configured to perform bucketing processing on second training data in the second training data set, and perform image encoding on the second training data by using a variational autoencoder;

[0113] an initialization unit, configured to load weights of the stable diffusion model to initialize the control model, and initialize the convolution layer to 0;

[0114] a loop iteration unit, configured to perform loop iteration training on the image encoded second training data by using the control model until a loop iteration requirement is met;

[0115] a second decoding unit, configured to decode a result of the loop iteration into a corresponding image by using a decoder of the variational autoencoder.

[0116] In an embodiment, the model training unit further comprises:

[0117] a second updating unit, configured to construct a second objective function according to the following formula, and perform parameter updating on the control model by using the second objective function:

[0118] ;

[0119] wherein, denotes the second objective function, denotes a latent variable without adding a time step, is a text constraint, is a control graph constraint.

[0120] Since the embodiments of the device part correspond to the embodiments of the method part, the embodiments of the device part are described in the description of the embodiments of the method part, and will not be described here.

[0121] The embodiments of the present application also provide a computer readable storage medium, which has a computer program stored thereon, and the computer program can implement the steps provided by the above embodiments when executed. The storage medium can include: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0122] The embodiment of the present application further provides a computer device, which can comprise a memory and a processor, the memory has a computer program stored therein, and the processor can realize the steps provided by the above embodiment when calling the computer program in the memory. Of course, the computer device can further comprise various network interfaces, power supplies and other components.

[0123] The various embodiments are described in a progressive manner in the specification, and each embodiment focuses on the difference from other embodiments. The same or similar parts of each embodiment can be referred to each other. For the system disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the related parts can be referred to the method part. It should be pointed out that, for ordinary skilled in the art, without departing from the principles of the present application, the present application can be improved and modified, and these improvements and modifications also fall within the protection scope of the claims of the present application.

[0124] It should also be noted that in the specification, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.

Claims

1. A method for generating a drone performance pattern, the method comprising: receiving a user input of a performance pattern; and generating a drone performance pattern based on the user input. The method comprises the following steps: collecting historical UAV performance patterns and text-labeling the historical UAV performance patterns to construct a training data set; training a text-to-image model and a corresponding control model using the training data set to construct a pattern generation model; generating a corresponding UAV performance pattern based on a set pattern generation requirement using the pattern generation model; the collecting historical UAV performance patterns and text-labeling the historical UAV performance patterns to construct a training data set comprises: dividing the historical UAV performance patterns into different levels based on UAV performance episodes; wherein the levels include minimal, simple, medium, and complex; generating text labels for the historical UAV performance patterns through a picture understanding model and combining the levels to construct a first training data set; obtaining original images corresponding to the historical UAV performance patterns and converting the original images into depth maps using a deep learning model; combining the depth maps, levels, and text labels to construct a second training data set; the training a text-to-image model and a corresponding control model using the training data set to construct a pattern generation model comprises: training the text-to-image model using the first training data set and training the control model using the second training data set; building the pattern generation model based on the trained text-to-image model and control model. 2.The method of claim 1, wherein, the training the text-to-image model using the first training data set comprises: inputting first training data in the first training data set into a stable diffusion model; encoding the first training data into latent space variables using a variational autoencoder in the stable diffusion model; iteratively diffusing the latent space variables based on the latent space until an iterative diffusion requirement is met to obtain iteratively denoised latent space variables; decoding the iteratively denoised latent space variables into corresponding images using a decoder of the variational autoencoder. 3.The method of claim 2, wherein, the training the text-to-image model using the first training data set further comprises: constructing a first objective function according to the following formula and updating parameters of the text-to-image model based on the first objective function: ; wherein L LDM represents an objective function, represents an encoder of the VAE, x represents the first training data, and y represents a text label, is noise added to the image, represents an encoding model of the text label, represents a U-net model structure, represents a latent variable to which t time step noise is added, and t represents a time step. 4.The method of claim 3, wherein, the training the control model using the second training data set comprises: setting a model structure and training hyperparameters of the control model; wherein the model structure of the control model includes an encoder block and an intermediate block in the stable diffusion model, and the control model adds a convolutional layer before each block; performing bucketing processing on second training data in the second training data set and image-encoding the second training data through a variational autoencoder; loading weights of the stable diffusion model to initialize the control model and initializing the convolutional layer to 0; performing cyclic iteration training on the image-encoded second training data using the control model until a cyclic iteration requirement is met; decoding the result of the cyclic iteration into corresponding images using a decoder of the variational autoencoder. 5.The method of claim 4, wherein, the training the control model using the second training data set further comprises: A second objective function is constructed according to the following formula, and the control model is updated in parameters using the second objective function: ; wherein, represents a second objective function, represents a latent variable without added time steps, is a text constraint, is a control chart constraint. 6.A drone performance pattern generation apparatus, characterized by, The method comprises the steps of: The data acquisition unit is configured to acquire historical UAV performance patterns and mark the historical UAV performance patterns with texts to construct a training data set; The model construction unit is configured to train a text-to-image model and a corresponding control model using the training data set to construct a pattern generation model; The pattern generation unit is configured to generate a corresponding UAV performance pattern using the pattern generation model based on a set pattern generation requirement; The data acquisition unit comprises: The grade division unit is configured to divide the historical UAV performance patterns into different grades based on UAV performance times; the grades include minimal, simple, medium, and complex; The first construction unit is configured to generate text labels for the historical UAV performance patterns using a picture understanding model and construct a first training data set in combination with the grades; The image conversion unit is configured to acquire original images corresponding to the historical UAV performance patterns and convert the original images into depth maps using a deep learning model; The second construction unit is configured to construct a second training data set in combination with the depth maps, grades, and text labels; The model construction unit comprises: The model training unit is configured to train the text-to-image model using the first training data set and train the control model using the second training data set; The model construction unit is configured to construct the pattern generation model based on the trained text-to-image model and control model.

7. A computer device, comprising: The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the UAV performance pattern generation method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the UAV performance pattern generation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for generating performance picture of cluster unmanned aerial vehicles, product, storage medium and electronic equipment

    CN112596536A

  • Method, device, equipment and medium for generating drone performance patterns

    CN118967855B

  • Pentograph model training method and device, equipment and storage medium

    CN117173504A

  • Fine adjustment method and device for text graph model, electronic equipment and storage medium

    CN119312842A