Unmanned aerial vehicle performance pattern generation method and device, computer equipment and storage medium
By constructing a training data set and utilizing the cultural graph model and control model, drone performance patterns are automatically generated, solving the problems of low generation efficiency and poor quality in existing technologies and achieving efficient and accurate pattern generation.
Patent Information
- Application Number
- CN202511255210.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-04
AI Technical Summary
When generating drone performance patterns, existing technologies have difficulty in efficiently generating patterns that meet style requirements and match the density of lines with the performance sorties, and require a lot of manual adjustments.
By collecting historical drone performance patterns, building a training dataset, and using the cultural graph model and control model for training, drone performance patterns that meet the set requirements are generated.
It reduces manual intervention, shortens the generation cycle, improves generation efficiency and pattern quality, and ensures that the pattern meets style requirements and that the density of lines matches the performance sequence.
Smart Images

Figure CN120747288A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer software technology, and in particular to a method and device for generating a drone performance pattern, a computer device, and a storage medium. Background Art
[0002] As drone light show technology continues to advance, the number of performances continues to increase, now reaching tens of thousands. The design of the performance pattern has become increasingly crucial within the overall drone light show. Currently, there are two main ways to create a drone light show pattern: one is direct design by designers; the other is to use design software to outline the performance pattern based on client-specified images, and then use discrete positioning to determine the position of the drone to form a complete drone show. However, as the complexity and number of performance images increase, the workload of manually obtaining performance patterns is increasing. At the same time, existing models have shortcomings in generating drone performance patterns: some models are difficult to meet the style requirements of drone performance patterns, and some models cannot generate patterns with line density that matches the performance sorties. Even if carefully designed text prompts are used, it usually takes multiple generations to obtain a pattern that is close to the requirements in terms of style and line density, and often a lot of modifications are required. In addition, related existing technologies, such as CN118967855B, are mainly used to solve problems such as position alignment and automatic discrete point distribution of existing performance patterns. It involves related devices, equipment and media, but does not mention the method for generating performance patterns; CN112596536A solves the problem of discretely distributing existing performance patterns to form performance images based on parameters such as drone performance sorties. Like CN118967855B, it does not involve the method for generating performance patterns. Summary of the Invention
[0003] Embodiments of the present invention provide a method, apparatus, computer device, and storage medium for generating a drone performance pattern, aiming to improve the generation efficiency and effect of the drone performance pattern.
[0004] In a first aspect, an embodiment of the present invention provides a method for generating a drone performance pattern, comprising: Collecting historical drone performance patterns and performing text labeling on the historical drone performance patterns to construct a training dataset; Using the training data set to train the Wensheng graph model and its corresponding control model, thereby constructing a pattern generation model; Based on the set pattern generation requirements, the pattern generation model is used to generate a corresponding drone performance pattern.
[0005] In a second aspect, an embodiment of the present invention provides a device for generating a drone performance pattern, comprising: A data collection unit, configured to collect historical drone performance patterns and perform text labeling on the historical drone performance patterns to construct a training data set; A model building unit, configured to train a Wensheng graph model and its corresponding control model using the training data set, thereby building a pattern generation model; The pattern generation unit is used to generate a corresponding drone performance pattern using the pattern generation model based on the set pattern generation requirements.
[0006] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for generating a drone performance pattern as described in the first aspect is implemented.
[0007] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for generating a drone performance pattern as described in the first aspect is implemented.
[0008] Embodiments of the present invention provide a method, apparatus, computer device, and storage medium for generating drone performance patterns. The method comprises: collecting historical drone performance patterns and performing text annotation on the patterns to construct a training dataset; using the training dataset to train a Vincent graph model and its corresponding control model to construct a pattern generation model; and using the pattern generation model to generate corresponding drone performance patterns based on set pattern generation requirements. The embodiments of the present invention collect historical drone performance patterns and perform text annotation on them to construct a training dataset. The dataset is then used to train the Vincent graph model and its corresponding control model to construct a pattern generation model. Finally, based on the set pattern generation requirements, the model is used to generate corresponding drone performance patterns. This reduces the need for manual design or repeated adjustments. Through automated model generation, the intensity of manual intervention is reduced, significantly shortening the drone performance pattern generation cycle and improving generation efficiency. Furthermore, because the training dataset is constructed based on historical drone performance patterns and is collaboratively trained with the control model, the generated patterns can better align with the style requirements of the drone performances and more accurately match the density of lines with the performances. This reduces the workload of subsequent modifications and improves the generation quality and applicability of the drone performance patterns. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0010] Figure 1 A schematic diagram of a flow chart of a method for generating a drone performance pattern provided by an embodiment of the present invention; Figure 2 A schematic block diagram of a device for generating a drone performance pattern according to an embodiment of the present invention; Figure 3 A schematic diagram of a data set in a method for generating a drone performance pattern provided by an embodiment of the present invention; Figure 4 A schematic diagram of another data set in a method for generating a drone performance pattern provided by an embodiment of the present invention; Figure 5 A network architecture diagram of a privacy release model in a method for generating a drone performance pattern provided by an embodiment of the present invention; Figure 6 A network architecture diagram of a control model in a method for generating a drone performance pattern provided by an embodiment of the present invention; Figure 7 This is a first generation effect diagram of a method for generating a drone performance pattern provided by an embodiment of the present invention; Figure 8 A second generation effect diagram of a method for generating a drone performance pattern provided by an embodiment of the present invention; Figure 9 This is a third generation effect diagram of a method for generating a drone performance pattern provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0012] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0013] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.
[0014] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0015] See below Figure 1 , an embodiment of the present invention provides a method for generating a drone performance pattern, specifically including: steps S101~S103.
[0016] Step S101: Collect historical drone performance patterns and perform text labeling on the historical drone performance patterns to construct a training data set; Step S102: using the training data set to train the Wensheng graph model and its corresponding control model to construct a pattern generation model; Step S103: Based on the set pattern generation requirements, the pattern generation model is used to generate a corresponding drone performance pattern.
[0017] In this embodiment, a training dataset is first constructed by collecting historical drone performance patterns and annotating them with text. This dataset is then used to train a text-based graph model and its control model, creating a pattern generation model. Finally, the model is used to generate corresponding drone performance patterns based on the specified pattern generation requirements. This reduces the need for manual design or repeated adjustments. Through automated model generation, the intensity of manual intervention is reduced, significantly shortening the generation cycle for drone performance patterns and improving generation efficiency. Furthermore, because the training dataset is constructed based on historical drone performance patterns and is collaboratively trained with the control model, the generated patterns can better align with the stylistic requirements of drone performances and more accurately match the density of lines with the performance schedule. This reduces the workload for subsequent revisions and improves the quality and applicability of the generated drone performance patterns.
[0018] In one embodiment, step S101 includes: Based on the number of drone performances, the historical drone performance patterns are divided into different levels; wherein the levels include minimalist, simple, medium and complex; Generating text labels for the historical drone performance patterns using an image understanding model, and constructing a first training dataset based on the levels; Obtaining an original image corresponding to the historical drone performance pattern, and converting the original image into a depth map using a deep learning model; A second training dataset is constructed by combining the depth map, level and text labels.
[0019] Because drone performance patterns primarily consist of simple drawings and line drawings, which differ significantly from sketches, they primarily feature clean black lines against a white background, facilitating subsequent discrete placement of lines to create the drone performance image. Furthermore, the density of lines in the performance pattern is closely related to the number of performances. Therefore, when text labeling the performance patterns, it is necessary to distinguish between the different line densities associated with the number of performances in the text description. Therefore, this embodiment categorizes the data into four levels based on the number of performances: minimal, simple, moderate, and intricate. An image understanding model and manual correction are then used to generate a text label describing each performance pattern.
[0020] Based on the divided levels, the first training data set for the text-based image model training is first generated. This data set is mainly used to train the model, input text, and generate drone performance patterns. In actual applications, the descriptive words of the above four levels are placed before the text label to synthesize the text label of each performance pattern. The entire data set has a total of 60,000 data with different aspect ratios, and each data consists of a performance pattern and a text label. As the complexity of the performance picture increases, the number of drone performances also increases. Therefore, the text labels can be generated by the image understanding model Qwen2.5-VL-7B-Instruct, and combined with manual correction and other methods to ensure that the text annotation is accurate and reliable. Among them, the four words minimal, simple, moderate, and intricate are used as prompt words to generate performance patterns with different line densities.
[0021] Combine Figure 3 The data in the first training dataset include performance numbers, performance patterns, and text labels. Here, because the text labels contain more content, they are replaced by characters in the figure (the same below). At the same time, because the English encoder is used, the English description data is used (the same below). Figure 3In the image, text label A means “minimal, line drawing style, white background, a blooming flower, with the petals occupying the main part, consists of five petals, each of which presents a soft curve and natural shape, with slightly curled edges and slight overlap between the petals. In the center of the flower, there is a cluster of slender stamens, dotted with small dots at the top.”; text label B means “simple, line drawing style, white background, a scene depicts a group of lotus flowers and leaves. Several leaves of varying sizes are neatly arranged in various locations, each with radiating veins and slightly wavy edges. Onelarge leaf is nearly circular in shape; another also features delicateveining. A blooming lotus flower and an unopened bud are all connected byslender stems.", that is, "a simple, line drawing style with a white background, depicting a scene of a group of lotus flowers and leaves.Lotus leaves of varying sizes are scattered throughout the painting, each with radiating veins and gently curled edges. One large leaf is nearly round, while another also displays delicate veining. A blooming lotus and a bud, about to bloom, are connected by a slender stem. "; Text label C indicates "moderate, line drawing style, white background, three blooming flowers and green leaves, with layers of petals and delicately carved stamens, show complexity and beauty. The leaves are of different shapes and have serrated edges, adding a sense of layering. The stems are interwoven to form a base." "; text label D means "intricate, line drawing style, white background, a steaming bowl of noodles, filled with various toppings, is lifted from the bowl with chopsticks, revealing a soft and springy texture. The bowl is filled with a variety of ingredients, including chunks of vegetables and slices of meat, and the surface is dotted with a few drops of broth."
[0022] A second training dataset is then generated for control model training. This dataset is primarily used to train the control model. Control maps and text are input, and drone performance patterns with the same structure as the existing images are generated. Existing control models convert the original image into line drawings, depth maps, normal maps, pose maps, semantic segmentation maps, and other components, and then use these converted images for image control. However, line drawings introduce background noise, making the generated image overly complex. Normal maps, pose maps, and semantic segmentation maps, among others, cause significant information loss in the original image, making it difficult to generate performance patterns with the same composition as the original. Only depth maps can provide accurate control information without significant information loss and are adaptable to all original image scenarios, making them highly suitable for performance pattern generation. Therefore, this embodiment uses depth maps as control maps. In practice, a training dataset of 60,000 records with varying aspect ratios was created. Each record consists of a depth map, a performance pattern, and text labels. The depth model converts the original image into a depth map, while the image understanding model generates text, combining line density terms to form text labels.
[0023] Combine Figure 4 The data in the second training dataset include performance numbers, depth maps, performance patterns and text labels. Figure 4In the image, the text label E means “minimal, line drawing style, white background, atruck. The truck is composed of simple geometric shapes. The front sectionfeatures a rectangular cab with the outlines of windows and doors. The bodyis a rectangular cargo box with a horizontal and vertical line on the siderepresenting the doors. The wheels are represented by three circles, eachwith a small dot inside to represent the wheel hub.”; the text label F means “simple, linedrawing style, white background, three chrysanthemums of different shapes anda few leaves, each flower outlined with simple and smooth lines, the petalsclearly layered, showing the states of bloom and bud. The flowers are connected by a slender stem, "A simple, line-drawn style, set against a white background, depicts three chrysanthemums of various shapes and a few leaves. Each flower is outlined with simple, flowing lines, and the petals are distinctly layered, showing both the blooming and budding states. The flowers are connected by a slender stem, dotted with several leaves of various shapes, with naturally curved edges.""; the text label G means "moderate, line drawing style, white background, a cluster of figs and their leaves, containing three ripe, oval figs with distinct longitudinalstripes, appearing plump and textured. Each fig is attached to a slenderbranch dotted with several broad leaves with wavy edges and clearly visibleveins, showcasing a natural texture.", that is, "moderate, line drawing style, white background, a cluster of figs and their leaves, containing three ripe, oval figs with distinct longitudinalstripes, appearing plump and textured. Each fig is attached to a slenderbranch dotted with several broad leaves with wavy edges and clearly visibleveins, showcasing a natural texture." "; the text tag H stands for "intricate, line drawing style,white background, a scene depicts an ancient sailing vessel sailing on thesea. Its streamlined hull, with distinct wavy lines on its underside,symbolizes its waterborne navigation. The vessel boasts three decks, eachwith neatly arranged windows and railings. The sails are hoisted high, thesails inflated, and a flag flutters from the mast.", that is, "intricate, line drawing style,white background, a scene depicts an ancient sailing vessel sailing on the sea. The streamlined hull is decorated with eye-catching wave patterns, symbolizing the characteristics of water navigation. The hull has three decks, each with neatly arranged portholes and railings. The sails are hoisted high, under the bulging sails, a flag flutters from the mast."
[0024] In one embodiment, step S102 includes: Using the first training data set to train the Vincent graph model, and using the second training data set to train the control model; Based on the trained Wensheng graph model and the control model, the pattern generation model is constructed.
[0025] This embodiment trains the text graph model and control model using a first training dataset and a second training dataset, respectively. This allows the text graph model to generate corresponding drone performance patterns based on input text labels, while the control model can generate drone performance patterns with the same structure as existing images based on input control graphs and text. This not only improves the efficiency and accuracy of drone performance pattern generation, but also ensures that the generated patterns better meet user expectations and needs. With the trained text graph model and control model, users only need to input the corresponding text labels or control graphs and text labels to quickly generate the desired drone performance patterns.
[0026] In one embodiment, the training of the text graph model using the first training dataset includes: inputting first training data in the first training data set into a stable diffusion model; Encoding the first training data into latent space variables using a variational autoencoder in a stable diffusion model; Iteratively diffusing the latent space variables based on the latent space until iterative diffusion requirements are met, thereby obtaining iteratively denoised latent space variables; The decoder of the variational autoencoder is used to decode the iteratively denoised latent space variables into the corresponding images.
[0027] Furthermore, the training of the culture graph model using the first training data set further includes: According to the following formula, a first objective function is constructed, and parameters of the culture graph model are updated based on the first objective function: ; Among them, L LDM represents the objective function, represents the encoder of VAE, x represents the first training data, y represents the text label, is the noise added to the image, An encoding model representing text labels, Represents the U-net model structure, represents a latent variable with t time-step noise added, where t represents the time step.
[0028] This embodiment uses Stable Diffusion (SD) as the text-to-image model. Stable Diffusion is a latent space text-to-image diffusion model. The image is encoded into a latent space variable through a Variational Autoencoder (VAE), and then an iterative diffusion process is performed in the latent space. During inference, the latent space variable obtained after iterative denoising is decoded into an image through VAE. At the same time, the specific image content is generated by constraining the image such as text. The whole process is as follows: Figure 5 shown. Figure 5 In the figure, the red area on the left represents the pixel space (PixelSpace), the green area in the middle represents the latent space (Latent Space), and the right represents the constraint (Conditioning). x is the image, represents the encoder of VAE, z is the latent variable in the latent space, is a latent variable with T time-step noise added, is the noise added to the image, t represents the time step, Represents the U-Net model structure, and the subscript θ represents the parameters of the neural network. is the decoder of VAE, is the decoded image, It is an encoding model of the constraints, which is used to map the constraints into an intermediate representation, and then map it to the intermediate layer of U-Net through the cross attention layer.
[0029] To support training with different aspect ratios, improve text-conforming capabilities, and balance generated image diversity and speed, an improved stable diffusion model, the Stable Diffusion XL model, can be used. In practical applications, to accommodate images of varying aspect ratios encountered in real-world scenarios, bucketed training can be performed using images of varying aspect ratios. Specifically, image sizes include the following: [768*1408, 768*1344, 768*1280, 896*1152, 960*1024, 1024*1024, 1024*960, 1152*896, 1280*768, 1344*768, 1408*768]. Furthermore, for faster convergence, existing model parameters can be used to initialize the model structure to be trained.
[0030] The specific training process of the stable diffusion model is as follows: (1) Determine the model structure to be trained and the training hyperparameters, including batch size, maximum number of rounds, optimizer, learning rate, learning rate schedule and other key hyperparameters; (2) The training data is bucketed according to the aspect ratio, and the images are encoded into latent space representations in advance using VAE; (3) Load the existing weights as the initial weights for the training model; (4) Iterate and loop to obtain the data of each batch, that is, the previously encoded latent space representation and text; (5) The text label is encoded by the text encoder and input into the U-Net structure together. The U-Net outputs the predicted noise residual , the predicted noise residual and the added noise Substitute the above objective function to calculate the loss value; (6) Back propagation calculates gradients and updates model parameters; (7) Repeat the above process until convergence or the maximum number of rounds is reached.
[0031] In one embodiment, the training of the control model using the second training data set includes: Setting a model structure and training hyperparameters of the control model; wherein the model structure of the control model includes an encoder block and an intermediate block in the stable diffusion model, and the control model has a convolutional layer added before each block; Performing bucket processing on the second training data in the second training data set, and performing image encoding on the second training data through a variational autoencoder; Loading the weights of the stable diffusion model to initialize the control model and initializing the convolutional layer to 0; Performing cyclic iterative training on the image-encoded second training data using the control model until cyclic iteration requirements are met; The decoder of the variational autoencoder is used to decode the result of the loop iteration into the corresponding image.
[0032] Furthermore, the training of the control model using the second training data set further includes: The second objective function is constructed according to the following formula, and the parameters of the control model are updated using the second objective function: ; in, represents the second objective function, represents the latent variable without adding time steps, is a text constraint, are control chart constraints.
[0033] In order to extract the existing image and the UAV performance pattern, this embodiment requires a control model to control the composition of the existing image. The control model (controlnet) can be used as an adaptation model of the SD model to generate image content by inputting text and controlling images. Figure 6 The control model copies the encoder block and middle block of the U-Net structure of the existing SD model and adds a 1×1 convolution layer with an initial weight of 0 before each block. The output of the control model is added to the middle block and decoder block of the U-Net structure through skip connections, thereby playing a control role. Only the 1×1 convolution layer, the copied encoder block and the middle block need to be trained, and the other parts do not participate in the training. During training, the control graph constraint obtained by encoding the constraint image by the control model , the text constraints obtained by the text encoder , combining each time step t and the corresponding latent variable , are input into the U-Net structure of SD to predict the noise residual, and then combined with the added noise , are substituted into the objective function of the control model (i.e., the second objective function) to calculate the loss. Backpropagation is then used to calculate the gradient, and the control model parameters are updated based on the gradient until the model converges. During inference, the control model encodes the constrained image and, along with the text constraints, is input into the U-Net structure of the SD model to control image generation. After gradual denoising and iteration, latent variables are generated that are identically distributed to the training image. These are then decoded into pixel space by the VAE decoder to produce the image.
[0034] To adapt to images with different aspect ratios in real application scenarios, we can use images with different aspect ratios for bucket training. The specific image sizes include the following: [768*1408, 768*1344, 768*1280, 896*1152, 960*1024, 1024*1024, 1024*960, 1152*896, 1280*768, 1344*768, 1408*768]. The specific training process of the control model is as follows: (1) Determine the control model structure to be trained and the training hyperparameters, including batch size, maximum number of rounds, optimizer, learning rate, learning rate schedule and other key hyperparameters; (2) The training data is bucketed according to the aspect ratio, and the images are encoded into latent space representations in advance using VAE; (3) Load the weights of the SDXL model as the initial weights of the training model, initialize the convolutional layer parameters to 0, and freeze the SDXL model parameters; (4) Iterate and loop to obtain the data of each batch, that is, the previously encoded latent space representation and text; (5) The text label is encoded by the text encoder and input into the U-Net structure together. The U-Net outputs the predicted noise residual. The predicted noise parameters and the added noise are substituted into the above objective function to calculate the loss value. (6) Back propagation calculates gradients and updates control model parameters; (7) Repeat the above process until convergence or the maximum number of rounds is reached.
[0035] In addition, combined with the above content, it can be seen that the control model needs to match the structure of the Wensheng graph model. Therefore, when SDXL is used in this embodiment, the above-trained SDXL model is used as the SDXL model of the control model, but its model parameters are not updated. Only the encoder block and intermediate block of its U-Net structure are copied as the control model, and its weights are used for initialization.
[0036] Based on the drone performance pattern generation method provided by this embodiment, a drone performance pattern that meets the user's expectations can be quickly generated based on the input text labels and control diagrams. In addition, the drone performance pattern generation method provided by this embodiment is also highly flexible and scalable. Users can freely adjust the content of text labels and control diagrams according to their actual needs, thereby generating a variety of drone performance patterns. Figure 7-Figure 9 , Figure 7 When the input texts are a1, a2, a3 and a4 respectively, the generated drone performance pattern is as follows Figure 7As shown in the figure, a1 is “minimal, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.”, that is, “minimal, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.”; a2 is “simple, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.”, that is, “simple, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.”; a3 is “moderate, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.”, that is, “moderate, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.”; a4 is “intricate, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.”, that is, “complex, line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.”; Figure 8 In the example, a drone performance pattern is generated based on the input text and the original image. The input text b1 is “line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.”, which means “line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.” The input text b2 is “line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.”, which means “line drawing style, white background, a classical car, round wheels, two windows, handlebars, side view.” Figure 9In the example, the corresponding drone performance pattern is generated based on the input text and the original image. The input text c1 is “line drawing style, white background, a cyberpunk panda, mechanical armor.”, that is, “line drawing style, white background, cyberpunk panda, mechanical armor.”; the input text c2 is “line drawing style, white background, acyberpunk cat, mechanical armor.”, that is, “line drawing style, white background, cyberpunk cat, mechanical armor.”; the input text c3 is “line drawing style, white background, a blooming lotus.”, that is, “line drawing style, white background, a blooming lotus.”; the input text c4 is “line drawing style, white background, an airplane in the sky.”, that is, “line drawing style, white background, an airplane in the sky.”
[0037] Figure 2 This is a schematic block diagram of a drone performance pattern generation device 200 provided in an embodiment of the present invention. The device 200 includes: The data collection unit 201 is used to collect historical drone performance patterns and perform text labeling on the historical drone performance patterns to construct a training data set; A model building unit 202 is used to train the Wensheng graph model and its corresponding control model using the training data set to build a pattern generation model; The pattern generation unit 203 is configured to generate a corresponding drone performance pattern using the pattern generation model based on a set pattern generation requirement.
[0038] In one embodiment, the data acquisition unit 201 includes: A level classification unit, configured to classify the historical drone performance patterns into different levels based on the number of drone performance sorties; wherein the levels include minimalist, simple, medium, and complex; A first construction unit is configured to generate text labels for the historical drone performance patterns using an image understanding model, and construct a first training dataset based on the levels; An image conversion unit, configured to obtain an original image corresponding to the historical drone performance pattern, and convert the original image into a depth map using a deep learning model; The second construction unit is configured to construct a second training data set by combining the depth map, the level, and the text label.
[0039] In one embodiment, the model building unit 202 includes: a model training unit, configured to train the Wensheng graph model using the first training data set, and to train the control model using the second training data set; The model building unit is used to build the pattern generation model based on the trained text graph model and the control model.
[0040] In one embodiment, the model training unit includes: a data input unit, configured to input first training data in the first training data set into a stable diffusion model; a data encoding unit, configured to encode the first training data into latent space variables using a variational autoencoder in a stable diffusion model; an iterative diffusion unit, configured to iteratively diffuse the latent space variables based on the latent space until iterative diffusion requirements are met, thereby obtaining iteratively denoised latent space variables; The first decoding unit is configured to decode the iteratively denoised latent space variable into a corresponding image using a decoder of the variational autoencoder.
[0041] In one embodiment, the model training unit further includes: The first updating unit is configured to construct a first objective function according to the following formula, and update parameters of the culture graph model based on the first objective function: ; Among them, L LDM represents the objective function, represents the encoder of VAE, x represents the first training data, y represents the text label, is the noise added to the image, An encoding model representing text labels, Represents the U-net model structure, represents a latent variable with t time-step noise added, where t represents the time step.
[0042] In one embodiment, the model training unit further includes: A model setting unit, configured to set a model structure and training hyperparameters of the control model; wherein the model structure of the control model includes an encoder block and an intermediate block in the stable diffusion model, and the control model has a convolutional layer added before each block; a data bucketing unit, configured to perform bucketing processing on the second training data in the second training data set, and perform image encoding on the second training data through a variational autoencoder; an initialization unit, configured to load the weights of the stable diffusion model to initialize the control model and initialize the convolutional layer to 0; a loop iteration unit, configured to perform loop iterative training on the image-encoded second training data using the control model until loop iteration requirements are met; The second decoding unit is used to decode the result of the loop iteration into a corresponding image using the decoder of the variational autoencoder.
[0043] In one embodiment, the model training unit further includes: The second updating unit is configured to construct a second objective function according to the following formula, and update the parameters of the control model using the second objective function: ; in, represents the second objective function, represents the latent variable without adding time steps, is a text constraint, are control chart constraints.
[0044] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and they will not be repeated here.
[0045] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed, the computer program can implement the steps provided in the above embodiment. The storage medium may include a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or other medium capable of storing program code.
[0046] The present invention also provides a computer device that may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, the steps provided in the above embodiment can be implemented. Of course, the computer device may also include various network interfaces, a power supply, and other components.
[0047] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
[0048] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
Claims
1. A method for generating a drone performance pattern, characterized in that: include: Collecting historical drone performance patterns and performing text labeling on the historical drone performance patterns to construct a training dataset; Using the training data set to train the Wensheng graph model and its corresponding control model, thereby constructing a pattern generation model; Based on the set pattern generation requirements, the pattern generation model is used to generate a corresponding drone performance pattern.
2. The method for generating a drone performance pattern according to claim 1, wherein: The collecting of historical drone performance patterns and text labeling of the historical drone performance patterns to construct a training dataset includes: Based on the number of drone performances, the historical drone performance patterns are divided into different levels; wherein the levels include minimalist, simple, medium and complex; Generating text labels for the historical drone performance patterns using an image understanding model, and constructing a first training dataset based on the levels; Obtaining an original image corresponding to the historical drone performance pattern, and converting the original image into a depth map using a deep learning model; A second training dataset is constructed by combining the depth map, level and text labels.
3. The method for generating a drone performance pattern according to claim 2, wherein: The method of using the training data set to train the Wensheng graph model and its corresponding control model to construct a pattern generation model includes: Using the first training data set to train the Vincent graph model, and using the second training data set to train the control model; Based on the trained Wensheng graph model and the control model, the pattern generation model is constructed.
4. The method for generating a drone performance pattern according to claim 3, wherein: The training of the culture graph model using the first training data set includes: inputting first training data in the first training data set into a stable diffusion model; Encoding the first training data into latent space variables using a variational autoencoder in a stable diffusion model; Iteratively diffusing the latent space variables based on the latent space until iterative diffusion requirements are met, thereby obtaining iteratively denoised latent space variables; The decoder of the variational autoencoder is used to decode the iteratively denoised latent space variables into the corresponding images.
5. The method for generating a drone performance pattern according to claim 4, wherein: The training of the culture graph model using the first training data set further includes: According to the following formula, a first objective function is constructed, and parameters of the culture graph model are updated based on the first objective function: ; Among them, L LDM represents the objective function, represents the encoder of VAE, x represents the first training data, y represents the text label, is the noise added to the image, An encoding model representing text labels, Represents the U-net model structure, represents a latent variable with t time-step noise added, where t represents the time step.
6. The method for generating a drone performance pattern according to claim 5, wherein: The using the second training data set to train the control model includes: Setting a model structure and training hyperparameters of the control model; wherein the model structure of the control model includes an encoder block and an intermediate block in the stable diffusion model, and the control model has a convolutional layer added before each block; Performing bucket processing on the second training data in the second training data set, and performing image encoding on the second training data through a variational autoencoder; Loading the weights of the stable diffusion model to initialize the control model and initializing the convolutional layer to 0; Performing cyclic iterative training on the second training data after image encoding using the control model until the cyclic iteration requirement is met; The decoder of the variational autoencoder is used to decode the result of the loop iteration into the corresponding image.
7. The method for generating a drone performance pattern according to claim 6, wherein: The training of the control model using the second training data set further includes: The second objective function is constructed according to the following formula, and the parameters of the control model are updated using the second objective function: ; in, represents the second objective function, represents the latent variable without adding time steps, is a text constraint, are control chart constraints.
8. A device for generating a drone performance pattern, characterized in that: include: A data collection unit, configured to collect historical drone performance patterns and perform text labeling on the historical drone performance patterns to construct a training data set; A model building unit, configured to train a Wensheng graph model and its corresponding control model using the training data set, thereby building a pattern generation model; The pattern generation unit is used to generate a corresponding drone performance pattern using the pattern generation model based on the set pattern generation requirements.
9. A computer device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for generating a drone performance pattern according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating a drone performance pattern according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method for generating performance picture of cluster unmanned aerial vehicles, product, storage medium and electronic equipment
CN112596536A
Method, device, equipment and medium for generating drone performance patterns
CN118967855B
Pentograph model training method and device, equipment and storage medium
CN117173504A
Fine adjustment method and device for text graph model, electronic equipment and storage medium
CN119312842A
Figure graph model fine adjustment method, device and equipment, storage medium and vehicle
CN120182737A
Cited By
Unmanned aerial vehicle performance picture generation method and device, equipment and medium
CN121505088A
Unmanned aerial vehicle performance picture generation method and device, equipment and medium
CN121505088B