Universal adaptive system and method for controlling virtual camera parameters in ai video generation system

A universal adaptive system with a GUI simplifies virtual camera control in generative AI systems by enabling automatic syntax adaptation and semantic enrichment, addressing the complexity and inefficiency of current methods, and improving content quality and personalization.

RU2865624C1Active Publication Date: 2026-07-07ПЕРЕГУДОВ ПАВЕЛ ГЕОРГИЕВИЧ
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
RU · RU
Patent Type
Patents
Current Assignee / Owner
ПЕРЕГУДОВ ПАВЕЛ ГЕОРГИЕВИЧ
Filing Date
2025-12-29
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Current virtual camera control in generative AI systems requires users to learn and remember unique syntax for each neural network model, leading to inefficient and complex interaction due to the lack of a unified abstraction layer and personalization mechanisms.

Method used

A universal adaptive system with a graphical user interface (GUI) that allows users to control virtual camera parameters through a single interface, utilizing a command conversion module for automatic syntax adaptation and a semantic enhancement module to enrich commands based on user feedback and neural network algorithms, along with an iterative learning module for personalized customization.

Benefits of technology

Simplifies user interaction by eliminating the need for manual syntax adaptation and enhances the quality and personalization of generated content through semantic enrichment and iterative learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_ABST
    Figure 00000001_ABST
Patent Text Reader

Abstract

FIELD: information technology.SUBSTANCE: technical result consists of simplifying and unifying the process of user interaction with various AI video generation models, improving the quality and personalization of generated video content. The system comprises a registry of AI models for the user to select a target model. The command conversion module receives signals from the GUI corresponding to user commands and automatically generates text prompts. This module includes a syntax adaptation module that extracts a base token (e.g., "pedestal_up") according to the syntax of the selected model, and a semantic enhancement module that semantically enriches this token. The system also optionally includes an iterative learning module that collects user feedback on video quality, evaluates it using a reward model, and uses this evaluation to further train the semantic enhancement module, thereby adapting the system to user preferences.EFFECT: universal adaptive control system comprising a GUI with visual camera control elements.10 cl, 6 dwg, 1 tbl
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Technical field

[0002] The invention relates to computer technology and human-machine interaction, specifically to graphical user interfaces (GUIs) and artificial intelligence (AI) systems. It can be used for unified control of virtual camera parameters (including movement, optics, lighting, and styling) in heterogeneous software environments and aggregators of neural network video generation models, such as Kling AI, Runway ML, Sora, Qwen, and others.

[0003] Technology Level

[0004] Currently, virtual camera movement control in generative artificial intelligence (AI) systems (hereinafter referred to as neural network models) is implemented primarily through text commands (prompts). Each neural network model uses its own unique syntax, set of commands, and parameters. For example, as shown in Table 1, the “Move up” command may require the syntax “pedestal_up” (Runway ML), “camera boom up” (Luma), or “camera move up” (Kling AI). This creates a significant barrier for users, who are forced to learn and remember the specific terminology and syntax for each individual neural network model. Existing graphical interfaces, such as those in Kling AI (Fig. 4), Runway ML (Fig. 5), or Higgsfield (Fig. 6), do not address this problem, providing either unsystematized icons or sliders that still require in-depth user knowledge to compose the basic text prompt.

[0005] Closest analogs include Runway LM and the Comfy UI graphical interface, which have similar camera control modules, but they lack a mechanism for additional learning based on user feedback, which makes the configuration less predictable and personalized (based on preliminary search and subjective assessments: Runway ML requires several hours and 5-10 generation iterations to understand the impact of commands, while the proposed solution provides more accurate and faster customization).

[0006] The closest known analogue (prototype) is the "CinePreGen" system, described in the article by Yiran Chen et al., "CinePreGen: Camera-Controlled Video Previsualization via Engine-Powered Diffusion." The prototype is a system for visual video pre-rendering that combines a 3D engine and a diffusion AI model. The prototype system includes a graphical interface with a storyboard and camera editor, which allows the user to visually specify camera parameters and trajectories (e.g., pan, tilt, push, dolly zoom). Based on this data, the system generates video footage using its own diffusion rendering pipeline, driven by data from the 3D engine (e.g., depth and pose maps).

[0007] The prototype's shortcomings include, firstly, its narrow specialization and monolithic architecture. The CinePreGen system doesn't address the problem of unifying controls for various third-party neural network models (Kling, Sora, etc.), as its GUI is tightly coupled only to its own internal rendering engine. Secondly, the prototype doesn't provide personalization or self-learning based on user preferences. It uses groundtruth data from the 3D engine, but lacks a mechanism for collecting user feedback (SF) on the final generated video, and it lacks an iterative learning module (e.g., RLHF) for adapting and improving the prompt generation process based on this SF.

[0008] Current solutions do not provide a unified abstraction layer that would allow any neural network model to be controlled through a single interface with automatic syntax adaptation and intelligent command enrichment.

[0009] Disclosure of invention

[0010] The technical problem of the invention is the low efficiency and complexity of managing the parameters of a virtual camera in heterogeneous AI video generation systems, due to the need to study and adapt a specific syntax for each neural network model and scene.

[0011] The main technical result is the simplification and unification of the process of user interaction with various neural network video generation models by eliminating the need for manual input and syntax adaptation.

[0012] An additional technical result is an increase in the quality and personalization of generated content due to the semantic enrichment of commands using neural network algorithms, as well as the provision of the possibility of custom system configuration through iterative further training of the control system based on the user's OS.

[0013] The specified technical results are achieved due to the fact that in a universal adaptive system for controlling the parameters of a virtual camera in AI video generation systems, containing a GUI with visual control elements for setting conceptual commands (parameters and scenarios) for controlling a virtual camera and an execution module for generating video content based on the specified parameters, according to the invention, the system additionally contains at least:

[0014] • a registry of neural network models, designed to allow the user to select a target neural network model from a list of supported ones;

[0015] • a command conversion module configured to receive a signal corresponding to at least one command from at least one visual GUI element and automatically convert the said command into at least one text fragment of a prompt based on the said signal, wherein the command conversion module comprises:

[0016] - a syntax adaptation module containing a database of rules and tokens, and configured to extract from said database at least one token corresponding to said signal and the target neural network model selected by the user, and to generate a basic text fragment of a prompt based on said token;

[0017] - and a semantic enhancement module, implemented in the form of at least one large language model, and configured to receive the specified basic text fragment of the prompt and semantically enrich it by adding context-dependent descriptions to form the final text fragment of the prompt and analyze the context of the animated image if the image 2 video conversion script is used.

[0018] The system operates in a manner that includes the user specifying at least one command for the spatial movement of a virtual camera through a graphical user interface (GUI) and generating video content, characterized in that it additionally includes the steps of:

[0019] • provide the user with a register of neural network models in the GUI and receive the user’s choice of the target neural network model;

[0020] • receive a signal corresponding to at least one command from a GUI visual element;

[0021] • form the specified command into the final text fragment of the prompt based on the specified signal by at least the following sub-steps:

[0022] - forming a basic text fragment of the prompt by extracting at least one token from a database of rules and tokens, wherein the token corresponds to the specified signal and the selected target neural network model;

[0023] - semantically enrich the specified (through the use of neural network algorithms for context analysis and knowledge bases) basic text fragment of the prompt using at least one large language model to form the final text fragment of the prompt;

[0024] - passes the specified final text fragment of the prompt to the selected target neural network model for generating video content.

[0025] Brief description of drawings

[0026] Fig. 1 shows the general architectural block diagram of the claimed system.

[0027] Fig. 2 shows a conceptual diagram of the graphical user interface with camera controls.

[0028] Fig. 3 shows a block diagram of the system operation algorithm (the main stages of the method).

[0029] Fig. 4-6, the interfaces mentioned in the "Background Art" section.

[0030] Table 1. Comparison of syntaxes of different AI models.

[0031] Table 1 is presented as an illustrative example demonstrating the fundamental differences in the syntax of neural network models for the basic set of camera control commands. The list of commands and corresponding tokens is not exhaustive and can be expanded for any of the supported models within the framework of the extensible system architecture described in the description.

[0032] Detailed description of the invention (essence and implementation)

[0033] Terminology and definitions: For the purposes of this description and claims, the following terms have an expanded interpretation:

[0034] • "Neural network technologies" (or "Neural network algorithms for semantic analysis"): For the purposes of this application, this term covers a wide range of artificial intelligence architectures capable of natural language processing and semantic analysis. This includes, but is not limited to, large language models (LLMs) on the Transformer architecture (e.g., GPT, BERT, LLaMA), recurrent neural networks (RNNs), generative adversarial networks (GANs), diffusion models, and any hybrid neural network architectures capable of generating or modifying text descriptions based on input data.

[0035] • "Virtual camera parameters": The term includes not only spatial movement, but also optical characteristics (focal length, aperture, bokeh), lighting parameters (exposure, color temperature), temporal characteristics (speed, interpolation), and stylistic features of camera work (e.g., "camera shake", "cinematic fly-by").

[0036] Description of the system in a static state (Fig. 1)

[0037] The claimed universal adaptive system for controlling the parameters of a virtual camera (hereinafter referred to as the system) is a software and hardware complex that includes at least the following functionally interconnected modules:

[0038] 1. Graphical user interface (GUI) 1 is implemented with a set of visual control elements 1.1 (Fig. 2) for setting conceptual commands (parameters and scenarios) for controlling the virtual camera and its operating modes. Set of visual control elements 1.1 contains a central element that sketches out the camera and visual interactive controls in the form of directional arrows and graphical icons for specifying commands corresponding to spatial movements (e.g., pedestal_up, pedestal_down, horizontal movement: truck_left and truck_right, depth movement: pull_out and pull_in, panning: pan_left and pan_right, tilt: tilt_up and tilt_down, orbital movement: orbit_left and orbit_right, rotation: roll_left and roll_right, zoom_in, zoom_out) and modes (static, follow_object). Moreover, the system architecture is open and provides the ability to dynamically add any other virtual camera control commands through the dynamic architecture expansion mechanism.The visual control set 1.1 can be modified and dynamically generated for new commands using the Semantic Enhancement Module 5 (LLM Enhancer) based on semantic processing of the incoming image (function: provides adaptability to unknown scenarios, impact on TR: increases intuitiveness and versatility, as in the example with "orbit_left" for a forest landscape).

[0039] In one embodiment, the set of visual control elements 1.1 for creating camera movement scenarios is implemented in the form of a timeline interface, allowing the user to place command icons (from set 1.1) at specific time marks, thus setting their sequence and activation duration throughout the entire generated video fragment.

[0040] In one embodiment, the set of visual controls 1.1 may include:

[0041] • Expandable visual elements for changing the camera's optical parameters (e.g. focal length, aperture, depth of field);

[0042] • Lighting control (e.g. exposure, white balance);

[0043] • Time parameters (eg speed, duration);

[0044] • Composition control (e.g. framing, symmetry, placement of objects in the frame, placement of the camera in space);

[0045] • Special camera modes (e.g. static mode, object following mode, autofocus mode, anti-shake mode).

[0046] In one embodiment, to provide additional customization, GUI 1 may include:

[0047] • a separate text field 14 for manual input of tokens and descriptive parameters of camera movement (for example, “very smooth, meditative camera movement” or “jerky and aggressive shaking of the manual camera along the horizontal axis”), which will be taken into account by the semantic enhancement module 5;

[0048] • Input module 15, connected to semantic enhancement module 5 to analyze not only camera commands but also the content to be animated.

[0049] 2. The registry of neural network models 6, implemented, for example, in the form of a drop-down list in the GUI 1, is functionally linked to the command conversion module 2 and is configured to store identifiers (API keys) and syntax rules for various third-party neural network models (e.g., Kling AI, Runway ML, Sora, Qwen), and provides the user with the choice of the target neural network model 8 (they can be conventionally designated, for example, as Model A, Model B, Model C). The list of models is expandable via API for any, including future ones

[0050] 3. Command conversion module 2, which is the core of the system, structurally includes at least three submodules:

[0051] 1) Syntax adaptation module 3 is linked to the rules and tokens database 4, which stores information on the mapping of signals from visual controls 1.1 to the basic text tokens for each model from the neural network model registry 6 (similar to Table 1). It converts an icon click (signal) into a basic token for a specific model, for example, the "Up" signal into the "pedestal_up" token for Runway or "camera move up" for Kling.

[0052] 2) Semantic Enhancement Module 5 (LLM Enhancer), implemented using neural network technologies, such as at least one large language model (LLM). Semantic Enhancement Module 5 is connected via input to Syntax Adaptation Module 3 and via output to Execution Module 7. Semantic Enhancement Module 5 also has a two-way connection with Iterative Learning Module 9 to receive control signals (weight updates). Semantic Enhancement Module 5 implements a sequential process of enriching the base text fragment of the prompt:

[0053] • In the first step, basic camera motion control tokens are extracted from the rules and tokens database 4 according to the syntax of the selected neural network model;

[0054] • In the second stage, a description of the nature of the camera movement is added to these tokens. The sources of context are: the initial data input module 15 (analyzes the reference image or source frame) and the manual input field 14;

[0055] • In the third stage, LLM performs semantic enrichment by integrating the master prompt, extracted tokens, the user's description of the motion pattern, and the semantic analysis of the input image to be animated. This process ensures the formation of the final prompt text fragment with increased predictability of camera behavior (function: generating context-sensitive prompts; impact on the technical result: improving the quality and cinematic quality of the generated video content, as in the example of enriching the "orbit_left" command for a forest landscape by adding parallax and lighting descriptions).

[0056] Example of work: base token "orbit right" + context (forest) --> Final prompt "camera move cinematic orbit right through pine trees, parallax effect".

[0057] 3) Execution Module 7 (API Aggregator) – a module that receives the generated prompt, which is linked to Command Conversion Module 2 and has the ability (e.g., via a network interface and API) to interact (send a request) with the user-selected external target neural network model 8. Execution Module 7 generates video content based on the resulting prompt. Execution Module 7 passes the resulting prompt to the external model, which then returns the generated video.

[0058] Command Conversion Module 2 Operation:

[0059] Receive signal: The user interacts with a visual element (e.g. the Fly Over icon).

[0060] Syntax adaptation: the module accesses the database and finds the token corresponding to the selected model (e.g. for Model A - "orbit right", for Model B - "arc cw").

[0061] • Semantic Enhancement: The base token is passed to Semantic Enhancement Module 5. This module, powered by neural network technologies, analyzes additional inputs (scene context, manual user input, image analysis). The neural network transforms the dry technical token into an enriched prompt. Example: instead of simply "orbit right," the neural network, having analyzed that the scene is a "forest landscape," generates: "smooth cinematic orbit right through pine branches, parallax effect, dynamic lighting."

[0062] 4. Optionally, to implement the additional training function, the system contains:

[0063] Iterative Learning Module 9 provides the Reinforcement Learning from Human Feedback (RLHF) function, and is structurally integrated with:

[0064] • Feedback data collection module (FDC) 10, configured to receive data from the GUI 1 (for example, user ratings “yes / no” or on a scale of 1-5 stars for each parameter).

[0065] • The reward model 11, which evaluates the effectiveness of the prompt, is connected to the data collection module OS 10 and trained on the user's preferences (initial training sources: paired comparisons from the knowledge base; function: evaluation of the effectiveness of the prompt, iterative improvement).

[0066] • Parameter update module 12, connected to reward model 11 and having an output connected to semantic enhancement module 5 to adjust its weights (algorithms: DPO or RWR; function: adjustment based on team ranking, retraining of LLM enhancer).

[0067] Control Module 13 (Agent Management) – Allows the user to save, load, or disable pre-trained versions of the Semantic Enhancement Module 5 as a personalized profile (agent). For managing such independent system components as the Iterative Learning Module 9 (RLHF) and the Semantic Enhancement Module 5 (LLM Enhancer), each of which can be independently activated or deactivated by the user, dedicated interface elements are provided, including a switch for complete deactivation, a weight slider for adjusting the degree of influence on scene processing, a name input field for saving the current version of the LLM Enhancer, and a text field 14 for additional information about the nature of camera movement and operator style.Retrained versions of the LLM enhancer are saved as personalized profiles, referred to as "agents," which can be commercialized as presets of camera styles and placed in the system's shared library for exchange between users (function: enabling batch customization and reuse; increasing the predictability of video content generation; and creating additional commercial value through the ability to monetize personalized models). An "agent" is a stored set of weights or instructions for the semantic enhancement module, adapted to a specific style. The system provides the ability to export and import these profiles via data serialization for exchange between users, enabling the creation of remote libraries of camera styles without revealing the internal data architecture of the profile itself.

[0068] A field for manual input of tokens and descriptive parameters of camera movement, functionally linked to GUI 1 and to Semantic Enhancement Module 5: data from this field is fed directly to LLM to enrich the prompt.

[0069] It's important to note that the proposed system is designed to control virtual camera parameters and complements the user's primary creative intent, expressed in the main text prompt. The system generates prompt fragments related to movement, camera angle, optical parameters, and technical specifications, which are then combined with the user's primary prompt to form the final query for the neural network model.

[0070] Implementation of the method, operation of the system (Fig. 3)

[0071] The system operates in a manner that includes the following steps.

[0072] Stage 1 (initialization) is the stage of preparing the system for operation, at which the user in GUI 1 selects from the register of neural network models 6 the target neural network model 8 with which he intends to work (for example, “Runway ML”).

[0073] Step 2 (selecting the target neural network model): the user in GUI 1 activates (sets) one or more visual elements of the set of visual controls 1.1 (Fig. 2) corresponding to the required commands for controlling the virtual camera, for example, presses the arrow corresponding to the command “pedestal_up”.

[0074] Stage 3 (Syntax Adaptation): GUI 1 transmits a signal (e.g., "cmd_pedestal_up") corresponding to the selected command to command translation module 2. Syntax adaptation module 3, having received this signal and information about the target model ("Runway ML") selected in stage 1 from the registry of neural network models 6, accesses the rule and token database 4. It extracts the corresponding base token from the rule and token database 4, for example, "pedestal_up" for the model selected in stage 1. If "Kling AI" had been selected, syntax adaptation module 3 would extract "camera move up". Thus, syntax adaptation module 3 forms the base text fragment of the prompt. At this stage, the user does not need to know the syntax (i.e., the specific token "pedestal_up"), since syntax adaptation module 3 does this for him.

[0075] Stage 4 (semantic enrichment): The neural network analyzes the basic text fragment of the prompt (containing the token "pedestal_up") and the context (image, text) to create a detailed, cinematic description, which is fed to Semantic Enhancement Module 5 (LLM Enhancer). Semantic Enhancement Module 5, being an LLM, analyzes this basic prompt and performs multi-stage enrichment.

[0076] It combines:

[0077] 1. Specified base prompt("pedestal_up");

[0078] 2. Optional tokens from the manual input field 14 descriptive parameters of camera movement (e.g. "smooth movement");

[0079] 3. Contextual information about the scene from Input Module 15, obtained either from the main prompt / reference image or by semantic analysis of the image to be animated itself (e.g., "forest landscape").

[0080] Based on this comprehensive analysis, Semantic Enhancement Module 5 semantically enriches the command, generating a final prompt text fragment, such as: "pedestal_up, cinematic, smooth camera movement through forest canopy".

[0081] Stage 5 (Execution): The final prompt is passed to execution module 7, which forwards it via API to target neural network model 8 ("Runway ML"). Target neural network model 8 generates video content and returns it to the user.

[0082] Optional steps 6-9 – iterative learning mode, the inclusion of which is decided by the user:

[0083] Step 6 (OS data collection): The system displays the generated video, and the user views the result and, through the OS data collection module 10 in GUI 1, rates the parameters involved in the generation (e.g., binary “yes / no” or a “1-5 star” scale for each parameter; function: generating team ratings).

[0084] Stage 7 (Reward Evaluation): The OS data is fed to reward model 11. This model, pre-trained on pairwise comparisons, evaluates how “good” the resulting prompt text fragment generated in stage 4 was, since it resulted in such a user rating (feature: prioritizing effective commands).

[0085] Stage 8 (retraining): Parameter update module 12, using the estimate from reward model 11 and algorithms such as direct preference optimization (DPO) or weighted reward regression (RWR), calculates the gradient and adjusts (retrains) the weights of the large language model of semantic boosting module 5 (function: preference adaptation).

[0086] Stage 9 (Personalization): The updated Semantic Enhancement Module 5 will generate semantically enriched prompts that better match the user's preferences in the future, for example, by using "smooth camera movement" more often if that token has received high ratings. The user can also save this trained version of Semantic Enhancement Module 5 as an "agent" via Management Module 13 for future projects, including leasing "camera styles" (function: commercialization, batch customization).

[0087] Thus, the stated technical result of "simplification and unification of interaction" is achieved through a GUI (eliminating textual terminology entry), a registry of neural network models connected to the system, and a syntax adaptation module (eliminating the need to learn and memorize the syntax of different models). The user interacts with a single unified GUI, and the adaptation module automatically inserts the correct base tokens (commands) for any selected neural network model, eliminating the need to study Table 1.

[0088] The technical result of "improved quality and personalization" is achieved through the use of Semantic Enhancement Module 5 (LLM Enhancer) and (optionally) Iterative Learning Module 9. Semantic Enhancement Module 5 enriches "dry" basic tokens (e.g., "camera move up") with context (e.g., "...smoothly moves upward, flying through the branches of the spruce trees..."), making the prompt more effective and enhancing the predictability and cinematic quality of each individual result in accordance with the creative task regarding camerawork. This result is additionally enhanced by Iterative Learning Module 9, which (if present) analyzes user ratings ("yes / no" or a "1-5 stars" scale for each parameter) and further trains the LLM Enhancer, which "remembers" successful semantic constructions, thereby personalizing the generation to the individual style and preferences of the user in each individual project.

[0089] The technical result of "enabling customization and further training" is achieved through the iterative training module 9, which allows the user to forcefully influence the weights of the LLM Enhancer, save its trained versions ("agents") and use them in other projects or transfer them to other users as presets of operator styles in the library.

Claims

1. A universal adaptive system for controlling the parameters of a virtual camera in AI video generation systems, comprising a graphical user interface (GUI) with at least one visual control element for specifying at least one command to control the virtual camera and an execution module for generating video content, characterized in that the system additionally contains at least a register of neural network models, designed to allow the user to select a target neural network model; a command conversion module configured to receive a signal corresponding to at least one command from at least one visual GUI element and automatically convert said command into at least one text fragment of a prompt based on said signal, wherein the command conversion module comprises: - a syntax adaptation module containing a database of rules and tokens and configured to extract from said database at least one token corresponding to said signal and the selected target neural network model, and to form on its basis a basic text fragment of the prompt; - and a semantic enhancement module, implemented on the basis of neural network technologies, configured to receive the specified basic text fragment of the prompt and its semantic enrichment by adding context-dependent descriptions to form the final text fragment of the prompt.

2. The system according to claim 1, characterized in that the visual control elements correspond to commands selected from a group including: spatial movement of the camera, changing optical parameters, controlling scene lighting, controlling frame composition parameters.

3. The system according to claim 1, characterized in that the semantic enhancement module is additionally configured to obtain contextual information about the scene from at least one source selected from the group: a reference image, the user's main text prompt, data from the semantic analysis of the image to be animated, and use this information in semantic enrichment.

4. The system according to claim 1, characterized in that the GUI additionally contains a separate text field for manual input by the user of descriptive parameters or tokens, wherein the semantic enhancement module is configured to use said descriptions or tokens when generating the final text fragment of the prompt.

5. The system according to claim 1, characterized in that the GUI additionally contains graphical elements for controlling camera movement scenarios, allowing the user to set the sequence and duration of activation of several camera control commands within the generated video.

6. The system according to claim 1, characterized in that it further comprises an iterative training module configured to receive feedback data from the user on the quality of the video content generated based on the final prompt and to use this data for further training of at least the semantic enhancement module, wherein the iterative training module comprises: a module for collecting feedback data for the user to provide quality assessments; a reward model trained on the user's preferences; and a parameter updating module configured to adjust the parameters of the semantic enhancement module based on the assessments from the reward model.

7. The system according to claim 6, characterized in that the iterative training module additionally contains a control module configured to save the retrained version of the semantic enhancement module in the form of a personalized profile, as well as export and import said profile for exchange between users.

8. The system according to claim 1, characterized in that the semantic enhancement module is implemented in the form of at least one large language model.

9. A method for controlling the parameters of a virtual camera in an AI video generation system, comprising the user specifying at least one command to control the parameters of a virtual camera through a GUI and generating video content, characterized in that the following steps are additionally performed: provide the GUI user with a register of neural network models and obtain a choice of the target neural network model; receive a signal from a visual GUI element corresponding to a user command; form the specified command into the final text fragment of the prompt based on the specified signal by at least the following sub-steps: forming a basic text fragment of the prompt by extracting at least one token from a database of rules and tokens, wherein the token corresponds to the specified signal and the selected target neural network model; semantically enrich the base text fragment using a semantic enhancement module based on neural network technologies to form the final text fragment of the prompt; pass the final text fragment of the prompt to the selected target neural network model to generate video content.

10. The method according to paragraph 9, characterized in that it additionally includes the steps of: receive feedback from the user about the quality of the generated video content; evaluate the effectiveness of the specified final text fragment of the prompt using a reward model trained on the user's preferences; and adjust the parameters of the semantic enhancement module based on the estimates from the reward model.