Prompt word optimization method, system and equipment for video generation and medium
By integrating multi-source information and combining spatial matching and reinforcement learning, accurate video generation prompt words are generated, which solves the problem of insufficient information integration in the existing technology, and achieves high-quality and accurate video generation to meet users' needs for high-quality videos.
Patent Information
- Application Number
- CN202510519510.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
AI Technical Summary
The existing video generation technology cannot effectively integrate multi-source information such as pictures uploaded by users, original prompt words, and motion pictures drawn by users, resulting in deviations from users' expectations, insufficient fineness, insufficient dynamic effects and details, and poor consistency between the input information, resulting in inconsistency or contradictions in the logic and visual aspects of the generated video.
By integrating multi-source information such as images, semantic text and motion maps, combined with spatial matching, reinforcement learning, physical model analysis and other technologies, an accurate prompt word containing scene adaptation parameters is generated, specifically including the user drawing motion maps of the dynamic mask layer and motion trajectory vector layer in the image editing interface, conducting thermal map analysis of contradictory areas, using reinforcement learning strategies for intention analysis, and reconstructing the action sequence through parameterized physical models to analyze environmental constraints and kinematic decoupling algorithms, and finally generating motion description parameters that conform to the domain characteristics.
It significantly improves the quality and accuracy of video generation, ensures that the generated video content is highly consistent with user expectations, and performs excellent dynamic effects and details, eliminates logical and visual contradictions, improves the overall quality and viewing of the video, and meets users' needs for high-quality videos.
Smart Images

Figure CN120449439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a prompt word optimization method, system, device and medium for video generation. Background Art
[0002] With the rapid development of artificial intelligence (AI), AI-powered image-based video technology has become a crucial tool in content creation. This technology can automatically generate dynamic video content based on user-provided raw images and simple prompts, significantly enriching the means and forms of content creation. In recent years, the rise of multimodal Large Language Model (LLM) prompt assistants has further advanced video generation technology, enabling users to quickly generate corresponding video content by inputting prompts and / or images.
[0003] However, with the continuous advancement of video generation technology, users are demanding higher levels of sophistication and control over generated videos. To meet these demands, cutting-edge technologies allow users to more precisely describe desired video motion by drawing motion regions and trajectories using motion graphs. By allowing users to draw dynamic regions and motion trajectories on input images, these technologies enable more precise control over the video generation process, resulting in video content that better meets user expectations.
[0004] While these new technologies offer users greater creative freedom, existing prompt word generation assistants fail to effectively integrate this complex input information. Specifically, existing prompt word generation assistants primarily focus on generating video content based on user-entered prompt words and / or images, while neglecting the user-drawn motion map as a crucial source of information. This can lead to inconsistencies between the input map, motion information, and prompt words during the video generation process, ultimately resulting in the generated video quality and accuracy failing to meet the high standards of users.
[0005] Specifically, the existing prompt word generation assistant has the following shortcomings:
[0006] 1. Insufficient information integration: The system cannot effectively integrate multiple sources of information, such as user-uploaded pictures, original prompt words, and user-drawn motion diagrams, resulting in a deviation between the generated video content and user expectations.
[0007] 2. Insufficient precision: Due to the lack of fine control of motion graphics, the generated videos often fail to meet user expectations in terms of dynamic effects and details.
[0008] 3. Poor consistency: Inconsistency between the input image, motion information, and prompt words may cause the generated video to be logically and visually incoherent or contradictory. Summary of the Invention
[0009] The purpose of the present invention is to provide a prompt word optimization method, system, device and medium for video generation. By integrating multi-source information such as images, semantic text and motion graphs, and combining technologies such as spatial matching, reinforcement learning, and physical model analysis, accurate prompt words containing scene adaptation parameters are generated, which significantly improves the quality and accuracy of video generation, thereby solving at least one of the above-mentioned existing technical problems.
[0010] In a first aspect, the present invention provides a method for optimizing prompt words for video generation, the method specifically comprising:
[0011] After the user uploads the initial image, the system receives the semantic description text entered by the user and guides the user to draw a motion map including a dynamic mask layer and a motion trajectory vector layer in the image editing interface;
[0012] Spatial matching is performed between the verb elements in the semantic description text and the motion trajectories in the motion map, and an attention mechanism is used to generate a heat map of the conflicting areas.
[0013] When conflicting areas are detected in the conflict area heat map, the reinforcement learning strategy is used to analyze the intention and generate a multi-dimensional optimization solution that includes trajectory direction correction, speed classification, and physical constraint prompts.
[0014] According to different scenario types, environmental constraints are analyzed through parameterized physical models, action sequences are reconstructed based on kinematic decoupling algorithms, and structural failure propagation simulations are performed using continuum mechanics models to generate motion description parameters that meet domain characteristics.
[0015] The motion graph is corrected according to the multi-dimensional optimization scheme and motion description parameters, and the corrected motion graph is iteratively verified. If the verification passes, the initial image and multiple sets of motion graphs are input into the multimodal large language model to generate the final prompt word containing the scene adaptation parameters.
[0016] In a second aspect, the present invention provides a prompt word optimization system for video generation, the system specifically comprising:
[0017] The first generation module is used to receive semantic description text input by the user after the user uploads the initial image, and guide the user to draw a motion map including a dynamic mask layer and a motion trajectory vector layer in the image editing interface;
[0018] The second generation module is used to spatially match the verb elements in the semantic description text with the motion trajectories in the motion map, and generate a heat map of the conflict area by combining the attention mechanism;
[0019] The third generation module is used to analyze the intentions of the conflicting areas in the conflict area heat map through reinforcement learning strategies, and generate a multi-dimensional optimization plan including trajectory direction correction, speed classification and physical constraint prompts;
[0020] The fourth generation module is used to analyze environmental constraints based on different scenario types through parameterized physical models, reconstruct action sequences based on kinematic decoupling algorithms, and simulate structural failure propagation using a continuum mechanics model to generate motion description parameters that meet domain characteristics.
[0021] The fifth generation module is used to modify the motion graph according to the multi-dimensional optimization scheme and motion description parameters, and iteratively verify the modified motion graph. If the verification passes, the initial image and multiple sets of motion graphs are input into the multimodal large language model to generate the final prompt word containing the scene adaptation parameters.
[0022] In a third aspect, the present invention provides a computer device comprising: a memory and a processor and a computer program stored in the memory, wherein when the computer program is executed on the processor, the prompt word optimization method for video generation as described in any one of the above methods is implemented.
[0023] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for optimizing prompt words for video generation as described in any one of the above methods is implemented.
[0024] Compared with the prior art, the present invention has at least one of the following technical effects:
[0025] 1. This invention integrates multi-source information such as images, semantic text and motion graphics, and combines technologies such as spatial matching, reinforcement learning, and physical model analysis to generate precise prompt words containing scene adaptation parameters, significantly improving the quality and accuracy of video generation.
[0026] 2. The present invention ensures that the generated video content is highly consistent with user expectations by integrating multi-source information such as pictures uploaded by users, original prompt words, and motion pictures drawn by users.
[0027] 3. The present invention utilizes the motion graph drawn by the user to finely control the video generation process, so that the generated video has better dynamic effects and details, meeting the user's demand for high-quality video.
[0028] 4. The present invention eliminates logical and visual contradictions in the generated video by ensuring the consistency among the input image, motion information and prompt words, thereby improving the overall quality and viewing experience of the video.
[0029] 5. The present invention achieves accurate drawing of motion graphs by dynamically drawing guide frames, real-time recording of vertex coordinates and generating dynamic mask layers using a second-order B-spline interpolation algorithm, and constructing motion trajectory vector layers using a cubic B-spline curve fitting algorithm, providing a reliable foundation for subsequent optimization.
[0030] 6. The present invention uses a natural language processing model to extract verb elements, combines it with an attention matching model to calculate the regional matching degree, and generates a conflict area heat map through a conflict intensity amplification factor, effectively identifying potential conflicts between semantic descriptions and motion trajectories.
[0031] 7. The present invention uses an attention matching model to comprehensively consider factors such as the attention weight of verb elements, direction encoding vectors, and trajectory direction change rates to accurately calculate the regional matching degree, providing a scientific basis for the generation of conflict area heat maps.
[0032] 8. The present invention uses reinforcement learning strategies to perform intent analysis and generate a multi-dimensional optimization solution that includes trajectory direction correction, speed classification, and physical constraint prompts, effectively resolving the conflict between semantic description and motion trajectory.
[0033] 9. According to different scene types, the present invention analyzes environmental constraints through parameterized physical models, reconstructs action sequences, and simulates the propagation of structural failures to generate motion description parameters that conform to domain characteristics, ensuring the authenticity and rationality of video generation.
[0034] 10. The present invention modifies the motion graph based on a multi-dimensional optimization scheme and motion description parameters, and performs iterative verification through a spatiotemporal consistency verification function, ultimately generating accurate prompt words containing scene adaptation parameters, significantly improving the quality of video generation and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0036] Figure 1 This is a flow chart of a method for optimizing prompt words for video generation provided by one embodiment of the present invention;
[0037] Figure 2 This is a structural diagram of a prompt word optimization system for video generation provided by one embodiment of the present invention;
[0038] Figure 3 It is a structural diagram of a computer device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0039] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0040] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0041] It will also be understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0042] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0043] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0044] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.
[0045] In the embodiments of the present application, the execution subject of the process includes a terminal device, which includes but is not limited to: a server, a computer, a smart phone, a tablet computer, and other devices capable of executing the method disclosed in the present application. Figure 1 A flow chart of a prompt word optimization method for video generation disclosed in the first embodiment of the present invention is shown, and is described in detail as follows:
[0046] S101 , after a user uploads an initial image, a semantic description text input by the user is received, and the user is guided to draw a motion image including a dynamic mask layer and a motion trajectory vector layer in an image editing interface.
[0047] In this embodiment, in existing AI image-generated video technology, users usually generate video content by uploading an initial image and inputting semantic description text (i.e., prompt words). However, as users' demand for the precision and control of video generation continues to increase, relying solely on prompt words and image input can no longer meet the requirements of high-quality video generation. To solve this problem, this embodiment guides users to draw a motion map containing a dynamic mask layer and a motion trajectory vector layer to enhance user control and generation effect during the video generation process.
[0048] Specifically, the video generation system's user interface includes an "Upload Image" button. Clicking this button prompts a file selection dialog box, allowing the user to select and upload an initial image from their local device. The system supports common image formats, such as JPEG, PNG, and BMP, ensuring that uploaded images can be correctly recognized and processed. Before entering subsequent processing steps, uploaded images undergo preprocessing, including but not limited to image resizing and color space conversion, to ensure that image quality meets subsequent processing requirements.
[0049] After the user uploads an image, the system displays a text input box, prompting the user to enter a semantically descriptive text related to the image content. The system has no specific requirements for the text format, but recommends using concise and clear sentences to describe the desired video content, such as "a person moves from left to right" or "an object rotates and gradually grows larger." The system preprocesses the user's text, including word segmentation and part-of-speech tagging, to facilitate subsequent spatial matching with the motion image.
[0050] After the user enters the semantic description text, the system automatically switches to the image editing interface, which includes the display area of the initial image and the toolbar for drawing motion diagrams.
[0051] The "Draw Mask" tool is available in the toolbar. Clicking it allows you to draw a dynamic mask layer on the original image. Using a mouse or stylus, you can draw a closed or open mask area on the image, which will define the portion of the video that needs to be dynamically changed. After drawing the mask, the system allows you to set its properties, such as transparency and color, to more clearly demonstrate the dynamic effect in subsequent video generation.
[0052] The "Draw Track" tool is available in the toolbar. Clicking this tool allows you to draw a motion track vector layer on a mask layer. Using a mouse or stylus, you can draw a motion track on the mask layer, which will define the movement of objects within the masked area. After drawing the track, you can set its properties, such as speed, acceleration, and direction, to more precisely control the object's movement in subsequent video generation.
[0053] Once the user has completed the motion map, the system saves the map (including the dynamic mask layer and the motion trajectory vector layer) along with the initial image as input data for subsequent video generation. The system then verifies the saved motion map to ensure that it is correctly formatted, has appropriate attributes, and matches the semantic description text entered by the user. If verification fails, the system prompts the user to redraw or adjust the motion map.
[0054] In this embodiment, after uploading an initial image and entering semantic description text, users can more accurately describe the desired video motion by drawing a motion map consisting of a dynamic mask layer and a motion trajectory vector layer. This method not only enhances user control over the video generation process, but also improves the quality and accuracy of the generated video, meeting user demands for high-quality video content.
[0055] S102, spatially match the verb elements in the semantic description text with the motion trajectories in the motion map, and generate a heat map of the conflict area by combining the attention mechanism.
[0056] In this embodiment, during the video generation process, the semantic description text (prompt words) provided by the user and the drawn motion map are two important sources of input information. However, due to the differences in expression and information granularity between the two, directly combining them for video generation may lead to inconsistencies in the input information, thereby affecting the quality and accuracy of the generated video. To solve this problem, this embodiment spatially matches the verb elements in the semantic description text with the motion trajectories in the motion map, and combines the attention mechanism to generate a heat map of conflicting areas, so as to discover and correct inconsistencies in the input information before video generation.
[0057] Specifically, first, the semantic description text input by the user is preprocessed, including steps such as word segmentation and part-of-speech tagging, so as to accurately identify the verb elements in the text. Based on the part-of-speech tagging results, the system identifies the verbs in the text and extracts them as key elements for video generation. For example, in "the character moves from left to right", the verb "move" is identified. In order to facilitate subsequent matching with the motion trajectory, the extracted verb elements are normalized, such as unifying the tense and voice of the verbs. The motion graph drawn by the user is preprocessed, including steps such as image denoising and edge detection, so as to accurately extract the motion trajectory. Based on the preprocessed motion graph, the system extracts the features of the motion trajectory, such as the starting point, end point, direction, speed, etc. of the trajectory. The extracted trajectory features are converted into vector form for subsequent spatial matching with the verb elements.
[0058] A spatial mapping relationship is established between the verb element and the motion trajectory. This means mapping the motion intent described by the verb element to the specific trajectory in the motion graph. For example, "move" is mapped to a trajectory from left to right. Based on the spatial mapping relationship, the matching degree between the verb element and the motion trajectory is calculated. The matching degree can be achieved by comparing the motion characteristics described by the verb element with the actual characteristics of the motion trajectory. Based on the matching degree calculation results, the degree of matching between the verb element and the motion trajectory is evaluated. If the matching degree is lower than the preset threshold, it is considered that there is inconsistency.
[0059] In order to more accurately detect inconsistencies in the input information, this embodiment introduces an attention mechanism. The attention mechanism can dynamically focus on important parts of the input information and assign different attention weights according to their importance. Based on the attention mechanism, the system identifies areas where there are inconsistencies between verb elements and motion trajectories, namely, contradictory areas. These areas may manifest as a mismatch between the movement intention described by the verb element and the actual characteristics of the motion trajectory. The identified contradictory areas are visualized in the form of a heat map. Different colors or brightness in the heat map represent different degrees of contradictory areas, so that users can intuitively understand the inconsistencies in the input information.
[0060] The generated heat map of conflicting areas is analyzed to identify the specific causes of inconsistencies. For example, this could be due to inaccurate descriptions of verb elements or imprecise motion trajectory drawing. Based on the conflicting area analysis results, the system generates corresponding correction suggestions. These correction suggestions may include modifying the description of verb elements or adjusting the motion trajectory drawing. The correction suggestions are fed back to the user, and the input information is revised based on the user's feedback. The corrected input information will serve as the final input for subsequent video generation.
[0061] In this embodiment, inconsistencies in the input information can be effectively discovered and corrected, improving the quality and accuracy of the generated video. Furthermore, by visually displaying a heat map of conflicting areas, users can more intuitively understand the problems in the input information and make corresponding adjustments based on the correction suggestions.
[0062] S103, when a conflict area is detected in the conflict area heat map, intention analysis is performed through a reinforcement learning strategy to generate a multi-dimensional optimization solution including trajectory direction correction, speed classification, and physical constraint prompts.
[0063] In this embodiment, during video generation or animation production, conflicting areas between the semantic description text and the motion trajectory map may cause the generated content to be inconsistent with the user's intent. When conflicting areas are detected in the conflict area heat map, an intelligent method is needed to analyze and correct the conflict. This embodiment improves the quality and accuracy of video generation by generating a multi-dimensional optimization solution that includes trajectory direction correction, speed grading, and physical constraint prompts.
[0064] Specifically, the generated conflict region heat map is first analyzed to identify conflicting areas. Conflicting areas typically appear as areas with unusual colors or excessive brightness in the heat map, indicating a significant inconsistency between the semantic description text and the motion trajectory map. Features of the conflicting areas are then extracted, including the location, size, and intensity of the conflict. These features serve as input to the subsequent reinforcement learning strategy.
[0065] Define the state space for reinforcement learning, including the characteristics of the current conflict area, the verb elements of the semantic description text, the characteristics of the motion trajectory, etc. The state space needs to contain enough information so that the reinforcement learning agent can fully understand the current conflict situation.
[0066] Define the action space for reinforcement learning, including operations such as trajectory direction correction, speed adjustment, and physical constraints. The action space should cover all possible optimization scenarios so that the agent can choose the best correction strategy.
[0067] Design a reward function to evaluate the effectiveness of the agent's actions in resolving the conflict. This reward function can be based on metrics such as the reduction of the conflict area and the alignment of the generated content with the user's intent.
[0068] Initialize the reinforcement learning agent, including its policy network, value network, and other parameters. The agent must possess a certain level of learning ability to continuously optimize its policy during training. Prepare training data, including a large number of heat maps of conflicting areas and their corresponding optimization solutions. The training data should be diverse so that the agent can learn the optimal correction strategy for various conflict situations. Input the training data into the reinforcement learning environment. The agent selects actions based on its current state and receives rewards from the environment. Through continuous iterative training, the agent gradually learns the optimal correction strategy, namely how to choose the most appropriate trajectory direction correction, speed classification, and physical constraint prompts for different conflict situations.
[0069] When a conflicting area is detected in a new conflict region heatmap, the features of the current conflicting area are fed into a trained reinforcement learning agent. Based on its learned policy, the agent performs intent analysis on the conflicting area to understand the user's true intent. Based on the intent analysis results, the agent generates a multi-dimensional optimization plan that includes trajectory direction correction, speed grading, and physical constraint hints. For example, the agent might recommend changing the direction of a trajectory segment from horizontal to vertical or adjusting the speed of a trajectory segment to match the verb elements in the semantic description text. The generated optimization plan is evaluated to ensure that it effectively resolves the current conflict and improves the alignment of the generated content with the user's intent. The evaluation results are fed back to the agent for further optimization of its strategy. The generated optimization plan is applied to the original motion trajectory map or semantic description text for correction and adjustment. The conflict region heatmap is regenerated to verify whether the corrected input information still contains conflicts. If the conflict is resolved, the optimization plan is effective; otherwise, further adjustments to the optimization plan or re-intention analysis are required.
[0070] In this example, a reinforcement learning strategy is used to automatically learn the optimal correction strategy, eliminating the need for human intervention. The optimization plan encompasses multiple dimensions, including trajectory direction, speed classification, and physical constraints, enabling comprehensive conflict resolution. Through intent analysis and optimization plan generation, conflicts can be quickly located and resolved, improving the efficiency of video generation or animation production.
[0071] S104, according to different scene types, analyze environmental constraints through parameterized physical models, reconstruct action sequences based on kinematic decoupling algorithms, and use continuous medium mechanics models to simulate structural failure propagation to generate motion description parameters that meet domain characteristics.
[0072] In this embodiment, existing video generation technologies use user-drawn motion diagrams (such as dynamic mask layers and motion trajectories) as a key source of information for describing desired motion patterns. However, existing prompt word generation assistants fail to effectively integrate motion diagrams with scene type information, resulting in the generated videos failing to meet user requirements in terms of dynamic effects, detailed expression, and logical consistency. To address this issue, a technical solution is needed that can dynamically analyze environmental constraints based on scene type, reconstruct motion sequences, and simulate the propagation of structural failures.
[0073] Specifically, the scene types are classified according to the initial images and semantic description texts uploaded by users. For example, natural scenes include forests, rivers, mountains, etc.; urban scenes include streets, buildings, transportation, etc.; and science fiction scenes include laboratories, alien environments, and future cities.
[0074] Corresponding parametric physical models are constructed for different scene types. For example, for natural scenes, environmental parameters such as wind, gravity, and water flow are introduced; for urban scenes, parameters such as traffic rules, building structure, and crowd density are introduced; for science fiction scenes, parameters such as anti-gravity, energy field, and light and shadow effects are introduced.
[0075] Decompose the motion trajectory drawn by the user into multiple independent action units (such as "jump", "rotation", and "translation"), and analyze the spatiotemporal relationship between each action unit.
[0076] Map environmental parameters in the parametric physics model to action units. For example, in the "Forest" scene, introduce a "wind influence parameter" to the "Jump" action unit to adjust the offset of the jump trajectory. In the "Laboratory" scene, introduce an "energy field constraint parameter" to the "Rotation" action unit to limit the rotation range.
[0077] Based on the decoupled action units and scene constraints, we reconstruct action sequences that conform to physical laws. For example, in a "street" scene, if the user draws a "vehicle driving" trajectory, the curvature and speed of the trajectory are adjusted according to traffic regulation parameters. In an "alien environment" scene, if the user draws a "hover movement" trajectory, the trajectory's hovering height and stability are adjusted according to anti-gravity parameters.
[0078] Define the key points in the motion diagram that may cause structural failure (such as "object collision points" and "trajectory fracture points"). Use the continuous medium mechanics model to simulate the propagation process of structural failure in the scene, specifically including: material parameter mapping: map the material of objects in the scene (such as "metal", "concrete", "glass") to the mechanics model, and define its elastic modulus, fracture toughness and other parameters. Failure mode analysis: analyze the propagation paths of different failure modes (such as "bending failure", "shear failure", and "tensile failure") in the scene. For example, in the "building" scene, if the motion trajectory drawn by the user causes a "wall collision", the mechanical model is used to simulate the crack propagation and collapse process of the wall under the action of the collision force. In the "mechanical device" scene, if the motion trajectory drawn by the user causes a "gear jam", the mechanical model is used to simulate the impact of gear failure on the overall device.
[0079] Based on the failure propagation simulation results, adjust the trajectory parameters in the motion diagram to avoid structural failure. For example, if the simulation results indicate a "wall collapse," adjust the motion trajectory to avoid the wall. If the simulation results indicate a "gear jam," adjust the motion trajectory to reduce friction between the gears.
[0080] Integrate the scene type, action sequence reconstruction results and failure propagation simulation results to generate motion description parameters that meet the domain characteristics. The generated motion description parameters are converted into natural language prompts, such as: "Scene type: forest, wind impact: wind direction southeast, wind speed 5m / s", "Action sequence: jump, wind offset: 0.5m".
[0081] In this embodiment, a parameterized physical model is used to ensure that motion description parameters are highly aligned with the scene type. Action sequence reconstruction and failure propagation simulation are used to improve the precision of motion description parameters. Multi-source information integration is used to eliminate inconsistencies between input images, motion information, and prompt words.
[0082] S105: The motion graph is modified according to the multi-dimensional optimization scheme and motion description parameters, and the modified motion graph is iteratively verified. If the verification passes, the initial image and multiple sets of motion graphs are input into the multimodal large language model to generate the final prompt word containing the scene adaptation parameters.
[0083] In this embodiment, the generation of motion graphs (e.g., character motion trajectories and object motion paths) in animation, virtual reality, game development, and other fields must balance physical plausibility, user intent, and scene adaptability. Traditional methods rely on single models or manual adjustments, making it difficult to handle the multi-dimensional constraints in complex scenes (e.g., conflicting motion trajectories, speed mismatches, and scene parameter conflicts), resulting in inconsistencies between the generated motion graph and the scene.
[0084] Specifically, it receives a multi-dimensional optimization plan, which includes the following three types of parameters: trajectory direction correction parameters, such as path smoothness and direction constraints (to avoid collisions and comply with physical laws); speed classification parameters, such as acceleration curves and speed thresholds (to avoid speeding or stalling); physical constraint prompts, such as friction, gravity, air resistance and other environmental parameters.
[0085] The initial motion graph is modified according to the optimization scheme. (1) Trajectory correction: The motion trajectory is adjusted to conform to the trajectory direction correction parameters, such as smoothing the trajectory through spline interpolation or recalculating the trajectory based on the physics engine. (2) Speed adjustment: The speed curve in the motion graph is adjusted according to the speed grading parameters, such as inserting acceleration transition segments between keyframes. (3) Physical constraint application: Physical constraint hints are mapped to the motion graph, such as introducing gravity parameters in the simulation to make the object motion conform to the real physical laws.
[0086] Verification metrics are defined to evaluate the corrected motion graph: (1) physical rationality: such as whether the trajectory conforms to Newtonian mechanics and whether the velocity is continuous; (2) scene adaptability: such as whether the motion graph is coordinated with other elements in the scene (such as obstacles and lighting); and (3) semantic consistency: such as whether the motion graph conforms to the description input by the user (such as "slow movement" and "fast jump").
[0087] The corrected motion graph is fed into the physics simulation engine to check for physical plausibility. The motion graph is then overlaid with the scene model to check for collisions, occlusions, and other issues. A natural language processing (NLP) model is used to calculate semantic similarity between the motion graph description and the user input. If validation fails, the optimization solution parameters are adjusted based on the failure criteria (e.g., increasing trajectory smoothness, lowering the velocity threshold), and the motion graph is re-corrected.
[0088] After verification, the initial image, multiple sets of motion images and user input descriptions are input into the multimodal large language model. The multimodal large language model integrates image, motion image and text information to understand the semantics and physical characteristics of the scene. Based on the understanding results, it generates scene adaptation parameters, such as lighting parameters: such as light intensity, color temperature, and direction (based on the interactive analysis of images and motion images), material parameters: such as reflectivity and roughness (based on collision information in motion images), and special effects parameters: such as particle effects and shadow effects (based on motion trajectories and scene descriptions).
[0089] The generated scene adaptation parameters are converted into natural language prompts, such as: "Scene lighting: warm colors, light intensity 500 lux, direction of illumination from 45 degrees from the upper left", "Material: the ground is rough concrete with a reflectivity of 0.3; the wall is smooth metal with a reflectivity of 0.8".
[0090] In this embodiment, the rationality of the motion graph is improved through comprehensive optimization of trajectory, speed, and physical constraints. An iterative verification process ensures a high degree of adaptability between the motion graph and the scene. A multimodal large language model is used to achieve semantic unification of images, motion graphs, and text, generating more accurate scene adaptation parameters.
[0091] In some embodiments, in the above step S101, after the user uploads the initial image, receiving the semantic description text input by the user, and guiding the user to draw a motion image including a dynamic mask layer and a motion trajectory vector layer in the image editing interface, specifically includes:
[0092] After receiving the initial image uploaded by the user, the user's input semantic description text is received through the human-computer interaction interface, and a dynamic drawing guide frame is generated on the image editing interface based on the semantic description text;
[0093] When a user touch operation is detected, the polygon vertex coordinate sequence is recorded in real time, and a second-order B-spline interpolation algorithm is used to generate a dynamic mask layer with gradual transparency;
[0094] After the mask layer is drawn, the trajectory anchor point set input by the user is captured, and the motion trajectory vector layer is constructed using the cubic B-spline curve fitting algorithm.
[0095] In this embodiment, existing video generation technologies require users to manually draw complex motion diagrams (such as dynamic mask layers and motion trajectories), but lack semantic guidance and intelligent assistance, resulting in low drawing efficiency and insufficient accuracy. To address this issue, an interactive drawing method that combines user semantic descriptions with intelligent algorithms is needed.
[0096] Specifically, the user uploads an initial image through a human-computer interaction interface (such as a touch screen, mouse, or voice input) and enters a semantic description text (such as "a person runs from left to right"). The system uses natural language processing (NLP) technology to parse the semantic description text, extracting key action information (such as "running"), direction information (such as "from left to right"), and target area information (such as "person").
[0097] Based on the analysis results, a dynamic drawing guide frame is generated in the image editing interface. For example, the action guide frame: the dynamic area that the user needs to draw (such as "character outline") is marked with a dotted frame, the direction guide line: the movement direction is marked with an arrow (such as "from left to right"), and the anchor point prompt: the preset anchor point position (such as "character joint point") is displayed around the target area to assist the user in positioning.
[0098] The user draws a polygonal mask area within the guide frame using touch gestures (such as sliding a finger or dragging a mouse). The system records the coordinate sequence of the polygon vertices along the user's touch path (e.g., [(x1, y1), (x2, y2), ..., (xn, yn)]) in real time. A second-order B-spline interpolation algorithm is used to smooth the vertex sequence, generating a continuous mask boundary curve. The transparency of the mask layer is dynamically adjusted based on the distance from the mask boundary to the center of the target area, for example, lower transparency near the boundary and higher transparency near the center.
[0099] After the user completes the mask layer drawing, they can enter a set of trajectory anchor points (e.g., "character joint motion trajectory") around the target area by clicking or dragging. The system records the anchor point coordinate sequence (e.g., [(x1, y1), (x2, y2), ..., (xm, ym)]) in real time. A cubic B-spline curve fitting algorithm is used to smooth the anchor point sequence, generating a continuous motion trajectory curve. The trajectory curve is then converted into a vector layer format (e.g., SVG path data), with direction arrows and speed annotations added.
[0100] During the drawing process, the system generates a real-time preview of the mask layer and the trajectory vector layer overlay for user verification. If the user finds any deviations, they can correct them by dragging anchor points, adjusting vertices, or re-entering the semantic description text. The system records the user's corrections and updates the semantic description and drawing guide based on the results, achieving iterative optimization.
[0101] In this embodiment, a drawing guide frame is automatically generated by parsing the user's semantic description, reducing the drawing difficulty. A B-spline interpolation algorithm is used to generate a smooth mask layer and trajectory vector layer, enhancing the visual effect. Dynamically adjust the transparency based on the position of the mask layer to enhance the realism of the dynamic effect.
[0102] In some embodiments, in the above step S102, spatially matching the verb elements in the semantic description text with the motion trajectories in the motion map and generating a conflict area heat map in combination with the attention mechanism specifically includes:
[0103] The natural language processing model is used to parse the semantic description text entered by the user, extract the verb element set, and generate the corresponding direction encoding vector;
[0104] Based on the continuous direction field in the motion trajectory vector layer, equidistant sampling is performed in the image space domain to obtain a discrete direction vector set;
[0105] When it is detected that the direction encoding vector is spatially associated with the discrete direction vector set, the attention matching model is called to calculate the regional matching degree;
[0106] The regional matching degree is transformed nonlinearly by the contradiction intensity amplification factor to obtain the thermal value distribution function;
[0107] The color coding is output according to the thermal value distribution function to generate a heat map of the conflict area.
[0108] In this example, existing video generation technologies can create semantic contradictions between the user-entered semantic description and the actual motion trajectory (e.g., "the object rises" versus "the trajectory goes down"), but lack effective means for detecting and visualizing these contradictions. To address this issue, it is necessary to develop an intelligent detection method that combines semantic parsing and motion trajectory analysis.
[0109] Specifically, the user submits a semantic description text, such as "the object moves from left to right and rises" through a human-computer interaction interface (such as text input or voice input). The semantic description text entered by the user (such as "the car drives from left to right and accelerates") is parsed through a natural language processing (NLP) model (such as BERT, GPT, etc.). A set of verb elements (such as "drive" and "accelerate") is extracted and the corresponding direction encoding vector is generated (such as "left→right" corresponds to the direction vector [1,0], and "accelerate" corresponds to the dynamic change vector [Δx,Δy]). The verb elements are mapped to the direction encoding vector to form a semantic-direction association dataset (such as "drive→[1,0]" and "accelerate→dynamic change features") for subsequent matching analysis.
[0110] Based on a motion trajectory vector layer (e.g., "car trajectory"), perform equidistant sampling in the image space domain to obtain a discrete set of direction vectors (e.g., [(1,0),(0.8,0.6),...,(0.2,-0.1)]). A continuous direction field is constructed using the rate of change of direction in the trajectory vector layer (e.g., "trajectory curvature"). Uniform or non-uniform sampling is performed within the direction field to obtain a discrete set of direction vectors.
[0111] When a spatial correlation (e.g., angular deviation < 30°) is detected between a direction encoding vector (e.g., "[1,0]") and a set of discrete direction vectors (e.g., [(1,0), (0.9, 0.1), ...]), the attention matching model is invoked. An attention mechanism (e.g., a multi-head attention model) is used to calculate the regional matching degree (e.g., the cosine similarity between the direction encoding vector and the discrete direction vector).
[0112] The regional matching degree is nonlinearly transformed using a contradiction intensity amplification factor (e.g., 1 + α * matching degree 2, where α is the amplification factor) to generate a thermal value distribution function H(x, y) = f(matching degree, contradiction intensity). The thermal value distribution function H(x, y) is used to output a color code (e.g., red = high contradiction, blue = low contradiction).
[0113] The heat map of conflicting areas is generated by outputting color codes (e.g., red → yellow → blue) based on the heat distribution function. The heat map can be superimposed on the original image to intuitively display the conflicting areas.
[0114] In this embodiment, the user's semantic description is parsed and motion trajectories are automatically matched, reducing the difficulty of conflict detection. A heat map is used to visually display conflicting areas, making it easier for users to correct them. Users can adjust the semantic description or motion trajectory based on the heat map, achieving iterative optimization.
[0115] Furthermore, the attention matching model satisfies
[0116]
[0117] Among them, M(x,y) represents the region matching degree, x and y represent the trajectory direction, ω k represents the attention weight of the k-th verb element, K represents the number of verb elements, represents the set of direction encoding vectors, represents a discrete direction vector set, θ v Indicates the standard direction angle corresponding to the verb element, Indicates the rate of change of trajectory in x direction, s i represents the tangential unit vector at the trajectory point;
[0118] The thermal value distribution function satisfies
[0119] H(x,y)=σ(α·(1-M(x,y)))
[0120] Where H(x,y) represents the value of the thermal value distribution function, σ(·) represents the Sigmoid activation function, and α represents the contradiction intensity amplification factor.
[0121] In this embodiment, attention weights are used to measure the contribution of the kth verb element to the match. Weights are dynamically assigned through an attention mechanism (such as the multi-head attention in the Transformer). For example, the verb "rise" has a higher weight for vertical trajectories, while "move" has a higher weight for horizontal trajectories. This can solve the problem of multiple verb interference (such as "rotate + move") and ensure that the dominant verb dominates the matching results.
[0122] The set of directional encoding vectors is used to map verb elements into directional vectors (e.g., "rise" → [0, 1]). Pre-trained models (such as BERT) are used to encode the semantic features of verbs into directional vectors, enabling a semantic-to-spatial mapping. This unifies the dimensions of semantics and trajectory, providing a foundation for subsequent matching.
[0123] The set of discrete direction vectors represents the local direction of the trajectory at a point (x, y). They are sampled at the trajectory vector layer (e.g., every 5 pixels) and reflect the actual direction of the trajectory. They provide spatial features of the trajectory and are matched with the semantic direction vector.
[0124] The standard direction angle represents the expected direction corresponding to the verb element (e.g., "left" → 180°). Verbs are converted to angles using a predefined direction dictionary (e.g., "left" → 180°, "right" → 0°). This provides a reference direction for matching and quantifies the deviation between semantics and trajectory.
[0125] The x-rate of change of a trajectory reflects the velocity or acceleration of the trajectory in the x-direction and is calculated from the trajectory's derivative. It helps determine whether the trajectory meets the dynamic characteristics of a verb (e.g., "accelerate" corresponds to an increase in the rate of change).
[0126] The tangent unit vector represents the tangent direction of the trajectory at point (x, y). It is calculated based on the geometry of the trajectory. It is used to provide local directional information of the trajectory and perform a dot product with the semantic direction vector to calculate similarity.
[0127] The Sigmoid activation function maps the regional matching degree to the range [0, 1], generating a normalized thermal value. This nonlinear transformation compresses the matching degree value. For example, a matching degree of 0.8 results in a thermal value of 0.73 (σ(0.8)≈0.73). This resolves the issue of inconsistent matching degree ranges and facilitates visualization.
[0128] The conflict intensity magnification factor is used to amplify or reduce the thermal value of the conflict area. This enhances the visibility of the conflict area and helps users quickly locate the problem.
[0129] In some embodiments, in step S103, when a conflict area is detected in the conflict area heat map, an intention analysis is performed using a reinforcement learning strategy to generate a multi-dimensional optimization solution including trajectory direction correction, speed classification, and physical constraint prompts, specifically including:
[0130] When a conflicting area is detected in the conflict area heat map, a state vector containing the heat value distribution, trajectory direction field, velocity field and physical constraint parameters is constructed;
[0131] The Actor-Critic strategy network is used to evaluate the current state and generate an action vector that includes trajectory direction correction, speed adjustment, and constraint parameter compensation.
[0132] Setting a reward function, wherein the reward function includes a heat map contradiction resolution term, a direction alignment term, and a constraint stability term;
[0133] Based on the state vector, action vector and reward function, the reinforcement learning algorithm is used to calculate the strategy optimization direction and output a multi-dimensional optimization plan.
[0134] In this example, in a video generation task, there may be a contradiction between the user's textual intent (e.g., "a car accelerates from left to right") and the actual generated trajectory (e.g., the trajectory curves downward). Existing techniques use heat maps to visualize the conflicting areas, but lack intelligent means to correct the contradiction, resulting in limited video generation quality.
[0135] Specifically, when a conflict area is detected in the conflict area heat map, a state vector is constructed, which includes the following dimensions: (1) Thermal value distribution: The thermal value matrix of the conflict area is extracted and normalized to the range [0,1]. (2) Trajectory direction field: The tangential unit vector of the trajectory at each point is calculated to generate the direction field matrix. (3) Velocity field: The velocity of each point is calculated based on the timestamp data of the trajectory to generate the velocity field matrix. (4) Physical constraint parameters: Including the acceleration limit of the trajectory, collision detection results, etc., encoded as scalars or vectors.
[0136] The Actor-Critic strategy network is used to evaluate the state vector and generate an action vector, which includes the following dimensions: (1) Trajectory direction correction: the direction adjustment angle of each point (such as -10° to +10°). (2) Speed adjustment: the speed is divided into three levels: low speed, medium speed, and high speed, and the speed adjustment coefficient of each point is output (such as 0.8, 1.0, 1.2). (3) Constraint parameter compensation: adjust the physical constraint parameters (such as increasing the acceleration limit to 10m / s 2 ).
[0137] The reward function satisfies
[0138] R=λ1·R1+λ2·R2+λ3·R3
[0139]
[0140] Among them, R represents the reward function value, R1 represents the contradiction elimination term of the heat map, R2 represents the direction alignment term of the heat map, R3 represents the constraint stability term of the heat map, λ1, λ2 and λ3 represent weight coefficients, N represents the number of regional points of the heat map, H i Represents the heat value of the i-th region point. The higher the heat value of the conflicting region, the greater the penalty. represents the target direction angle, represents the century direction angle. The higher the degree of alignment, the greater the reward. M represents the number of physical constraints. j It represents the penalty value of the j-th physical constraint. The higher the degree of constraint satisfaction, the greater the reward.
[0141] Based on the state vector S, the action vector A, and the reward function R, a reinforcement learning algorithm (such as PPO) is used to calculate the policy optimization direction and output a multi-dimensional optimization solution. Specifically, the state-action-reward triple (S, A, R) is stored. The parameters of the actor and critic networks are optimized based on the reward function R. After iterative optimization, the final action vector A is output, which includes trajectory direction correction, velocity classification, and physical constraint information.
[0142] In this example, reinforcement learning is used to automatically generate a multi-dimensional optimization solution to improve the consistency of trajectory with user intent. The design of the state vector and reward function reduces computational complexity and is suitable for real-time video generation tasks. It can be extended to other physically constrained scenarios (such as collision avoidance and trajectory smoothing).
[0143] In some embodiments, in step S104, according to different scene types, environmental constraints are analyzed using a parameterized physical model, the action sequence is reconstructed based on a kinematic decoupling algorithm, and a continuum mechanics model is used to simulate structural failure propagation to generate motion description parameters that conform to domain characteristics. Specifically, the following steps are performed:
[0144] Identify the scene type characteristics of the target scene, and determine the environmental objects, main objects, and whether there are structural damage conditions based on the scene type characteristics;
[0145] Constructing an environmental dynamics equation based on the object pose vector of the environmental object and the corresponding constraint conditions, wherein the environmental dynamics equation is used to output environmental constraint parameters;
[0146] Based on Lie group decomposition theorem, the continuous motion of the subject object is decomposed into a sequence of discrete motion primitives;
[0147] When structural failure conditions are detected, stress propagation data are calculated using the governing equations of continuum mechanics;
[0148] Environmental constraint parameters, discrete motion primitive sequences, and stress propagation data are integrated to generate motion description parameters that include multi-physics field coupling features.
[0149] In this embodiment, scene type features of the target scene are extracted, including scene type (such as "city streets", "natural landscapes", "industrial scenes"), environmental objects (such as vehicles, trees, pedestrians, buildings), and subject objects (such as robots, characters). Based on the object pose vectors (position, posture) of the environmental objects and corresponding constraints (such as collision detection, friction model, gravity influence), the environmental dynamics equation is constructed, and environmental constraint parameters (such as maximum displacement, maximum speed, and force range) are output. Based on the Lie group decomposition theorem, the continuous motion of the subject object is decomposed into a sequence of discrete motion primitives. The user-drawn motion graph (dynamic area and trajectory) is parsed and converted into the constraints of the motion primitive sequence. When structural damage conditions are detected (such as robotic arm overload or building collapse), the continuum mechanics model is triggered. The continuum mechanics model (such as elastic mechanics equations) is used to calculate stress propagation data (such as stress distribution and crack propagation path). The environmental constraint parameters (such as maximum displacement of 1.5 meters), the discrete motion primitive sequence (such as "move-grab-place"), the stress propagation data (such as crack propagation path), and the user-drawn motion graph (such as trajectory) are integrated. Generate motion description parameters containing multi-physics field coupling features, including: environmental constraint parameters: maximum displacement, maximum velocity, force range; motion primitive sequence: time series and spatial coordinates of discrete actions; structural failure parameters: stress distribution, crack propagation path; user motion diagram parameters: dynamic areas and trajectories drawn by the user.
[0150] In this embodiment, multi-physics information, such as environmental constraints, kinematics, and structural failures, is integrated to generate prompts that conform to actual physical laws. By leveraging user-drawn motion graphs and kinematic decoupling technology, the precision and dynamic effects of generated video are enhanced. Consistency between the input graph, motion information, and prompts is ensured to avoid logical and visual inconsistencies.
[0151] In some embodiments, in step S105, the motion graph is modified according to the multi-dimensional optimization scheme and the motion description parameters, and the modified motion graph is iteratively verified. If the verification passes, the initial image and multiple sets of motion graphs are input into the multimodal large language model to generate the final prompt word containing the scene adaptation parameters, which specifically includes:
[0152] Based on the trajectory correction amount and motion description parameters in the multi-dimensional optimization scheme, the corrected motion trajectory is generated through the weighted fusion equation to obtain the correction result;
[0153] The correction result is iteratively verified through the spatiotemporal consistency verification function, and a verification pass signal is triggered when the function value of the spatiotemporal consistency verification function is greater than the preset function threshold;
[0154] The verified initial image and multiple sets of motion images are input into the multimodal large language model to generate prompt words containing scene adaptation parameters.
[0155] In this embodiment, existing video generation technologies often make it difficult to directly use user-drawn motion graphs (dynamic areas and trajectories) for video generation. The main reasons include: User-drawn trajectories may not conform to physical laws (e.g., sudden speed changes, path conflicts); The corrected motion graphs are not rigorously verified, which may result in logical or visual inconsistencies in the generated video; The generated prompt words do not take into account scene characteristics, resulting in a mismatch between the video content and the scene.
[0156] Specifically, based on the user-drawn motion diagram and semantic description text, the system identifies unreasonable parts of the trajectory (such as sudden speed changes and path conflicts). It then integrates trajectory corrections (such as reducing the speed to 2 m / s) and motion description parameters (such as a maximum displacement of 1.5 meters) into a multi-dimensional optimization solution. Based on the trajectory corrections and motion description parameters in the multi-dimensional optimization solution, a weighted fusion equation is used to generate a corrected motion trajectory. The corrected motion trajectory is then output, including parameters such as position, velocity, and acceleration.
[0157] A spatiotemporal consistency verification function is constructed to evaluate whether the modified motion graph meets the following conditions: (1) Temporal consistency: whether the temporal sequence of the motion trajectory is reasonable (e.g., whether the speed changes smoothly). (2) Spatial consistency: whether the motion trajectory has no conflicts with the environmental objects in the scene (e.g., obstacles, other characters).
[0158] The modified motion graph is iteratively verified until the function value of the spatiotemporal consistency verification function exceeds the preset function threshold. The preset function threshold is set to 0.8 (the function value range is 0 to 1, with 1 indicating complete consistency). For example, the initial function value is 0.6, which increases to 0.7 after the first iteration and to 0.85 after the second iteration, triggering a verification pass signal. When the function value exceeds the preset threshold, a verification pass signal is triggered, indicating that the modified motion graph meets the spatiotemporal consistency requirements.
[0159] The verified initial image and multiple sets of motion graphs are fed into a multimodal large language model. Based on the initial image and motion graphs, the multimodal large language model generates prompt words containing scene adaptation parameters. The final prompt words, containing the scene adaptation parameters, are then output for video generation.
[0160] In this embodiment, a multi-dimensional optimization solution corrects unreasonable user-drawn trajectories to ensure that the motion trajectory conforms to physical laws. A spatiotemporal consistency verification function ensures that the corrected motion graph meets temporal and spatial consistency requirements. The generated prompts include scene adaptation parameters to ensure that the video content closely matches the scene characteristics.
[0161] Reference Figure 2 An embodiment of the present invention provides a prompt word optimization system for video generation, the system specifically comprising:
[0162] The first generation module is used to receive semantic description text input by the user after the user uploads the initial image, and guide the user to draw a motion map including a dynamic mask layer and a motion trajectory vector layer in the image editing interface;
[0163] The second generation module is used to spatially match the verb elements in the semantic description text with the motion trajectories in the motion map, and generate a heat map of the conflict area by combining the attention mechanism;
[0164] The third generation module is used to analyze the intentions of the conflicting areas in the conflict area heat map through reinforcement learning strategies, and generate a multi-dimensional optimization plan including trajectory direction correction, speed classification and physical constraint prompts;
[0165] The fourth generation module is used to analyze environmental constraints based on different scenario types through parameterized physical models, reconstruct action sequences based on kinematic decoupling algorithms, and simulate structural failure propagation using a continuum mechanics model to generate motion description parameters that meet domain characteristics.
[0166] The fifth generation module is used to modify the motion graph according to the multi-dimensional optimization scheme and motion description parameters, and iteratively verify the modified motion graph. If the verification passes, the initial image and multiple sets of motion graphs are input into the multimodal large language model to generate the final prompt word containing the scene adaptation parameters.
[0167] It is understandable that if Figure 1 The contents of the embodiment of the prompt word optimization method for video generation shown in the figure are all applicable to the embodiment of the prompt word optimization system for video generation. The functions specifically implemented by the embodiment of the prompt word optimization system for video generation are similar to those in the embodiment of the figure. Figure 1 The embodiment of the prompt word optimization method for video generation shown is the same as that shown in FIG. Figure 1 The beneficial effects achieved by the embodiment of the prompt word optimization method for video generation shown are also the same.
[0168] It should be noted that the information interaction, execution process and other contents between the above-mentioned systems are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0169] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0170] Reference Figure 3 The embodiment of the present invention further provides a computer device 3, comprising: a memory 302, a processor 301, and a computer program 303 stored in the memory 302. When the computer program 303 is executed on the processor 301, the prompt word optimization method for video generation as described in any one of the above methods is implemented.
[0171] The computer device 3 may be a desktop computer, a notebook computer, a PDA, a cloud server or other computing devices. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art will understand that Figure 3 This is merely an example of the computer device 3 and does not constitute a limitation on the computer device 3 . The computer device 3 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the computer device 3 may also include input and output devices, network access devices, etc.
[0172] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor.
[0173] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may also be an external storage device of the computer device 3, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 3. Furthermore, the memory 302 may include both an internal storage unit of the computer device 3 and an external storage device. The memory 302 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program. The memory 302 may also be used to temporarily store data that has been output or is about to be output.
[0174] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for optimizing prompt words for video generation as described in any one of the above methods is implemented.
[0175] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process of the above-mentioned method embodiment by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can at least include: any entity or device capable of carrying computer program code to the camera / terminal device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, mobile hard drive, magnetic disk, or optical disk. In some jurisdictions, based on legislation and patent practice, computer-readable media cannot be electric carrier signals or telecommunication signals.
[0176] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0177] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0178] In the embodiments disclosed in the present application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0179] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
Claims
1. A prompt word optimization method for video generation, characterized in that: The method specifically includes: After the user uploads the initial image, the system receives the semantic description text entered by the user and guides the user to draw a motion map including a dynamic mask layer and a motion trajectory vector layer in the image editing interface; Spatial matching is performed between the verb elements in the semantic description text and the motion trajectories in the motion map, and an attention mechanism is used to generate a heat map of the conflicting areas. When conflicting areas are detected in the conflict area heat map, the reinforcement learning strategy is used to analyze the intention and generate a multi-dimensional optimization solution that includes trajectory direction correction, speed classification, and physical constraint prompts. According to different scenario types, environmental constraints are analyzed through parameterized physical models, action sequences are reconstructed based on kinematic decoupling algorithms, and structural failure propagation simulations are performed using continuum mechanics models to generate motion description parameters that meet domain characteristics. The motion graph is corrected according to the multi-dimensional optimization scheme and motion description parameters, and the corrected motion graph is iteratively verified. If the verification passes, the initial image and multiple sets of motion graphs are input into the multimodal large language model to generate the final prompt word containing the scene adaptation parameters.
2. The method according to claim 1, characterized in that After the user uploads the initial image, the user inputs a semantic description text, and the user is guided to draw a motion map including a dynamic mask layer and a motion trajectory vector layer in the image editing interface, specifically including: After receiving the initial image uploaded by the user, the user's input semantic description text is received through the human-computer interaction interface, and a dynamic drawing guide frame is generated on the image editing interface based on the semantic description text; When a user touch operation is detected, the polygon vertex coordinate sequence is recorded in real time, and a second-order B-spline interpolation algorithm is used to generate a dynamic mask layer with gradual transparency; After the mask layer is drawn, the trajectory anchor point set input by the user is captured, and the motion trajectory vector layer is constructed using the cubic B-spline curve fitting algorithm.
3. The method according to claim 1, characterized in that The process of spatially matching the verb elements in the semantic description text with the motion trajectories in the motion map and generating a heat map of conflicting regions by combining the attention mechanism specifically includes: The natural language processing model is used to parse the semantic description text entered by the user, extract the verb element set, and generate the corresponding direction encoding vector; Based on the continuous direction field in the motion trajectory vector layer, equidistant sampling is performed in the image space domain to obtain a discrete direction vector set; When it is detected that the direction encoding vector is spatially associated with the discrete direction vector set, the attention matching model is called to calculate the regional matching degree; The regional matching degree is transformed nonlinearly by the contradiction intensity amplification factor to obtain the thermal value distribution function; The color coding is output according to the thermal value distribution function to generate a heat map of the conflict area.
4. The method according to claim 3, characterized in that The attention matching model satisfies Among them, M(x,y) represents the region matching degree, x and y represent the trajectory direction, ω k represents the attention weight of the k-th verb element, K represents the number of verb elements, represents the set of direction encoding vectors, represents a discrete direction vector set, θ v Indicates the standard direction angle corresponding to the verb element, Indicates the rate of change of trajectory in x direction, s i represents the tangential unit vector at the trajectory point; The thermal value distribution function satisfies H(x,y)=σ(α·(1-M(x,y))) Where H(x,y) represents the value of the thermal value distribution function, σ(·) represents the Sigmoid activation function, and α represents the contradiction intensity amplification factor.
5. The method according to claim 1, wherein When a conflict area is detected in the conflict area heat map, the intention analysis is performed through the reinforcement learning strategy to generate a multi-dimensional optimization solution including trajectory direction correction, speed classification and physical constraint prompts, including: When a conflicting area is detected in the conflict area heat map, a state vector containing the heat value distribution, trajectory direction field, velocity field and physical constraint parameters is constructed; The Actor-Critic strategy network evaluates the current state and generates an action vector that includes trajectory direction correction, speed adjustment, and constraint parameter compensation. Setting a reward function, wherein the reward function includes a heat map contradiction resolution term, a direction alignment term, and a constraint stability term; Based on the state vector, action vector and reward function, the reinforcement learning algorithm is used to calculate the strategy optimization direction and output a multi-dimensional optimization plan.
6. The method according to claim 1, characterized in that According to different scenario types, the environmental constraints are analyzed through parameterized physical models, the action sequence is reconstructed based on the kinematic decoupling algorithm, and the structural failure propagation simulation is performed using the continuum mechanics model to generate motion description parameters that meet the characteristics of the domain. Specifically, Identify the scene type characteristics of the target scene, and determine the environmental objects, main objects, and whether there are structural damage conditions based on the scene type characteristics; Constructing an environmental dynamics equation based on the object pose vector of the environmental object and the corresponding constraint conditions, wherein the environmental dynamics equation is used to output environmental constraint parameters; Based on Lie group decomposition theorem, the continuous motion of the subject object is decomposed into a sequence of discrete motion primitives; When structural failure conditions are detected, stress propagation data are calculated using the governing equations of continuum mechanics; Environmental constraint parameters, discrete motion primitive sequences, and stress propagation data are integrated to generate motion description parameters that include multi-physics field coupling features.
7. The method according to any one of claims 1 to 6, characterized in that The motion graph is modified according to the multi-dimensional optimization scheme and motion description parameters, and the modified motion graph is iteratively verified. If the verification passes, the initial image and multiple sets of motion graphs are input into the multimodal large language model to generate the final prompt word containing the scene adaptation parameters, which specifically includes: Based on the trajectory correction amount and motion description parameters in the multi-dimensional optimization scheme, the corrected motion trajectory is generated through the weighted fusion equation to obtain the correction result; The correction result is iteratively verified through the spatiotemporal consistency verification function, and a verification pass signal is triggered when the function value of the spatiotemporal consistency verification function is greater than the preset function threshold; The verified initial image and multiple sets of motion images are input into the multimodal large language model to generate prompt words containing scene adaptation parameters.
8. A prompt word optimization system for video generation, characterized in that: The system specifically includes: The first generation module is used to receive semantic description text input by the user after the user uploads the initial image, and guide the user to draw a motion map including a dynamic mask layer and a motion trajectory vector layer in the image editing interface; The second generation module is used to spatially match the verb elements in the semantic description text with the motion trajectories in the motion map, and generate a heat map of the conflict area by combining the attention mechanism; The third generation module is used to analyze the intentions of the conflicting areas in the conflict area heat map through reinforcement learning strategies, and generate a multi-dimensional optimization plan including trajectory direction correction, speed classification and physical constraint prompts; The fourth generation module is used to analyze environmental constraints based on different scenario types through parameterized physical models, reconstruct action sequences based on kinematic decoupling algorithms, and simulate structural failure propagation using a continuum mechanics model to generate motion description parameters that meet domain characteristics. The fifth generation module is used to modify the motion graph according to the multi-dimensional optimization scheme and motion description parameters, and iteratively verify the modified motion graph. If the verification passes, the initial image and multiple sets of motion graphs are input into the multimodal large language model to generate the final prompt word containing the scene adaptation parameters.
9. A computer device, characterized in that: include: A memory, a processor, and a computer program stored in the memory, which, when executed on the processor, implements the prompt word optimization method for video generation according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the method for optimizing prompt words for video generation according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Iterative self-optimization text video method and system based on knowledge enhancement
CN122138025A
A knowledge-enhanced iterative self-optimizing text-generated video method and system
CN122138025B