A method and system for generating a picture-text-video based on cross-modal semantic mapping

By constructing a hierarchical structured semantic tree and a partitioned attention coordination mechanism, the problem of refined expression of attribute adjectives and scene adverbs in existing technologies is solved, generating time-coherent dynamic videos and improving the product display effect in e-commerce advertisements.

CN120730138BActive Publication Date: 2026-06-26BEIJING HUIFENG RUNDA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING HUIFENG RUNDA TECH CO LTD
Filing Date
2025-06-18
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively express the nuanced details of attribute adjectives and scene adverbs when generating product demonstration videos, leading to visual confusion, misaligned sequence of function demonstrations, material or style conflicts, and insufficient visual continuity.

Method used

By constructing a hierarchical structured semantic tree, extracting core object nouns, attribute adjectives, and scene adverbs, and employing a partitioned coordinated attention mechanism and dynamic calibration technology, a temporally coherent dynamic video is generated to ensure accurate mapping of materials, styles, and scenes.

Benefits of technology

It enables refined expression of attribute adjectives and scene adverbs, avoiding visual confusion and temporal misalignment, and improving the semantic fit and content generation efficiency of video generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120730138B_ABST
    Figure CN120730138B_ABST
Patent Text Reader

Abstract

The application provides a kind of text and video generation method and system based on cross-modal semantic mapping, it is related to data processing technical field, the method includes: step 1, input product description text, execute hierarchical semantic decoupling, extract core object noun, attribute adjective and scene adverb, and construct hierarchical structured semantic tree;Step 2, based on hierarchical structured semantic tree, execute the area exploration of fine-grained modifier semantics, identify the associated area of attribute adjective or scene adverb, generate semantic adaptation correction factor for each associated area.The application realizes the automatic generation of product description text to semantic precision, time sequence coherent dynamic video by hierarchical semantic decoupling, area semantic mapping, cross-modal feature fusion and dynamic space-time calibration, ensures that visual effect is consistent with text semantics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for generating text and video based on cross-modal semantic mapping. Background Technology

[0002] In fields such as e-commerce advertising, the technology of generating product demonstration videos based on text descriptions continues to develop, but existing systems still face the following challenges in the implementation process:

[0003] Current mainstream methods (such as some end-to-end generative models) tend to focus on the visual representation of core object nouns (such as "dress") during semantic mapping, while there is room for optimization in the refined expression of attribute adjectives (such as "lace patchwork" and "elegant and intellectual") and scene adverbs (such as "raising an arm during exercise"). In some complex scenarios, visual confusion may occur between the "elegant and intellectual" style and the "sweet" style; when the correlation between scene adverbs and action logic is insufficient, it may lead to misalignment of the function display sequence (such as "heart rate monitoring" and the arm-raising action not being fully synchronized).

[0004] Some existing solutions (such as fusion methods based on global splicing) may cause material or style conflicts when dealing with complex semantics. When the text description involves multiple attributes in the same area (such as "chiffon" and "cotton"), it may generate an unnatural transition effect. There may be a lack of visual continuity between the differentiated style expressions of adjacent areas (such as the elegant design of the cuffs and the lively decoration of the skirt). Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and system for generating text and video based on cross-modal semantic mapping, so as to achieve pixel-level semantic alignment between text description and dynamic video.

[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0007] Firstly, a method for generating text-to-image videos based on cross-modal semantic mapping, the method comprising:

[0008] Step 1: Input product description text, perform hierarchical semantic decoupling, extract core object nouns, attribute adjectives and scenario adverbs, and construct a hierarchical structured semantic tree;

[0009] Step 2: Based on the hierarchical structured semantic tree, perform fine-grained semantic modification region exploration, identify the associated regions of attribute adjectives or scene adverbs, and generate semantic adaptation correction factors for each associated region;

[0010] Step 3: Using a partitioned coordinated attention mechanism, the semantic adaptation correction factor is injected into the feature fusion link of the corresponding visual region to generate a cross-modal feature vector, and a key frame sequence is generated based on the cross-modal feature vector, wherein each frame contains the associated region segmentation label.

[0011] Step 4: Based on the keyframe sequence, perform dynamic calibration, extract scene adverb semantic units from the hierarchical structured semantic tree, construct spatiotemporal constraint relationships, and inject spatiotemporal constraint relationships into the associated region segmentation identifiers through a dynamic correction and transmission process, outputting a temporally coherent dynamic video.

[0012] Furthermore, the product description text is input, hierarchical semantic decoupling is performed, core object nouns, attribute adjectives, and scenario adverbs are extracted, and a hierarchical structured semantic tree is constructed, including:

[0013] Perform syntactic dependency parsing on the product description text to extract the core object nouns as the root node of the semantic tree; identify the direct dependent adjectives that modify the core object nouns and bind them as the first-level child nodes of the root node; identify the adverbial phrases associated with verbs and map them as the second-level time constraint child nodes of the root node;

[0014] Based on the syntactic distance between the first-level child nodes and the second-level time-constrained child nodes and the root node of the semantic tree, a hierarchical index is constructed to generate a hierarchical structured semantic tree.

[0015] Furthermore, based on a hierarchical structured semantic tree, fine-grained semantic region exploration is performed to identify associated regions of attribute adjectives or scene adverbs, and semantic adaptation correction factors are generated for each associated region, including:

[0016] Based on the semantics of the first-layer child nodes, the visual space segmentation of the core object is driven, generating the texture segmentation region corresponding to the material class child nodes and the outline segmentation region corresponding to the style class child nodes.

[0017] For texture segmentation regions, detect whether there are multiple material class child nodes in the same region. If so, generate a weighted fusion material correction factor.

[0018] For the contour segmentation region, the deviation value between the geometric features and the semantics of the style class child nodes is detected. If the deviation value exceeds the preset threshold, a style correction factor with directional constraints is generated.

[0019] Furthermore, based on the semantics of the first-layer child nodes, the visual space segmentation of the core object is driven, generating texture segmentation regions corresponding to material-type child nodes and contour segmentation regions corresponding to style-type child nodes, including:

[0020] The semantics of material class adjectives in the first-level child nodes are parsed to drive the modeling of the surface material distribution of the core object and generate the initial material region; boundary continuity optimization is performed on the initial material region, and adjacent sub-regions with the same material properties are merged to generate texture segmentation regions;

[0021] The semantics of style-class adjectives in the first-level child nodes are parsed to drive the geometric contour modeling of the core object and generate the initial contour region; topological closure optimization is performed on the initial contour region to generate the contour segmentation region;

[0022] Establish a spatial mapping relationship between texture segmentation regions and contour segmentation regions, and identify the contour partition to which each texture region belongs.

[0023] Furthermore, a partitioned coordinated attention mechanism is employed to inject semantic adaptation correction factors into the feature fusion link of the corresponding visual region, generating cross-modal feature vectors. Based on these cross-modal feature vectors, a keyframe sequence is generated, where each frame contains associated region segmentation markers, including:

[0024] The material correction factor and style correction factor are spatially encoded according to the coordinates of the texture segmentation region and the contour segmentation region to generate an indexed correction matrix.

[0025] Retrieve region feature vectors that match spatial coordinates from the visual basic feature library, and perform affine transformation on the correction matrix and the region feature vectors to generate cross-modal feature vectors.

[0026] Based on cross-modal feature vectors, a sequence of key frames is rendered, and texture segmentation region coordinates, contour segmentation region coordinates, and region temporal identifiers are embedded in each frame to form a group of associated region segmentation identifier frames.

[0027] Furthermore, region feature vectors matching the spatial coordinates are retrieved from the visual feature library. An affine transformation is then performed on the correction matrix and the region feature vectors to generate cross-modal feature vectors, including:

[0028] Material correction factors are spatially mesh-encoded according to the coordinates of texture segmentation regions to generate a material correction matrix; style correction factors are spatially mesh-encoded according to the coordinates of contour segmentation regions to generate a style correction matrix; based on the coordinates of texture segmentation regions, material feature vectors matching the coordinates are retrieved from the visual basic feature library; based on the coordinates of contour segmentation regions, contour feature vectors matching the coordinates are retrieved from the visual basic feature library.

[0029] An affine transformation is performed between the material correction matrix and the material feature vector to generate optimized material features; an affine transformation is performed between the style correction matrix and the contour feature vector to generate optimized contour features.

[0030] By fusing optimized material features and optimized contour features, a cross-modal feature vector is generated.

[0031] Furthermore, based on the keyframe sequence, dynamic calibration is performed to extract scene adverb semantic units from the hierarchical structured semantic tree, construct spatiotemporal constraint relationships, and inject spatiotemporal constraint relationships into the associated region segmentation identifiers through a dynamic correction and transmission process, outputting a temporally coherent dynamic video, including:

[0032] Extract the scene adverb semantic units from the second-level time constraint sub-nodes in the semantic tree, parse the time adverbs to generate the starting time period, parse the action adverbs to generate the target action parameters, and construct an action-time period mapping table;

[0033] In the associated region segmentation identifier frame group, the material correction domain is located based on the texture segmentation region coordinates; the motion correction domain is located based on the contour segmentation region coordinates.

[0034] When the material similarity between adjacent frames in the material correction domain is lower than the threshold, a transition frame is inserted based on the time period marker in the mapping table; when the motion parameters in the motion correction domain deviate from the constraint values ​​in the mapping table beyond the tolerance, the motion phase is redirected based on the regional temporal identifier, and a temporally coherent dynamic video is output.

[0035] Secondly, a text-to-video generation system based on cross-modal semantic mapping includes:

[0036] The semantic decoupling module is used to input product description text, perform hierarchical semantic decoupling, extract core object nouns, attribute adjectives and scenario adverbs, and construct a hierarchical structured semantic tree;

[0037] The semantic exploration module is used to perform fine-grained semantic region exploration based on a hierarchical structured semantic tree, identify the associated regions of attribute adjectives or scene adverbs, and generate semantic adaptation correction factors for each associated region.

[0038] The feature fusion module is used to inject semantic adaptation correction factors into the feature fusion link of the corresponding visual region using a partitioned coordinated attention mechanism, generate cross-modal feature vectors, and generate keyframe sequences containing related region segmentation labels based on the cross-modal feature vectors.

[0039] The dynamic generation module is used to perform dynamic calibration based on key frame sequences, extract scene adverb semantic units from hierarchical structured semantic trees, construct spatiotemporal constraint relationships, inject spatiotemporal constraint relationships into the segmentation identifiers of associated regions through dynamic correction and transmission processes, and output temporally coherent dynamic video.

[0040] Thirdly, a computing device includes:

[0041] One or more processors;

[0042] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0043] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0044] The above-described solution of the present invention has at least the following beneficial effects:

[0045] By constructing a structured semantic tree through hierarchical semantic decoupling, we can achieve hierarchical extraction of core object nouns, attribute adjectives, and scene adverbs, avoiding the loss or misjudgment of attribute / scene semantics in traditional methods. For example, for "elegant and intellectual lace dress," we can accurately distinguish the style boundary between "intellectual" and "sweet," reducing visual performance deviations. By binding scene adverbs with action triggering logic (such as associating "heart rate monitoring during exercise" with a hand-raising action), we can avoid the timing mismatch of function display. Utilizing fine-grained semantic region exploration and partitioned attention coordination mechanisms, we inject material / style correction factors into corresponding visual regions, replacing the traditional global feature splicing strategy. For example, we use weighted fusion correction factors for mutually exclusive attributes in the same region ("chiffon" and "cotton") to avoid texture mixing; for opposing styles in adjacent regions ("elegant" cuffs and "sweet" hem), we eliminate visual disjointedness and improve the refinement of feature fusion through contour segmentation and transition constraints.

[0046] Based on semantic tree-based scene adverb semantic units, spatiotemporal constraints are constructed, and the discontinuity problem of traditional temporal modeling is solved by dynamically correcting the transmission process. For example, the display period of the "blood oxygen monitoring during sleep" function is limited by time adverbs to avoid erroneous triggering of non-corresponding action frames; for decorative elements (such as "bows"), the action phase is redirected through regional temporal identifiers to eliminate visual state jumps and ensure the temporal consistency between video dynamic effects and text semantics. The closed-loop design of the entire process from semantic decoupling to spatiotemporal calibration achieves accurate mapping from text description to visual content. The generated videos can accurately convey the complex attributes of products such as material, style, and usage scenarios, which is especially suitable for the dynamic display of complex products in e-commerce advertising, reducing manual design costs while improving content generation efficiency and semantic fit. Attached Figure Description

[0047] Figure 1 This is a flowchart illustrating a text-to-video generation method based on cross-modal semantic mapping, provided by an embodiment of the present invention.

[0048] Figure 2 This is a schematic diagram of a text and video generation system based on cross-modal semantic mapping provided by an embodiment of the present invention. Detailed Implementation

[0049] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0050] like Figure 1 As shown, an embodiment of the present invention proposes a text-to-video generation method based on cross-modal semantic mapping, the method comprising the following steps:

[0051] Step 1: Input product description text, perform hierarchical semantic decoupling, extract core object nouns, attribute adjectives and scenario adverbs, and construct a hierarchical structured semantic tree;

[0052] Step 2: Based on the hierarchical structured semantic tree, perform fine-grained semantic modification region exploration, identify the associated regions of attribute adjectives or scene adverbs, and generate semantic adaptation correction factors for each associated region;

[0053] Step 3: Using a partitioned coordinated attention mechanism, the semantic adaptation correction factor is injected into the feature fusion link of the corresponding visual region to generate a cross-modal feature vector, and a key frame sequence is generated based on the cross-modal feature vector, wherein each frame contains the associated region segmentation label.

[0054] Step 4: Based on the keyframe sequence, perform dynamic calibration, extract scene adverb semantic units from the hierarchical structured semantic tree, construct spatiotemporal constraint relationships, and inject spatiotemporal constraint relationships into the associated region segmentation identifiers through a dynamic correction and transmission process, outputting a temporally coherent dynamic video.

[0055] In this embodiment of the invention, by extracting core objects, attributes, and scene adverbs, the text description is transformed into a hierarchical semantic tree, avoiding the mixing of semantic information in traditional methods. The structured semantic tree can systematically capture multi-dimensional information such as material, style, and action in the text. Through semantically driven region segmentation, attributes directly correspond to the texture partitions of the 3D model, avoiding visual deviations caused by global material mixing. For example, in a running shoe model, the boundary between the leather area and the mesh area is precisely defined using semantic correction factors, resulting in a natural texture transition. For composite attributes such as "gradient color + frosted texture," the generated semantic adaptation correction factor can simultaneously adjust the color gradient parameters and roughness, ensuring that multiple modifications are visually presented in a coordinated manner, avoiding semantic breaks caused by single-dimensional corrections in traditional methods.

[0056] The partitioned coordinated attention mechanism associates semantic correction factors such as material and contour with visual feature vectors to generate feature vectors containing cross-modal information. For example, after fusing the contour curvature parameter of "retro style" with the texture parameter of "aged leather," the model simultaneously presents a synergistic effect of retro shape and aged material. A spatiotemporal mapping table is constructed based on scene adverbs to ensure strict synchronization between material changes (such as mesh transparency) and actions (shoe sole compression) on the timeline, avoiding the misalignment of actions and functions in traditional rendering. A dynamic calibration mechanism automatically detects abrupt material changes (such as color jumps) or action deviations (such as insufficient compression displacement), ensuring smooth playback and semantic conflicts by inserting transition frames or redirecting action phases. For example, when the material similarity is below a threshold, a transition frame is automatically generated to make color changes smooth and natural.

[0057] In a preferred embodiment of the present invention, step 1 above, which involves inputting product description text, performing hierarchical semantic decoupling, extracting core object nouns, attribute adjectives, and scene adverbs, and constructing a hierarchical structured semantic tree, may include:

[0058] Step 100: Perform syntactic dependency parsing on the product description text and extract the core object nouns as the root node of the semantic tree;

[0059] Identify the direct dependent adjectives that modify the nouns of the core object and bind them as the first-level child nodes of the root node; identify the adverbial phrases associated with verbs and map them as the second-level time-constrained child nodes of the root node;

[0060] Step 101: Based on the syntactic distance between the first-level child nodes and the second-level time-constrained child nodes and the root node of the semantic tree, construct a hierarchical index to generate a hierarchical structured semantic tree.

[0061] In this embodiment of the invention, the input text (e.g., "This lightweight and breathable summer sports vest can quickly wick away sweat when running") is segmented into independent lexical units (e.g., "this", "lightweight", "breathable", "summer", "sports vest", "when running", "can", "quickly", "wick away sweat"). Stop words (e.g., "this", "of", "can") are removed, and core semantic words are retained to obtain a preprocessed text sequence. A pre-trained dependency syntax model (e.g., spaCy's en_core_web_lg) is used to analyze the grammatical relationships between words and generate a dependency tree. For example, "sports vest" is the subject of the verb "wick away sweat" (nsubj relation) and plays the role of the core argument in the sentence. "When running" is associated with the verb "wick away sweat" through an "advcl" (adverbial clause) relation, indicating a temporal condition.

[0062] Traverse the dependency tree to find nodes that satisfy the following conditions:

[0063] Directly involved in the core action of the sentence (such as the subject of "sweating");

[0064] It has the highest centrality (by calculating the betweenness centrality of a node in the dependency tree, i.e., the frequency with which it acts as a bridge for the shortest path);

[0065] Exclude modifiers (such as "summer" as a time modifier, not as the core word alone);

[0066] Finally, "sports vest" was determined as the root node.

[0067] Binding logic for first-level child nodes (attribute adjectives):

[0068] In the dependency tree, search for words that are directly connected to the root node "sports vest" through an "amod" (adjective modification) relationship. For example, "lightweight" and "breathable" both directly modify "sports vest" through an "amod" relationship and are extracted as candidate adjectives.

[0069] The selected adjectives are semantically categorized:

[0070] Material type: such as "cotton" or "velvet";

[0071] Functional features: such as "lightweight" and "breathable";

[0072] Style categories: such as "elegant" and "sweet".

[0073] Categorization is achieved by querying a predefined semantic dictionary; for example, "lightweight" is categorized as a functional attribute. The categorized adjectives are then grouped by semantic type and bound to the first-level child nodes of the root node. For example:

[0074] 1.1: Lightweight and thin (functional);

[0075] 1.2: Breathable (functional).

[0076] Mapping rules for second-level child nodes (time-constrained adverb phrases):

[0077] In the dependency tree, look for phrases that are connected to verbs (such as "sweat") via "advmod" (adverb modification) or "advcl" (adverbial clause). For example:

[0078] "While running" modifies "sweating" using the "advcl" relationship, indicating a time condition;

[0079] The word “quickly” modifies “sweating” through the “advmod” relationship, indicating the manner of the action.

[0080] Spatiotemporal constraint classification is performed on the identified adverbial phrases:

[0081] Time constraints: such as "while running" or "during exercise";

[0082] Method constraints: such as "fast" or "slow";

[0083] Location constraints: such as "outdoor" or "indoor".

[0084] The classification is based on the semantic features and dependency relationships of words. For example, "while running" is classified as a time-constrained phrase because it contains both action and time markers. The classified adverbial phrases are mapped to second-level child nodes of the root node, establishing indirect connections through intermediate verbs. For example:

[0085] 2.1: While running (time constraint, related verb "sweat");

[0086] 2.2: Fast (method constraint, related verb "sweating").

[0087] Step 101: Starting from the root node, traverse the dependency tree using Breadth-First Search (BFS), initializing the distance of the root node to 0. For each child node, its distance is the distance of its parent node plus 1. For example: "Sports Vest" (distance 0) → "Sweater" (distance 1, via the nsubj relationship) → "Running" (distance 2, via the advcl relationship). If multiple paths connect the same node, take the distance of the shortest path. For example, "Fast" indirectly connects "Sports Vest" via "Sweater," the path is "Sports Vest → Sweater → Fast," and the distance is 2.

[0088] Generation of hierarchical indexes and construction of semantic trees:

[0089] Based on the calculated syntactic distance, the nodes are divided into different levels:

[0090] Distance 0: Root node level;

[0091] Distance 1: First-level child nodes (attribute adjective);

[0092] Distance 2: Second-level child node (time-constrained adverb phrase).

[0093] The encoding method adopted is "hierarchical.sequential".

[0094] Root node: 0;

[0095] First-level child nodes: After grouping by semantic type, each group is sorted according to its position in the dependency tree, such as 1.1 (first of the material class) and 1.2 (first of the function class);

[0096] Second-level child nodes: grouped and sorted by constraint type, such as 2.1 (time constraint first) and 2.2 (method constraint first).

[0097] Organize all nodes into a tree structure according to their index relationships.

[0098] By using grammatical dependency parsing, the system clearly distinguishes between core object nouns (e.g., "dress"), attribute adjectives (e.g., "lace patchwork," "elegant and intellectual"), and scene adverbs (e.g., "raising an arm while exercising"), avoiding the loss or confusion of modifying semantics. For example, "elegant and intellectual," as a first-level child node, directly modifies the core object, accurately driving visual style generation and preventing misjudgments of the "sweet" style. The identification and layering of adverbial phrases (e.g., "while exercising") ensures that scene logic (e.g., action triggering conditions) is fully captured, laying the foundation for subsequent dynamic temporal calibration. The hierarchical index (root node → first-level attribute → second-level scene) transforms textual semantics into a tree structure, visualizing the relationship between "object-attribute-scene." For example, the material attribute of "leather sofa" (first level) and the scene constraint of "in the living room" (second level) clearly define the subordinate relationship through the tree structure, driving region segmentation during visual generation (material corresponds to the texture area, scene corresponds to the background setting). Quantifying syntactic distance (e.g., 1 for attribute adjectives and 2 for scene adverbs) clarifies semantic priority and avoids feature fusion chaos caused by spreading out multi-dimensional information.

[0099] For scenarios with multiple modifiers (such as "blue striped cotton shirt"), hierarchical binding ensures that "blue," "striped," and "cotton" are all first-level child nodes, corresponding to independent visual areas of color, pattern, and material, respectively, thus solving the texture mixing problem caused by attribute superposition in traditional methods. Hierarchical mapping of adverbial phrases (such as "while running" and "quickly" as second-level child nodes in "quickly sweats while running") clarifies the time and manner constraints of actions, avoiding misalignment of function display timing (such as the sweating effect being out of sync with the running action).

[0100] In a preferred embodiment of the present invention, step 2 above, based on a hierarchical structured semantic tree, performs fine-grained semantic region exploration to identify associated regions of attribute adjectives or scene adverbs, and generates a semantic adaptation correction factor for each associated region, may include:

[0101] Step 200: Based on the semantics of the first-layer child nodes, drive the visual space segmentation of the core object, generating the texture segmentation region corresponding to the material class child nodes and the contour segmentation region corresponding to the style class child nodes, specifically including:

[0102] Step 2020: Parse the semantics of material class adjectives in the first-layer child nodes to drive the modeling of the surface material distribution of the core object and generate the initial material region; perform boundary continuity optimization on the initial material region, merge adjacent sub-regions with the same material properties, and generate texture segmentation regions;

[0103] Step 2021: Parse the semantics of style-class adjectives in the first-level child nodes to drive the geometric contour modeling of the core object and generate the initial contour region; perform topological closure optimization on the initial contour region to generate the contour segmentation region;

[0104] Step 2022: Establish the spatial mapping relationship between the texture segmentation region and the contour segmentation region, and identify the contour partition to which each texture region belongs;

[0105] Step 201: For the texture segmentation region, detect whether there are multiple material class child nodes in the same region. If so, generate a weighted fusion material correction factor.

[0106] Step 202: For the contour segmentation region, detect the deviation value between the geometric features and the semantics of the style class child nodes. If the deviation value exceeds the preset threshold, generate a style correction factor with directional constraints.

[0107] In this embodiment of the invention, taking the input text "dark gray wool blend hooded coat with durable leather trim at the elbows" as an example, the material-related adjectives ("wool blend" and "leather") in the first layer of sub-nodes are parsed using NLP tools (such as spaCy). After extracting keywords, a pre-trained Word2Vec model is called to calculate word vectors, which are then matched with a material feature library (containing parameters such as wool fiber density and leather tensile strength). For example, "wool blend" matches the parameter set in the feature library with wool fiber diameter of 18-25μm and a blend ratio of 50% wool + 50% polyester fiber. Combined with the 3D template model of the core object "coat" (which has been bound to an ergonomic mesh), the material semantics are transformed into surface attribute distributions using the material node system of 3D software such as Blender. For example, "wool blend" corresponds to the main fabric area, with a diffuse reflectance coefficient of 0.7 and a roughness of 0.6; "leather" corresponds to the elbow area, with a reflectance of 0.4 and a roughness of 0.3.

[0108] Initial material region generation and boundary optimization:

[0109] Based on material distribution parameters, an initial mask is generated on the UV unwrapping map of the 3D model: through threshold segmentation (e.g., regions with a material probability > 0.6 are marked as valid), binarization is performed using Python's OpenCV library to obtain the initial contours of the "wool blend region" and the "cortical region". For example, the elbow cortical region is presented as an independent polygonal selection in the UV map.

[0110] Boundary continuity optimization involves two steps:

[0111] Morphological operations: First, perform a dilation operation on the 3×3 convolution kernel (expanding the region edge by 2 pixels), then perform an erosion operation (shrinking the edge by 1 pixel) to smooth the jagged edges;

[0112] Adjacent region merging: Calculate the similarity of material features of adjacent sub-regions (e.g., extract texture direction histograms and use Bach distance to calculate differences), set a threshold of 0.8 (distance < 0.8 is considered similar), merge the wool blend body and cuff areas of the same material, and eliminate small fragment partitions.

[0113] Step 2021: Analyze the style-related adjectives (such as "hooded design" and "rugged lines of durable leather at the elbows") in the first-layer child nodes, and extract geometric feature keywords using the rule engine: "hooded" corresponds to a rounded top structure with a preset radius of curvature of 15cm; "rugged lines" corresponds to straight edges, allowing vertex angle deviation ≤10°. Based on the vertex coordinates of the outer coat's 3D model, use Maya's deformer tool for parametric adjustments:

[0114] For the vertices of the hooded area, move the top vertex up 2cm along the Z-axis, and at the same time adjust the weights of the adjacent vertices to increase the brim curvature from the initial R=10cm to R=15cm;

[0115] The edge vertices of the elbow cortical region were adjusted from rounded corners (curvature 0.5) to right angles (curvature 0.1), and the normal vector directions were modified in batches using the vertex group editing mode.

[0116] Topology checks, using mesh analysis tools for 3D models (such as Blender's mesh repair function), detected three non-closed edges at the connection between the hood and collar (with notch lengths of 2mm and 1.5mm, respectively). For each notch edge vertex, the average normal vector of adjacent vertices was calculated to generate a new triangular facet; for example, for notch edge ABC, a new vertex D was inserted in the ABC plane, connecting AD, BD, and CD to form a closed facet; Laplacian smoothing was iterated five times, each time shifting the vertex 10% towards the centroid of adjacent vertices to ensure the curvature continuity of the contour edges, ultimately generating the complete hood contour, the independent elbow cortex contour, and other partitions.

[0117] Step 2022: Unify the UV coordinates (2D) of the texture segmentation with the world coordinates of the 3D contour model. Extract the vertex coordinates (ua, vb) of the UV map and convert them into 3D coordinates (x, y, z) using the UV mapping matrix of the model (exported by the 3D software, including scaling and rotation parameters). For example, the vertex (ua = 0.3, vb = 0.4) of the elbow cortex region in the UV map is used to obtain its actual position in 3D space through the matrix operation [x, y, z] = M × [ua, vb, 1]. A ray perpendicular to the UV plane is emitted from the center (u0, v0) of a texture region in the UV map and converted into a 3D ray using the inverse mapping matrix. The ray intersects with the contour-segmented 3D model, and the contour partition where the nearest intersection point is located is recorded. For example, if the ray intersects with the elbow cortex contour model, the texture region is identified as the "elbow cortex partition".

[0118] Step 201, taking a "burgundy silk and off-white cotton dress" as an example, firstly, semantic parsing determines that the first-layer child nodes associated with the texture segmentation region (such as the skirt hem) are "silk" and "cotton". Detection logic: Traverse the list of material adjectives corresponding to this region. When ≥2 different material labels appear, trigger the correction process. On the UV unwrapping map of the 3D model, the conflict region (i.e., the boundary between silk and cotton) is marked by the material distribution probability map, which is usually represented as a transition zone with a width of 5-8 pixels (corresponding to a physical distance of 1-2 cm in the 3D model).

[0119] Weighted fusion calculation of material parameters (generating correction factors):

[0120] The word "silk" appears twice in the text, totaling 10 words, with a frequency of 0.2; the word "cotton" appears once, with a frequency of 0.1. Assume that out of 100 descriptions of dresses in the domain document library, 20 contain "silk". 50 articles containing the word "cotton" The weights for "silk" and "cotton" are normalized, i.e. Ensure that the sum of the weights of multiple materials is 1 to avoid excessive bias in color / roughness towards a certain material due to weight imbalance when merging parameters.

[0121] Diffuse color correction factor:

[0122] Silk base color RGB (180, 30, 80), cotton base color RGB (240, 220, 200), calculate the corrected color value for each pixel in the transition area:

[0123] R = 180 × normalized silk weight + 240 × normalized cotton weight;

[0124] G = 30 × normalized silk weight + 220 × normalized cotton weight;

[0125] B = 80 × normalized silk weight + 200 × normalized cotton weight;

[0126] The generated color correction factor is RGB(180×normalized silk weight + 240×normalized cotton weight, 30×normalized silk weight + 220×normalized cotton weight, 80×normalized silk weight + 200×normalized cotton weight). This factor is used to cover the original pixel color and achieve a gradient from silk red to cotton off-white.

[0127] Roughness correction factor:

[0128] Silk roughness is 0.2 (smooth), cotton roughness is 0.6 (rough). The roughness correction factor after correction is equal to the silk roughness × normalized silk weight + the cotton roughness × normalized cotton weight. This factor directly affects the roughness parameter of the pixel, so that the surface graininess gradually increases when the transition area changes from smooth silk to rough cotton.

[0129] Step 202: Taking a "retro rounded neckline (required radius R = 5cm)" as an example, the actual contour radius is measured to be R = 3cm using 3D modeling software, and the deviation value is calculated as 5 - 3 = 2cm. The preset threshold is 1cm (i.e., correction is triggered when the deviation exceeds 1cm). If the actual radius is less than the target value, it is considered "insufficient radius," and the contour needs to be expanded outward; if the actual radius is greater than the target value, it is considered "excessive radius," and the contour needs to be contracted inward. The normal vector (the vector perpendicular to the surface where that point is located) of each vertex of the neckline contour determines the offset direction. For insufficient radius, the normal vector points outward (away from the center), and the offset direction is outward along the normal vector; if excessive radius, the normal vector points to the center, and the offset direction is inward.

[0130] For each vertex, calculate the displacement correction factor (i.e., the distance and direction of the required offset):

[0131] The displacement reflects the length a vertex needs to move. Displacement = distance from vertex to center of circle × (target radius - actual radius) ÷ target radius. For example, if a vertex is 4cm from the center, the radius R corresponding to the current actual radius is 3cm, and the target radius R determined by the style class child node semantics is 5cm, then the displacement = 4 × (5 - 3) ÷ 5 = 1.6cm. This indicates that the vertex needs to move 1.6cm in a specific direction to make the contour meet the target radius requirement. The direction of displacement is determined by the vertex's normal vector, which is a vector perpendicular to the surface on which the vertex is located. It defines the direction of vertex movement. For example, if the vertex is located on an arc surface, its normal vector points outwards, and the vertex will move along this outward direction; if the surface is concave, the normal vector points inwards, and the vertex will move inwards.

[0132] Based on the calculated displacement and determined displacement direction, a vertex displacement correction factor is generated. This factor is presented as a three-dimensional vector containing components in the X, Y, and Z dimensions. If the vertex normal vector is along the positive X-axis and the displacement is 1.6cm, then the generated vertex displacement correction factor is (1.6cm, 0, 0); if the normal vector is along the positive Y-axis, then the displacement correction factor is (0, 1.6cm, 0). This three-dimensional vector directly acts on the vertex coordinates, driving the vertex to move a corresponding distance along the specified direction, thereby changing the contour shape to meet the target curvature requirement.

[0133] Material / style semantic-driven segmentation transforms abstract descriptions into concrete texture regions and contour partitions, avoiding the "chiffon and cotton texture mixing" problem caused by global feature mixing in traditional methods. Weighted fusion of material correction factors addresses the transition problem between multiple materials in the same region, while directionally constrained style factors correct contour deviations, enhancing visual realism. The established texture-contour space mapping provides accurate coordinate indices for subsequent feature fusion, reducing semantic misalignment during feature injection. For text containing multiple modifiers, hierarchical segmentation and correction factor generation simultaneously handle the visual mapping of material, style, and function, ensuring accurate expression of semantics across all dimensions.

[0134] In a preferred embodiment of the present invention, step 3 above employs a partitioned coordinated attention mechanism to inject a semantic adaptation correction factor into the feature fusion link of the corresponding visual region, generating a cross-modal feature vector, and generating a keyframe sequence based on the cross-modal feature vector, wherein each frame contains an associated region segmentation identifier, which may include:

[0135] Step 300: Spatial encoding of the material correction factor and style correction factor according to the coordinates of the texture segmentation region and the contour segmentation region to generate an indexed correction matrix;

[0136] Step 301: Retrieve region feature vectors matching the spatial coordinates from the visual basic feature library, and perform an affine transformation on the correction matrix and the region feature vectors to generate cross-modal feature vectors, specifically including:

[0137] Step 3010: Encode the material correction factor using spatial grids based on the coordinates of the texture segmentation region to generate a material correction matrix; encode the style correction factor using spatial grids based on the coordinates of the contour segmentation region to generate a style correction matrix; retrieve the material feature vector matching the coordinates from the visual basic feature library based on the coordinates of the texture segmentation region; retrieve the contour feature vector matching the coordinates from the visual basic feature library based on the coordinates of the contour segmentation region.

[0138] Step 3011: Perform an affine transformation on the material correction matrix and the material feature vector to generate optimized material features; perform an affine transformation on the style correction matrix and the contour feature vector to generate optimized contour features;

[0139] Step 3012: Fuse optimized material features and optimized contour features to generate cross-modal feature vectors;

[0140] Step 302: Based on cross-modal feature vectors, render keyframe sequences and embed texture segmentation region coordinates, contour segmentation region coordinates and region temporal identifiers in each frame to form a group of associated region segmentation identifier frames.

[0141] In this embodiment of the invention, the generated material correction factor (including parameters such as color and roughness) and style correction factor (including parameters such as contour offset and curvature adjustment value) are obtained, and the coordinate information of the texture segmentation region and the contour segmentation region is read simultaneously. Wherein:

[0142] The coordinates of the texture segmentation region are based on the two-dimensional coordinates (u, v) of the UV unfolded image. For example, the coordinate range of a certain region is u∈[0.2, 0.6], v∈[0.3, 0.7]. The coordinates of the contour segmentation region are the vertex coordinates in three-dimensional space. The texture segmentation region is divided into a grid (e.g., a 100×100 grid) according to a preset precision. Each grid cell corresponds to the actual physical size of the three-dimensional model surface (e.g., 0.5cm×0.5cm). For each grid cell, it is determined whether it belongs to the texture segmentation region based on its center coordinates. If it does, a material correction factor (e.g., color RGB value, roughness value) is filled into the cell to form a material correction grid. The contour segmentation region is divided into a voxel grid (e.g., a 30×30×30 voxel) in three-dimensional space. Each voxel corresponds to a cube in three-dimensional space (e.g., 1cm×1cm×1cm). A ray casting method is used to determine whether the voxel contains contour vertices. If it does, a style correction factor (e.g., vertex displacement, normal direction) is filled into the voxel to form a style correction grid. Convert 2D material correction meshes and 3D style correction meshes into matrix structures:

[0143] The material correction matrix M_texture has dimensions of [number of mesh rows × number of mesh columns × number of correction parameters], for example, 100×100×4 (R, G, B, roughness); the style correction matrix M_shape has dimensions of [number of voxel layers × number of voxel rows × number of voxel columns × number of correction parameters]. Add a composite index to each element in the matrix, with the index format "region type_coordinate segmentation_parameter type".

[0144] Step 3010: Normalize each grid cell parameter in M_texture to the [0, 1] interval (e.g., divide the color value by 255, and use the roughness directly), and divide it into blocks according to the coordinates of the texture region (e.g., each 10×10 grid is a block), generating multiple sub-matrices (e.g., 100 sub-matrices of 10×10×4). Cluster each voxel parameter in M_shape according to the coordinates of the contour region (e.g., divide the vertices into 5 clusters based on the K-means algorithm), with each cluster corresponding to a sub-matrix (e.g., 5 sub-matrices of 6×6×6×3).

[0145] The UV coordinates of the texture segmentation region are converted into index keys of the feature library (e.g., "uv_0.2-0.6_0.3-0.7"), and a matching material feature vector is searched in the feature library. The feature vector includes a basic color vector (e.g., [0.3, 0.4, 0.5] corresponding to RGB normalized values), a texture mode vector (e.g., [0.6, 0.2, 0.1] representing the proportion of fabric texture), etc. The 3D coordinates of the contour segmentation region are converted into bounding box coordinates (e.g., "x10-20_y5-15_z0-10"), and a matching contour feature vector is searched in the feature library. The feature vector includes a shape curvature vector (e.g., [0.8, 0.3] representing the degree of roundness), a vertex distribution vector (e.g., [0.5, 0.5] representing symmetry), etc.

[0146] Step 3011: Perform a weighted fusion of the submatrices of the material correction matrix and the material eigenvectors:

[0147] For each color parameter in a submatrix, the color vector of the eigenvector is multiplied element-wise (e.g., the submatrix color value [0.8, 0.2, 0.1] is multiplied by the eigenvector [0.3, 0.4, 0.5], and then an offset (e.g., [0.1, 0.1, 0.1]) is added to generate the optimized color vector.

[0148] The roughness parameter is linearly combined with the texture pattern vector of the feature vector (e.g., the correction value 0.7 and the pattern vector 0.6 are combined with weights of 0.7 and 0.3 to obtain 0.7×0.7+0.3×0.6).

[0149] Perform a spatial transformation between the submatrices of the style correction matrix and the contour feature vectors:

[0150] For each submatrix, the displacement parameters are adjusted in relation to the curvature vector of the eigenvector (e.g., a displacement of 1cm is amplified to 1.2cm based on the curvature vector of 0.8 to make the contour fit the curvature requirements better).

[0151] For the vertex distribution vector, rotate it according to the direction of the normal vector in the correction matrix (e.g., rotate the vertex distribution vector [0.5, 0.5] 15 degrees along the normal vector to ensure contour symmetry).

[0152] Step 3012 unifies the optimized material features (such as color vectors and roughness vectors) and contour features (such as curvature vectors and displacement vectors) into vectors of the same dimension. For example, if the material features are 10-dimensional (3-dimensional color + 1-dimensional roughness + 6-dimensional texture pattern) and the contour features are 8-dimensional (2-dimensional curvature + 3-dimensional vertex distribution + 3-dimensional displacement), both are transformed into 12-dimensional vectors through dimensionality reduction. Based on the partitioned coordinated attention mechanism, the weights of the material features and contour features are calculated: features with a larger area proportion in the texture segmentation region are assigned higher weights (e.g., material feature weight 0.6, contour feature weight 0.4).

[0153] Feature vectors are merged by weight: such as material vector V_texture = [v1, v2, ... v12], contour vector V_shape = [s1, s2, ... s12], and cross-modal vector V_cross = 0.6 × V_texture + 0.4 × V_shape, to ensure the organic integration of material and contour information.

[0154] Step 302: Material features drive texture mapping generation, such as converting the color vector [0.24, 0.08, 0.05] to RGB (61, 20, 13), and generating a texture with a roughness value of 0.67. Contour features drive the displacement of model vertices, such as applying a displacement of 1.2cm after adjusting the curvature vector to the vertex coordinates to generate a new contour shape. Render keyframes according to a preset animation sequence (e.g., 100 frames in total).

[0155] Frames 1-20 showcase material feature changes, frames 21-60 demonstrate contour deformation, and frames 61-100 display the coordinated changes in material and contour. Cross-modal feature vectors are invoked in real-time during each frame's rendering to ensure dynamic consistency between material and contour. Three types of identifiers are embedded in the metadata of each frame's image:

[0156] Texture segmentation region coordinates: Store the UV coordinate range of the texture region in JSON format (e.g., {"uv":[[0.2, 0.3], [0.6, 0.7]]});

[0157] Contour segmentation region coordinates: Stores the three-dimensional coordinate set of contour vertices (e.g., {"vertices":[[x1, y1, z1], [x2, y2, z2]]});

[0158] Region time sequence identifier: Records the time interval of this region in the video (e.g., {"start_frame":21, "end_frame":60}).

[0159] Generate associated region segmentation identifier frame groups.

[0160] By employing a partitioned attention coordination mechanism, material and style correction factors are precisely injected into corresponding visual regions, avoiding texture conflicts or contour deviations caused by global feature mixing. Region segmentation markers (coordinates + timing) embedded in keyframes provide clear guidance for dynamic calibration. For example, when displaying strap deformation, the system can precisely control the starting position and time of deformation based on the contour region coordinates and timing markers, preventing asynchrony between function display and action. For products containing multiple materials and styles, cross-modal feature vectors can simultaneously fuse material, shape, and color information, ensuring coherent expression of semantics across dimensions in dynamic videos, enhancing visual realism and user experience. The design of indexed correction matrices and region markers allows the system to flexibly adapt to different product descriptions; when adding materials or styles, only the correction factors and coordinate information need to be updated to quickly generate corresponding visual content.

[0161] In a preferred embodiment of the present invention, step 4 above, which involves dynamic calibration based on the keyframe sequence, extracting scene adverb semantic units from the hierarchical structured semantic tree, constructing spatiotemporal constraint relationships, and injecting spatiotemporal constraint relationships into the associated region segmentation identifier through a dynamic correction and transmission process, and outputting a temporally coherent dynamic video, may include:

[0162] Step 400: Extract the scene adverb semantic units of the second-level time constraint sub-nodes in the semantic tree, parse the time adverbs to generate the starting time period, parse the action adverbs to generate the target action parameters, and construct the action-time period mapping relationship table;

[0163] Step 401: In the associated region segmentation identifier frame group, locate the material correction domain according to the texture segmentation region coordinates; locate the motion correction domain according to the contour segmentation region coordinates;

[0164] Step 402: When the material similarity between adjacent frames in the material correction domain is lower than the threshold, a transition frame is inserted based on the time period marker in the mapping table; when the motion parameters in the motion correction domain deviate from the constraint values ​​in the mapping table beyond the tolerance, the motion phase is redirected based on the regional temporal identifier, and a temporally coherent dynamic video is output.

[0165] In this embodiment of the invention, scene adverb semantic units are extracted from the second-level time constraint sub-nodes of the hierarchical structured semantic tree using part-of-speech tagging algorithms (such as rule-based adverb recognition). For example, for the text "When running outdoors, the breathable mesh upper of the shoe quickly wicks away sweat, and when jumping and landing, the cushioning module of the sole is significantly compressed," the system will locate "when running outdoors," "when jumping and landing" (time adverb phrases), and "quickly" and "significantly" (action adverbs) in the time constraint sub-nodes. The extraction of semantic units needs to be combined with syntactic dependency relations. For example, "quickly" is used as an adverbial of "wicking away sweat," and the dependency tree is used to determine the core action word "wicking away sweat" that it modifies.

[0166] Based on the preset video duration or total number of keyframes, convert time adverbs into specific time segments. Assume the total video duration is 10 seconds (25 frames / second, 250 frames total):

[0167] "Running outdoors" is interpreted as a continuous motion and mapped to frames 100-200 (corresponding to 4-8 seconds);

[0168] The "jump landing" is interpreted as an instantaneous action, mapped to frames 220-225 (corresponding to 8.8-9 seconds).

[0169] Time conversion rules: Phrases containing "time" or "in process" are mapped to duration segments, while phrases containing "instant" or "time" are mapped to instantaneous actions with a frame range of ≤5 frames.

[0170] Action adverbs are converted into quantization parameters using predefined mapping rules:

[0171] The "Fast" setting corresponds to the dynamic display speed of the perspiration function, set to a material transparency change of ≥0.05 per frame (reduced from 0.8 to 0.3).

[0172] "Significant" corresponds to the displacement of the sole compression, which is set to a compression distance ≥ 1.5cm (compressed from an initial thickness of 3cm to 1.5cm).

[0173] Complex motion parameters need to be combined with domain knowledge. For example, "compression of the damping module" needs to be associated with the pressure-deformation curve to convert "significant" into deformation exceeding 60% of the static deformation.

[0174] The action-time mapping table uses a three-dimensional structure:

[0175] Line dimension: Action scenario (e.g., "shoe upper sweat-wicking" "sole compression");

[0176] Column dimension: Time period (start frame - end frame, e.g., 100-200 frames);

[0177] Depth dimension: Target parameter set (e.g., [transparency change rate 0.05 / frame, compression displacement 1.5cm]).

[0178] Step 401: Read the texture segmentation coordinates from the associated region segmentation identifier frame group and mark the material region in each frame image. The texture coordinates adopt a UV mapping relationship. For example, the coordinate range of the mesh area of ​​the shoe upper in the UV map is (u=0.2, v=0.3)-(u=0.8, v=0.7). Through the UV-screen coordinate transformation matrix, generate the corresponding two-dimensional mask in each frame (e.g., in frame 100, the corresponding screen coordinates of this area are x=100-500px, y=200-400px). If a frame contains "mesh + leather" splicing, generate two masks respectively, and the material correction domain is the union of the masks. Based on the three-dimensional coordinates of the contour segmentation region, locate the dynamic region in the motion-related frame.

[0179] The three-dimensional coordinate set of the sole outline is [(x1, y1, z1), (x2, y2, z2), ...]. It is converted into screen coordinates through the camera projection matrix. A dynamic mask of the sole area is generated in frames 220-225 (adjusted according to motion deformation). The motion correction domain of the sole outline is activated only in the time period specified by the mapping table (such as frames 220-225). Other frames are regarded as static areas.

[0180] Step 402: Locate the material correction domain (e.g., a texture partition of a certain area of ​​the product) and obtain the material feature parameters of two adjacent frames. For example, in frame n and frame n+1, extract the color values ​​(specific values ​​of the red, green, and blue channels), transparency percentage (e.g., 80% and 30%), and surface roughness quantification values ​​(e.g., 0.6 and 0.9) of that area. Calculate the numerical differences of the color channels (e.g., red changing from 200 to 180, a difference of 20), the drop in transparency from 80% to 30% (50%), and the increase in roughness from 0.6 to 0.9 (0.3), and determine whether these changes exceed the visually acceptable range of a natural transition.

[0181] Preset similarity thresholds for different materials (e.g., color difference for fabric materials is allowed to be ≤15, transparency difference ≤20%, roughness change ≤0.15; thresholds for metal materials can be relaxed). When a parameter changes beyond the corresponding threshold, a transition frame insertion mechanism is triggered.

[0182] The number of frames to be inserted is calculated based on the degree to which the parameters exceed the threshold. If the color difference exceeds 5 (20-15), 1 frame is inserted for every 5 differences. If the transparency difference exceeds 30% (50%-20%), 1 frame is inserted for every 10% difference. After comprehensive consideration, 4 transition frames are inserted.

[0183] Perform linear interpolation on each material parameter:

[0184] The colors change from RGB (200, 220, 240) to (180, 200, 220). During the 4-frame transition, each channel decreases by 5 (20 ÷ 4 = 5) in each frame, generating transition frame colors of (195, 215, 235), (190, 210, 230), (185, 205, 225), and (180, 200, 220) respectively.

[0185] The transparency decreases from 80% to 30% by 12.5% ​​per frame (50% ÷ 4 = 12.5%), with transition frames having transparency of 67.5%, 55%, 42.5%, and 30%.

[0186] The roughness ranges from 0.6 to 0.9, increasing by 0.075 per frame (0.3 ÷ 4 = 0.075), with transition frame roughnesses of 0.675, 0.75, 0.825, and 0.9.

[0187] Between frame n and frame n+1, four transition frames are inserted sequentially according to the calculated parameters, so that the original frame order becomes n→transition frame 1→transition frame 2→transition frame 3→transition frame 4→n+1, ensuring that the material parameters change smoothly frame by frame and eliminating visual breaks.

[0188] Real-time monitoring and deviation calculation of motion parameters:

[0189] In the motion correction domain (such as the motion partition of a product's outline), extract the actual motion data. For example, a certain action requires a displacement of 1.5 cm within 5 frames (0.3 cm per frame), but the actual displacement is 0.5 cm in the first frame, 0.3 cm in the second frame, and 0.2 cm in the third frame, totaling 1.0 cm. The target cumulative displacement should be 0.9 cm (0.3 × 3 = 0.9), resulting in a deviation of +0.1 cm. Compare this deviation to a preset tolerance (such as ±0.05 cm / frame) to determine if the deviation exceeds the allowable range. If it does, phase redirection needs to be triggered.

[0190] Motion phase adjustment and parameter reallocation:

[0191] Redistribute the remaining displacement: The total target is 1.5 cm, 1.0 cm has been completed, and 0.5 cm needs to be completed in the remaining 2 frames. Adjust the displacement to 0.25 cm per frame (0.5 ÷ 2 = 0.25).

[0192] Smoothing speed changes: Check the speed difference between adjacent frames (if the original speed of frame 3 is 0.2 cm / frame, and frame 4 directly changes to 0.25 cm / frame, the change rate is 25%). If it exceeds the preset speed fluctuation threshold (e.g., 20%), further refine the adjustment. For example, shift 0.23 cm in frame 4 and 0.27 cm in frame 5 to keep the speed change rate within 15%. Fill the adjusted displacement into frames 4 and 5 to ensure that the displacement of each frame is within the tolerance range and the total displacement meets the standard.

[0193] Play back the adjusted motion frames to check if the total displacement is 1.5 cm, if the displacement changes uniformly in each frame, and if it is synchronized with other dynamic effects such as material changes. For example, ensure that the change in the transparency of the shoe upper material ends synchronously when the shoe sole compression action is completed to avoid timing misalignment. Merge the inserted material transition frames and the adjusted motion frames into the original keyframe sequence in chronological order to form a new frame queue. For example, if the original sequence is ...n, n+1, ...k, k+1..., after insertion it becomes ...n, transition 1, transition 2, transition 3, transition 4, n+1, ...k (adjusted), k+1...

[0194] Iterate through all frames and check the time synchronization between the material correction domain and the motion correction domain: for example, whether the frame where the material transparency begins to change is consistent with the frame where the motion begins. If there is an offset (e.g., the material change is two frames earlier than the motion), adjust the insertion position of the transition frame to ensure coordinated multi-dimensional dynamic effects. Perform image enhancement (e.g., anti-aliasing, sharpening) on ​​all frames, add motion blur to motion frames (e.g., edge blurring when the sole of a shoe is compressed), synthesize the video at the set resolution (e.g., 1920×1080) and frame rate (25fps), and embed metadata (correction records, timing identifiers) to ensure that the output video has a coherent time sequence and smooth visuals.

[0195] By using an action-time period mapping table, the system ensures that scene semantics are accurately presented in the video according to preset times and parameters, avoiding discrepancies between actions and time. It automatically detects and corrects abrupt material changes and action deviations, eliminating screen flicker or motion stuttering by inserting transition frames and redirecting action phases, thus improving the video viewing experience. For products with multiple actions and material changes, it can simultaneously handle dynamic calibration of materials and actions, ensuring that every detail conforms to semantic logic. No manual frame-by-frame adjustments are required; the system can automatically generate spatiotemporal constraints and perform calibration based on the semantic tree. When new scene adverbs or action parameters are added, only the mapping table needs to be updated for quick adaptation, reducing video production costs.

[0196] like Figure 2 As shown, embodiments of the present invention also provide a text-to-video generation system based on cross-modal semantic mapping, comprising:

[0197] The semantic decoupling module is used to input product description text, perform hierarchical semantic decoupling, extract core object nouns, attribute adjectives and scenario adverbs, and construct a hierarchical structured semantic tree;

[0198] The semantic exploration module is used to perform fine-grained semantic region exploration based on a hierarchical structured semantic tree, identify the associated regions of attribute adjectives or scene adverbs, and generate semantic adaptation correction factors for each associated region.

[0199] The feature fusion module is used to inject semantic adaptation correction factors into the feature fusion link of the corresponding visual region using a partitioned coordinated attention mechanism, generate cross-modal feature vectors, and generate keyframe sequences containing related region segmentation labels based on the cross-modal feature vectors.

[0200] The dynamic generation module is used to perform dynamic calibration based on key frame sequences, extract scene adverb semantic units from hierarchical structured semantic trees, construct spatiotemporal constraint relationships, inject spatiotemporal constraint relationships into the segmentation identifiers of associated regions through dynamic correction and transmission processes, and output temporally coherent dynamic video.

[0201] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0202] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0203] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0204] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for generating text-to-image video based on cross-modal semantic mapping, characterized in that, The method includes: Step 1: Input product description text, perform hierarchical semantic decoupling, extract core object nouns, attribute adjectives and scenario adverbs, and construct a hierarchical structured semantic tree; Step 2: Based on the hierarchical structured semantic tree, perform fine-grained semantic modification region exploration, identify the associated regions of attribute adjectives or scene adverbs, and generate semantic adaptation correction factors for each associated region; Step 3: Using a partitioned coordinated attention mechanism, the semantic adaptation correction factor is injected into the feature fusion link of the corresponding visual region to generate a cross-modal feature vector, and a key frame sequence is generated based on the cross-modal feature vector, wherein each frame contains the associated region segmentation label. Step 4: Based on the keyframe sequence, perform dynamic calibration, extract scene adverb semantic units from the hierarchical structured semantic tree, construct spatiotemporal constraint relationships, and inject spatiotemporal constraint relationships into the associated region segmentation identifiers through a dynamic correction and transmission process, outputting a temporally coherent dynamic video.

2. The image-text video generation method based on cross-modal semantic mapping according to claim 1, characterized in that, Input product description text, perform hierarchical semantic decoupling, extract core object nouns, attribute adjectives, and scenario adverbs, and construct a hierarchical structured semantic tree, including: Perform syntactic dependency parsing on the product description text to extract the core object nouns as the root node of the semantic tree; identify the direct dependent adjectives that modify the core object nouns and bind them as the first-level child nodes of the root node; identify the adverbial phrases associated with verbs and map them as the second-level time constraint child nodes of the root node; Based on the syntactic distance between the first-level child nodes and the second-level time-constrained child nodes and the root node of the semantic tree, a hierarchical index is constructed to generate a hierarchical structured semantic tree.

3. The image-text video generation method based on cross-modal semantic mapping according to claim 2, characterized in that, Based on a hierarchical structured semantic tree, fine-grained semantic region exploration is performed to identify associated regions of attribute adjectives or scene adverbs, and semantic adaptation correction factors are generated for each associated region, including: Based on the semantics of the first-layer child nodes, the visual space segmentation of the core object is driven, generating the texture segmentation region corresponding to the material class child nodes and the outline segmentation region corresponding to the style class child nodes. For texture segmentation regions, detect whether there are multiple material class child nodes in the same region. If so, generate a weighted fusion material correction factor. For the contour segmentation region, the deviation value between the geometric features and the semantics of the style class child nodes is detected. If the deviation value exceeds the preset threshold, a style correction factor with directional constraints is generated.

4. The image-text video generation method based on cross-modal semantic mapping according to claim 3, characterized in that, Based on the semantics of the first-layer child nodes, the visual space segmentation of the core object is driven, generating texture segmentation regions corresponding to material-type child nodes and contour segmentation regions corresponding to style-type child nodes, including: The semantics of material class adjectives in the first-level child nodes are parsed to drive the modeling of the surface material distribution of the core object and generate the initial material region; boundary continuity optimization is performed on the initial material region, and adjacent sub-regions with the same material properties are merged to generate texture segmentation regions; The semantics of style-class adjectives in the first-level child nodes are parsed to drive the geometric contour modeling of the core object and generate the initial contour region; topological closure optimization is performed on the initial contour region to generate the contour segmentation region; Establish a spatial mapping relationship between texture segmentation regions and contour segmentation regions, and identify the contour partition to which each texture region belongs.

5. The image-text video generation method based on cross-modal semantic mapping according to claim 4, characterized in that, A partitioned coordinated attention mechanism is employed to inject semantic adaptation correction factors into the feature fusion link of the corresponding visual region, generating cross-modal feature vectors. Based on these cross-modal feature vectors, a keyframe sequence is generated, where each frame contains associated region segmentation markers, including: The material correction factor and style correction factor are spatially encoded according to the coordinates of the texture segmentation region and the contour segmentation region to generate an indexed correction matrix. Retrieve region feature vectors that match spatial coordinates from the visual basic feature library, and perform affine transformation on the correction matrix and the region feature vectors to generate cross-modal feature vectors. Based on cross-modal feature vectors, a sequence of key frames is rendered, and texture segmentation region coordinates, contour segmentation region coordinates, and region temporal identifiers are embedded in each frame to form a group of associated region segmentation identifier frames.

6. The image-text video generation method based on cross-modal semantic mapping according to claim 5, characterized in that, Retrieve region feature vectors matching spatial coordinates from the visual feature library, perform affine transformation on the correction matrix and region feature vectors to generate cross-modal feature vectors, including: Material correction factors are spatially mesh-encoded according to the coordinates of texture segmentation regions to generate a material correction matrix; style correction factors are spatially mesh-encoded according to the coordinates of contour segmentation regions to generate a style correction matrix; based on the coordinates of texture segmentation regions, material feature vectors matching the coordinates are retrieved from the visual basic feature library; based on the coordinates of contour segmentation regions, contour feature vectors matching the coordinates are retrieved from the visual basic feature library. An affine transformation is performed between the material correction matrix and the material feature vector to generate optimized material features; an affine transformation is performed between the style correction matrix and the contour feature vector to generate optimized contour features. By fusing optimized material features and optimized contour features, a cross-modal feature vector is generated.

7. The image-text video generation method based on cross-modal semantic mapping according to claim 6, characterized in that, Based on keyframe sequences, dynamic calibration is performed to extract scene adverb semantic units from a hierarchical structured semantic tree, constructing spatiotemporal constraint relationships. Through a dynamic correction and propagation process, these spatiotemporal constraint relationships are injected into the segmentation identifiers of associated regions, outputting a temporally coherent dynamic video, including: Extract the scene adverb semantic units from the second-level time constraint sub-nodes in the semantic tree, parse the time adverbs to generate the starting time period, parse the action adverbs to generate the target action parameters, and construct an action-time period mapping table; In the associated region segmentation identifier frame group, the material correction domain is located based on the texture segmentation region coordinates; the motion correction domain is located based on the contour segmentation region coordinates. When the material similarity between adjacent frames in the material correction domain is lower than the threshold, a transition frame is inserted based on the time period marker in the mapping table; when the motion parameters in the motion correction domain deviate from the constraint values ​​in the mapping table beyond the tolerance, the motion phase is redirected based on the regional temporal identifier, and a temporally coherent dynamic video is output.

8. A text-to-video generation system based on cross-modal semantic mapping, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The semantic decoupling module is used to input product description text, perform hierarchical semantic decoupling, extract core object nouns, attribute adjectives and scenario adverbs, and construct a hierarchical structured semantic tree; The semantic exploration module is used to perform fine-grained semantic region exploration based on a hierarchical structured semantic tree, identify the associated regions of attribute adjectives or scene adverbs, and generate semantic adaptation correction factors for each associated region. The feature fusion module is used to inject semantic adaptation correction factors into the feature fusion link of the corresponding visual region using a partitioned coordinated attention mechanism, generate cross-modal feature vectors, and generate keyframe sequences containing related region segmentation labels based on the cross-modal feature vectors. The dynamic generation module is used to perform dynamic calibration based on key frame sequences, extract scene adverb semantic units from hierarchical structured semantic trees, construct spatiotemporal constraint relationships, inject spatiotemporal constraint relationships into the segmentation identifiers of associated regions through dynamic correction and transmission processes, and output temporally coherent dynamic video.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video generation method and device, electronic equipment and storage medium

    CN117676277A

  • Method for generating video from text based on feature decoupling enhancement

    CN118658106A