Intelligent transcoding method and system for UI design manuscript fused with multi-modal large model

By using a multimodal large model for cross-modal parsing and bidirectional synchronization of UI design drafts, the problem of insufficient understanding of multimodal information in UI design draft to code technology is solved, and high-quality, maintainable cross-platform code generation and development efficiency are achieved.

CN120973379APending Publication Date: 2025-11-18SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511100502.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing UI design draft to code conversion technologies suffer from insufficient understanding of multimodal design information, difficulty in integrating visual styles, interaction logic and business semantics, poor layout adaptability and maintainability of generated code, inability to dynamically optimize component structure, and poor compatibility with design draft version iterations, resulting in low development efficiency.

Method used

A multimodal large model is used for cross-modal joint analysis to generate platform-independent intermediate representations. An improved Cassowary algorithm is used to solve the layout constraint equations, a two-way synchronization mechanism between the design draft and the code is established, incremental updates are achieved by combining the SimHash algorithm, and visual fidelity testing and performance optimization are carried out through the debugging and optimization center.

Benefits of technology

It achieves pixel-level style restoration and precise mapping of business logic, improves code generation quality and maintainability, supports bidirectional traceability between design drafts and code, reduces full-scale refactoring, and improves development efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120973379A_ABST
    Figure CN120973379A_ABST
Patent Text Reader

Abstract

The invention relates to the cross technical field of artificial intelligence and software development, in particular to a UI design draft intelligent transcoding method and system fused with a multi-modal large model, and the method comprises the following steps: carrying out the cross-modal joint analysis of a design draft through a multi-modal analysis engine, and extracting visual features, semantic association and interaction logic; generating a platform-independent intermediate representation containing style description, layout constraint and data binding; generating a multi-platform response type code based on a dynamic layout adapter, and solving a layout constraint equation by adopting an improved CasSowary algorithm; establishing a bidirectional synchronization mechanism of a design draft and codes, and realizing incremental updating through a SimHash algorithm; the method has the beneficial effects that a visual-semantic-code cross-modal cognitive link is creatively constructed, and accurate mapping of pixel-level style reduction and business logic is realized aiming at complex scenes such as multilayer nested components and dynamic interaction logic which are difficult to process by a traditional tool.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence and software development, in particular to a UI design draft intelligent conversion code method and system fusing a multi-modal large model. BACKGROUND

[0002] Currently, the traditional UI design draft conversion code technology mainly relies on manual annotation and rule engine to realize the recognition and code generation of design elements. The existing method usually extracts the control position information by using an image segmentation algorithm, identifies the text content by combining with the OCR technology, and generates the basic layout code through template matching. However, such technology has the following significant defects: firstly, the understanding ability of multi-modal design information is insufficient, and it is difficult to effectively fuse the visual style (such as gradient background), the interaction logic (such as animation trigger condition) and the business semantics (such as data binding relationship); secondly, the layout adaptability and maintainability of the generated code are poor, and the component structure cannot be dynamically optimized according to the target platform characteristics (such as mobile end responsive layout); thirdly, the semantic gap between the design draft and the code leads to low style restoration degree, and complex visual effects (such as shadow superposition, irregular graphics) often need manual secondary adjustment.

[0003] Although the current deep learning-based solution has improved the element detection accuracy, it still faces the following key challenges: the cross-modal alignment technology is not mature, and it is difficult to establish the accurate mapping relationship between the visual features of design elements and the code semantics; the dynamic layout generation algorithm lacks the business context perception ability, resulting in redundant code structure and poor scalability; the multi-platform adaptation mechanism is rigid, and the iOS / Android / Web three-end code needs to be generated separately, which has high maintenance cost. In addition, the existing system has poor compatibility for design draft version iteration, and slight style modification often causes large-scale code reconstruction, which seriously restricts the development efficiency. SUMMARY

[0004] The purpose of the present application is to provide a UI design draft intelligent conversion code method and system fusing a multi-modal large model to solve the problems raised in the background technology.

[0005] To achieve the above purpose, the present application provides the following technical scheme: a UI design draft intelligent conversion code method fusing a multi-modal large model, comprising the following steps:

[0006] 1) performing cross-modal joint analysis on the design draft by using a multi-modal analysis engine to extract visual features, semantic associations and interaction logic;

[0007] 2) generating a platform-independent intermediate representation containing style description, layout constraint and data binding;

[0008] 3) generating multi-platform responsive code based on a dynamic layout adapter, and solving the layout constraint equation by using an improved Cassowary algorithm.

[0009] 4) Establish a two-way synchronization mechanism between design and code, and realize incremental update through SimHash algorithm;

[0010] 5) Perform visual fidelity testing and performance optimization through a debugging optimization center, and output cross-platform code that meets industrial standards.

[0011] Preferably, the multi-modal analysis engine in step 1) includes: a visual feature extraction module that uses a cascade convolutional network to analyze vector graphic features and extract mixed mode and mask path parameters; a semantic association modeling module that generates a semantic dependency graph containing interactive intent based on the RoBERTa architecture; a cross-modal alignment module that optimizes the cross-modal embedding space of the CLIP-ViT model through contrastive learning and uses a two-way loss function to train positive and negative samples.

[0012] Preferably, the intermediate representation in step 2) includes: a semantic type classification field that distinguishes between button and list view component types; a set of style rules injected with CSS variables; a layout constraint equation array; and a data binding configuration of data source address and format converter.

[0013] Preferably, the multi-platform responsive code generated based on the dynamic layout adapter in step 3) includes: SwiftUI declarative code integrated with the Combine framework for the iOS platform; Jetpack Compose component tree containing Coroutine scope for the Android platform; and React functional components using CSS Module to isolate styles for the Web platform.

[0014] Preferably, the two-way synchronization mechanism in step 4) includes: updating design parameters inversely through AST analysis of code modifications; triggering local regeneration when the Hamming distance exceeds a threshold; and supporting version history backtracking and difference comparison functions;

[0015] The debugging optimization center in step 5) includes a visual fidelity testing module based on the structural similarity algorithm, with a dynamic threshold set to SSIM ≥ 0.92; a real-time rendering pipeline monitoring module that automatically optimizes component structure when FPS is below 60; and a platform-specific Lint rule checker that enforces code compliance with SwiftUI / Compose / React best practices.

[0016] A system for a UI design draft intelligent code conversion method based on a multi-modal large model, comprising:

[0017] A multi-modal analysis engine that builds a visual semantic joint analysis channel, including visual feature extraction, semantic association modeling, and cross-modal alignment.

[0018] The intelligent generation engine constructs a two-stage code generation system to achieve accurate conversion from design semantics to platform code, including UI meta description generation and multi-platform code conversion;

[0019] The dynamic layout adapter introduces a hybrid optimization strategy to solve the responsiveness problem of multi-terminal layout, including a constraint solving engine, a breakpoint decision mechanism, and visual focus optimization.

[0020] The bidirectional synchronization module establishes a version collaboration mechanism between design and code, supports closed-loop iteration, uses the SimHash algorithm to generate semantic fingerprints, triggers local regeneration when the Hamming distance exceeds the threshold, and updates the design parameters by reversing code modification through AST parsing.

[0021] The debugging and optimization center builds a full-chain quality assurance system, including visual fidelity testing, performance optimizer, and code style check functions.

[0022] In preferred multimodal parsing engines:

[0023] Visual feature extraction uses a cascaded convolutional network to process the design draft input, analyzes vector graphic features from Sketch / Figma source files, and establishes a visual feature matrix that includes professional parameters such as blending mode and mask path.

[0024] Semantic association modeling is based on the RoBERTa architecture to build a bidirectional attention mechanism, which parses the annotated text in the design draft and generates a semantic dependency graph containing interactive intent. Nodes represent component functions and edge weights represent event triggering priorities.

[0025] Cross-modal alignment optimizes the cross-modal embedding space of the CLIP-ViT model through contrastive learning. A dual-path loss function is designed, with positive samples coming from manually labeled code snippets and negative samples from erroneous codes generated by random perturbation.

[0026] Preferably, in the intelligent generation engine: UI meta description generation encodes multimodal features into platform-independent intermediate representations, and multi-platform code conversion generates target code based on parameterized templates. Typical conversion strategies include: iOS platform: generating SwiftUI declarative code and automatically integrating the Combine framework to achieve data binding; Android platform: building a Jetpack Compose component tree with built-in Coroutine scope management; Web platform: outputting React functional components and using CSSModules to isolate style conflicts.

[0027] Preferably, in the dynamic layout adapter:

[0028] The constraint solving engine uses the improved Cassowary algorithm to solve the layout equations and supports weight priority settings, such as: Button.width==0.8*Parent.width{strength:REQUIRED}; Visual focus optimization combines eye-tracking datasets to assign layout flexibility coefficients to core operation areas, and the calculation formula is: FlexWeight=1 / (1+e^(-(AttentionScore-0.5))).

[0029] Preferably, in the debugging and optimization center: visual fidelity testing uses the structural similarity algorithm (SSIM) to compare the rendering results with the design draft, and sets a dynamic threshold: if(ssimScore<0.92)triggerRedraw(); the performance optimizer monitors the rendering pipeline in real time, and automatically applies optimization strategies when the FPS is below 60; code style checks are integrated with the platform's proprietary Lint rule library to ensure that the generated code conforms to best practices.

[0030] Compared with the prior art, the beneficial effects of the present invention are:

[0031] This invention proposes a method and system for intelligently converting UI design drafts into code by integrating a multimodal large model. It creatively constructs a cross-modal cognitive link of "visual-semantic-code," achieving pixel-level style restoration and precise mapping of business logic for complex scenarios such as multi-layered nested components and dynamic interaction logic that are difficult for traditional tools to handle. By introducing the context-aware capabilities of the multimodal large model, the system significantly improves code generation quality and maintainability. More groundbreakingly, it supports bidirectional traceability between design drafts and code; when developers modify the size of the bottom navigation bar icons, the system automatically locates the associated layout constraints and style files, enabling partial updates rather than a full refactoring. Attached Figure Description

[0032] Figure 1 This is a flowchart of the method of the present invention;

[0033] Figure 2 This is a system block diagram of the present invention. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the present invention clear and complete, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only some, not all, embodiments of the present invention, and are merely illustrative of the embodiments of the present invention. They are not intended to limit the embodiments of the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Example 1, please refer to Figure 1This invention provides a technical solution: a method for intelligently converting UI design drafts into code by integrating multimodal large models, taking the development scenario of e-commerce App product detail page as an example, including the following steps:

[0036] (1) Construction of cross-domain datasets

[0037] We scraped 100,000 publicly available design drafts from the Figma community, covering e-commerce, social media, and utility applications, including layer grouping, style parameters (RGBA values, corner radius), and interaction annotations (click / swipe events).

[0038] Match the Android / iOS / Web code in the GitHub open source project and extract the component implementations that correspond to the design draft (such as the mapping relationship between RecyclerView and UICollectionView).

[0039] The manual annotation team adds semantic tags to complex interactive scenarios (such as nested scrolling views and dynamic data binding), and the annotation accuracy is ensured to be ≥98% through cross-validation.

[0040] (2) Pre-training of multimodal alignment model

[0041] Based on the DeCLIP architecture, a three-stage training process is designed:

[0042] Visual-Text Alignment: Input a screenshot of the design draft and the labeled text (such as "the bottom navigation bar contains homepage, message, and personal center icons"), and reduce the feature distance through comparison and learning.

[0043] Visual-code alignment: Input the design draft and target code (SwiftUI / Compose / React) into a dual-tower network to calculate the cosine similarity at the component level.

[0044] Full-modal joint optimization: An adversarial training strategy is introduced, the discriminator network distinguishes the distribution differences between generated code and manually written code, and the generator network is continuously iterated and optimized.

[0045] (3) Domain adaptation and dynamic fine-tuning

[0046] 1. Knowledge Injection

[0047] Inject e-commerce design guidelines (human-computer interaction guidelines) to force the generated model to have a button size ≥10mm and a font contrast ≥4.5:1. Build e-commerce-specific components and add an adaptation matrix to the MLP layer of the visual encoder using LoRA.

[0048] 2. Dynamically optimize the process

[0049] When the system detects that a user has modified the layout of the generated e-commerce promotion page three times in a row, it first automatically extracts the differences in modification (e.g., the spacing is adjusted from 8dp to 12dp, and the title font is made bold). At the same time, it performs gradient backpropagation in the latent space, updates the adapter parameters, and continuously improves the generation accuracy of similar components after the update.

[0050] (4) Deployment and Verification

[0051] 1. Design Draft Input Analysis

[0052] Input the design draft of the e-commerce promotional activity page (including dynamic price tags, countdown components, and 3D product display area).

[0053] The system automatically identifies the binding path between the price tag and the backend data, and the display area relies on the Three.js engine to generate WebGL context configuration parameters.

[0054] 2. Code generation and optimization

[0055] Generates SwiftUI + Combine declarative code for iOS; Jetpack Compose + Kotlin Flow code for Android; and React Hooks + CSS Module combined code for Web. It also generates a responsive constraint optimizer, such as automatically adding responsive breakpoints when the width of a 3D container exceeds the safe zone in the mobile viewport.

[0056] 3. End-to-end verification

[0057] Visual reproduction test: Use a pixel-level comparison tool (PhantomCSS) to detect the differences between the generated page and the design draft, requiring the SSIM value of key areas to be greater than 0.9.

[0058] Performance stress test: Simulate tens of thousands of concurrent accesses and optimize the rendering pipeline through the Chrome Performance panel.

[0059] Cross-platform verification: The generated code for iOS was statically analyzed using Xcode (0 warnings), and the code for Android was tested on a Pixel 6 Pro.

[0060] Example 2, based on Example 1, proposes a system for the intelligent code-to-code method of UI design drafts fused with multimodal large models as described in claim 5, comprising:

[0061] 1. Multimodal parsing engine

[0062] The multimodal parsing engine constructs a joint visual-semantic parsing channel, bridging the semantic gap between design drafts and code. This module mainly includes three parts: visual feature extraction, semantic association modeling, and cross-modal alignment.

[0063] Visual feature extraction: CascadeCNN is used to process the design draft input. Vector graphics features are parsed from Sketch / Figma source files to establish a visual feature matrix that includes professional parameters such as blend mode and mask path.

[0064] Semantic association modeling: Based on the RoBERTa architecture, a bidirectional attention mechanism is built to parse the annotated text in the design draft (such as "click to jump to product details") and generate a semantic dependency graph (SDG) containing the interaction intent. Nodes represent component functions, and edge weights represent the event triggering priority.

[0065] Cross-modal alignment: The cross-modal embedding space of the CLIP-ViT model is optimized through contrastive learning. A dual-path loss function is designed, where positive samples come from manually labeled code snippets and negative samples are erroneous codes generated by random perturbation.

[0066] 2. Intelligent Generation Engine

[0067] The intelligent generation engine module mainly constructs the following two-stage code generation system to achieve accurate conversion from design semantics to platform code.

[0068] 1. UI Meta Description Generation: Encodes multimodal features into a platform-independent intermediate representation (UI-MetaJSON). Key data structures include:

[0069] interface UIComponent{

[0070] semanticType: "Button"|"ListView"; / / Semantic category

[0071]

[0072] 2. Multi-platform code conversion: Generate target code based on parameterized templates. Typical conversion strategies include:

[0073] iOS platform: Generates SwiftUI declarative code and automatically integrates the Combine framework for data binding.

[0074] @ObservedObject var productModel:ProductViewModel

[0075] Android platform: Building Jetpack Compose component trees with built-in Coroutine scope management.

[0076] val scope=rememberCoroutineScope()

[0077] Web platform: Outputs React functional components, using CSS Modules to isolate style conflicts.

[0078] import styles from'. / ProductCard.module.css'

[0079] 3. Dynamic Layout Adapter

[0080] Dynamic layout adapters introduce hybrid optimization strategies to address the responsiveness challenges of multi-platform layouts:

[0081] Constraint Solving Engine: Employs an improved Cassowary algorithm to solve layout equations, supporting weight priority settings.

[0082] Button.width==0.8*Parent.width{strength:REQUIRED}

[0083] Breakpoint decision mechanism: Dynamically select layout scheme based on device feature library, such as when an iOS device is detected:

[0084]

[0085] }

[0086] Visual focus optimization: By combining eye-tracking datasets, layout flexibility coefficients are assigned to core operational areas (such as shopping cart buttons).

[0087] FlexWeight = 1 / (1+e)

[0088] -(AttentionScore-0.5) )

[0090] 4. Two-way synchronization module

[0091] The bidirectional synchronization module establishes a version collaboration mechanism between design and code, while also supporting closed-loop iteration. It uses the SimHash algorithm to generate semantic fingerprints for the design draft and code. When the Hamming distance exceeds a threshold (ΔH ≥ 0.15), it triggers local regeneration to achieve incremental updates. Secondly, it uses AST parsing to analyze code modifications and then updates the design draft parameters in reverse.

[0092] 5. Debugging and Optimization Center

[0093] The debugging and optimization center has built a full-chain quality assurance system to ensure the industrial-grade reliability of delivered code, including functions such as visual fidelity testing, performance optimizer, and code style check.

[0094] 1. Visual fidelity test: The rendered result is compared with the design draft using the Structural Similarity Algorithm (SSIM), and a dynamic threshold is set.

[0095] if(ssimScore<0.92)triggerRedraw()

[0096] 2. Performance Optimizer: Monitors the rendering pipeline in real time and automatically applies optimization strategies when FPS drops below 60, such as:

[0097] / / Detect list scrolling stuttering

[0098] if (scrollFPS < 45) {

[0099] replaceVStackWithLazyVStack()

[0100] }

[0101] 3. Code style check: Integrates the platform's proprietary Lint rule library to ensure generated code conforms to best practices, such as mandating that Swift code follow these guidelines.

[0102] guard let safeValue=optionalValue else{return}.

[0103] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for intelligently converting UI design drafts into code by integrating multimodal large models, characterized in that: Includes the following steps: 1) The design draft is jointly analyzed across modalities using a multimodal parsing engine to extract visual features, semantic relationships, and interaction logic; 2) Generate a platform-independent intermediate representation that includes style descriptions, layout constraints, and data binding; 3) Generate multi-platform responsive code based on dynamic layout adapters, and use an improved Cassowary algorithm to solve the layout constraint equations; 4) Establish a two-way synchronization mechanism between design drafts and code, and implement incremental updates through the SimHash algorithm; 5) Visual fidelity testing and performance optimization are performed through the debugging and optimization center to output cross-platform code that conforms to industry standards.

2. The method for intelligently converting UI design drafts into code based on a multimodal large model as described in claim 1, characterized in that: The multimodal parsing engine described in step 1) includes: a visual feature extraction module, which uses a cascaded convolutional network to parse vector graphics features and extract blending mode and mask path parameters; a semantic association modeling module, which generates a semantic dependency graph containing interactive intent based on the RoBERTa architecture; and a cross-modal alignment module, which optimizes the cross-modal embedding space of the CLIP-ViT model through contrastive learning and uses a dual-path loss function for training positive and negative samples.

3. The method for intelligently converting UI design drafts into code based on a multimodal large model as described in claim 2, characterized in that: The intermediate representation described in step 2) includes: a semantic type classification field to distinguish between button and list view component types; a set of style rules for CSS variable injection; an array of layout constraint equations; and data binding configuration for data source address and format converter.

4. The method for intelligently converting UI design drafts into code based on a multimodal large model as described in claim 3, characterized in that: Step 3) generates multi-platform responsive code based on the dynamic layout adapter. Specifically, this includes: generating SwiftUI declarative code integrating the Combine framework on the iOS platform; generating a Jetpack Compose component tree containing Coroutine scope on the Android platform; and outputting React functional components with CSS Module isolation styles on the Web platform.

5. The method for intelligently converting UI design drafts into code based on a multimodal large model as described in claim 4, characterized in that: The bidirectional synchronization mechanism in step 4) includes: modifying the reverse update design parameters through AST parsing code; triggering local regeneration when the Hamming distance exceeds a threshold; and supporting version history backtracking and difference comparison functions. The debugging and optimization center mentioned in step 5) includes a visual fidelity testing module based on structural similarity algorithm, with a dynamic threshold set to SSIM≥0.92; a real-time rendering pipeline monitoring module that automatically optimizes component structure when FPS is below 60; and a platform-specific Lint rule checker that forces code to conform to SwiftUI / Compose / React best practices.

6. A system for the intelligent code-to-code method of UI design drafts merging multimodal large models according to claim 5, characterized in that: include: The multimodal parsing engine constructs a joint visual-semantic parsing channel, comprising three parts: visual feature extraction, semantic association modeling, and cross-modal alignment. The intelligent generation engine constructs a two-stage code generation system to achieve accurate conversion from design semantics to platform code, including UI meta description generation and multi-platform code conversion; The dynamic layout adapter introduces a hybrid optimization strategy to solve the responsiveness problem of multi-terminal layout, including a constraint solving engine, a breakpoint decision mechanism, and visual focus optimization. The bidirectional synchronization module establishes a version collaboration mechanism between design and code, supports closed-loop iteration, uses the SimHash algorithm to generate semantic fingerprints, triggers local regeneration when the Hamming distance exceeds the threshold, and updates the design parameters by reversing code modification through AST parsing. The debugging and optimization center builds a full-chain quality assurance system, including visual fidelity testing, performance optimizer, and code style check functions.

7. The system according to claim 6, characterized in that: In the multimodal parsing engine: Visual feature extraction uses a cascaded convolutional network to process the design draft input, analyzes vector graphic features from Sketch / Figma source files, and establishes a visual feature matrix that includes professional parameters such as blending mode and mask path. Semantic association modeling is based on the RoBERTa architecture to build a bidirectional attention mechanism, which parses the annotated text in the design draft and generates a semantic dependency graph containing interactive intent. Nodes represent component functions and edge weights represent event triggering priorities. Cross-modal alignment optimizes the cross-modal embedding space of the CLIP-ViT model through contrastive learning. A dual-path loss function is designed, with positive samples coming from manually labeled code snippets and negative samples from erroneous codes generated by random perturbation.

8. The system according to claim 7, characterized in that: In the intelligent generation engine: UI meta description generation encodes multimodal features into platform-independent intermediate representations. Multi-platform code conversion generates target code based on parameterized templates. Typical conversion strategies include: iOS platform: generating SwiftUI declarative code and automatically integrating the Combine framework to achieve data binding; Android platform: building a Jetpack Compose component tree with built-in Coroutine scope management; Web platform: outputting React functional components and using CSS Modules to isolate style conflicts.

9. A system according to claim 8, characterized in that: In the dynamic layout adapter: The constraint solving engine uses the improved Cassowary algorithm to solve the layout equations and supports weight priority settings, such as: Button.width==0.8*Parent.width{strength:REQUIRED}; Visual focus optimization combines eye-tracking datasets to assign layout flexibility coefficients to core operation areas, and the calculation formula is: FlexWeight=1 / (1+e^(-(AttentionScore-0.5))).

10. A system according to claim 9, characterized in that: In the debugging and optimization center: the visual fidelity test uses the Structural Similarity Algorithm (SSIM) to compare the rendered results with the design draft, and sets a dynamic threshold: if (ssimScore < 0.92) triggerRedraw(); the performance optimizer monitors the rendering pipeline in real time, and automatically applies optimization strategies when the FPS is below 60; the code style check integrates the platform's proprietary Lint rule library to ensure that the generated code conforms to best practices.

Citation Information

Cited By

  • Near storage processing unit for data cleaning and data lake cleaning system

    CN122173027A

  • Near-memory processing unit and data lake cleansing system oriented to data cleansing

    CN122173027B