Low-code intelligent development assisting method, system and equipment based on multi-mode perception

The multi-modal perception-based low-code system addresses high development barriers and inefficiencies by converting diverse user inputs into structured data for intelligent code generation and collaboration, improving efficiency and code quality.

CN120315686APending Publication Date: 2025-07-15SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510394025.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing low-code platform has high development thresholds, low development efficiency and code quality, low collaboration efficiency and insufficient intelligence.

Method used

Receive user needs through multimodal perception technology, generate structured data, generate code based on Transformer, optimize code with graph neural network, and improve collaboration efficiency through context perception and conflict detection algorithms.

Benefits of technology

Significantly lower the development threshold, improve development efficiency and code quality, improve team collaboration efficiency, and reduce manual code writing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120315686A_ABST
    Figure CN120315686A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of software development, and discloses a low-code intelligent development assisting method, system and equipment based on multi-modal perception, and the method comprises the steps: receiving a demand input by a user through voice, a hand-drawn flow chart or natural language description, and converting the input demand into structured data; generating a corresponding code based on the structured data; and based on the development guidance information of the user, optimizing the code, and recommending the component. According to the method, by combining the multi-mode perception technology and the low-code development platform, intelligent code generation, optimization and cooperation can be realized, the development threshold is remarkably reduced, and the development efficiency is improved. Through the method, a developer can complete a low-code development task more efficiently, the manual code writing time is shortened, the code quality is improved, and meanwhile the cooperation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of software development technologies, for example, to a low-code intelligent development assistance method, system, and device based on multimodal perception. Background Art

[0002] With the acceleration of digital transformation, low-code development platforms have gradually become the preferred tools for enterprises to quickly build application programs due to their high efficiency and ease of use. However, the existing low-code platforms still have the following problems:

[0003] High development threshold: Although low-code platforms reduce the programming difficulty, for non-professional developers, certain technical backgrounds and comprehension abilities are still required.

[0004] Functional limitations: Existing platforms usually rely on predefined components and templates, making it difficult to meet complex or personalized development requirements.

[0005] Low collaboration efficiency: In scenarios of multi-developer collaboration, code conflicts and communication costs are relatively high, affecting development efficiency.

[0006] Insufficient intelligence: Existing platforms lack intelligent assistance functions such as automatic code generation, optimization, and recommendation, resulting in limited improvement in development efficiency.

[0007] Therefore, the existing low-code platforms have problems of high development threshold, low development efficiency, and low code quality.

[0008] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of this application. Summary of the Invention

[0009] To have a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Instead, it serves as a preface to the subsequent detailed description.

[0010] The low-code intelligent development assistance method, system, and device based on multimodal perception in this disclosure solve the problems of high development threshold, low development efficiency, and low code quality in existing low-code platforms.

[0011] Embodiments of this disclosure provide a low-code intelligent development assistance method based on multimodal perception, which includes:

[0012] Receiving the requirements input by the user through voice, hand-drawn flowcharts, or natural language descriptions, and converting the input requirements into structured data;

[0013] Generating corresponding code based on the structured data;

[0014] Optimize the code and recommend components based on the user's development guidance information.

[0015] In some embodiments, receive the requirements input by the user through voice, hand-drawn flowcharts, or natural language descriptions, and convert the input requirements into structured data, including:

[0016] For the requirements input by the user through voice, use a preset speech recognition algorithm to convert the requirements input by the user through voice into text, and combine with natural language processing algorithms to extract key information and convert it into instructions for the low-code platform;

[0017] For the requirements input by the user through a hand-drawn flowchart, use a preset computer vision recognition algorithm to recognize the requirements input by the user through the hand-drawn flowchart, obtain the corresponding graphics and text, and perform relationship modeling on the recognized graphics and text through a preset graph neural network to convert the recognition result into structured data;

[0018] For the requirements input by the user through natural language descriptions, use a preset language model to analyze the requirements input by the user through natural language descriptions, extract key information and convert it into instructions for the low-code platform.

[0019] In some embodiments, generate corresponding code based on the structured data, including:

[0020] According to the code generation model of Transformer, based on the structured data and combined with the characteristics of the low-code platform, generate components or modules that conform to the platform specifications;

[0021] According to the template matching algorithm of graph embedding, combined with the graph neural network, based on the structured data, determine the corresponding template, and based on the determined template, expand the generated components or modules that conform to the platform specifications to obtain the expanded code;

[0022] Optimize the expanded code based on a preset static code analysis tool and dynamic code analysis algorithm to generate the corresponding code.

[0023] In some embodiments, optimize the code and recommend components based on the user's development guidance information, including:

[0024] Based on the user's development guidance information, determine the user's development habits and preferences;

[0025] Based on the user's development habits and preferences, adjust the code and recommend components.

[0026] In some embodiments, adjust the code and recommend components based on the user's development habits and preferences, including:

[0027] Adjust the code based on the user's development habits and preferences;

[0028] Based on the user's development habits and preferences, according to the structured data, and combined with the component relationships in the preset knowledge graph, recommend the corresponding components.

[0029] In some embodiments, the method further includes:

[0030] Obtain the user's development guidance information in real time;

[0031] Predict the user's behavior based on the real-time obtained development guidance information;

[0032] Based on the predicted user behavior, adjust the generated code in real time and recommend components.

[0033] In some embodiments, the method further includes:

[0034] Real-time synchronize the development collaboration of multiple users, and detect the collaboration conflicts of users based on the conflict detection algorithm of the graph neural network;

[0035] Based on the rule-based merging strategy or machine learning model, merge the user's collaboration conflicts or generate multiple solutions.

[0036] The embodiments of the present disclosure provide a low-code intelligent development assistance system based on multimodal perception. The system includes:

[0037] A multimodal perception module, configured to receive the requirements input by the user through voice, hand-drawn flowcharts or natural language descriptions, and convert the input requirements into structured data;

[0038] An intelligent code generation module, configured to generate corresponding code based on the structured data;

[0039] A context awareness module, configured to optimize the code and recommend components based on the user's development guidance information.

[0040] The embodiments of the present disclosure provide an electronic device, which includes at least one processor;

[0041] And a memory communicatively connected to at least one processor;

[0042] The memory stores instructions executable by at least one processor. The instructions are executed by at least one processor so that at least one processor can execute the above-mentioned low-code intelligent development assistance method based on multimodal perception.

[0043] An embodiment of the present disclosure provides a storage medium storing program instructions that, when running, execute the above-mentioned low-code intelligent development assistance method based on multi-modal perception.

[0044] The low-code intelligent development assistance method, system, device, and storage medium based on multi-modal perception provided by the embodiments of the present disclosure can achieve the following technical effects:

[0045] By combining multi-modal perception technology and a low-code development platform, the embodiments of the present disclosure can achieve intelligent code generation, optimization, and collaboration, significantly reducing the development threshold and improving development efficiency. Through this method, developers can complete low-code development tasks more efficiently, reduce the time for manually writing code, improve code quality, and at the same time enhance collaboration efficiency.

[0046] The above general description and the following description are merely exemplary and explanatory and are not used to limit this application. Description of the Drawings

[0047] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations and the drawings do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and among them:

[0048] Figure 1 is a flowchart of a low-code intelligent development assistance method based on multi-modal perception provided by an embodiment of the present disclosure;

[0049] Figure 2 is a structural diagram of a low-code intelligent development assistance system based on multi-modal perception provided by an embodiment of the present disclosure;

[0050] Figure 3 is a structural diagram of a low-code intelligent development assistance device based on multi-modal perception provided by an embodiment of the present disclosure. Detailed Embodiments

[0051] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below in conjunction with the drawings. The attached drawings are only for reference and illustration purposes and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, multiple details are provided to provide a full understanding of the disclosed embodiments. However, one or more embodiments can still be implemented without these details. In other cases, well-known structures and systems can be shown in a simplified manner.

[0052] In the embodiments of the present disclosure, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so as to implement the embodiments of the present disclosure described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions.

[0053] Unless otherwise specified, the term "plurality" means two or more.

[0054] In the embodiments of the present disclosure, the character " / " indicates that the front and rear objects are in an "or" relationship. For example, A / B means: A or B.

[0055] The term "and / or" is an associative relationship describing an object, indicating that three relationships can exist. For example, A and / or B means: A or B, or, the three relationships of A and B.

[0056] The term "correspond to" may refer to an associative relationship or a binding relationship. A corresponding to B means that there is an associative relationship or a binding relationship between A and B.

[0057] Next, a low-code intelligent development assistance method, system, device, and storage medium based on multi-modal perception provided by the embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0058] Figure 1 It is a schematic flowchart of a low-code intelligent development assistance method based on multi-modal perception provided by the embodiments of the present disclosure.

[0059] Combined with Figure 1 As shown, the low-code intelligent development assistance method based on multi-modal perception may include:

[0060] S101, receiving the requirements input by the user through voice, hand-drawn flowchart or natural language description, and converting the input requirements into structured data;

[0061] S102, generating corresponding code based on the structured data;

[0062] S103, optimizing the code based on the user's development guidance information and recommending components.

[0063] In some embodiments, receiving the requirements input by the user through voice, hand-drawn flowchart or natural language description, and converting the input requirements into structured data may include:

[0064] For the requirements input by the user through voice, using a preset speech recognition algorithm to convert the requirements input by the user through voice into text, and combining natural language processing algorithms to extract key information and convert it into instructions for the low-code platform;

[0065] For the requirements input by the user through a hand-drawn flowchart, use a preset computer vision recognition algorithm to recognize the requirements input by the user through the hand-drawn flowchart, obtain the corresponding graphics and text, and perform relationship modeling on the recognized graphics and text through a preset graph neural network to convert the recognition result into structured data;

[0066] For the requirements input by the user through natural language description, use a preset language model to analyze the requirements input by the user through natural language description, extract key information and convert it into instructions for the low-code platform.

[0067] In some embodiments, generating corresponding code based on the structured data may include:

[0068] According to the code generation model of Transformer, based on the structured data and combined with the characteristics of the low-code platform, generate components or modules that meet the platform specifications;

[0069] According to the template matching algorithm of graph embedding, combined with the graph neural network, based on the structured data, determine the corresponding template, and based on the determined template, expand the generated components or modules that meet the platform specifications to obtain the expanded code;

[0070] Based on a preset static code analysis tool and dynamic code analysis algorithm, optimize the expanded code to generate the corresponding code.

[0071] In some embodiments, optimizing the code and recommending components based on the user's development guidance information may include:

[0072] Based on the user's development guidance information, determine the user's development habits and preferences;

[0073] Based on the user's development habits and preferences, adjust the code and recommend components.

[0074] In some embodiments, optimizing the code and recommending components based on the user's development habits and preferences may include:

[0075] Based on the user's development habits and preferences, adjust the code;

[0076] Based on the user's development habits and preferences, according to the structured data, combined with the component relationships in the preset knowledge graph, recommend the corresponding components.

[0077] In some embodiments, Figure 1 the method in [[ ]] may further include:

[0078] Obtain the user's development guidance information in real time;

[0079] Predict user behavior based on the development guidance information obtained in real time;

[0080] Based on the predicted user behavior, adjust the generated code in real time and recommend components.

[0081] In some embodiments, Figure 1 the method in may further include:

[0082] Synchronize the development collaboration of multiple users in real time, and detect the collaboration conflicts of users based on the conflict detection algorithm of the graph neural network;

[0083] Based on the rule-based merging strategy or machine learning model, merge the collaboration conflicts of users or generate multiple solutions.

[0084] By combining multi-modal perception technology and a low-code development platform, the present disclosure can achieve intelligent code generation, optimization, and collaboration, significantly reducing the development threshold and improving the development efficiency. Through this method, developers can complete low-code development tasks more efficiently, reduce the time for manually writing code, improve the code quality, and at the same time improve the team collaboration efficiency.

[0085] Compared with Figure 1 the low-code intelligent development assistance method based on multi-modal perception in, the present disclosure also provides a low-code intelligent development assistance system based on multi-modal perception. It can be understood that it can be based on Figure 2 the function description of the corresponding system module in, and further explain Figure 1 the low-code intelligent development assistance method based on multi-modal perception in. As Figure 2 shown, the system may specifically include:

[0086] A multi-modal perception module 201, configured to receive the requirements input by the user through voice, hand-drawn flowcharts, or natural language descriptions, and convert the input requirements into structured data;

[0087] An intelligent code generation module 202, configured to generate corresponding code based on the structured data;

[0088] A context awareness module 203, configured to optimize the code and recommend components based on the user's development guidance information.

[0089] In some embodiments, receiving the requirements input by the user through voice, hand-drawn flowcharts, or natural language descriptions and converting the input requirements into structured data includes:

[0090] For the requirements input by the user through voice, use the preset speech recognition algorithm to convert the requirements input by the user through voice into text, and combine with the natural language processing algorithm to extract key information and convert it into instructions for the low-code platform;

[0091] For the requirements input by the user through a hand-drawn flowchart, use the preset computer vision recognition algorithm to recognize the requirements input by the user through the hand-drawn flowchart, obtain the corresponding graphics and text, and perform relationship modeling on the recognized graphics and text through the preset graph neural network to convert the recognition result into structured data;

[0092] For the requirements input by the user through natural language description, use the preset language model to analyze the requirements input by the user through natural language description, extract key information and convert it into instructions for the low-code platform.

[0093] In some embodiments, based on the structured data, generate corresponding code, including:

[0094] According to the code generation model of Transformer, based on the structured data and combined with the characteristics of the low-code platform, generate components or modules that meet the platform specifications;

[0095] According to the template matching algorithm of graph embedding, combined with the graph neural network, based on the structured data, determine the corresponding template, and based on the determined template, expand the generated components or modules that meet the platform specifications to obtain the expanded code;

[0096] Based on the preset static code analysis tool and dynamic code analysis algorithm, optimize the expanded code to generate the corresponding code.

[0097] In some embodiments, based on the user's development guidance information, optimize the code and recommend components, including:

[0098] Based on the user's development guidance information, determine the user's development habits and preferences;

[0099] Based on the user's development habits and preferences, adjust the code and recommend components.

[0100] In some embodiments, based on the user's development habits and preferences, adjust the code and recommend components, including:

[0101] Based on the user's development habits and preferences, adjust the code;

[0102] Based on the user's development habits and preferences, according to the structured data, combined with the component relationships in the preset knowledge graph, recommend the corresponding components.

[0103] In some embodiments, the context awareness module 203 can also be used to obtain the user's development guidance information in real time;

[0104] Predict the user's behavior based on the development guidance information obtained in real time;

[0105] Based on the predicted user behavior, adjust the generated code in real time and recommend components.

[0106] In some embodiments, Figure 2 The low-code intelligent development assistance system based on multimodal perception shown can also include a conflict resolution module, which can be used to synchronize the development collaboration of multiple users in real time and detect the collaboration conflicts of users based on the conflict detection algorithm of the graph neural network;

[0107] Based on a rule-based merging strategy or a machine learning model, merge the collaboration conflicts of users or generate multiple solutions.

[0108] The above development process can be divided into an input stage, a processing stage, an optimization stage, and a collaboration stage.

[0109] Input stage: The developer inputs requirements through voice, hand-drawn flowcharts, or natural language descriptions.

[0110] Processing stage: The multimodal perception module converts the input into structured data, and the intelligent code generation module generates the corresponding code or low-code components.

[0111] Optimization stage: The context awareness module optimizes the generated code according to the developer's context and recommends relevant components.

[0112] Collaboration stage: Multiple developers collaborate in real time, and the system automatically detects and resolves code conflicts.

[0113] In the above system, the multi-modal perception module serves as the input layer of the system, aiming to convert the developer's requirements into structured instructions recognizable by the low-code platform through various means such as speech, hand-drawn flowcharts, and natural language descriptions. For speech recognition and natural language understanding, an improved Transformer architecture (such as Wav2Vec 2.0 + Transformer) is adopted, combined with an adaptive attention mechanism to improve the accuracy and robustness of speech recognition. For the low-code development scenario, the NLU model is optimized to enable it to recognize specific development terms and instructions. This means introducing multi-language support, allowing developers to use different languages for speech input and performing semantic alignment through cross-language models (such as mBERT) to ensure the consistency of multi-language input. For hand-drawn flowcharts, a CNN architecture combined with a generative adversarial network (GAN) and an attention mechanism is used to improve the recognition accuracy of hand-drawn flowcharts. The relationships between the recognized graphics and text are modeled through a graph neural network (GNN) to generate a more accurate logical flow. At the same time, the graphics and text in the hand-drawn flowchart are recognized through a convolutional neural network (CNN) and converted into an executable logical flow. This enables real-time feedback and correction of hand-drawn flowcharts, allowing developers to view the recognition results in real-time during the hand-drawing process and make adjustments. For natural language description processing, an encoder-decoder architecture based on Transformer (such as T5) is adopted, combined with domain adaptation technology to perform semantic parsing on natural language descriptions. A context-aware pre-trained model (such as RoBERTa) is introduced to dynamically adjust the generated instructions or parameters. Support for multi-turn conversations in natural language descriptions is achieved, allowing developers to gradually refine the requirement descriptions through multi-turn interactions, and the system dynamically adjusts the functions of the generated low-code components according to each turn of the conversation.

[0114] The application scenarios of this module are extensive. For example, non-professional developers can describe their requirements through speech, and the system automatically generates corresponding low-code components, or developers can design business processes through hand-drawn flowcharts, and the system converts them into executable workflows. Existing multi-modal inputs are usually limited to a single application scenario. This patent integrates speech, images, and natural language descriptions deeply into the low-code development platform through the multi-modal perception module, supports multi-language input and multi-turn conversations. Its core value lies in reducing the development threshold, enabling users without a technical background to easily use the low-code platform, while significantly improving development efficiency and reducing the time for manual input and configuration.

[0115] As the core processing layer of the system, the intelligent code generation module is responsible for automatically generating low-code components or code snippets based on the output of the multimodal perception module and providing optimization suggestions. This module is based on the code generation model of Transformer (such as Codex, CodeGen), generates code snippets according to the input, introduces a multi-task learning framework, enabling the code generation model to learn both code generation and code optimization tasks simultaneously. Combining the characteristics of the low-code platform and reinforcement learning, while generating components or modules that comply with the platform specifications, it dynamically adjusts the quality and efficiency of the generated code. This can support the adaptive optimization of code generation and dynamically adjust the complexity and performance of the generated code according to project requirements and context. At the same time, through the semantic matching algorithm, the most suitable low-code template is selected according to the input requirements. Adopting a template matching algorithm based on graph embedding and combining it with a graph neural network (GNN) to improve the accuracy and efficiency of template matching. Introducing dynamic parameter filling technology to support developers' custom extensions based on templates. In addition, use static code analysis tools (such as ESLint, Pylint) to check the generated code, and combine static code analysis tools and dynamic code analysis techniques to perform multi-dimensional optimization on the generated code. Introduce a machine learning-based code quality evaluation model to automatically evaluate the quality of the generated code and provide optimization suggestions.

[0116] Its application scenarios include that developers input functional requirements and the system automatically generates corresponding low-code components, or developers select predefined templates and the system dynamically generates code according to the input data. Existing code generation tools usually rely on predefined templates and lack the ability of dynamic optimization. This patent realizes the adaptive optimization of code generation through multi-task learning and reinforcement learning, supports the automated testing of code optimization, and its core value lies in significantly improving development efficiency, reducing the time of manual code writing, while improving code quality and ensuring that the generated code meets project requirements.

[0117] As the intelligent layer of the system, the context-aware module aims to dynamically adjust the generated code or recommend relevant components by analyzing the developer's context (such as project type, historical development habits, etc.), thereby improving development efficiency and code quality. By introducing deep reinforcement learning algorithms, it dynamically learns the developer's development habits and preferences (such as frequently used components, code styles, etc.), supports real-time learning and prediction of developer behavior, combines knowledge graph technology, semantically models the developer's context, improves the accuracy and depth of context awareness, and dynamically adjusts the generated code according to the current project type (such as web applications, mobile applications, etc.). It adopts a recommendation algorithm based on graph neural networks (GNNs), combines the component relationships in the knowledge graph, and provides more accurate component recommendations. By introducing multi-modal recommendation technology, it supports multi-modal interaction for component recommendations, combines voice, image, and text inputs, provides richer recommendation results, and further optimizes the recommendations through multi-round conversations. In addition, it dynamically generates best practice guidelines by combining domain knowledge graphs and project types. By introducing a reinforcement learning-based recommendation system, it supports the dynamic update of best practices and personalized recommendations, and dynamically adjusts the recommendation strategy according to project requirements and developer feedback. For example, for projects with high security requirements, the system recommends using encryption libraries or security frameworks. Its application scenarios include when a developer needs to implement a certain function, the system recommends the most suitable low-code components, or the system dynamically adjusts the generated code according to the project type and developer habits.

[0118] Existing context-aware tools are usually limited to simple historical data analysis. The core value of this module lies in realizing real-time behavior learning and prediction of developers through deep reinforcement learning and knowledge graph technology, providing more accurate component recommendations and best practice recommendations. It further improves code quality, ensures that the generated code meets project requirements and best practices, while enhancing development efficiency and reducing the time for developers to search for and select components.

[0119] The collaboration and conflict resolution module, as the collaboration layer of the system, aims to support real-time collaboration among multiple developers and automatically detect and resolve code conflicts through intelligent algorithms, thereby improving the efficiency of team collaboration. This module utilizes real-time synchronization technologies (such as Operational Transformation or CRDT) to support multiple developers in simultaneously editing the same project and synchronizes developers' modifications in real-time through an event-driven architecture. Meanwhile, it uses difference comparison algorithms (such as the Diff algorithm) to detect code conflicts. For example, when multiple developers modify the same line of code simultaneously, the system automatically detects the conflict and prompts the developers. Then, combined with deep learning algorithms and static code analysis tools, it conducts multi-dimensional detection of code conflicts. By introducing a conflict detection algorithm based on graph neural networks (GNN), the accuracy and efficiency of conflict detection are improved. This method supports real-time detection and visualization of code conflicts, and developers can intuitively understand the conflict locations and reasons through the visualization interface. In addition, it automatically resolves conflicts through intelligent algorithms (such as rule-based merging strategies or machine learning models), and adopts a conflict resolution algorithm based on reinforcement learning to dynamically learn conflict resolution strategies. Combining natural language processing technology, it automatically generates explanations and suggestions for conflict resolution. For example, the system automatically merges the modifications of two developers or provides multiple solutions for developers to choose from. Its application scenarios include multiple developers simultaneously editing a low-code page, the system synchronizing their modifications in real-time, or the system automatically resolving code conflicts to reduce communication costs.

[0120] Existing collaboration tools usually lack intelligent conflict resolution capabilities. This patent realizes real-time detection and automated resolution of code conflicts through deep learning algorithms and blockchain technology. The core value of this module lies in significantly improving the efficiency of team collaboration, reducing communication costs and conflict resolution time, while ensuring code consistency and integrity.

[0121] In a specific example, the core function of the multi-modal perception module is to convert various input methods (such as voice, hand-drawn flowcharts, natural language descriptions, etc.) into structured instructions recognizable by the low-code platform. Specifically, it can include the following steps:

[0122] 1) Speech recognition

[0123] Using advanced speech recognition technologies (such as Google Speech-to-Text, Whisper, etc.), convert voice input into text. Combining natural language processing (NLP) technology, extract key information and convert it into instructions for the low-code platform. This step improves the accuracy and robustness of speech recognition through an adaptive attention mechanism, supporting multi-language input and multi-round conversations.

[0124] Example: The developer's voice input is "Create a user login page", and the system generates the corresponding low-code component after recognition.

[0125] 2) Image Recognition (Hand-drawn Flowchart)

[0126] Use computer vision technologies (such as OpenCV, YOLO, etc.) to recognize the graphics and text in the hand-drawn flowchart. Through the graph neural network (GNN), model the relationships between the recognized graphics and text, and convert the recognition results into structured data (such as flowchart nodes, connection relationships, etc.). This step supports real-time feedback and correction of the hand-drawn flowchart. Developers can view the recognition results in real time during the hand-drawing process and make adjustments.

[0127] Example: The developer hand-draws a flowchart, and the system automatically converts it into a workflow and generates the corresponding code.

[0128] 3) Natural Language Processing

[0129] Use pre-trained language models (such as GPT, BERT, etc.) to understand natural language descriptions. Through the encoder-decoder architecture based on Transformer (such as T5), combined with domain adaptation technologies, perform semantic parsing on natural language descriptions. This step supports multi-round conversations. Developers can gradually improve the requirement descriptions through multi-round interactions, and the system dynamically adjusts the generated low-code components according to each round of conversations.

[0130] Example: The developer inputs "Add a button that jumps to the home page when clicked", and the system generates the corresponding low-code component.

[0131] The core function of the intelligent code generation module is to automatically generate low-code components or code snippets based on the output of the multimodal perception module. The specific implementation method can be as follows:

[0132] 1) Code Generation Model

[0133] Adopt a code generation model based on Transformer (such as Codex, CodeGen, etc.), combined with the characteristics of the low-code platform, to generate components or modules that conform to the platform specifications. Through a multi-task learning framework, enable the code generation model to learn both code generation and code optimization tasks simultaneously, and combine reinforcement learning to dynamically adjust the quality and efficiency of the generated code.

[0134] Example: Input "Create a user login page", and the system generates the corresponding HTML, CSS, and JavaScript code.

[0135] 2) Template Matching

[0136] Use a template matching algorithm based on graph embedding, combined with the graph neural network (GNN), to improve the accuracy and efficiency of template matching. Introduce dynamic parameter filling technology to support developers' custom extensions based on templates.

[0137] Example: When the input is "Create a data table", the system selects a predefined data table template and fills in the data.

[0138] 3) Code Optimization

[0139] Combine static code analysis tools (such as ESLint, Pylint, etc.) and dynamic code analysis techniques to perform multi-dimensional optimization on the generated code. Introduce a code quality evaluation model based on machine learning to automatically evaluate the quality of the generated code and provide optimization suggestions. Support automated testing for code optimization. By automatically generating test cases, verify the performance and stability of the optimized code in different scenarios.

[0140] Example: The system automatically optimizes the generated code, removing redundant parts or fixing potential errors.

[0141] The core function of the context-aware module is to analyze the developer's context (such as project type, historical development habits, etc.) and dynamically adjust the generated code or recommend relevant components. The specific implementation methods can be as follows:

[0142] 1) Context Analysis

[0143] Use deep reinforcement learning algorithms to dynamically learn the developer's development habits and preferences. Combine knowledge graph technology to perform semantic modeling on the developer's context, improving the accuracy and depth of context awareness. Support real-time learning and prediction of developer behavior. The system can dynamically adjust the generated code and recommended components according to the developer's real-time behavior.

[0144] Example: If a developer often uses a certain UI component library, the system preferentially recommends components from that library.

[0145] 2) Component Recommendation

[0146] Adopt a recommendation algorithm based on graph neural networks (GNN), combine the component relationships in the knowledge graph, and provide more accurate component recommendations. Introduce multi-modal recommendation technology, combine voice, image, and text inputs, and provide richer recommendation results.

[0147] Example: When a developer needs to implement a chart function, the system recommends the most suitable chart component.

[0148] 3) Best Practice Recommendation

[0149] Combine the domain knowledge graph and project type to dynamically generate best practice guidelines. Introduce a recommendation system based on reinforcement learning, and dynamically adjust the recommendation strategy according to project requirements and developer feedback. Support dynamic updates and personalized recommendations for best practices. The system can dynamically adjust the recommended content according to project progress and developer feedback.

[0150] Example: For projects with high security requirements, the system recommends using encryption libraries or security frameworks.

[0151] The core functions of the collaboration and conflict resolution module are to support real-time collaboration among multiple developers and automatically detect and resolve code conflicts through intelligent algorithms. The specific implementation methods can be as follows:

[0152] 1) Real-time collaboration

[0153] Use real-time synchronization technologies (such as Operational Transformation or CRDT), combined with blockchain technology, to ensure the transparency and traceability of the collaboration process. Support real-time collaboration among multiple developers in different network environments, and the system can dynamically adjust the synchronization strategy according to the network conditions.

[0154] Example: Multiple developers simultaneously edit a low-code page, and the system synchronizes their modifications in real time.

[0155] 2) Conflict detection

[0156] Use difference comparison algorithms (such as Diff algorithm) combined with deep learning algorithms and static code analysis tools to perform multi-dimensional detection of code conflicts. Introduce a conflict detection algorithm based on graph neural network (GNN) to improve the accuracy and efficiency of conflict detection. Support real-time detection and visualization of code conflicts, and developers can intuitively understand the conflict locations and reasons through the visualization interface.

[0157] Example: Two developers simultaneously modify the same line of code, and the system detects the conflict and prompts the developers.

[0158] 3) Conflict resolution

[0159] Use intelligent algorithms (such as rule-based merging strategies or machine learning models) to automatically resolve conflicts. The system can provide multiple solutions for developers to choose from according to the type and context information of the conflicts.

[0160] Example: The system automatically merges the modifications of two developers or provides multiple solutions for developers to choose from.

[0161] Describe the above development process in combination with some specific usage scenarios.

[0162] Scenario 1: The developer describes the requirement by voice: "Create a user login page." The system automatically generates the corresponding low-code components and recommends relevant APIs.

[0163] Scenario 2: The developer hand-draws a flowchart, and the system automatically converts it into a workflow and generates the corresponding code.

[0164] Scenario 3: Multiple developers edit the same module simultaneously, and the system automatically merges the code and resolves conflicts.

[0165] In the low-code intelligent development assistance system based on multi-modal perception disclosed in this application, it aims to achieve intelligent code generation, optimization, and collaboration by combining multi-modal perception technology and a low-code development platform, significantly reducing the development threshold and improving development efficiency. The system includes four core modules: a multi-modal perception module, an intelligent code generation module, a context awareness module, and a collaboration and conflict resolution module. The multi-modal perception module supports various input methods such as voice, hand-drawn flowcharts, and natural language descriptions, and converts them into structured instructions recognizable by the low-code platform; the intelligent code generation module automatically generates low-code components or code snippets based on the input and provides optimization suggestions; the context awareness module analyzes the developer's context (such as project type, historical development habits, etc.) and dynamically adjusts the generated code or recommends relevant components; the collaboration and conflict resolution module supports real-time collaboration among multiple developers and automatically detects and resolves code conflicts through intelligent algorithms. The core innovation points of this invention lie in multi-modal perception-driven intelligent code generation, context-aware intelligent code optimization, low-code component recommendation based on a knowledge graph, and real-time collaboration and intelligent conflict resolution. Through this system, developers can complete low-code development tasks more efficiently, reduce the time of manually writing code, improve code quality, and at the same time enhance team collaboration efficiency. This invention is applicable to various low-code development scenarios and has broad application prospects and commercial value.

[0166] The innovation points of the embodiments of this disclosure include:

[0167] Multi-modal perception-driven intelligent code generation: Automatically generate low-code modules or code snippets through various input methods such as voice, hand-drawn flowcharts, and natural language descriptions.

[0168] Context-aware intelligent code optimization: Dynamically optimize the generated code according to the developer's context and recommend best practices.

[0169] Low-code component recommendation based on a knowledge graph: Use the knowledge graph to intelligently recommend relevant components and provide usage examples.

[0170] Real-time collaboration and intelligent conflict resolution: Support real-time collaboration among multiple developers and automatically resolve code conflicts through intelligent algorithms.

[0171] The present disclosure deeply integrates multimodal perception technology with a low-code development platform to create an intelligent development assistant system. Through various input methods such as voice, hand-drawn flowcharts, and natural language descriptions, the system converts the developer's requirements into instructions recognizable by the low-code platform and automatically generates code or low-code components. At the same time, the system can dynamically optimize the generated code according to the developer's context, recommend best practices and relevant components, significantly improving development efficiency and code quality. In addition, the system supports real-time collaboration among multiple developers and automatically detects and resolves code conflicts through intelligent algorithms, reducing communication costs in team collaboration. These technical features together constitute an efficient and intelligent low-code development solution applicable to various development scenarios.

[0172] It solves the problems of insufficient intelligence and high development threshold of existing platforms. By introducing multimodal perception technology, the system can support more natural input methods, enabling non-professional developers to easily use the low-code platform. The intelligent code generation and optimization functions reduce the time for manual code writing and improve code quality. The context-aware module can dynamically adjust the generated code according to the developer's habits and project requirements, further enhancing development efficiency. The collaboration and conflict resolution module significantly improves team collaboration efficiency through real-time synchronization and intelligent algorithms. These improvements make the low-code platform more user-friendly and efficient, and better able to meet complex and personalized development needs.

[0173] Combined Figure 3 As shown, the embodiment of the present disclosure also provides a low-code intelligent development assistance device 300 based on multimodal perception, including a processor 304 and a memory 301. Optionally, the system may further include a communication interface 302 and a bus 303. Among them, the processor 304, the communication interface 302, and the memory 301 can communicate with each other through the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call the logical instructions in the memory 301 to execute the low-code intelligent development assistance method based on multimodal perception in the above embodiment.

[0174] In addition, when the logical instructions in the above-mentioned memory 301 are implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium.

[0175] The memory 301, being a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as the program instructions / modules corresponding to the methods in the embodiments of the present disclosure. The processor 304 executes functional applications and data processing by running the program instructions / modules stored in the memory 301, that is, implements the low-code intelligent development assistance method based on multi-modal perception in the above embodiments.

[0176] The memory 301 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device and the like. In addition, the memory 301 may include high-speed random access memory and may also include non-volatile memory.

[0177] The embodiments of the present disclosure provide a computer-readable storage medium storing computer-executable instructions, and the computer-executable instructions are configured as the low-code intelligent development assistance method based on multi-modal perception.

[0178] The above computer-readable storage medium may be a transient computer-readable storage medium or a non-transient computer-readable storage medium.

[0179] The technical solution of the embodiments of the present disclosure can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of the embodiments of the present disclosure. The foregoing storage medium may be a non-transient storage medium, including: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes, or may also be a transient storage medium.

[0180] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. Embodiments merely represent possible variations. Unless explicitly required, individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. As used in the description of the embodiments, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groupings of these. Without further limitation, an element defined by the statement "comprising an..." does not preclude the presence of additional identical elements in the process, method, or apparatus including the element. Herein, each embodiment may focus on the differences from other embodiments, and the same or similar parts among the embodiments may be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, the relevant parts may refer to the description of the method part.

[0181] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner may depend on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0182] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units can be merely a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components can be or can not be physically separated. The components shown as units can be or can not be physical units, that is, they can be located in one place or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. Additionally, in the embodiments of the present disclosure, the various functional units can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit.

[0183] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks can also occur in a different order than that disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, which can depend on the functions involved. Each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0184] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0185] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.

[0186] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0187] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0188] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0189] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, or a server of a distributed system, or a server incorporating a blockchain.

[0190] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this disclosure can be achieved, and this is not limited herein.

[0191] The above specific implementation manners do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A low-code intelligent development assistance method based on multimodal perception, characterized in that The method includes: Receiving the requirements input by the user through voice, hand-drawn flowcharts, or natural language descriptions, and converting the input requirements into structured data; Generating corresponding code based on the structured data; Optimizing the code based on the user's development guidance information and recommending components.

2. The method according to claim 1, wherein The step of receiving the requirements input by the user through voice, hand-drawn flowcharts, or natural language descriptions and converting the input requirements into structured data includes: For the requirements input by the user through voice, using a preset speech recognition algorithm to convert the requirements input by the user through voice into text, and combining with natural language processing algorithms to extract key information and convert it into instructions for the low-code platform; For the requirements input by the user through hand-drawn flowcharts, using a preset computer vision recognition algorithm to recognize the requirements input by the user through hand-drawn flowcharts, obtaining the corresponding graphics and text, and performing relationship modeling on the recognized graphics and text through a preset graph neural network to convert the recognition results into structured data; For the requirements input by the user through natural language descriptions, using a preset language model to analyze the requirements input by the user through natural language descriptions, extracting key information and converting it into instructions for the low-code platform.

3. The method according to claim 1, characterized in that, The step of generating corresponding code based on the structured data includes: According to the code generation model of Transformer, based on the structured data, combined with the characteristics of the low-code platform, generating components or modules that conform to the platform specifications; According to the template matching algorithm of graph embedding, combined with the graph neural network, based on the structured data, determining the corresponding template, and based on the determined template, expanding the generated components or modules that conform to the platform specifications to obtain the expanded code; Based on a preset static code analysis tool and dynamic code analysis algorithm, optimizing the expanded code to generate the corresponding code.

4. The method according to claim 1, wherein The step of optimizing the code based on the user's development guidance information and recommending components includes: Determining the user's development habits and preferences based on the user's development guidance information; Adjusting the code based on the user's development habits and preferences and recommending components.

5. The method according to claim 4, wherein The step of adjusting the code based on the user's development habits and preferences and recommending components includes: Adjusting the code based on the user's development habits and preferences; Based on the user's development habits and preferences, according to the structured data, combined with the component relationships in the preset knowledge graph, recommending the corresponding components.

6. The method according to claim 4, characterized in that, The method further includes: Obtaining the user's development guidance information in real time; Predicting the user's behavior based on the real-time obtained development guidance information; Based on the predicted user behavior, adjusting the generated code in real time and recommending components.

7. The method according to claim 1, characterized in that, The method further includes: Real-time synchronizing the development collaboration of multiple users, and detecting the collaboration conflicts of users based on the conflict detection algorithm of the graph neural network; Merging the collaboration conflicts of users or generating multiple solutions based on a rule-based merging strategy or a machine learning model.

8. A low-code intelligent development assistance system based on multimodal perception, characterized in that, The system includes: A multi-modal perception module, configured to receive the requirements input by the user through voice, hand-drawn flowcharts or natural language descriptions, and convert the input requirements into structured data; An intelligent code generation module, configured to generate corresponding code based on the structured data; A context awareness module, configured to optimize the code and recommend components based on the user's development guidance information.

9. An electronic device, characterized in that, Comprising: At least one processor; And a memory communicatively connected to the at least one processor; Wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.