Code automatic generation method and system based on multi-modal large model and medium

By using a multimodal large model-based approach, the semantic alignment and context-aware bias issues of multimodal input features are resolved, enabling high-accuracy generation of complex business logic code, improving the accuracy and compatibility of code generation, and reducing repetitive work for developers.

CN120929064APending Publication Date: 2025-11-11SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510839729.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing code generation technologies suffer from deviations in semantic alignment and context awareness of multimodal input features, resulting in incomplete understanding of requirements and low accuracy in generating complex business logic code.

Method used

The method adopts a multimodal large model-based approach, which converts multimodal data into feature vectors through a preset processing algorithm, obtains the spatial mapping relationship of feature vectors using a preset multimodal model, constructs a shared vector space, and constructs environmental constraint vectors through historical project code context awareness, and finally generates running code that conforms to preset technical specifications and architectural features.

Benefits of technology

It achieves accurate fusion and contextual understanding of multimodal inputs, improves the accuracy and compatibility of code generation, reduces repetitive coding work, and enhances development efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929064A_ABST
    Figure CN120929064A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic code generation method and system based on a multi-modal large model and a medium, and mainly relates to the technical field of automatic code generation. The method and the device are used for solving the problem of incomplete demand understanding caused by deviation between semantic alignment of multi-modal input features and context sensing and the problem of low accuracy of complex business logic code generation in the existing scheme. Comprising the following steps: constructing an environment constraint vector through history project code context sensing; wherein the environment constraint vector comprises dependency scanning data, technology stack analysis data and code style matching data; according to the multi-dimensional fusion feature and the environment constraint vector, generating an operation code conforming to a preset technical specification and a preset architecture feature; and performing pre-training on the operation code based on the pre-training code generation model, calculating a multi-dimensional reward function, performing reinforcement learning on the operation code, obtaining a feedback optimization code, and reversely injecting the optimization code into the pre-training code generation model to obtain a final operation code.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automatic code generation technology, and in particular to a method, system and medium for automatic code generation based on a multimodal large model. Background Technology

[0002] With the rapid development of artificial intelligence and automation technologies, AI applications are becoming increasingly widespread. Especially in the field of software development, AI technology has permeated every stage of code writing. In recent years, AI-based code generation tools have emerged continuously. These tools, through natural language processing technology, have initially achieved the function of converting user-described requirements in natural language into executable code, thereby improving development efficiency.

[0003] However, existing code generation technologies still have significant shortcomings in the fusion and processing of multimodal input features and the generation of code for complex business logic. Specifically, on the one hand, when the input information includes multiple modalities such as text, images, and speech, the system struggles to accurately align the semantics between features of different modalities, leading to deviations in the contextual understanding of user needs. On the other hand, when faced with complex business scenarios involving multiple conditional judgments and nested loops, the generated code often falls short of expectations in terms of logical accuracy and functional completeness. These technical bottlenecks severely restrict the application effectiveness of AI code generation tools in real-world production environments.

[0004] Therefore, there is an urgent need for a method, system, and medium for automatic code generation based on multimodal large models to solve the following problems: the incomplete understanding of requirements due to the deviation between the semantic alignment and context awareness of multimodal input features, and the low accuracy of generating complex business logic code. Summary of the Invention

[0005] This application provides a method, system, and medium for automatic code generation based on a multimodal large model, in order to solve the problems of incomplete understanding of requirements due to deviations in semantic alignment and context awareness of multimodal input features in existing solutions, and low accuracy in generating complex business logic code.

[0006] Firstly, this application provides a method for automatic code generation based on a multimodal large model, the method comprising: Obtain multimodal data and historical project code corresponding to the current project requirements; the multimodal data includes text data, audio data, and image data; convert the multimodal data into feature vectors using a preset processing algorithm; use a preset multimodal model to obtain the spatial mapping relationship between the feature vectors corresponding to the multimodal data, and set a shared vector space for the feature vectors corresponding to different modal data. Feature vector fusion is performed on feature vectors in a shared vector space to obtain multidimensional fused features; By leveraging historical project code context awareness, an environment constraint vector is constructed; this vector includes dependency scanning data, technology stack analysis data, and code style matching data. Based on multi-dimensional fusion features and environmental constraint vectors, generate runtime code that conforms to preset technical specifications and preset architectural features; The code is pre-trained using the CodeT5 pre-trained code generation model, and a multi-dimensional reward function is calculated. Then, the code is reinforced through learning to obtain optimized code based on feedback. Finally, the optimized code is back-injected into the CodeT5 pre-trained code generation model to obtain the final code.

[0007] In one implementation of this application, a preset processing algorithm is used to convert multimodal data into feature vectors, specifically including: Audio data is converted into text data using a speech recognition model; Image information is identified from image data using the improved YOLOv5 algorithm; The CLIP model is used for text encoders and image encoders to convert text data and image information into feature vectors, respectively.

[0008] In one implementation of this application, the pre-defined multimodal model employs a contrastive learning framework, and the loss function is: Calculate the loss function ; in, For the preset boundary hyperparameters, For text feature vectors, To match image feature vectors, The feature vector of the negative sample image; For the matched positive sample pairs, it represents the text. and images Similarity between them For mismatched negative sample pairs, representing text and images The similarity between them; The value of i ranges from [1, n], and the values ​​of j and k range from [1, m]. n represents the total number of text feature vectors, and m represents the total number of image feature vectors.

[0009] In one implementation of this application, an environment constraint vector is constructed through historical project code context awareness, specifically including: Parse the configuration files of historical project code, extract technical components and version information, establish a technical ecosystem graph of project component relationships, and use technical components, version information, and technical ecosystem graph as dependency scanning data; Based on the technology ecosystem map, compatibility rules of components within the technology stack are retrieved, and the code structure and features in historical project code are identified through AST (Abstract Syntax Tree) to verify the degree of matching between the technology stack and the architecture pattern. The degree of matching between code structure, technology stack and architecture pattern is obtained as technology stack analysis data. By learning the naming conventions and architectural features of the project's historical code using an LSTM model, code style matching data can be obtained.

[0010] In one implementation of this application, the calculation of the multidimensional reward function specifically includes: Through the formula: Calculate the multidimensional reward function; in, As a performance reward for the currently running code, As a safety reward, As a readability bonus, These are the preset parameters.

[0011] In one implementation of this application, after generating runtime code that conforms to preset technical specifications and preset architectural characteristics, the method further includes: It connects to the enterprise project technical specification library to retrieve knowledge of the running code, and performs code API calls and syntax standardization checks in the running code.

[0012] Secondly, this application provides an automatic code generation system based on a multimodal large model, the system comprising: The multimodal information input module is used to obtain multimodal data and historical project code corresponding to the current project requirements; the multimodal data includes text data, audio data, and image data. The multimodal information processing module is used to convert multimodal data into feature vectors through a preset processing algorithm; to obtain the spatial mapping relationship between the feature vectors corresponding to the multimodal data using a preset multimodal model, and to set a shared vector space for the feature vectors corresponding to different modal data; and to perform feature vector fusion on the feature vectors in the shared vector space to obtain multidimensional fused features. The context-aware module is used to construct environment constraint vectors by understanding the context of historical project code; these environment constraint vectors include dependency scanning data, technology stack analysis data, and code style matching data. The code generation module is used to generate running code that conforms to preset technical specifications and preset architectural features based on multi-dimensional fusion features and environmental constraint vectors. The running code is pre-trained based on the CodeT5 pre-trained code generation model, multi-dimensional reward function is calculated, and then reinforcement learning is performed on the running code to obtain feedback optimized code. The optimized code is then back-injected into the CodeT5 pre-trained code generation model to obtain the final running code.

[0013] In one implementation of this application, the multimodal information processing module includes a data conversion unit. Used to convert audio data into text data through a speech recognition model; Image information is identified from image data using the improved YOLOv5 algorithm; The CLIP model is used for text encoders and image encoders to convert text data and image information into feature vectors, respectively.

[0014] In one implementation of this application, the context-aware module includes a constraint acquisition unit. Configuration files used to parse historical project code, extract technical components and version information, establish a technical ecosystem graph of project component relationships, and use technical components, version information, and technical ecosystem graph as dependency scanning data; Based on the technology ecosystem map, compatibility rules of components within the technology stack are retrieved, and the code structure and features in historical project code are identified through AST (Abstract Syntax Tree) to verify the degree of matching between the technology stack and the architecture pattern. The degree of matching between code structure, technology stack and architecture pattern is obtained as technology stack analysis data. By learning the naming conventions and architectural features of the project's historical code using an LSTM model, code style matching data can be obtained.

[0015] Thirdly, this application provides a non-volatile computer storage medium storing computer instructions, which, when executed, implement an automatic code generation method based on a multimodal large model as described above.

[0016] The online learning method provided in this application embodiment, As can be seen from the above technical solutions, this application has the following advantages: By employing a multimodal large-scale model-based automatic code generation method, which integrates multi-dimensional data such as natural language, visual images, and structured code, multimodal input compatibility and parsing accuracy are achieved. Furthermore, by reconstructing the software development paradigm through context awareness and code generation models, and by learning the unified technical specifications and architectural features of a project, automated code generation is achieved, improving the accuracy and compatibility of code generation. This not only frees developers from repetitive coding tasks, allowing them to focus on efficiently generating code automatically and researching higher-value technical architectures and innovative algorithms, but also solves the problems of incomplete understanding of requirements due to deviations in the semantic alignment and context awareness of existing multimodal input features, as well as the low accuracy of generating code for complex business logic. Attached Figure Description

[0017] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a method for automatically generating code based on a multimodal large model provided in an embodiment of this application.

[0019] Figure 2 This is a schematic diagram of the internal structure of a code automatic generation system based on a multimodal large model provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Those skilled in the art should understand that the embodiments described below are merely preferred embodiments of this disclosure and do not imply that this disclosure can only be implemented through these preferred embodiments. These preferred embodiments are merely used to explain the technical principles of this disclosure and are not intended to limit the scope of protection of this disclosure. Based on the preferred embodiments provided by this disclosure, all other embodiments obtained by those skilled in the art without creative effort should still fall within the scope of protection of this disclosure.

[0022] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0023] The technical solutions proposed in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0024] The embodiment provides a method for automatic code generation based on a multimodal large model, such as Figure 1 As shown in the embodiments of this application, the method mainly includes the following steps: Step 110: Obtain the multimodal data and historical project code corresponding to the current project requirements.

[0025] Multimodal data includes text data, audio data, and image data.

[0026] It should be noted that different methods of acquiring multimodal data can include manual input.

[0027] Those skilled in the art will understand that obtaining multimodal data (including text, audio, and image data) and historical project code corresponding to the current project requirements in this step has the following beneficial effects: Improving model training quality: Multimodal data can provide richer contextual information for code generation models, making up for the limitations of single-modal data in semantic expression, thereby improving the model's ability to understand complex requirements.

[0028] Enhance the completeness of requirements analysis: text data can be directly linked to code logic, audio data may contain supplementary explanations from developers, and image data (such as flowcharts and interface designs) can help understand business logic, making requirements analysis more comprehensive.

[0029] Adapting to diverse input scenarios: Supporting manual input increases the flexibility of data collection, facilitates the integration of data from different sources (such as meeting recordings, design documents, and code comments), and optimizes the model's generalization ability.

[0030] Reduce data acquisition costs: Reusing code from historical projects reduces reliance on new data and shortens the model training cycle.

[0031] Step 120: Convert multimodal data into feature vectors using a preset processing algorithm; use a preset multimodal model to obtain the spatial mapping relationship between the feature vectors corresponding to the multimodal data, and set a shared vector space for the feature vectors corresponding to different modal data; perform feature vector fusion on the feature vectors in the shared vector space to obtain multidimensional fused features.

[0032] In some embodiments, multimodal data is converted into feature vectors using a preset processing algorithm, specifically including: Audio data is converted into text data using a speech recognition model; Image information is identified from image data using the improved YOLOv5 algorithm; The CLIP model is used for text encoders and image encoders to convert text data and image information into feature vectors, respectively.

[0033] The preset multimodal model uses a contrastive learning framework, and the loss function is: Calculate the loss function ; in, For the preset boundary hyperparameters, For text feature vectors, To match image feature vectors, The feature vector of the negative sample image; For the matched positive sample pairs, it represents the text. and images Similarity between them For mismatched negative sample pairs, representing text and images The similarity between them; The value of i ranges from [1, n], and the values ​​of j and k range from [1, m]. n represents the total number of text feature vectors, and m represents the total number of image feature vectors.

[0034] Based on the above description, those skilled in the art will understand that this step converts multimodal data (text, audio, and images) into feature vectors using a preset processing algorithm, effectively preserving the semantic information of each modality. Specifically: Audio data: Speech recognition models (such as Whisper or Wav2Vec) convert speech into text, which facilitates subsequent unified encoding and avoids the complexity of directly processing audio signals.

[0035] Image data: The improved YOLOv5 algorithm optimizes target detection efficiency, accurately identifies key information in images (such as UI components and flowchart elements), and reduces noise interference.

[0036] Text and Image Encoding: The CLIP model's text encoder and image encoder extract feature vectors respectively, ensuring comparability between modalities and laying the foundation for subsequent cross-modal alignment.

[0037] The pre-set multimodal model involved in this step adopts a contrastive learning framework, which optimizes the spatial distribution of feature vectors through a loss function, specifically as follows: Semantic alignment enhancement: Contrastive learning improves cross-modal semantic consistency by distinguishing between positive sample pairs (matching text-image pairs) and negative sample pairs (non-matching text-image pairs), reducing the distance between relevant feature vectors and increasing the distance between irrelevant feature vectors.

[0038] Shared vector space optimization: By controlling the similarity threshold of positive and negative sample pairs through boundary hyperparameters (such as), the model is prevented from overfitting, and the feature vectors of different modalities are comparable in the shared space.

[0039] This step fuses feature vectors in the shared vector space (e.g., through weighted concatenation or attention mechanisms) to obtain multi-dimensional fused features. The technical effects include: Enhanced context awareness: The fusion feature integrates semantic information from text, visual information from images, and supplementary descriptions from audio, enabling the model to more comprehensively understand complex requirements.

[0040] Enhanced ability to express complex logic: Multidimensional features compensate for the limitations of single modality. For example, flowcharts in images can assist business logic described in text, improving the accuracy of code generation.

[0041] The design of the loss function involved in this step (such as the contrastive loss based on InfoNCE) further optimizes the feature space distribution and reduces intermodal bias by maximizing the similarity of positive sample pairs and minimizing the similarity of negative sample pairs.

[0042] Step 130: Construct an environment constraint vector by leveraging the historical project code context awareness.

[0043] The environment constraint vector includes dependency scanning data, technology stack analysis data, and code style matching data. By leveraging the context awareness of historical project code, an environment constraint vector is constructed, specifically including: Parse the configuration files of historical project code, extract technical components and version information, establish a technical ecosystem graph of project component relationships, and use technical components, version information, and technical ecosystem graph as dependency scanning data; Based on the technology ecosystem map, compatibility rules of components within the technology stack are retrieved, and the code structure and features in historical project code are identified through AST (Abstract Syntax Tree) to verify the degree of matching between the technology stack and the architecture pattern. The degree of matching between code structure, technology stack and architecture pattern is obtained as technology stack analysis data. By learning the naming conventions and architectural features of the project's historical code using an LSTM model, code style matching data can be obtained.

[0044] Based on the above description, those skilled in the art will understand that: this step, by parsing historical project configuration files (such as pom.xml / package.json), extracting component version information, and constructing a technology ecosystem graph, can identify potential dependency conflicts. The establishment of the technology ecosystem graph automatically avoids known incompatible combinations when generating new code. Technology stack adaptability optimization, combined with AST analysis and ecosystem graph retrieval, can quantitatively evaluate the matching degree between the technology stack and the architectural pattern, avoiding the generation of code structures that do not conform to the target technology stack. Code style consistency maintains the LSTM model's learning of historical code naming conventions and architectural characteristics, ensuring that the generated code maintains a consistent style with existing projects. Dynamic adaptation of constraints: the multi-dimensional composition of constraint vectors (dependencies / technology stack / style) can be updated with project iterations. In a continuous integration environment, each commit triggers constraint vector optimization, forming a positive feedback loop.

[0045] Step 140: Generate running code that conforms to preset technical specifications and preset architecture features based on multi-dimensional fusion features and environmental constraint vectors; pre-train the running code based on the CodeT5 pre-trained code generation model, calculate the multi-dimensional reward function, and then reinforce the running code to obtain the feedback optimized code. Then, the optimized code is back-injected into the CodeT5 pre-trained code generation model to obtain the final running code.

[0046] In some embodiments, calculating the multidimensional reward function specifically includes: Through the formula: Calculate the multidimensional reward function; in, As a performance reward for the currently running code, As a safety reward, As a readability bonus, These are the preset parameters.

[0047] After generating runtime code that conforms to preset technical specifications and preset architectural characteristics, the method also includes: It connects to the enterprise project technical specification library to retrieve knowledge of the running code, and performs code API calls and syntax standardization checks in the running code.

[0048] Based on the above description, those skilled in the art will understand that this step constructs a closed-loop feedback system by combining a CodeT5 pre-trained model with a multi-dimensional reward function (performance / security / readability). This scheme calculates code quality metrics in real time during the model inference phase and dynamically adjusts the generation strategy through reinforcement learning. The introduction of environmental constraint vectors ensures that each optimization iteration conforms to preset technical specifications, forming a continuous improvement chain of "generation-evaluation-optimization".

[0049] By combining the real-time retrieval mechanism of the multidimensional constraint fusion technical specification library with AST analysis, the following can be achieved: Architecture feature verification: Matching preset architecture patterns through abstract syntax trees; API compliance check: Dynamically comparing calling standards in the enterprise's technical specification library; Syntax standardization verification: Style constraints based on project historical data.

[0050] The optimized code reverse injection mechanism has been implemented: Domain knowledge accumulation: solidifying project experience into model parameters; Long-tail scenario coverage: Improve the model's ability to handle edge cases by continuously injecting and optimizing samples; Technical debt prevention: Automatically avoid anti-patterns identified in historical projects.

[0051] In addition, this application Figure 2 This application provides an embodiment of an automatic code generation system based on a multimodal large model. For example... Figure 2 As shown in the embodiments of this application, the system mainly includes: The multimodal information input module 210 is used to obtain multimodal data and historical project codes corresponding to the current project requirements; the multimodal data includes text data, audio data, and image data.

[0052] The multimodal information processing module 230 is used to convert multimodal data into feature vectors through a preset processing algorithm; to obtain the spatial mapping relationship between the feature vectors corresponding to the multimodal data using a preset multimodal model, and to set a shared vector space for the feature vectors corresponding to different modal data; and to perform feature vector fusion on the feature vectors in the shared vector space to obtain multidimensional fused features.

[0053] The multimodal information processing module 230 includes a data conversion unit. Used to convert audio data into text data through a speech recognition model; Image information is identified from image data using the improved YOLOv5 algorithm; The CLIP model is used for text encoders and image encoders to convert text data and image information into feature vectors, respectively.

[0054] The context-aware module 230 is used to construct an environment constraint vector by being aware of the historical project code context; the environment constraint vector includes dependency scanning data, technology stack analysis data, and code style matching data.

[0055] Context-aware module 230 includes a constraint acquisition unit. Configuration files used to parse historical project code, extract technical components and version information, establish a technical ecosystem graph of project component relationships, and use technical components, version information, and technical ecosystem graph as dependency scanning data; Based on the technology ecosystem map, compatibility rules of components within the technology stack are retrieved, and the code structure and features in historical project code are identified through AST (Abstract Syntax Tree) to verify the degree of matching between the technology stack and the architecture pattern. The degree of matching between code structure, technology stack and architecture pattern is obtained as technology stack analysis data. By learning the naming conventions and architectural features of the project's historical code using an LSTM model, code style matching data can be obtained.

[0056] The code generation module 240 is used to generate runtime code that conforms to preset technical specifications and preset architectural features based on multi-dimensional fusion features and environmental constraint vectors. It pre-trains the runtime code using the CodeT5 pre-trained code generation model, calculates a multi-dimensional reward function, and then uses reinforcement learning to obtain optimized code. This optimized code is then back-injected into the CodeT5 pre-trained code generation model to obtain the final runtime code. In addition, embodiments of this application also provide a non-volatile computer storage medium storing executable instructions, which, when executed, implement the above-described method for automatic code generation based on a multimodal large model.

[0057] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for automatic code generation based on a multimodal large model, characterized in that, The method includes: Retrieve multimodal data and historical project code corresponding to the current project requirements; the multimodal data includes text data, audio data, and image data. Multimodal data is converted into feature vectors using a pre-defined processing algorithm; a pre-defined multimodal model is used to obtain the spatial mapping relationship between the feature vectors corresponding to the multimodal data, and a shared vector space is set for the feature vectors corresponding to different modal data; feature vector fusion is performed on the feature vectors in the shared vector space to obtain multidimensional fused features; By leveraging historical project code context awareness, an environment constraint vector is constructed; this vector includes dependency scanning data, technology stack analysis data, and code style matching data. Based on multidimensional fusion features and environmental constraint vectors, running code that conforms to preset technical specifications and preset architectural features is generated; the running code is pre-trained based on the CodeT5 pre-trained code generation model, multidimensional reward function is calculated, and then reinforcement learning is performed on the running code to obtain feedback optimized code, which is then back-injected into the CodeT5 pre-trained code generation model to obtain the final running code.

2. The automatic code generation method based on a multimodal large model according to claim 1, characterized in that, The multimodal data is converted into feature vectors using a pre-defined processing algorithm, specifically including: Audio data is converted into text data using a speech recognition model; Image information is identified from image data using the improved YOLOv5 algorithm; The CLIP model is used for text encoders and image encoders to convert text data and image information into feature vectors, respectively.

3. The automatic code generation method based on a multimodal large model according to claim 1, characterized in that, The pre-defined multimodal model uses a contrastive learning framework, with the loss function being: Calculate the loss function ; in, For the preset boundary hyperparameters, For text feature vectors, To match image feature vectors, The feature vector of the negative sample image; For the matched positive sample pairs, it represents the text. and images Similarity between them For mismatched negative sample pairs, representing text and images The similarity between them; The value of i ranges from [1, n], and the values ​​of j and k range from [1, m]. n represents the total number of text feature vectors, and m represents the total number of image feature vectors.

4. The automatic code generation method based on a multimodal large model according to claim 1, characterized in that, By leveraging the historical project code context awareness, an environment constraint vector is constructed, specifically including: Parse the configuration files of historical project code, extract technical components and version information, establish a technical ecosystem graph of project component relationships, and use technical components, version information, and technical ecosystem graph as dependency scanning data; Based on the technology ecosystem map, compatibility rules of components within the technology stack are retrieved, and the code structure and features in historical project code are identified through AST (Abstract Syntax Tree) to verify the degree of matching between the technology stack and the architecture pattern. The degree of matching between code structure, technology stack and architecture pattern is obtained as technology stack analysis data. By learning the naming conventions and architectural features of the project's historical code using an LSTM model, code style matching data can be obtained.

5. The automatic code generation method based on a multimodal large model according to claim 1, characterized in that, Calculating the multidimensional reward function specifically includes: Through the formula: Calculate the multidimensional reward function; in, As a performance reward for the currently running code, As a safety reward, As a readability bonus, These are the preset parameters.

6. The automatic code generation method based on a multimodal large model according to claim 1, characterized in that, After generating runtime code that conforms to preset technical specifications and preset architectural characteristics, the method further includes: It connects to the enterprise project technical specification library to retrieve knowledge of the running code, and performs code API calls and syntax standardization checks in the running code.

7. A code automatic generation system based on a multimodal large model, characterized in that, The system includes: The multimodal information input module is used to obtain multimodal data and historical project code corresponding to the current project requirements; the multimodal data includes text data, audio data, and image data. The multimodal information processing module is used to convert multimodal data into feature vectors through a preset processing algorithm; to obtain the spatial mapping relationship between the feature vectors corresponding to the multimodal data using a preset multimodal model, and to set a shared vector space for the feature vectors corresponding to different modal data; and to perform feature vector fusion on the feature vectors in the shared vector space to obtain multidimensional fused features. The context-aware module is used to construct environment constraint vectors by understanding the context of historical project code; these environment constraint vectors include dependency scanning data, technology stack analysis data, and code style matching data. The code generation module is used to generate running code that conforms to preset technical specifications and preset architectural features based on multi-dimensional fusion features and environmental constraint vectors. The running code is pre-trained based on the CodeT5 pre-trained code generation model, multi-dimensional reward function is calculated, and then reinforcement learning is performed on the running code to obtain feedback optimized code. The optimized code is then back-injected into the CodeT5 pre-trained code generation model to obtain the final running code.

8. The code automatic generation system based on a multimodal large model according to claim 7, characterized in that, The multimodal information processing module includes a data conversion unit. Used to convert audio data into text data through a speech recognition model; Image information is identified from image data using the improved YOLOv5 algorithm; The CLIP model is used for text encoders and image encoders to convert text data and image information into feature vectors, respectively.

9. The code automatic generation system based on a multimodal large model according to claim 7, characterized in that, The context-aware module includes a constraint acquisition unit. Configuration files used to parse historical project code, extract technical components and version information, establish a technical ecosystem graph of project component relationships, and use technical components, version information, and technical ecosystem graph as dependency scanning data; Based on the technology ecosystem map, compatibility rules of components within the technology stack are retrieved, and the code structure and features in historical project code are identified through AST (Abstract Syntax Tree) to verify the degree of matching between the technology stack and the architecture pattern. The degree of matching between code structure, technology stack and architecture pattern is obtained as technology stack analysis data. By learning the naming conventions and architectural features of the project's historical code using an LSTM model, code style matching data can be obtained.

10. A non-volatile computer storage medium, characterized in that, It stores computer instructions, which, when executed, implement a method for automatic code generation based on a multimodal large model as described in any one of claims 1-6.

Citation Information

Cited By

  • Intelligent fusion development platform demand understanding and framework automatic generation method

    CN121387243A