Code conversion method and device, equipment, medium and computer program product
By extracting and splitting non-code information, combining multimodal knowledge graphs to generate initial codes, and optimizing and feedback verification, the problem of low accuracy and efficiency of data conversion codes in the prior art is solved, and more efficient and accurate code conversion is achieved.
Patent Information
- Application Number
- CN202510004229.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
AI Technical Summary
The accuracy and efficiency of data conversion codes in the prior art lead to low efficiency and low efficiency of scientific researchers when converting theoretical research data into practical applications and prone to understanding deviations or inconsistent implementation.
By extracting the non-code information to be converted, key text information is obtained and splitting it to determine the code relevance score for each text block. Based on these scores, code-oriented text blocks are filtered out, and initial code is generated based on the multimodal knowledge graph, and the initial code is optimized and feedback verified, and finally the converted code of multi-programming languages is obtained.
It improves the accuracy and efficiency of the data conversion code, reduces the time and energy investment of scientific researchers in interpreting and implementing theoretical research materials, and reduces the risk of understanding deviations and inconsistent implementation.
Smart Images

Figure CN119987732A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing technology, and in particular to a code conversion method, device, equipment, medium and computer program product. Background Art
[0002] In the current scientific research environment, abundant information resources have promoted the development of science and technology. However, abundant information resources also make how to efficiently transform theoretical research materials into practical applications a technical problem that needs to be solved urgently. Researchers usually need to spend a lot of time and energy to interpret and implement the complex algorithms and experimental contents described in theoretical research materials. This process is not only inefficient, but also prone to misunderstanding or inconsistent implementation.
[0003] Currently, there are some automatic code generation technologies based on machine learning, which mainly focus on establishing the association between text and code snippets, and then outputting the code snippets corresponding to the text through this association. However, the existing methods can only obtain sporadic code snippets, and relevant personnel are required to compile the generated code later. In addition, the accuracy of code conversion by the existing methods is low, and it cannot improve the efficiency of data conversion code. Summary of the invention
[0004] The present invention provides a code conversion method, device, equipment, medium and computer program product, which are used to solve the defects of low accuracy and low efficiency of data conversion code in the prior art, and improve the accuracy and efficiency of data conversion code.
[0005] The present invention provides a code conversion method, comprising the following steps.
[0006] Extract the non-code information to be converted to obtain key text information; Splitting the key text information into multiple text blocks, and determining a code relevance score for each of the text blocks; Based on the code relevance score, determining a code-oriented text block; Generate initial code based on a multimodal knowledge graph and the code-oriented text block; the multimodal knowledge graph is constructed based on historical text information and historical code information; The initial code is optimized and feedback-verified to obtain conversion codes in multiple programming languages.
[0007] According to a code conversion method provided by the present invention, the non-code information to be converted is extracted to obtain key text information, which includes: The non-code information to be converted is subjected to sequential analysis and content block annotation to obtain component labels and component position information; the component labels include text labels of text components and non-text labels of non-text components; Performing text recognition on the text component and the non-text component based on the component position information to obtain a recognized text; Based on the recognized text and the component labels, key text information is determined.
[0008] According to a code conversion method provided by the present invention, the determining key text information based on the recognized text and the component labels includes: Based on the component labels, code-irrelevant text in the identified text is screened out to obtain a primary screened text; Matching the primary screening text based on a preset regular expression to obtain a secondary screening text; The secondary screening text is matched based on a preset key text list to obtain key text information.
[0009] According to a code conversion method provided by the present invention, the key text information is split into multiple text blocks, and the code relevance score of each text block is determined, which includes: Splitting the key text information based on the component position information to obtain multiple text blocks; Determine the adjacent blocks of each text block, input each text block and the adjacent blocks into a text analysis model, and obtain a code relevance score for each text block; the text analysis model is constructed based on historical text blocks, historical code scores and prompt information; the prompt information is determined based on each text block and the adjacent blocks.
[0010] According to a code conversion method provided by the present invention, the code conversion method further includes: Performing entity extraction on the historical text information to obtain multiple text entities; Determine the entity relationship between each of the text entities and other text entities, and construct a text knowledge graph based on the entity relationship; Constructing a code knowledge graph based on the historical code information; Based on the text knowledge graph and the code knowledge graph, a multimodal knowledge graph is constructed.
[0011] According to a code conversion method provided by the present invention, the optimization and feedback verification of the initial code to obtain the conversion code of multiple programming languages includes: Testing the logical consistency of the initial code to obtain a test result; Based on the detection result, the initial code is optimized for performance, readability and compatibility to obtain an optimization result; Based on the test results of the test cases, the initial code is feedback verified to obtain conversion codes in multiple programming languages; the test cases are generated based on the optimization results.
[0012] The present invention also provides a code conversion device, comprising the following modules: A key text information extraction module is used to extract the non-code information to be converted to obtain key text information; A code relevance score determination module, used to split the key text information into multiple text blocks, and determine the code relevance score of each text block; A code-oriented text block determination module, used to determine a code-oriented text block based on the code relevance score; An initial code generation module, used to generate initial code based on a multimodal knowledge graph and the code-oriented text block; the multimodal knowledge graph is constructed based on historical text information and historical code information; The code conversion module is used to optimize and feedback-verify the initial code to obtain conversion codes in multiple programming languages.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-mentioned code conversion methods when executing the computer program.
[0014] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the code conversion method described in any one of the above is implemented.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the code conversion method described above is implemented.
[0016] The code conversion method, device, equipment, medium and computer program product provided by the present invention extract key text information from non-code information, the key text information is integrated and split to obtain multiple text blocks, and the code relevance score of each text block is further determined, and the code-oriented text blocks are screened out through the code relevance score; based on the code-oriented text blocks and the created multimodal knowledge graph, the initial code is generated, and finally the initial code is optimized and feedback verified to obtain the conversion code of multiple programming languages. The present invention automatically generates the corresponding code by parsing the code logic in the non-code information, thereby improving the efficiency and accuracy of the data conversion code. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 This is one of the flow charts of the code conversion method provided by the present invention.
[0019] Figure 2 This is the second flow chart of the code conversion method provided by the present invention.
[0020] Figure 3 It is a structural schematic diagram of the code conversion device provided by the present invention.
[0021] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] Combine the following Figure 1-Figure 4 The code conversion method, device, apparatus, medium and computer program product of the present invention are described.
[0024] Figure 1 This is one of the flow charts of the code conversion method provided by the present invention, such as Figure 1 As shown, the method includes the following: Step 100: extract the non-code information to be converted to obtain key text information; Specifically, the code conversion method provided by the present invention first extracts text information from the non-code information to be converted (hereinafter referred to as research data instead of explanation) to obtain key text information.
[0025] The specific process includes S1, research data information extraction. The specific content of step S1 includes: S1.1, Layout analysis: Analyze the layout of the research materials to obtain the various components of the research materials and record the position of each component in the research materials; then mark the various components of the research materials and label each component with corresponding labels, such as title, abstract, paragraph, table, and picture, etc. Step S1.1 finally obtains the various components of the research materials, the labels of the components, and the positions in the research materials.
[0026] S1.2, text extraction: extract the research material components labeled as titles and paragraphs obtained in S1.1, and sort these components according to the logical structure of the research material.
[0027] S1.3, text recognition: According to the position information of each component in S1.2, the corresponding positions of the non-text information are intercepted in turn, and sent to the text detection model to extract the text information.
[0028] S1.4 Text cleaning and content extraction: Remove irrelevant information from the text, such as references, acknowledgments, author information, and institutional affiliation. Extract key parts of the text, such as the abstract, introduction, related work, methods, and experimental results.
[0029] Step 200: split the key text information into multiple text blocks, and determine the code relevance score of each text block; Specifically, the code conversion method provided by the present invention further includes S2, semantic analysis and code information extraction, and the specific content of step S2 includes: S2.1. Divide the text structure of the research material obtained in the above step S1 into multiple text blocks.
[0030] S2.2. Analyze the code relevance of text blocks: Identify which text blocks contain information that can be converted into code. The basis for identification is the code relevance score of each text block.
[0031] The specific contents of the above step S2.2 include: S2.2.1. Determine prompt information: Set prompt information to guide the large language model to analyze the code relevance of each text block.
[0032] S2.2.2. Code relevance score: Determine the code relevance score for each text block based on the results output by the model after the prompt.
[0033] Step 300: Determine a code-oriented text block based on the code relevance score; Specifically, the specific contents of the above step S2.2 also include: S2.2.3. Determine the code-oriented text blocks in the text blocks based on the code relevance score obtained in the above step S2.2.2.
[0034] In order to analyze the code relevance of each text block, this application proposes a scoring rule framework, which includes five main aspects: keyword and phrase identification, contextual coherence and depth of detail, technical difficulty and complexity, implementation intent and practicality, and the presence of code examples. Each aspect has a corresponding scoring standard, which includes high relevance (e.g., 4 to 5 points), medium relevance (e.g., 2 to 3 points), and low relevance (e.g., 0 to 1 points). The final total score is obtained by averaging the scores of each aspect. If the total score is greater than or equal to 3 points, the text block is considered to be code-oriented; conversely, if the total score is less than 3 points, the text block is considered not to be code-oriented.
[0035] Step 400: Generate initial code based on the multimodal knowledge graph and the code-oriented text block; the multimodal knowledge graph is constructed based on historical text information and historical code information; The multimodal knowledge graph provided in this application is a knowledge graph that can express algorithms, data structures, variables, functions and their interrelationships, using information generated during research materials and code generation. This graph includes not only text information, but also non-text information, such as pictures, tables and formulas, which can fully reflect the connection between research materials and their implementation codes.
[0036] The specific contents of the above step S2.2 also include: S2.2.4. Query and analyze the multimodal knowledge graph to identify all information related to the research material.
[0037] S2.2.5. Generate initial code: Generate initial code using the information determined in S2.2.4 above and the code-oriented text block determined in S2.2.3 above.
[0038] Step 500: Optimize and perform feedback verification on the initial code to obtain conversion codes in multiple programming languages.
[0039] Specifically, the code conversion method provided by the present invention further includes S3, code optimization, verification and feedback loop, and the specific content of step S3 includes: S3.1. Code logic check and optimization: including performance optimization, readability optimization and compatibility optimization.
[0040] S3.2. Automated verification: including designing test cases and implementing automated testing.
[0041] S3.3. Result evaluation and error analysis: including code quality evaluation, error analysis and marking of problem points.
[0042] S3.4, Feedback loop and model adjustment: Feedback the test and evaluation results to the large language model for adjustment and retraining, gradually improving the accuracy of code generation. Repeat the feedback loop until the generated code passes all verifications and tests.
[0043] S3.5. Multi-language support and platform compatibility: The code conversion method provided in this embodiment supports the generation of codes in multiple programming languages and cross-platform compatibility.
[0044] This embodiment extracts key text information from non-code information, and the key text information is integrated and split to obtain multiple text blocks, further determine the code relevance score of each text block, and screen out code-oriented text blocks through the code relevance score; based on the code-oriented text blocks and the created multimodal knowledge graph, generate initial code, and finally optimize and feedback verify the initial code to obtain conversion code of multiple programming languages. The present invention automatically generates corresponding code by parsing the code logic in non-code information, thereby improving the efficiency and accuracy of data conversion code.
[0045] Figure 2 This is the second flow chart of the code conversion method provided by the present invention, such as Figure 2 As shown, the method may also include: Step 110: performing sequence analysis and content block annotation on the non-code information to be converted to obtain component labels and component position information; the component labels include text labels of text components and non-text labels of non-text components; Step 120: performing text recognition on the text component and the non-text component based on the component position information to obtain a recognized text; Step 130: Determine key text information based on the recognized text and the component labels.
[0046] Specifically, the content of the above step S1.1 includes: Input research materials in a certain format and convert them into a series of pictures by page number.
[0047] Layout analysis: Use a pre-trained layout analysis model to analyze the images on each page, identify the various components (including titles, abstracts, paragraphs, tables, images, and formulas), and record their location information.
[0048] The content of the above step S1.2 includes: Extract titles and paragraphs: Filter out the parts labeled as titles and paragraphs from the results of layout analysis.
[0049] Logical structure sorting: Sort the extracted titles and paragraphs according to the position information of each component to restore the logical structure of the research material.
[0050] The content of the above step S1.3 includes: Position interception: Based on the position information obtained in S1.2 above, the areas of the corresponding components in the research data are intercepted in sequence.
[0051] Text detection and recognition: The captured component areas are sent to the pre-trained text detection model to identify the text information therein.
[0052] Text information integration: associate the recognized text information with the corresponding component labels to form a complete text content.
[0053] This embodiment obtains key text information by performing text extraction on the non-code information to be converted, thereby laying a data foundation for the completion of the code conversion method.
[0054] In one embodiment, the code conversion method provided in the embodiment of the present application may also include: Step 131: based on the component labels, filter out code-irrelevant text in the identified text to obtain a primary screening text; Step 132: matching the primary screening text based on a preset regular expression to obtain a secondary screening text; Step 133: Match the secondary screened text based on a preset key text list to obtain key text information.
[0055] Specifically, the content of the above step S1.4 includes: S1.4.1. Keyword matching: Identify and mark the key parts that need to be retained through keyword matching.
[0056] Keyword List: Create a list of common heading and paragraph starter markers.
[0057] Match and tag: Traverse the extracted text, find content that matches the keywords in the above keyword list, and tag the matched parts.
[0058] S1.4.2, Regular expression filtering: Use regular expressions to identify and remove unnecessary parts.
[0059] Regular Expression Mode: Write a regular expression to match the content that you don't want to filter.
[0060] Apply and Filter: Apply regular expressions to text to find and remove unwanted information.
[0061] S1.4.3, non-text information extraction: extract the location information of the pictures, tables and formulas saved in S1.1 above, and cut out the areas of the corresponding components in the pictures in turn. Send them to the visual macro model to obtain text information descriptions of the pictures, tables and formulas.
[0062] S1.4.4. Content organization: Organize and save the marked key parts.
[0063] Content classification: Based on the markings in step S1.4.1 above, divide the text into different sections, such as abstract, introduction, and methods.
[0064] Arrange the output: Arrange the classified key parts into a clear document structure, insert their text descriptions based on the information of the pictures, tables and formulas saved in S1.1 above, and ensure that each part is arranged in the order in which it appears in the original document.
[0065] Final output: Generate a document containing only the key parts and their related non-text information, removing all irrelevant information.
[0066] This embodiment further improves the accuracy of code conversion by filtering out information irrelevant to the code.
[0067] In one embodiment, the code conversion method provided in the embodiment of the present application may also include: Step 210: split the key text information based on the component position information to obtain multiple text blocks; Step 220, determine the adjacent blocks of each of the text blocks, input each of the text blocks and the adjacent blocks into a text analysis model, and obtain a code relevance score for each of the text blocks; the text analysis model is constructed based on historical text blocks, historical code scores and prompt information; the prompt information is determined based on each of the text blocks and the adjacent blocks.
[0068] Specifically, the code relevance score of each text block is obtained through a text analysis model. Each text block and its adjacent text blocks, that is, the information between each text block and its context, are input into the text analysis model to obtain the code relevance score of each text block.
[0069] The text analysis model is built based on historical text blocks, historical code scores, and prompt information. Using the information generated during the research data and code generation process, a multimodal knowledge graph is created that can express algorithms, data structures, variables, functions, and their relationships. This graph includes not only text information, but also non-text information such as pictures, tables, and formulas to fully reflect the connection between research data and its implementation code.
[0070] This embodiment determines the code relevance score of each text block by constructing a multimodal knowledge graph based on historical data.
[0071] In one embodiment, the code conversion method provided in the embodiment of the present application may also include: Step 10: extracting entities from the historical text information to obtain multiple text entities; Step 20: determine the entity relationship between each of the text entities and other text entities, and construct a text knowledge graph based on the entity relationship; Step 30: construct a code knowledge graph based on the historical code information; Step 40: construct a multimodal knowledge graph based on the text knowledge graph and the code knowledge graph.
[0072] Specifically, the embodiment of the present application defines an ontology, which includes entities (nodes) and relationships (edges). For example, the definition of an entity is as follows: Research Paper: A document describing the results of a scientific study. Task: The specific problem that the model is trying to solve. Model Architecture: The design structure of the model. Data Structure: The organization of the data. Dataset: The collection of data used to train and test the model. Evaluation Metrics: The criteria for measuring the performance of the model. Experimental Results: The experimental results and statistics obtained after the model is trained. Experimental Configuration: Details of the hardware and software environment required to conduct the experiment. Figures: Graphs or tables in the research paper. Formulas: Mathematical formulas in the research paper.
[0073] Relationship definition: Propose: connects research data and model architecture, indicating the model architecture proposed in the research data. Solve: connects model architecture and task, indicating the problem solved by the model. Preprocess: the relationship from dataset to preprocessing method, indicating the preprocessing performed on the dataset. Use dataset: connects model architecture and dataset, indicating the dataset used by the model. Evaluate: connects model architecture and evaluation indicator, indicating the evaluation criteria of model performance.
[0074] For entities related to research materials, attribute extraction can cover various types of information, such as the title, author, publication date, journal name, abstract, and keywords of the research materials; the type of model architecture, number of layers, activation function, loss function, optimizer, and learning rate; the data type, number of samples, number of labels, number of features, source, and license of the dataset; the type and value of the evaluation metric; the hardware, software, running time, hyperparameters, training data, validation data, and test data of the experimental configuration; and other attributes, such as code links, open source status, citation count, and download count. Description: Connect an entity with a graph (or formula) to indicate that the entity is described by a graph (or formula).
[0075] The code knowledge graph construction process also includes: Entity Recognition: Use pre-trained NER models to identify key entities in text, such as paper titles, model architectures, data structures, datasets, evaluation metrics, etc. Use OCR technology to identify values in charts or symbols in formulas.
[0076] Relation extraction: Identify relationships between entities, such as propose, solve, preprocess, use, and evaluate.
[0077] Attribute extraction: Extract the attributes of entities from text, such as the number of layers in the model architecture, activation functions, etc. Extract attribute information from charts and formulas, such as the data range in charts and parameters in formulas.
[0078] Graph construction: Use graph database technology to build a knowledge graph. The nodes in the graph represent entities, and the edges represent the relationships between entities.
[0079] This embodiment generates the initial code to be adjusted by constructing a multimodal knowledge graph.
[0080] In one embodiment, the code conversion method provided in the embodiment of the present application may also include: Step 510: Detect the logical consistency of the initial code to obtain a detection result; Step 520: Optimize the performance, readability and compatibility of the initial code based on the detection result to obtain an optimization result; Step 530: Based on the test results of the test cases, feedback verification is performed on the initial code to obtain conversion codes in multiple programming languages; the test cases are generated based on the optimization results.
[0081] Specifically, the specific contents of S3.1 above also include: Logical consistency check: Check whether the algorithm steps, conditional judgments, and loop structures are correct.
[0082] Performance optimization: Identify and replace inefficient data structures and algorithms. Reduce unnecessary calculations and memory usage.
[0083] Improve readability: Follow the coding standards of your chosen programming language. Add clear comments and documentation strings.
[0084] Compatibility optimization: Ensure that the code can run on different operating systems and compilers. For specific language features, check whether they are widely supported.
[0085] The specific contents of S3.2 above also include: Design test cases: Write test cases based on the experimental data and results in the paper.
[0086] Implement automated testing: Use an automated testing framework to perform multiple rounds of testing on the generated code to verify that its functions and outputs are consistent with the paper description.
[0087] The specific contents of S3.3 above also include: Assess code quality: Check whether the code meets the expected functional and performance requirements.
[0088] Error analysis: Compare the output of your code with the expected results in the paper.
[0089] Mark problem spots: Add markers or comments to your code to indicate areas that require further investigation.
[0090] This embodiment improves the accuracy of code conversion by optimizing the initial code and performing loop feedback.
[0091] The code conversion device provided by the present invention is described below. The code conversion device described below and the code conversion method described above can be referred to each other.
[0092] Please refer to Figure 3 The present invention also provides a code conversion device, comprising: The key text information extraction module 301 is used to extract the non-code information to be converted to obtain the key text information; A code relevance score determination module 302 is used to split the key text information into multiple text blocks and determine a code relevance score for each of the text blocks; A code-oriented text block determination module 303, used to determine a code-oriented text block based on the code relevance score; An initial code generation module 304 is used to generate initial code based on a multimodal knowledge graph and the code-oriented text block; the multimodal knowledge graph is constructed based on historical text information and historical code information; The code conversion module 305 is used to optimize and feedback-verify the initial code to obtain conversion codes in multiple programming languages.
[0093] Optionally, the key text information extraction module includes: A non-code information processing unit is used to perform sequential analysis and content block annotation on the non-code information to be converted, and obtain component labels and component position information; the component labels include text labels of text components and non-text labels of non-text components; A text recognition unit, used for performing text recognition on the text component and the non-text component based on the component position information to obtain a recognized text; The key text information determining unit is used to determine the key text information based on the identified text and the component labels.
[0094] Optionally, the key text information determining unit includes: A recognition text screening unit, used for screening out code-irrelevant text in the recognition text based on the component labels to obtain a primary screening text; A text primary screening unit, used for matching the primary screening text based on a preset regular expression to obtain a secondary screening text; The secondary screening text matching unit is used to match the secondary screening text based on a preset key text list to obtain key text information.
[0095] Optionally, the code relevance score determination module includes: A key text information splitting unit, used for splitting the key text information based on the component position information to obtain multiple text blocks; A code relevance score determination unit is used to determine the adjacent blocks of each of the text blocks, input each of the text blocks and the adjacent blocks into a text analysis model, and obtain a code relevance score for each of the text blocks; the text analysis model is constructed based on historical text blocks, historical code scores and prompt information; the prompt information is determined based on each of the text blocks and the adjacent blocks.
[0096] Optionally, the code conversion device further includes: An entity extraction module, used to extract entities from the historical text information to obtain multiple text entities; A text knowledge graph construction module, used to determine the entity relationship between each of the text entities and other text entities, and to construct a text knowledge graph based on the entity relationship; A code knowledge graph construction module, used to construct a code knowledge graph based on the historical code information; A multimodal knowledge graph construction module is used to construct a multimodal knowledge graph based on the text knowledge graph and the code knowledge graph.
[0097] Optionally, the code conversion module includes: An initial code detection unit, used to detect the logical consistency of the initial code and obtain a detection result; An initial code optimization unit, used to optimize the performance, readability and compatibility of the initial code based on the detection result to obtain an optimization result; The initial code feedback verification unit is used to perform feedback verification on the initial code based on the test results of the test cases to obtain conversion codes in multiple programming languages; the test cases are generated based on the optimization results.
[0098] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the code conversion method, which includes: extracting the non-code information to be converted to obtain key text information; splitting the key text information to obtain multiple text blocks, and determining the code relevance score of each text block; determining the code-oriented text block based on the code relevance score; generating the initial code based on the multimodal knowledge graph and the code-oriented text block; the multimodal knowledge graph is constructed based on the historical text information and the historical code information; optimizing and feedback-verifying the initial code to obtain the conversion code of multiple programming languages.
[0099] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0100] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the code conversion method provided by the above methods, which method includes: extracting non-code information to be converted to obtain key text information; splitting the key text information to obtain multiple text blocks, and determining the code relevance score of each text block; determining a code-oriented text block based on the code relevance score; generating initial code based on a multimodal knowledge graph and the code-oriented text block; the multimodal knowledge graph is constructed based on historical text information and historical code information; optimizing and feedback-verifying the initial code to obtain conversion code for multiple programming languages.
[0101] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the code conversion method provided by the above-mentioned methods, the method comprising: extracting non-code information to be converted to obtain key text information; splitting the key text information to obtain multiple text blocks, and determining a code relevance score for each of the text blocks; determining a code-oriented text block based on the code relevance score; generating initial code based on a multimodal knowledge graph and the code-oriented text block; the multimodal knowledge graph is constructed based on historical text information and historical code information; optimizing and feedback-verifying the initial code to obtain conversion code for multiple programming languages.
[0102] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0103] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A code conversion method, characterized in that: include: Extract the non-code information to be converted to obtain key text information; Splitting the key text information into multiple text blocks, and determining a code relevance score for each of the text blocks; Based on the code relevance score, determining a code-oriented text block; Generate initial code based on the multimodal knowledge graph and the code-oriented text block; The multimodal knowledge graph is constructed based on historical text information and historical code information; The initial code is optimized and feedback-verified to obtain conversion codes in multiple programming languages.
2. The code conversion method according to claim 1, characterized in that: The non-code information to be converted is extracted to obtain key text information including: The non-code information to be converted is subjected to sequential analysis and content block annotation to obtain component labels and component position information; the component labels include text labels of text components and non-text labels of non-text components; Performing text recognition on the text component and the non-text component based on the component position information to obtain a recognized text; Based on the recognized text and the component labels, key text information is determined.
3. The code conversion method according to claim 2, characterized in that: The determining of key text information based on the identified text and the component labels includes: Based on the component labels, code-irrelevant text in the identified text is screened out to obtain a primary screened text; Matching the primary screening text based on a preset regular expression to obtain a secondary screening text; The secondary screening text is matched based on a preset key text list to obtain key text information.
4. The code conversion method according to claim 2, characterized in that: The step of splitting the key text information to obtain a plurality of text blocks and determining the code relevance score of each of the text blocks comprises: Splitting the key text information based on the component position information to obtain multiple text blocks; Determine the adjacent blocks of each text block, input each text block and the adjacent blocks into a text analysis model, and obtain a code relevance score for each text block; the text analysis model is constructed based on historical text blocks, historical code scores and prompt information; the prompt information is determined based on each text block and the adjacent blocks.
5. The code conversion method according to claim 1, characterized in that: The code conversion method further includes: Performing entity extraction on the historical text information to obtain multiple text entities; Determine the entity relationship between each of the text entities and other text entities, and construct a text knowledge graph based on the entity relationship; Constructing a code knowledge graph based on the historical code information; Based on the text knowledge graph and the code knowledge graph, a multimodal knowledge graph is constructed.
6. The code conversion method according to claim 1, characterized in that: The optimizing and feedback verification of the initial code to obtain the conversion code of multiple programming languages comprises: Testing the logical consistency of the initial code to obtain a test result; Based on the detection result, the initial code is optimized for performance, readability and compatibility to obtain an optimization result; Based on the test results of the test cases, the initial code is feedback verified to obtain conversion codes in multiple programming languages; the test cases are generated based on the optimization results.
7. A code conversion device, characterized in that: include: A key text information extraction module is used to extract the non-code information to be converted to obtain key text information; A code relevance score determination module, used to split the key text information into multiple text blocks, and determine the code relevance score of each text block; A code-oriented text block determination module, used to determine a code-oriented text block based on the code relevance score; An initial code generation module, used to generate initial code based on the multimodal knowledge graph and the code-oriented text block; The multimodal knowledge graph is constructed based on historical text information and historical code information; The code conversion module is used to optimize and feedback-verify the initial code to obtain conversion codes in multiple programming languages.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the code conversion method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the code conversion method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the code conversion method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Software system development method based on multi-modal AI large model
CN120762638A
Code generation method and device based on artificial intelligence and electronic equipment
CN121349418A