Flight simulator interface code automatic generation method, system and device based on large language model and storage medium

By adopting an automatic interface code generation method based on a large language model, the problems of low efficiency and poor accuracy in interface code generation in flight simulation training systems are solved, achieving efficient and reliable automatic interface code generation and improving the system's adaptability and code quality.

CN120872322BActive Publication Date: 2025-12-09CHINA SOUTHERN TECHNOLOGY (GUANGDONG HENGQIN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511395860.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-12-09
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

Existing technologies in flight simulation training systems suffer from low efficiency and poor accuracy in interface code generation, as well as poor system adaptability. They struggle to handle complex scenarios such as multi-threaded synchronization and exception recovery, and the fixed code structure leads to insufficient maintainability.

Method used

A large language model-based approach is adopted. After acquiring interface data from external devices, noise filtering and standardization are performed. Then, deep learning and semantic understanding algorithms are used to generate interface code templates. These templates are instantiated by combining example data and data interaction rules. The model parameters are dynamically adjusted through test case verification and performance evaluation to achieve closed-loop automated generation.

Benefits of technology

It enables unified collection and management of interface data from multi-source heterogeneous devices, generates code with reasonable structure and accurate logic, improves the practicality and reliability of the code, ensures the correctness of the code's functions and the stability of its operation, forms a self-improving closed-loop system, and improves development efficiency and code quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872322B_ABST
    Figure CN120872322B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of large language models, and provides a large language model-based flight simulator interface code automatic generation method, system, device and storage medium, which solves the problems of low generation efficiency, poor accuracy and poor system adaptability of flight target interface code. The method comprises the following steps: acquiring interface data of a plurality of external devices of a flight target, filtering noise, extracting key information, and converting the interface data into structured input data; using a deep learning and semantic understanding algorithm to deeply analyze the data, extract key semantic information, and generate an interface code template; combining example data and interaction rules to instantiate the template content and generate target code; generating multi-scenario test cases based on a test case algorithm to verify the correctness, stability and performance of the code; comprehensively analyzing the test and performance results, dynamically adjusting the model parameters and training data, and realizing closed-loop automatic code generation. The application improves the generation efficiency, accuracy and system adaptability of the flight target interface code.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of large language models, and in particular to a large language model-based flight simulator interface code automatic generation method, system, device and storage medium. BACKGROUND

[0002] In a flight simulation training system, interface code undertakes the key task of data interaction between the simulator and various external simulation or real devices. Due to the differences in communication protocols and data formats adopted by different device manufacturers, interface development needs to continuously adapt to multiple heterogeneous protocols, and the generated code is required to have high real-time performance, high reliability and good maintainability to meet the strict requirements of flight simulation systems for data exchange accuracy and timeliness.

[0003] A current existing scheme adopts a template-based code generation technology. This technology predefines code templates for multiple common communication protocols, selects the corresponding template according to the protocol description file of the target device, and realizes automatic filling of part of the data fields through rule mapping, finally generating basic interface code. This scheme reduces the workload of completely manual coding to a certain extent.

[0004] However, this scheme still has obvious limitations in practical application. The protocol types covered by its code template library are limited, and it has poor adaptability when facing new or private protocols that are not predefined. The rule mapping method relies on the manually predefined mapping table, and has insufficient expression ability for complex interaction logic and state transition relationships in the protocol, which leads to logical deviations in the generated code when handling multi-thread synchronization, exception recovery and other scenarios. At the same time, the code structure generated by the template tends to be fixed, and there is still room for improvement in code efficiency and maintainability, making it difficult to fully meet the implementation requirements of high reliability and real-time code for flight simulation systems. SUMMARY

[0005] The present application provides a large language model-based flight simulator interface code automatic generation method, system, device and storage medium to solve the problems of low generation efficiency, poor accuracy and poor system adaptability of flight target interface code in the prior art.

[0006] To solve the above technical problems, in a first aspect, the present application provides a large language model-based flight simulator interface code automatic generation method, comprising:

[0007] Obtaining interface data provided by a plurality of external devices in a flight target, the interface data including communication protocols, data formats and example data;

[0008] The interface data is subjected to noise filtering processing, key information of a communication protocol and key information of a data format are extracted from the processed interface data, and the key information of the communication protocol and the key information of the data format are converted according to a preset standardized data format to form structured input data;

[0009] The structured input data is input into a pre-trained large language model, the structured input data is subjected to deep analysis by a deep learning algorithm in the large language model, key semantic information is extracted from the analysis result by a semantic understanding algorithm in the large language model, and an interface code template is generated based on the key semantic information;

[0010] Based on the example data and the data interaction rule, the variables, functions and communication interfaces in the interface code template are instantiated to generate a target interface code;

[0011] Based on the target interface code and the example data, a test case generation algorithm is used to generate test cases in multiple scenarios, the correctness, stability and performance of the interface are verified by executing the test cases, and a test verification result is obtained;

[0012] The target interface code is subjected to structural complexity analysis and running efficiency evaluation processing to obtain a performance analysis result, and the parameters and training data of the large language model are dynamically adjusted in combination with the test verification result and the performance analysis result to realize a closed-loop automatic interface code generation process.

[0013] Optionally, the structured input data is input into a pre-trained large language model, the structured input data is subjected to deep analysis by a deep learning algorithm in the large language model, key semantic information is extracted from the analysis result by a semantic understanding algorithm in the large language model, and an interface code template is generated based on the key semantic information, including:

[0014] The structured input data is input into a large language model, and the protocol definition content and the data format description content in the structured input data are jointly coded by a coding module of the large language model to form a model input sequence;

[0015] A deep learning algorithm provided by the large language model is used to perform multi-level semantic analysis on the model input sequence to obtain a deep analysis result;

[0016] A semantic understanding algorithm provided by the large language model is used to perform semantic relationship reasoning on the deep analysis result to obtain key semantic information, and the key semantic information includes protocol interaction logic and data mapping relationship;

[0017] Generate an interface code template based on the protocol interaction logic and the data mapping relationship.

[0018] Optionally, the semantic understanding algorithm provided by the large language model is used to perform semantic relationship reasoning on the deep parsing result to obtain key semantic information, including:

[0019] Perform dependency relationship analysis on the protocol syntax units in the deep parsing result to identify trigger conditions and timing constraint relationships between protocol commands and responses.

[0020] Perform semantic role labeling on the data structure definitions in the deep parsing result to analyze the functional attributes, value ranges, and encoding methods of data fields.

[0021] Establish a state transition model of message sequences based on the trigger conditions and the timing constraint relationships, and deduce a complete workflow of protocol interaction based on the state transition model.

[0022] Construct mapping rules and conversion relationships between data fields based on the functional attributes, value ranges, and encoding methods, and form standard templates for data parsing and encapsulation based on the mapping rules and conversion relationships.

[0023] Generate key semantic information based on the complete workflow and the standard templates.

[0024] Optionally, generating key semantic information based on the complete workflow and the standard templates includes:

[0025] Convert the complete workflow into a state transition rule set containing state transition conditions and message processing sequences.

[0026] Convert the standard templates into a data operation rule set containing field mapping relationships and data conversion methods.

[0027] Construct a semantic description file based on the state transition rule set and the data operation rule set, which defines protocol interaction logic and data mapping relationships.

[0028] Structurally encapsulate the semantic description file to generate key semantic information.

[0029] Optionally, instantiating variables, functions, and communication interfaces in the interface code template based on the example data and data interaction rules to generate target interface code includes:

[0030] Extract data samples with actual numerical values and corresponding data type descriptions from the example data.

[0031] According to the data transmission sequence and the verification requirement of the data interaction rule, the specific implementation mode of the variable and the function in the interface code template is determined;

[0032] The data sample, the data type description, and the corresponding data structure and communication interface in the interface code template are matched and bound;

[0033] According to the specific implementation mode, the message organization format and the interaction timing flow specified in the communication protocol are combined to fill the bound interface code template, and the target interface code is generated.

[0034] Optionally, based on the target interface code and the example data, a test case generation algorithm is used to generate multiple-scenario test cases, the correctness, stability, and performance of the interface are verified by executing the test cases, and test verification results are obtained, including:

[0035] According to the interface function implemented by the target interface code, normal data flow test scenarios and abnormal data flow test scenarios are generated according to the test requirements corresponding to the interface function;

[0036] Based on the example data, a test case generation algorithm is used to generate input test data sets corresponding to normal data processing scenarios and abnormal data processing scenarios, respectively;

[0037] The input test data set and the corresponding test scene configuration information are combined to form multiple-scenario test cases;

[0038] According to the message organization format and the interaction timing flow specified in the communication protocol, the test cases are executed in the corresponding scenarios to obtain interface output results;

[0039] The interface output results are compared with the preset expected behavior to verify the correctness of the interface under different load conditions, and to verify the stability and performance, and test verification results are obtained.

[0040] Optionally, the target interface code is subjected to structural complexity analysis and running efficiency evaluation processing to obtain performance analysis results, and the parameters and training data of the large language model are dynamically adjusted in combination with the test verification results and the performance analysis results, including:

[0041] The target interface code is subjected to structural complexity analysis to generate quantized indicators corresponding to the code logic complexity and the module dependency relationship, respectively;

[0042] The running efficiency of the target interface code is evaluated in a preset simulation data interaction environment to obtain resource occupation rate and data processing throughput;

[0043] generate a performance analysis result based on the quantization index, the resource occupancy rate, and the data processing throughput;

[0044] determine a code defect and a performance bottleneck type based on the test verification result and the performance analysis result;

[0045] dynamically adjust parameters and training data of the large language model based on the code defect and the performance bottleneck type.

[0046] In a second aspect, the present application provides a large language model-based flight simulator interface code automatic generation system, comprising:

[0047] An acquisition module is configured to acquire interface data provided by a plurality of external devices in a flight target, wherein the interface data comprises a communication protocol, a data format, and example data;

[0048] A filtering module is configured to perform noise filtering processing on the interface data, extract key information of the communication protocol and key information of the data format from the processed interface data, and convert the key information of the communication protocol and the key information of the data format into structured input data according to a preset standardized data format;

[0049] An input module is configured to input the structured input data into a pre-trained large language model, perform deep analysis on the structured input data through a deep learning algorithm in the large language model, extract key semantic information from the analysis result using a semantic understanding algorithm in the large language model, and generate an interface code template based on the key semantic information;

[0050] A generation module is configured to instantiate variables, functions, and communication interfaces in the interface code template based on the example data and data interaction rules, and generate a target interface code;

[0051] A verification module is configured to generate test cases for multiple scenarios using a test case generation algorithm based on the target interface code and the example data, verify the correctness, stability, and performance of the interface by executing the test cases, and obtain a test verification result;

[0052] An evaluation module is configured to perform structural complexity analysis and running efficiency evaluation processing on the target interface code, obtain a performance analysis result, combine the test verification result and the performance analysis result, and dynamically adjust parameters and training data of the large language model to realize a closed-loop automatic interface code generation process.

[0053] In a third aspect, the present application provides an electronic device, comprising:

[0054] A memory is configured to store a computer program;

[0055] The processor is configured to implement the steps of the method for automatically generating an interface code of a flight simulator based on a large language model according to the first aspect.

[0056] In a fourth aspect, the present application provides a computer readable storage medium, wherein a computer program is stored in the computer readable storage medium, and the computer program is configured to implement the steps of the method for automatically generating an interface code of a flight simulator based on a large language model according to the first aspect when executed by a processor.

[0057] The technical scheme provided by the present application has the following beneficial effects:

[0058] The present application realizes unified collection and centralized management of interface data of multiple source heterogeneous devices, and provides a complete data basis for subsequent automatic processing. The data quality and consistency are improved, the normativity and usability of input data are ensured, and a foundation is laid for efficient analysis and understanding of models. The strong learning and reasoning ability of the large language model is utilized to automatically understand the deep semantics of protocols and data formats, generate a code framework with reasonable structure and accurate logic, and reduce the difficulty and workload of code writing. Abstract code templates are converted into specific executable codes to ensure that the generated codes can accurately reflect the actual data interaction requirements and improve the practicality and accuracy of the codes. Comprehensive test cases are automatically generated and executed to verify the correctness, stability and performance of the generated codes, and the reliability of the codes is improved. The quality of the codes is quantitatively evaluated and continuously optimized to form a self-improving closed-loop system, and the overall quality and generation efficiency of the output codes are continuously improved.

[0059] Further, the structured input data after joint encoding is subjected to multi-level semantic analysis by a deep learning algorithm provided by the large language model, and the semantic relationship reasoning algorithm is used to analyze the results to extract key semantic information containing protocol interaction logic and data mapping relationship, and finally generate an interface code template based on the information.

[0060] Furthermore, the deep semantic understanding and reasoning ability of the large language model is utilized to automatically extract the key interaction logic and mapping relationship from the protocol and data format, and generate an interface code template with clear structure and accurate logic based on the extracted information, which greatly improves the automation level and intelligent degree of code generation, and ensures the high consistency between the code template and the protocol specification.

[0061] These and other aspects of the present application will become more apparent from the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings needed to be used in the description of the embodiments or the prior art will be briefly introduced. Obviously, the accompanying drawings described below are only some embodiments of the present application, and all other drawings obtained by those of ordinary skill in the art without creative labor based on these drawings belong to the protection scope of the present application.

[0063] Figure 1 A flowchart of a flight simulator interface code automatic generation method based on a large language model provided by an embodiment of the present application;

[0064] Figure 2 A data collection module data interaction flowchart of a flight simulator interface code automatic generation method based on a large language model provided by an embodiment of the present application;

[0065] Figure 3 A DeepSeek large model data processing and code generation flowchart of a flight simulator interface code automatic generation method based on a large language model provided by an embodiment of the present application;

[0066] Figure 4 A test case generation and verification flowchart of a flight simulator interface code automatic generation method based on a large language model provided by an embodiment of the present application;

[0067] Figure 5 A connection mode of a code analysis tool and a performance analysis tool and an interface code and a flow direction of analysis data of a flight simulator interface code automatic generation method based on a large language model provided by an embodiment of the present application;

[0068] Figure 6 System optimization and iteration of a flight simulator interface code automatic generation method based on a large language model provided by an embodiment of the present application;

[0069] Figure 7 A structure schematic diagram of a flight simulator interface code automatic generation system based on a large language model provided by an embodiment of the present application. DETAILED DESCRIPTION

[0070] In order to make the person skilled in the art better understand the present application, the present application will be further described in detail below in combination with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor belong to the protection scope of the present application.

[0071] The core of the present application is to provide a flight simulator interface code automatic generation method based on a large language model, and a specific embodiment thereof has a flowchart as shown in Figure 1 The method comprises the following steps:

[0072] Step 101: Obtain interface data provided by a plurality of external devices in a flight target, wherein the interface data comprises a communication protocol, a data format and example data.

[0073] In step 101, the flight target refers to the overall flight vehicle system that needs to be simulated and trained. The external device refers to a simulation or real component that interacts with the flight simulator, such as an avionics system, a power system, a control load system, etc. The interface data is a collection of all relevant information for realizing communication and data exchange between the flight simulator and the external device. The communication protocol is a part of the interface data, which specifies a set of rules for data transmission and exchange in the network, including the structure of the data packet, the start and end markers, the verification method, the correspondence between commands and responses, etc. The data format is a part of the interface data, which defines the specific organization form and meaning of the data content, such as whether the data type of a certain data field is an integer or a floating point number, the length of the data in bytes, the unit of the data, and the physical meaning represented by the data. The example data is a part of the interface data, which is a set of real recorded original data streams or data packet instances that meet the requirements of the communication protocol and the data format, and is used to specifically show the real appearance of data interaction.

[0074] In the embodiment of the present application, a data acquisition tool or program is used to establish a communication connection with each external device, and according to the interface document provided by each device, the data from these devices is actively read or passively received. These data include text or configuration files that specify communication rules (communication protocol), documents that define data organization methods (data format), and log files that record real data interaction processes (example data). All these collected raw materials together constitute the input basis for subsequent automated processing.

[0075] For example, the developer obtains a User Datagram Protocol (UDP) document, a format document defining 20 data structures, and a data log file recording 100 sets of actual transmitted and received data from the device manual and communication log of the navigation display unit (external device A) and the engine control unit (external device B) of the flight target. These documents and files together constitute the interface data obtained in step 101.

[0076] Step 102: Perform noise filtering on the interface data, extract key information of the communication protocol and key information of the data format from the processed interface data, and convert the key information of the communication protocol and the key information of the data format according to a preset standardized data format to form structured input data.

[0077] In step 102, noise filtering refers to the process of identifying and removing errors, redundancies or incomplete information from the original interface data, with the purpose of improving data quality. Key information of the communication protocol refers to protocol type, data transmission format, and data interaction rules. Key information of the data format refers to data structure and data type. Standardized data format is a pre-defined and unified data organization specification, which is used to ensure that all information from different sources can be processed by subsequent large language models in the same style. Structured input data is a clear and well-organized information set obtained after cleaning, extraction and conversion, which is easy for computer programs, especially large language models, to read and understand efficiently and accurately.

[0078] In the embodiments of the present application, first, noise filtering is performed on the original interface data obtained in step 101, such as removing garbled lines in logs, correcting spelling errors in documents, etc. Then, key information in the communication protocol, such as protocol type and important interaction rules, is extracted from the cleaned data, and key information in the data format, such as core data structure and data type, is extracted. Then, these extracted key information is converted and reorganized according to a pre-defined standardized data format. Finally, a structured input data that is neat and can be directly used by a large language model is output.

[0079] For example, the interface data obtained from devices A and B is cleaned, and 3 data records with format errors in the data log are removed. Then, the interaction rule "instruction 0xA1 needs to be responded within 5 milliseconds" is extracted from the communication protocol, and the data type "pitch angle data is 4 bytes unsigned integer" is extracted from the data format. The extracted key information is converted into structured input data in JavaScript Object Notation (JSON), and the JSON file clearly contains protocol rules and data format definitions.

[0080] Step 103: Input the structured input data into a pre-trained large language model, perform deep analysis on the structured input data through a deep learning algorithm in the large language model, extract key semantic information from the analysis result using a semantic understanding algorithm in the large language model, and generate an interface code template based on the key semantic information.

[0081] In step 103, the large language model is a pre-trained artificial intelligence model with strong text understanding and generation capabilities. The deep learning algorithm is a machine learning technique used in the large language model to automatically learn and extract complex patterns and features from data. Deep analysis refers to the process of in-depth analysis and understanding of structured data input by the large language model using its deep learning capabilities. The semantic understanding algorithm is a technique used in the large language model to understand the true meaning and association behind the text and data. Key semantic information is the core knowledge about how the protocol works and how the data is mapped, extracted after deep analysis and semantic understanding. The interface code template is a code draft generated based on key semantic information, containing the main framework and logic of the code but not yet filled with specific values.

[0082] In the embodiments of the present application, the structured input data generated in step 102 is input into the pre-trained large language model. The model first uses its internal deep learning algorithm to comprehensively and deeply analyze the data and understand its internal structure. Then, it uses its semantic understanding algorithm to reason about the analysis results and extract key semantic information such as protocol interaction logic and data mapping relationships. Finally, the model generates a structured interface code template based on these key semantic information, which defines the overall framework and main functions of the code.

[0083] For example, the structured input data in JSON format is input into the large language model. The model extracts key semantic information through analysis, including "the protocol interaction logic is to send 0xA1 instruction to the main system and wait for the received data packet" and "the data mapping relationship is that the data in bytes 12 to 15 of the received data packet needs to be processed and converted into the actual angle value". Based on this, the model generates an interface code template that contains a function framework for sending instructions and a function framework for parsing data packets.

[0084] Step 104: Based on the example data and data interaction rules, the variables, functions, and communication interfaces in the interface code template are instantiated to generate the target interface code.

[0085] In step 104, the data interaction rules are part of the "key information of the communication protocol". Instantiation refers to the process of filling the abstract framework and variables in the interface code template with specific, executable code content. The target interface code refers to the complete, directly compilable and executable program code obtained after instantiation.

[0086] In the embodiments of the present application, according to the example data obtained in step 101, the specific data type and value sample of the variable in the code template are determined. At the same time, according to the data interaction rules derived from the protocol, the specific implementation logic of the function in the template is determined. Then the example data is bound to the preset data structure in the template, and the assembly and parsing logic of the message is filled according to the data interaction rules. Finally, a complete and usable target interface code is generated.

[0087] For example, a set of real data is extracted from the initial 100 sets of example data, in which the pitch angle original value is 1234567. According to the check rules specified in the protocol, the calculation logic of the Cyclic Redundancy Check (CRC) is implemented in the check function framework of the interface code template. The real data is bound to the data structure in the template, and the message processing function is completely filled, and finally a target interface code is generated which can correctly send 0xA1 instruction and parse the reply data packet.

[0088] Step 105: Based on the target interface code and the example data, a test case generation algorithm is used to generate multiple scene test cases, and by executing the test cases, the correctness, stability and performance of the interface are verified, and a test verification result is obtained.

[0089] In step 105, the test case generation algorithm is a computer program or method that can automatically create test scenes and test data according to code functions and input data. The test case is a specific test scheme containing test input data and expected output results generated by the test case generation algorithm. The test verification result is a comprehensive evaluation report on the correctness, stability and performance of the target interface code after executing all test cases.

[0090] In the embodiments of the present application, according to the function implemented by the target interface code generated in step 104 and the example data in step 101, a test case generation algorithm is used to automatically design test cases in multiple scenarios, including normal and abnormal situations. Then automatically run these test cases, record the actual output of the code. By comparing the actual output with the expected result, the correctness of the code is verified. The stability is tested by long time running. The performance is evaluated by monitoring the processing speed. Finally, the test verification result is obtained.

[0091] For example, using the test case generation algorithm, 85 normal scenario test cases and 15 abnormal scenario test cases are created based on 100 sets of example data. After executing all test cases, it is found that the generated code will make an error when processing a specific abnormal data packet, and the memory usage will slowly increase after continuous running for 1 hour. These findings are recorded in the test verification result.

[0092] Step 106: Perform structural complexity analysis and running efficiency evaluation on the target interface code to obtain performance analysis results. Combine the test verification results and the performance analysis results to dynamically adjust the parameters and training data of the large language model, to realize the closed-loop automatic generation of interface code.

[0093] In step 106, structural complexity analysis refers to using tools to analyze the structure of the code and evaluate its complexity, usually focusing on metrics such as cyclomatic complexity. Running efficiency evaluation refers to running the code in a simulated real environment to measure its resource occupancy and processing speed, etc. Performance analysis results are quantitative reports on code quality and technical performance obtained by combining structural complexity and running efficiency evaluation. Closed-loop automation refers to feeding the final output results (test verification results and performance analysis results) of the entire process back to the large language model at the starting stage to guide its self-optimization, thereby forming an automatic cycle and continuously improving system.

[0094] In the embodiments of the present application, the code analysis tool is used to analyze the structural complexity of the target interface code generated in step 104, and calculate its cyclomatic complexity and other indicators. At the same time, the code is run in a simulated environment to evaluate the central processing unit usage (CPU) and data processing throughput. Combine the test verification results of step 105 and the performance analysis results of this step to locate the defects and performance bottlenecks in the code. According to these findings, dynamically adjust the parameters and training data of the large language model to optimize its code generation strategy, thereby realizing closed-loop automation and improving the quality of the code generated next time.

[0095] For example, the target interface code is analyzed and the cyclomatic complexity is measured to be 15. Under the simulated load of processing 1000 data packets per second, the CPU occupancy rate is monitored to be 12%, and the throughput is 950 data packets per second. Combining the defects found in the test verification results, it is determined that the problem is caused by the overly complex logic of a function in the code. Therefore, the error information and performance data are fed back to the large language model, the training data is adjusted, the weight of similar correct code samples is increased, and the model is optimized. When generating code for similar protocols next time, the model can generate code with simpler structure and more correct logic.

[0096] The application realizes intelligent generation and continuous optimization of flight simulator interface code through a series of automatic steps. The method can automatically process multi-source heterogeneous interface data, use a large language model to deeply understand protocol semantics and generate high-quality code templates, combine instantiation and comprehensive testing and verification to ensure code reliability, and finally continuously improve code generation quality and system adaptive ability through a closed-loop feedback mechanism, effectively improving development efficiency, code accuracy and overall system performance.

[0097] To solve how to make a large language model accurately understand a communication protocol and generate a high-quality interface code template, in some embodiments, step 103: the structured input data is input into a pre-trained large language model, the structured input data is deeply parsed by a deep learning algorithm in the large language model, key semantic information is extracted from the parsed results by a semantic understanding algorithm in the large language model, and an interface code template is generated based on the key semantic information, including:

[0098] Step 201: input the structured input data into the large language model, and jointly encode the protocol definition content and data format description content in the structured input data by the encoding module of the large language model to form a model input sequence.

[0099] In step 201, the encoding module is a component of the large language model responsible for converting different forms of input data into a series of digital identifiers understandable by the model. The protocol definition content and data format description content in the structured input data are extracted and converted from the information obtained after noise filtering and standardization processing of the original interface data; the protocol definition content includes protocol type, data transmission format and data interaction rules extracted from the communication protocol, and the data format description content includes data structure, data type and data constraint conditions extracted from the data format. Joint encoding refers to the process of uniformly converting the protocol definition content and data format description content as a whole. The model input sequence is an ordered digital sequence containing all semantic information of the protocol and data format obtained after joint encoding, which is used as the direct input of the large language model.

[0100] In the embodiments of the application, the large language model first receives the structured input data, and then the encoding module in the large language model starts to work, reads and analyzes the protocol definition content and data format description content contained in the data as a whole information, maps all these text and structure information into a continuous digital sequence rich in semantic relationship through specific conversion rules, and finally forms a complete model input sequence for subsequent analysis.

[0101] Step 202: using the deep learning algorithm provided by the large language model, the model input sequence is subjected to multi-level semantic analysis, and a deep analysis result is obtained.

[0102] In step 202, multi-level semantic analysis refers to the process of analyzing and understanding the input sequence from shallow to deep and from surface to connotation by the model. The deep analysis result is the internal deep representation formed by the model after multi-level semantic analysis, which contains complex structure and semantic information in the input data.

[0103] In the embodiments of the present application, the large language model uses the deep learning algorithm provided by it to deeply process the model input sequence generated in step 201. This processing process contains multiple levels. The model first analyzes the basic syntax structure of the sequence, then understands the interaction rules expressed by it, and finally understands the deep semantic association. Through this series of layer-by-layer progressive analysis, a deep analysis result that can fully reflect the essence of the protocol and data format is finally obtained.

[0104] Step 203: using the semantic understanding algorithm provided by the large language model, the deep analysis result is subjected to semantic relationship reasoning, and key semantic information is obtained, including protocol interaction logic and data mapping relationship.

[0105] In step 203, semantic relationship reasoning refers to the process of inferring implicit and not explicitly expressed semantic associations based on existing information using logic and knowledge. The protocol interaction logic is a component of the key semantic information, which describes the message exchange order and condition judgment rules that both parties must follow to complete a specific function. The data mapping relationship is another component of the key semantic information, which defines the correspondence between fields in different data formats and the necessary conversion rules.

[0106] In the embodiments of the present application, the large language model uses the semantic understanding algorithm provided by it to deeply analyze the deep analysis result obtained in step 202. The model analyzes the association between various elements in the analysis result, infers the complete interaction logic of the communication parties, accurately deduces how the data fields correspond and convert, and finally extracts the key semantic information containing the protocol interaction logic and the data mapping relationship.

[0107] Step 204: generating an interface code template based on the protocol interaction logic and the data mapping relationship.

[0108] In the embodiments of the present application, the model starts to construct the code framework according to the two parts of key semantic information of protocol interaction logic and data mapping relationship obtained in step 203. The protocol interaction logic is used to generate the control flow code structure, such as function call sequence and conditional judgment branch. The data mapping relationship is used to generate data structure definition and data conversion function prototype. Finally, these elements are combined into an interface code template with clear structure and complete logic.

[0109] The following is a specific example:

[0110] In the foregoing embodiments, after inputting the JSON format structured input data containing the rule "instruction 0xA1 needs to be responded within 5 milliseconds" and the type definition of "pitch angle data is 4 bytes of unsigned integer" into the large language model, the model first encodes the protocol definition content and the data format description content through its encoding module, converts the rule and structure described in text into a model input sequence containing 512 digital identifiers. The length of the sequence is determined by the preset input dimension of the model. Then, the large language model provides a deep learning algorithm to perform multi-level semantic analysis on the model input sequence. The model identifies the part representing the instruction type, the part representing the time constraint, and the part representing the data structure in the sequence in turn, and understands the dependency relationship between them. Finally, a deep analysis result containing complete semantics is obtained. Then, the large language model provides a semantic understanding algorithm to perform semantic relationship reasoning on the deep analysis result. The model infers that when receiving instruction 0xA1, a data packet must be replied within 5 milliseconds, and further analyzes that the 4 bytes starting from the 12th byte in the data packet together represent an unsigned integer. This integer needs to be converted to have physical meaning, thereby extracting the key semantic information containing the protocol interaction logic and the data mapping relationship. Finally, based on the key semantic information, an interface code template is generated, which includes a sending function framework for sending 0xA1 instruction, a receiving processing function framework for waiting and receiving data packet within 5 milliseconds, and a data analysis function framework for extracting the data from the 12th to 15th byte in the received data packet and converting it to the actual angle value.

[0111] In the embodiments of the present application, the above complete step scheme improves the accuracy and reliability of code generation by letting the large language model deeply understand the semantics of protocols and data, and automatically generating a code framework with excellent structure, effectively reducing the complexity and manual intervention requirements of subsequent development, and laying a solid foundation for quickly constructing high-quality interface code.

[0112] To solve how to accurately extract key semantic information understandable by machines from the deep parsing result, in some embodiments, step 203: using the semantic understanding algorithm provided by the large language model, performing semantic relationship reasoning on the deep parsing result to obtain key semantic information, including:

[0113] Step 301: performing dependency relationship analysis on the protocol syntax units in the deep parsing result to identify the trigger condition and timing constraint relationship between the protocol command and the response.

[0114] In step 301, the protocol syntax units in the deep parsing result are derived from the intermediate parsing data generated by the large language model after distributed representation learning on the structured input data, which means the smallest unit with independent syntax function identified in the communication protocol, such as protocol command, response code, parameter identifier, etc. Dependency relationship analysis is a method of analyzing the dependent, modifying or triggering relationship between these units. Trigger condition refers to the prerequisite or rule that a protocol command must meet to trigger a specific response. Timing constraint relationship specifies the sequence and delay requirements between the command and the response in time.

[0115] In the embodiments of the present application, the semantic understanding algorithm first performs dependency relationship analysis on the protocol syntax units in the deep parsing result to identify the dependency relationship between each command and response, and then accurately extracts the specific conditions required to trigger the response and the time sequence and delay constraints they must comply with.

[0116] Step 302: performing semantic role labeling on the data structure definition in the deep parsing result to analyze the functional attributes, value range and encoding method of the data field.

[0117] In step 302, the data structure definition in the deep parsing result is derived from the intermediate parsing data generated by the large language model after distributed representation learning on the structured input data, which means the formal specification about data organization method parsed from the data format description, including field composition, type system and memory layout, etc. Data field refers to the smallest data unit that constitutes the data structure, each field has a specific data type, length and semantic meaning. Semantic role labeling is a process of assigning a specific function label to each field in the data structure. Functional attributes describe the role or physical meaning represented by a data field. Value range defines the upper and lower limits of the numerical values that a data field can represent. Encoding method specifies the specific representation method of a data field in binary or character stream.

[0118] In the embodiments of the present application, the semantic understanding algorithm then performs semantic role labeling on the data structure definition in the deep parsing result, parses what the specific function of each data field is, what the valid value range is, and which encoding format is used for storage and transmission.

[0119] Step 303: According to the trigger condition and the timing constraint relationship, a state transition model of the message sequence is established, and a complete workflow of protocol interaction is derived based on the state transition model.

[0120] In step 303, the message sequence refers to a message set with a specific timing relationship formed according to the protocol specification, which is derived from the message exchange order model established according to the trigger condition and the timing constraint relationship between the protocol syntax units. The state transition model is an abstract representation for describing how the system switches from one state to another state under different message stimuli. The complete workflow of protocol interaction is a sequence diagram derived based on the state transition model, which clearly describes all possible paths from the beginning to the end of the interaction.

[0121] In the embodiments of the present application, according to the trigger condition and the timing constraint relationship identified in step 301, the semantic understanding algorithm begins to construct the state transition model of the message sequence, which clearly shows the state migration path of the system under various messages. Based on this model, the algorithm can derive the complete workflow of the entire protocol interaction.

[0122] Step 304: According to the function attribute, the value range, and the encoding method, a mapping rule and a conversion relationship between data fields are constructed, and a standard template for data parsing and packaging is formed based on the mapping rule and the conversion relationship.

[0123] In step 304, the mapping rule defines how different data fields correspond to and associate with each other. The conversion relationship describes the calculation or processing operation required to convert the data of one field to another field. The standard template for data parsing and packaging is formed based on the mapping rule and the conversion relationship, which is a standardized specification for guiding how to extract data from the original data stream and how to package the data into a transmission format.

[0124] In the embodiments of the present application, according to the data field function attribute, the value range, and the encoding method parsed in step 302, the semantic understanding algorithm constructs the mapping rule and the conversion relationship between these fields, clearly defines how the fields correspond and how the data is converted, and forms the standard template for data parsing and packaging based on these rules and relationships.

[0125] Step 305: Based on the complete workflow and the standard template, key semantic information is generated.

[0126] In the embodiments of the present application, based on the complete protocol interaction workflow derived in step 303 and the data parsing and encapsulation standard template formed in step 304, the semantic understanding algorithm integrates them together to generate the final key semantic information which completely defines the protocol interaction logic and the data mapping relationship.

[0127] The following is a specific example:

[0128] Based on the obtained deep parsing result, the semantic understanding algorithm first performs dependency analysis on the protocol syntax units in the deep parsing result, identifies the trigger condition between the protocol command instruction 0xA1 and the response data packet 0x55AA as that the instruction 0xA1 is the only command for requesting data, and the timing constraint relationship is that the response must be completed within 5 milliseconds. Then, semantic role labeling is performed on the data structure definition in the deep parsing result, and it is parsed that the data field function attribute corresponding to the 4 bytes starting from the 12th byte in the data packet is the pitch angle, the value range is 0 to 4294967295, and the encoding mode is 32-bit unsigned integer. Then, a state transition model of the message sequence is established according to the identified trigger condition and timing constraint relationship, the model includes three states, state one is a waiting instruction state, state two is a state of waiting for a response after receiving the instruction 0xA1, and state three is a data processing state after receiving the data packet 0x55AA, the model specifies that the condition for converting from state one to state two is to receive the instruction 0xA1, the condition for converting from state two to state three is to receive the data packet 0x55AA within 5 milliseconds, and if it is timed out, it returns to state one. Based on the state transition model, the complete protocol interaction workflow is derived as follows: after sending the instruction 0xA1, start the 5 millisecond timer, if the data packet 0x55AA is received before the timer is timed out, perform data processing, otherwise, it is considered as a timeout failure. Then, according to the data field function attribute pitch angle, the value range 0 to 4294967295 and the encoding mode 32-bit unsigned integer, a mapping rule between data fields is constructed, the rule specifies that the 12th to 15th bytes of the data packet are mapped to the pitch angle variable, and a conversion relationship is established, the conversion relationship is that the actual value of the pitch angle is equal to the original value divided by the conversion coefficient 1000, wherein the actual value of the pitch angle represents the final physical quantity in degrees, the original value represents the 4-byte unsigned integer extracted from the data packet, and the conversion coefficient 1000 is a dimensionless scaling factor. Based on the mapping rule and the conversion relationship, a standard template for data parsing and encapsulation is formed, which specifies how to extract the specified bytes from the data packet and complete the numerical conversion. Finally, based on the derived complete workflow and the formed standard template, the key semantic information is generated, which completely contains the protocol interaction logic and the data mapping relationship.

[0129] In the embodiments of the present application, the above complete step scheme is converted into clear, structured and machine understandable core semantic information through layer-by-layer semantic reasoning of complex protocol and data format description, which provides a solid and reliable logical basis for subsequent accurate generation of interface code, and improves the capture accuracy and depth of protocol intent and data meaning.

[0130] In order to convert the protocol workflow and data operation specification into machine-readable standardized semantic information, in some embodiments, step 305: generating key semantic information based on the complete workflow and the standard template, includes:

[0131] Step 401: converting the complete workflow into a state transition rule set containing state transition conditions and message processing sequences.

[0132] In step 401, the state transition condition refers to the event or judgment condition that must be met to trigger the system to switch from one state to another. The message processing sequence specifies the order in which different messages are processed and responded. The state transition rule set is a collection of multiple independent rules obtained by disassembling the complete workflow, each rule clearly describes the state change that will occur under what conditions and how to process the message.

[0133] In the embodiments of the present application, each state change link and message processing step described in the complete workflow is disassembled and refined, and an independent rule is created for each link, which clearly specifies the specific conditions required for state transition and the processing sequence of messages in the process. The collection of all these rules constitutes the state transition rule set.

[0134] Step 402: converting the standard template into a data operation rule set containing field mapping relationships and data conversion methods.

[0135] In step 402, the field mapping relationship indicates the correspondence between the fields of the source data and the target data. The data conversion method defines the specific calculation or processing steps required to convert the source data into the target data. The data operation rule set is a collection of multiple independent rules obtained by disassembling all data operation details contained in the standard template, each rule clearly describes how a group of fields is mapped and how data is converted.

[0136] In the embodiments of the present application, each group of field mapping relationships and each kind of data conversion method defined in the standard template is extracted and formatted, and an independent rule is created for each group of mapping and each kind of conversion, which clearly specifies how the fields correspond to each other and how the data is specifically converted. The collection of all these rules constitutes the data operation rule set.

[0137] Step 403: constructing a semantic description file according to the state transition rule set and the data operation rule set, the semantic description file defining protocol interaction logic and data mapping relationship.

[0138] In step 403, the semantic description file is a file written in a specific format for machine reading, which integrates all core information of protocol interaction and data operation.

[0139] In the embodiment of the present application, the state transition rule set obtained in step 401 and the data operation rule set obtained in step 402 are taken as inputs, and the contents in the two rule sets are systematically organized into a file according to a pre-defined file format, which finally constitutes the semantic description file, which completely defines the protocol interaction logic and data mapping relationship.

[0140] Step 404: structurally packaging the semantic description file to generate key semantic information.

[0141] In step 404, structurally packaging refers to a process of adding necessary metadata information to the semantic description file and packaging it into a standard format.

[0142] In the embodiment of the present application, the semantic description file generated in step 403 is processed, version number, creation time and other management information are added, and it is converted into a standard data exchange format which is compact and easy to parse, and the final achievement obtained after this processing is the key semantic information.

[0143] The following is a specific example:

[0144] According to the complete workflow and the formed standard template, the complete workflow is first converted into a state transition rule set containing three rules. Rule one specifies that the state transition condition from the waiting instruction state to the waiting response state is receiving instruction 0xA1, and the message processing sequence is to immediately send instruction 0xA1. Rule two specifies that the state transition condition from the waiting response state to the data processing state is receiving data packet 0x55AA within 5 milliseconds, and the message processing sequence is to start a 5-millisecond timer and process data after receiving the data packet. Rule three specifies that the state transition condition from the waiting response state to the waiting instruction state is 5-millisecond timeout, and the message processing sequence is to terminate the current session. Then the standard template is converted into a data operation rule set containing one rule. The rule specifies that the field mapping relationship is that bytes 12 to 15 of the data packet are mapped to the pitch angle variable, and the data conversion method is that the pitch angle actual value is equal to the original value divided by the conversion coefficient 1000, where the pitch angle actual value represents the final physical quantity in degrees, the original value represents the 4-byte unsigned integer extracted from the data packet, and the conversion coefficient 1000 is a dimensionless scaling factor. For example, when the original value is 1234567, the pitch angle actual value is equal to 1234567 divided by 1000, which is 1234.567, and the pitch angle actual value is 1234.567 degrees. Then, according to the state transition rule set containing three rules and the data operation rule set containing one rule, a semantic description file is constructed. The file is written in JSON format, where the state transition chapter completely describes the three state transition rules, and the data operation chapter clearly defines the field mapping relationship and the data conversion method. The file completely defines the protocol interaction logic and the data mapping relationship. Finally, the semantic description file is structurally packaged, metadata information such as file version number 1.0 and CRC32 algorithm calculation check code 0x76D8FDC3 is added, and the final key semantic information is generated. The calculation process of the check code is to perform cyclic redundancy check calculation on the file content, which is used to ensure the integrity of the information in the transmission process.

[0145] In the embodiments of the present application, the above complete step scheme converts the flow and template into a rule set and encapsulates it into a standard information package, so that the key protocol and data semantics can be clearly and unambiguously defined and transmitted, providing accurate and reliable input basis for subsequent automatic code generation, and improving the information processing efficiency and overall reliability of the system.

[0146] In order to convert the abstract interface code template into specific executable target code, in some embodiments, step 104: based on the example data and data interaction rules, the variables, functions and communication interfaces in the interface code template are instantiated to generate target interface code, including:

[0147] Step 501: Extracting data samples with actual numerical values and corresponding data type descriptions from the example data.

[0148] In step 501, the data sample refers to a real data instance containing specific values extracted from the example data. The data type description is information corresponding to the data sample, which explains the value category and format of the data sample.

[0149] In the embodiments of the present application, representative data records are selected from the provided example data, and data segments containing actual numerical values are extracted, and at the same time, the data type descriptions corresponding to these data segments are recorded.

[0150] Step 502: According to the data transmission sequence and verification requirements of the data interaction rule, determine the specific implementation of the variables and functions in the interface code template.

[0151] In step 502, the data transmission sequence and verification requirements in the data interaction rule are derived from the key semantic information extracted after deep analysis of the communication protocol, which means the arrangement and sending order rules of each field in the data packet, and the verification mechanism for ensuring data integrity and accuracy, such as CRC check, parity check, etc. The specific implementation refers to the specific code writing method for the variables and function logic in the interface code template that have not been defined according to the requirements of the data interaction rule.

[0152] In the embodiments of the present application, according to the data transmission sequence and data verification regulations that must be followed in the data interaction rule, it is determined how each variable in the interface code template should be defined and how each function should be written in its internal logic, so as to form a specific implementation scheme.

[0153] Step 503: Matching and binding the data samples, the data type descriptions, and the corresponding data structures and communication interfaces in the interface code template.

[0154] In step 503, matching and binding refers to the process of establishing a one-to-one correspondence between the data samples and their type descriptions and the abstract data structures and communication interfaces reserved in the interface code template.

[0155] In the embodiments of the present application, the data samples and data type descriptions extracted in step 501 are associated with the data structures and communication interfaces already defined in the interface code template, ensuring that real data can be connected to the correct abstract location in the code.

[0156] Step 504: According to the specific implementation, combining the message organization format and interaction timing flow specified in the communication protocol, filling the bound interface code template to generate the target interface code.

[0157] In step 504, the message organization format and the interactive timing flow specified in the communication protocol are derived from the specification description obtained by parsing the protocol definition, which respectively means the binary or character level organization structure of the message (such as the layout of the frame header, data field, and frame tail), and the time sequence and response relationship that must be followed in the message exchange process; the data interaction rule focuses on the order and verification at the data level, while the communication protocol specifies the message structure and interactive timing at a higher level, and the two are the relationship between the specific implementation and the overall framework, and the data interaction rule needs to follow the overall specification of the communication protocol.

[0158] In the embodiment of the present application, the variable and function specific implementation determined according to step 502, while strictly complying with the message specific organization style and interactive time sequence specified in the communication protocol, is used to fill in the details of the interface code template that has completed data binding, and finally generate the complete target interface code.

[0159] The following is a specific example:

[0160] With the generated interface code template and the existing example data, a set of data samples with actual values is first extracted from the 100 example data sets, which contains the pitch angle original value 1234567, and its corresponding data type description is recorded as 32-bit unsigned integer. Then, according to the data interaction rules, the data transmission sequence must be sent first and then received, and the CRC32 algorithm must be used for data integrity verification, the specific implementation of the variables and functions in the interface code template is determined, which includes defining a variable named pitch to store the pitch angle data, and implementing a function to calculate the CRC check code, which uses the standard CRC32 polynomial algorithm. Then the pitch angle original value 1234567 in the data sample and its data type description 32-bit unsigned integer are matched and bound with the pre-defined attitude data structure and serial communication interface in the interface code template to ensure that the data can be correctly stored and transmitted. Finally, according to the determined specific implementation, combined with the message organization format requirement in the communication protocol that each data packet must start with 0x55AA and end with 2-byte CRC check code, and the interaction timing flow requirement that the request response cycle must be completed within 5 milliseconds, the bound interface code template is completely filled, and the message assembly function is responsible for constructing the data packet starting with 0x55AA, the data sending function is responsible for sending the instruction 0xA1 through the serial port, the data receiving function is responsible for waiting and receiving the data packet within 5 milliseconds, and the data parsing function is responsible for extracting the data from the 12th to 15th byte of the data packet and converting it to the actual angle value, wherein the conversion formula is that the actual value of the pitch angle is equal to the original value divided by the conversion coefficient 1000, in which the actual value of the pitch angle represents the final physical quantity in degrees, the original value represents the 4-byte unsigned integer extracted from the data packet, and the conversion coefficient 1000 is a dimensionless scaling factor, for example, when the original value is 1234567, the formula calculation process is 1234567 divided by 1000 equals 1234.567, and the actual value of the pitch angle is 1234.567 degrees, and the CRC check function is responsible for verifying the received data packet, and finally a complete and usable target interface code is generated.

[0161] In the embodiments of the present application, the above complete step scheme successfully converts the general code framework into a special interface code that can be run by combining real data with abstract templates and filling in details according to rules, ensuring that the generated code can accurately reflect the protocol specification and handle real data, greatly improving the development efficiency and code reliability.

[0162] In order to comprehensively verify the quality and reliability of the generated interface code, in some embodiments, step 105: based on the target interface code and the example data, a multi-scene test case is generated using a test case generation algorithm, the correctness, stability and performance of the interface are verified by executing the test case, and a test verification result is obtained, including:

[0163] Step 601: According to the interface function implemented by the target interface code, the normal data flow test scene and the abnormal data flow test scene are generated according to the test requirements corresponding to the interface function.

[0164] In step 601, the interface function refers to the specific capabilities such as data transmission, protocol analysis, and message processing implemented by the target interface code, which is derived from the actual role identified after analyzing the generated interface code. The test requirements corresponding to the interface function refer to the test requirements for verifying whether each function of the interface works normally, including data transmission correctness, protocol compliance, and exception handling capability, which are derived from the analysis of the interface function and the requirements of the communication protocol specification. The normal data flow test scene refers to the test environment simulating the data interaction of the device in the standard working state. The abnormal data flow test scene refers to the test environment simulating the data interaction of the device in the non-standard or error state.

[0165] In the embodiments of the present application, first, it is analyzed which interface functions are implemented by the target interface code, then for each interface function, the test requirements it needs to meet are determined, and according to these test requirements, the test scene simulating the normal working condition and the test scene simulating various abnormal conditions are designed respectively.

[0166] Step 602: Based on the example data, an input test data set corresponding to the normal data processing scene and the abnormal data processing scene is generated using a test case generation algorithm.

[0167] In step 602, the input test data set refers to a set of input data generated for a test scene.

[0168] In the embodiments of the present application, based on the existing example data, a set of test data conforming to the normal data specification is generated for the normal data processing scene determined in step 601 using a test case generation algorithm, and a set of test data containing various error conditions is generated for the abnormal data processing scene.

[0169] Step 603: The input test data set and the corresponding test scene configuration information are combined to form a multi-scene test case.

[0170] In step 603, the test scene configuration information refers to setting information defining test environment, test parameters, expected results, and other test execution conditions, and is derived from configuration rules defined in advance for each scene according to test requirements.

[0171] In the embodiments of the present application, the input test data set generated in step 602 is combined with the corresponding test scene configuration information to form a complete test case containing specific test data and test conditions for each test scene.

[0172] In step 604, the interface output result refers to the actual response and data output generated by the target interface code after receiving the input data of the test case.

[0173] In step 604, the interface output result refers to the actual response and data output generated by the target interface code after receiving the input data of the test case.

[0174] In the embodiments of the present application, the test case formed in step 603 is executed in the corresponding test scene strictly according to the message organization format and interactive timing flow specified by the communication protocol, and all output results generated by the target interface code are recorded.

[0175] In step 605, the preset expected behavior refers to the correct output that the interface should generate under specific input, which is defined in advance according to the interface function and protocol specification. Different load conditions refer to various data flow intensity and processing pressure situations simulated during testing, including low load, normal load, peak load, and abnormal load, etc.

[0176] In step 605, the preset expected behavior refers to the correct output that the interface should generate under specific input, which is defined in advance according to the interface function and protocol specification. Different load conditions refer to various data flow intensity and processing pressure situations simulated during testing, including low load, normal load, peak load, and abnormal load, etc.

[0177] In the embodiments of the present application, the interface output result obtained in step 604 is compared in detail with the preset expected behavior, and the correctness of the interface response under various load conditions is checked. At the same time, the stability of the interface is verified through long-time running test, and the performance of the interface is evaluated through measuring response time and other indicators, and finally the complete test verification result is obtained.

[0178] The following is a specific example:

[0179] After the generated target interface code and the existing 100 sets of example data are obtained, first, according to the instruction sending and data parsing functions implemented by the target interface code, normal data flow test scenarios are generated according to the test requirements corresponding to the functions to simulate the correct instruction interaction process, and abnormal data flow test scenarios are generated to simulate the timeout response and data format error conditions. Based on the 100 sets of example data, 85 sets of input test data corresponding to the normal data processing scenarios are generated using the test case generation algorithm, and these data all conform to the protocol specification. At the same time, 15 sets of input test data corresponding to the abnormal data processing scenarios are generated, which contain error check codes or timeout trigger conditions. The 85 sets of normal test data and the normal scene configuration information are combined, and the 15 sets of abnormal test data and the abnormal scene configuration information are combined to form 100 test cases with multiple scenes. According to the organization format that the message must start with 0x55AA and end with a 2-byte CRC check code, and the interactive timing flow that the request response must be completed within 5 milliseconds, all 100 test cases are executed in the corresponding scene to obtain the interface output result. The interface output result is compared with the preset expected behavior, wherein the preset expected behavior is obtained according to the protocol specification, including correctly parsing the pitch angle data in the normal scene and returning an error code in the abnormal scene. Through comparison, it is found that the output result of the target interface code in processing the 73rd abnormal test case does not match the expectation, which simulates the check code error condition, and the code fails to correctly return the error code but performs error data parsing. At the same time, in the stability test after continuous running for 1 hour, it is monitored that the memory usage increases from the initial 45MB to 68MB, and in the performance test, under the load condition of processing 1000 data packets per second, the average response time is measured to be 1.05 milliseconds. These verification results are recorded in the test verification result, wherein the memory usage growth value is 68 minus 45, equal to 23MB, and the average response time is calculated by dividing the total response time by the number of test cases.

[0180] In the embodiments of the present application, the above complete step scheme generates multiple scene test cases automatically and performs comprehensive verification, ensuring multi-dimensional quality evaluation of the interface code function, stability and performance, providing accurate and reliable basis for subsequent code optimization, and improving test efficiency and verification integrity.

[0181] In order to realize self-optimization and continuous improvement of the code generation system, in some embodiments, step 106: the structure complexity analysis and running efficiency evaluation of the target interface code are performed to obtain performance analysis results, and the parameters and training data of the large language model are dynamically adjusted in combination with the test verification results and the performance analysis results, including:

[0182] Step 701: performing structural complexity analysis on the target interface code to generate quantized indicators corresponding to code logic complexity and module dependency, respectively.

[0183] In step 701, code logic complexity is a quantized indicator measuring the complexity of conditional judgment and loop structure in the code. Module dependency refers to the association between different functional modules in the code. The quantized indicator is a measurement result in numerical form.

[0184] In the embodiments of the present application, a special code analysis tool is used to perform structural complexity analysis on the target interface code. The number of conditional branches and the depth of loop nesting in the code are calculated to generate a numerical indicator representing code logic complexity. Meanwhile, a function call relationship graph is analyzed to generate a quantized indicator describing the dependency between modules.

[0185] Step 702: performing running efficiency evaluation on the target interface code in a preset simulation data interaction environment to obtain resource occupation rate and data processing throughput.

[0186] In step 702, the simulation data interaction environment refers to a software test environment constructed by simulating real device communication. The resource occupation rate refers to the proportion of computing resources used by the code during running. The data processing throughput refers to the amount of data that the code can process per unit time.

[0187] In the embodiments of the present application, the target interface code is run in a preset simulation data interaction environment. The processor and memory usage during code running are collected in real time by a monitoring tool to obtain the resource occupation rate. Meanwhile, the number of data packets successfully processed by the code within a specific time period is counted to obtain the data processing throughput.

[0188] Step 703: generating performance analysis results based on the quantized indicators, the resource occupation rate, and the data processing throughput.

[0189] In the embodiments of the present application, the code logic complexity and module dependency quantized indicators generated in step 701 are comprehensively analyzed with the resource occupation rate and data processing throughput obtained in step 702 to generate a performance analysis result that comprehensively reflects the code quality and performance.

[0190] Step 704: determining the types of code defects and performance bottlenecks according to the test verification results and the performance analysis results.

[0191] In step 704, code defects refer to functional errors or logical errors in the code. Performance bottlenecks refer to critical parts or modules that limit the running efficiency of the code.

[0192] In the embodiments of the present application, according to the function errors and abnormal conditions recorded in the test verification results, combined with the resource use efficiency and data processing capacity indicators displayed in the performance analysis results, the specific defect types and performance bottleneck types existing in the code are comprehensively analyzed and determined. The specific implementation process is: analyzing the error reports in the test verification results and the abnormal indicators in the performance analysis results, correlating and comparing the function errors, running crashes and other problems occurring in the test with the high complexity, low throughput and other abnormal indicators identified in the performance analysis, so as to determine the specific code defect type and performance bottleneck type. For example, when the test verification result shows that the data parsing error occurs, and the performance analysis result shows that the parsing function has a high cyclomatic complexity of 20 and a processing throughput lower than the standard value, it is determined that the code defect is a logic error and the performance bottleneck is that the function structure is too complex.

[0193] Step 705: dynamically adjusting the parameters and training data of the large language model according to the code defect and the performance bottleneck type.

[0194] In step 705, dynamic adjustment refers to the process of modifying the model configuration and training content in real time according to the feedback results.

[0195] In the embodiments of the present application, according to the code defect type and performance bottleneck type determined in step 704, the internal parameter settings of the large language model are adjusted in a targeted manner, and the training data content of the model is updated, the proportion of correct code samples is increased, and the code patterns leading to defects are reduced.

[0196] The following is a specific example:

[0197] With the generated target interface code and the existing test verification results, the structural complexity of the target interface code is analyzed first, and the code analysis tool is used to calculate the cyclomatic complexity of 15, which is a quantitative indicator of the complexity of the code logic. The value is calculated by counting the number of conditional judgment nodes and edges in the code. The specific calculation formula is cyclomatic complexity equal to the number of edges minus the number of nodes plus 2, where the number of edges represents the number of edges in the program control flow graph, and the number of nodes represents the number of nodes in the control flow graph. At the same time, the quantitative indicator of module dependency is 0.3, which is calculated by analyzing the ratio of the number of dependent connections between modules in the function call graph to the total possible number of connections. Then the running efficiency of the target interface code is evaluated in the preset simulation data interaction environment. The load condition of processing 1000 data packets per second is simulated, and the performance monitoring tool is used to get the central processing unit occupancy rate of 12%, the memory occupancy rate of 45MB, and the data processing throughput of 950 data packets per second. The throughput value is calculated by dividing the total number of successfully processed data packets 57000 by 60 seconds, i.e. 57000 divided by 60 equals 950. Then based on the quantitative indicators of cyclomatic complexity 15, module dependency 0.3, central processing unit occupancy rate 12%, memory occupancy rate 45MB and data processing throughput 950 data packets per second, the performance analysis result is generated. According to the test verification result recorded in the processing of specific abnormal data packet error and the slow growth of memory usage from 45MB to 68MB, combined with the high cyclomatic complexity and relatively low data processing throughput in the performance analysis result, it is determined that the code defect is abnormal processing logic error, and the performance bottleneck is the low memory usage efficiency caused by the high code structure complexity. Finally, according to the determined code defect type and performance bottleneck type, the parameter configuration of the large language model is dynamically adjusted, the generation weight of complex condition structure is reduced, and the code sample with high memory usage efficiency is added to the training data to optimize the model and improve the accuracy and efficiency of subsequent code generation.

[0198] In the embodiments of the present application, the above complete step scheme accurately identifies the quality problems and efficiency bottlenecks of the generated code through systematic code analysis and performance evaluation, and accordingly realizes the targeted optimization of the large language model, effectively improves the accuracy and efficiency of subsequent code generation, and forms a good self-improvement cycle.

[0199] The following is a specific embodiment:

[0200] 1. Data collection and preprocessing

[0201] Through the data acquisition module, the simulation parts / real parts communication protocols, data formats and example data provided by different manufacturers are collected (such as Figure 2As shown, the data collection module is in communication connection with the external simulation / real device, for obtaining relevant data. The collected data is preprocessed, the noise data is removed by data cleaning algorithm, the protocol type, data transmission format, data interaction rule and other key information in the communication protocol are extracted by data extraction algorithm, and the data structure, data type and other key information in the data format are extracted.

[0202] The data cleaning can be represented by the following formula: wherein, represents the cleaned data set, represents the original data set, represents the noise data set. When extracting data, it is assumed that the communication protocol information set is , the data format information set is , and the sets of extracted key information are and , and the extraction process can be represented as and , wherein and are the communication protocol key information extraction function and the data format key information extraction function, respectively. The key information is converted according to the preset standardized data format to form a data file convenient for DeepSeek large model processing.

[0203] 2. DeepSeek large model module

[0204] A pre-trained DeepSeek large model is introduced, which has strong learning and reasoning ability. The preprocessed standardized data is input into the DeepSeek large model, and the DeepSeek large model performs deep analysis on the communication protocol and data format through deep learning algorithm, and extracts key semantic information by semantic understanding algorithm. Assuming that the input data is X, the code template output after processing by the DeepSeek large model is T, the specific implementation details are D, and the finally generated interface code is C, the process can be represented as: wherein, represents the DeepSeek large model, and G represents a function for generating specific implementation details according to input information I (such as example data and data interaction rules, etc.). Based on the extracted key semantic information, the DeepSeek large model generates the corresponding code template, and then fills in the specific implementation details according to the example data and data interaction rules, etc., to finally form a complete interface code (in Figure 3 , the connection relationship between the DeepSeek large model and the data processing module and the code generation module, as well as the data flow and code generation process, is shown).

[0205] 3. Test case automatic generation and verification

[0206] The test case generation module receives the generated interface code and the sample data. According to the functional requirements of the interface code and the characteristics of the sample data, the test case generation module automatically generates test cases using a test case generation algorithm. Assuming that the set of functional requirements of the interface code is , the sample data is S, and the generated test case set is TC, the generation process can be represented as: , where Gen is the test case generation function. The generated test cases are input into the test execution module, and the test cases are executed. By comparing the output results of the interface code with the expected results of the sample data, the correctness of the interface code is verified; by running the interface code for a long time and monitoring its running state, the stability of the interface code is verified; by measuring the time and other indicators of the interface code processing data, the performance of the interface code is verified. Figure 4 The connection relationship and data interaction process between the test case generation module, the test execution module, and the interface code are shown.

[0207] 4. Code complexity and running efficiency analysis

[0208] The code analysis tool is introduced to analyze the complexity of the generated interface code, and the readability and maintainability of the code are evaluated from the aspects of cyclomatic complexity, code lines, function call depth, etc. At the same time, the performance analysis tool is used to evaluate the running efficiency of the interface code by simulating the actual data interaction scenario, and the response time, throughput, and other indicators of the interface code processing data are monitored to ensure that the interface code meets the real-time requirements. Figure 5 The connection mode of the code analysis tool and the performance analysis tool with the interface code and the flow direction of the analysis data are presented. The cyclomatic complexity calculation can use the McCabe method. Assuming that the number of decision nodes in the code is n, the number of connection edges is e, and the number of regions is r, the cyclomatic complexity V(G) calculation formula is: .

[0209] 5. System optimization and iteration

[0210] According to the test verification results and the performance analysis results, the DeepSeek large model is optimized. If the test verification finds errors in the interface code, the error information and related data are fed back to the DeepSeek large model, and the accuracy of code generation is improved by adjusting the parameters and training data of the model. If the performance analysis result shows that the running efficiency of the interface code does not meet the standard, the code performance bottleneck is analyzed, the code generation strategy of the DeepSeek large model is optimized, and the efficiency of code generation is improved to realize the optimization and iteration of the system. The flow chart is shown in Figure 6 .

[0211] Figure 7A structural schematic diagram of a flight simulator interface code automatic generation system based on a large language model provided by an embodiment of the present application is shown in the following figure, and the specific implementation part describes:

[0212] An acquisition module 71 is configured to acquire interface data provided by a plurality of external devices in a flight target, the interface data including a communication protocol, a data format, and example data.

[0213] A filtering module 72 is configured to perform noise filtering processing on the interface data, extract key information of the communication protocol and key information of the data format from the processed interface data, and convert the key information of the communication protocol and the key information of the data format according to a preset standardized data format to form structured input data.

[0214] An input module 73 is configured to input the structured input data to a pre-trained large language model, perform deep analysis on the structured input data through a deep learning algorithm in the large language model, extract key semantic information from the analysis result by using a semantic understanding algorithm in the large language model, and generate an interface code template based on the key semantic information.

[0215] A generation module 74 is configured to instantiate variables, functions, and communication interfaces in the interface code template based on the example data and data interaction rules to generate a target interface code.

[0216] A verification module 75 is configured to generate test cases in multiple scenarios by using a test case generation algorithm based on the target interface code and the example data, verify the correctness, stability, and performance of the interface by executing the test cases, and obtain a test verification result.

[0217] An evaluation module 76 is configured to perform structural complexity analysis and running efficiency evaluation processing on the target interface code to obtain a performance analysis result, dynamically adjust parameters and training data of the large language model in combination with the test verification result and the performance analysis result, and realize a closed-loop automatic interface code generation process.

[0218] The flight simulator interface code automatic generation system based on the large language model of the embodiment of the present application is used to realize the flight simulator interface code automatic generation method based on the large language model described above, and therefore the specific implementation part of the flight simulator interface code automatic generation system based on the large language model can be seen in the embodiment part of the flight simulator interface code automatic generation method based on the large language model described above. The specific implementation can be referred to the description of the corresponding embodiment part, and will not be repeated here.

[0219] The application further provides an electronic device, comprising a memory for storing a computer program, and a processor for executing the computer program to implement the steps of the method for automatically generating flight simulator interface code based on a large language model.

[0220] The application further provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to implement the steps of the method for automatically generating flight simulator interface code based on a large language model.

[0221] In an example embodiment, the computer-readable storage medium can include, but is not limited to, a U disk, a read-only memory, a random access memory, a mobile hard disk, a magnetic disk or an optical disk, and various media capable of storing a computer program.

[0222] The embodiments of the application further provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in the method for automatically generating flight simulator interface code based on a large language model.

[0223] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in general terms in the above description. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the application.

[0224] The above describes in detail the method, system, device and storage medium for automatically generating flight simulator interface code based on a large language model. The principles and implementation modes of the application are described by applying specific examples. The above description of the examples is only used to help understand the method and core idea of the application. It should be noted that, for those skilled in the art, without departing from the principles of the application, some improvements and modifications can be made to the application, and these improvements and modifications also fall within the protection scope of the application.

Claims

1. A large language model-based flight simulator interface code automatic generation method, characterized in that, The method comprises the following steps: acquiring interface data provided by a plurality of external devices in a flight target, the interface data comprising a communication protocol, a data format and example data; performing noise filtering processing on the interface data, extracting key information of the communication protocol and key information of the data format from the processed interface data, and converting the key information of the communication protocol and the key information of the data format into a structured input data according to a preset standardized data format; inputting the structured input data into a pre-trained large language model, performing deep analysis on the structured input data by a deep learning algorithm in the large language model, extracting key semantic information from the analysis result by a semantic understanding algorithm in the large language model, and generating an interface code template based on the key semantic information; instantiating variables, functions and communication interfaces in the interface code template based on the example data and data interaction rules to generate a target interface code; generating test cases in multiple scenarios by a test case generation algorithm based on the target interface code and the example data, verifying the correctness, stability and performance of the interface by executing the test cases, and obtaining test verification results; performing structural complexity analysis and running efficiency evaluation processing on the target interface code to obtain performance analysis results, and dynamically adjusting parameters and training data of the large language model in combination with the test verification results and the performance analysis results to realize a closed-loop automatic interface code generation process; the step of inputting the structured input data into a pre-trained large language model, performing deep analysis on the structured input data by a deep learning algorithm in the large language model, extracting key semantic information from the analysis result by a semantic understanding algorithm in the large language model, and generating an interface code template based on the key semantic information comprises the following steps: inputting the structured input data into the large language model, jointly encoding protocol definition content and data format description content in the structured input data by an encoding module of the large language model to form a model input sequence; performing multi-level semantic analysis on the model input sequence by a deep learning algorithm provided by the large language model to obtain a deep analysis result; performing semantic relationship reasoning on the deep analysis result by a semantic understanding algorithm provided by the large language model to obtain key semantic information, the key semantic information comprising protocol interaction logic and data mapping relationship; generating an interface code template based on the protocol interaction logic and the data mapping relationship; the step of performing semantic relationship reasoning on the deep analysis result by a semantic understanding algorithm provided by the large language model to obtain key semantic information comprises the following steps: performing dependency relationship analysis on protocol syntax units in the deep analysis result to identify trigger conditions and timing constraint relationships between protocol commands and responses; performing semantic role labeling on data structure definitions in the deep analysis result to analyze functional attributes, value ranges and encoding methods of data fields; According to the trigger condition and the timing constraint relationship, a state transition model of a message sequence is established, and a complete workflow of protocol interaction is derived based on the state transition model; According to the functional attribute, the value range and the encoding mode, a mapping rule and a conversion relationship between data fields are constructed, and a standard template for data parsing and packaging is formed based on the mapping rule and the conversion relationship; Based on the complete workflow and the standard template, key semantic information is generated. 2.The large language model-based flight simulator interface code automatic generation method according to claim 1, wherein, The generation of key semantic information based on the complete workflow and the standard template includes: Converting the complete workflow into a state transition rule set containing state transition conditions and message processing sequences; Converting the standard template into a data operation rule set containing field mapping relationships and data conversion methods; According to the state transition rule set and the data operation rule set, a semantic description file is constructed, which defines the protocol interaction logic and data mapping relationship; The semantic description file is structured and packaged to generate key semantic information. 3.The large language model based flight simulator interface code automatic generation method according to claim 1, wherein, The instantiation processing of variables, functions and communication interfaces in the interface code template based on the example data and data interaction rules to generate target interface code includes: Extracting data samples with actual numerical values and corresponding data type descriptions from the example data; According to the data transmission sequence and verification requirements of the data interaction rules, the specific implementation mode of variables and functions in the interface code template is determined; The data samples, the data type descriptions and the corresponding data structures and communication interfaces in the interface code template are matched and bound; According to the specific implementation mode, the message organization format and the interaction timing flow specified in the communication protocol are combined to fill the bound interface code template to generate the target interface code. 4.The method of claim 1, wherein, Based on the target interface code and the example data, a test case generation algorithm is used to generate multiple-scenario test cases, the correctness, stability and performance of the interface are verified by executing the test cases, and test verification results are obtained, including: According to the interface function implemented by the target interface code, normal data flow test scenarios and abnormal data flow test scenarios are generated according to the test requirements corresponding to the interface function; Based on the example data, a test case generation algorithm is used to generate input test data sets corresponding to normal data processing scenarios and abnormal data processing scenarios respectively; The input test data sets and the corresponding test scenario configuration information are combined to form multiple-scenario test cases; According to the message organization format and the interaction timing flow specified in the communication protocol, the test cases are executed in the corresponding scenarios to obtain interface output results; The interface output results are compared with the preset expected behavior to verify the correctness of the interface under different load conditions, and to verify the stability and performance, and to obtain test verification results. 5.The large language model based flight simulator interface code automatic generation method according to claim 1, wherein, The structural complexity analysis and running efficiency evaluation of the target interface code are performed to obtain performance analysis results, and the parameters and training data of the large language model are dynamically adjusted in combination with the test verification results and the performance analysis results, including: Performing structural complexity analysis on the target interface code to generate quantization indicators corresponding to code logic complexity and module dependency relationship respectively; In a preset simulation data interaction environment, the running efficiency of the target interface code is evaluated to obtain resource occupancy rate and data processing throughput; Based on the quantization indicators, the resource occupancy rate and the data processing throughput, performance analysis results are generated; According to the test verification results and performance analysis results, determine the code defect and performance bottleneck type; According to the code defect and the performance bottleneck type, dynamically adjust the parameters and training data of the large language model.

6. A large language model-based flight simulator interface code automatic generation system, characterized by, Comprise: The acquisition module is used for acquiring interface data provided by a plurality of external devices in a flight target, and the interface data includes communication protocol, data format and example data; The filtering module is used for noise filtering processing on the interface data, extracting key information of the communication protocol and key information of the data format from the processed interface data, and converting the key information of the communication protocol and the key information of the data format into a standardized data format according to a preset standardized data format to form structured input data; The input module is used for inputting the structured input data into a pre-trained large language model, performing deep analysis on the structured input data through a deep learning algorithm in the large language model, extracting key semantic information from the analysis result by using a semantic understanding algorithm in the large language model, and generating an interface code template based on the key semantic information; The generation module is used for instantiating variables, functions and communication interfaces in the interface code template based on the example data and data interaction rules to generate target interface code; The verification module is used for generating test cases in multiple scenarios based on the target interface code and the example data by using a test case generation algorithm, verifying the correctness, stability and performance of the interface by executing the test cases, and obtaining test verification results; The evaluation module is used for performing structural complexity analysis and running efficiency evaluation on the target interface code to obtain performance analysis results, and dynamically adjusting the parameters and training data of the large language model in combination with the test verification results and the performance analysis results, so as to realize a closed-loop automatic interface code generation process. The structured input data is input into a pre-trained large language model, the structured input data is deeply analyzed through a deep learning algorithm in the large language model, key semantic information is extracted from the analysis result by using a semantic understanding algorithm in the large language model, and an interface code template is generated based on the key semantic information, including: The structured input data is input into a large language model, and the protocol definition content and data format description content in the structured input data are jointly coded by a coding module of the large language model to form a model input sequence; The deep learning algorithm provided by the large language model is used to perform multi-level semantic analysis on the model input sequence to obtain a deep analysis result; The semantic understanding algorithm provided by the large language model is used to perform semantic relationship reasoning on the deep analysis result to obtain key semantic information, and the key semantic information includes protocol interaction logic and data mapping relationship; Based on the protocol interaction logic and the data mapping relationship, an interface code template is generated; The semantic understanding algorithm provided by the large language model is used to perform semantic relationship reasoning on the deep analysis result to obtain key semantic information, including: Dependency relationship analysis is performed on the protocol syntax unit in the deep analysis result to identify the trigger condition and timing constraint relationship between the protocol command and response; The data structure definition in the deep analysis result is subjected to semantic role labeling to analyze the functional attribute, value range and encoding method of the data field; According to the trigger condition and the timing constraint relationship, a state transition model of the message sequence is established, and a complete workflow of the protocol interaction is deduced based on the state transition model; According to the functional attribute, the value range and the encoding method, a mapping rule and a conversion relationship between data fields are constructed, and a standard template for data analysis and encapsulation is formed based on the mapping rule and the conversion relationship; Based on the complete workflow and the standard template, key semantic information is generated.

7. An electronic device, comprising: Including: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the flight simulator interface code automatic generation method based on the large language model according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the flight simulator interface code automatic generation method based on the large language model according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Visual intelligent programming method based on large language model

    CN119045806A

  • Internet of Things equipment access protocol component development method, equipment and storage medium

    CN119396365A