A code generation method and system based on RAG and framework-based COT
Through the large-model code generation method based on RAG and framework COT, the problem of low time efficiency and hardware utilization in the development of Verilog language is solved, and the efficiency and low power consumption of FPGA is achieved.
Patent Information
- Application Number
- CN202510453480.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
In the prior art, FPGA-based software radio system (SDR) has low time efficiency and hardware utilization during the development of Verilog language, which cannot meet the needs of FPGA to achieve high efficiency and low power consumption.
The large-model code generation method based on RAG and framework COT is adopted. By inputting the signal processing algorithm to be converted into the framework COT large model, the algorithm framework is built, and optimization is carried out according to user needs to generate FPGA executable optimization code.
It improves the execution efficiency of FPGA in hardware circuits, reduces delay and resource usage, and achieves higher time and space efficiency.
Smart Images

Figure CN119987743B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a code generation method and system based on RAG and framework-based COT. Background Art
[0002] The emergence of software-defined radio (SDR) systems has brought great changes to the field of wireless communication. It has flexibility and adaptability that cannot be achieved by hardware radio systems and is an important aspect of wireless network research. In recent years, due to the dynamic nature, high performance, and reconfigurability of field-programmable gate arrays (FPGAs) being able to well meet the requirements of SDRs, FPGAs have become a commonly used SDR platform. Different from traditional SDRs, FPGAs need to interact more directly with hardware through a hardware description language (HDL) (such as Verilog).
[0003] Currently, some scholars have verified that large language models (LLMs) have the potential to accelerate HDL development, such as the ability to generate 8-bit adder code, code detection and repair, and assisting users in learning HDL to generate simple HDL code for small computing tasks. Since many signal processing algorithms used in SDRs are relatively complex and consume a large amount of resources, although current LLMs can generate basic HDL code for signal processing algorithms, the underlying logic of these codes lacks in-depth knowledge in the professional field, resulting in codes often having insufficient scalability, occupying a large amount of logic resources, and being inefficient in terms of time resources. Therefore, existing LLMs can no longer meet the requirements of FPGAs for achieving high efficiency and low power consumption. Summary of the Invention
[0004] Based on this, embodiments of the present invention provide a code generation method and system based on RAG and framework-based COT, aiming to solve the problem of low FPGA time efficiency and hardware utilization rate in the process of developing Verilog language for FPGA-based SDRs in the prior art.
[0005] The first aspect of the embodiments of the present invention provides a code generation method based on RAG and framework-based COT, the method comprising:
[0006] Input the signal processing algorithm to be converted into a framework-based COT large model, and build an algorithm framework according to the operation process of the framework-based COT large model, wherein the signal processing algorithm to be converted is embodied in the form of matlab or C language;
[0007] Generate basic code according to the algorithm framework;
[0008] According to user requirements, through a retrieval-augmented generation method, focus on a preset part of the generated basic code, match professional terms, and optimize the preset part;
[0009] Output the optimized code, where both the basic code and the optimized code are FPGA-executable codes.
[0010] Furthermore, the steps of building an algorithm framework according to the operation process of the framework-based COT large model include:
[0011] Perform mathematical operations and algorithm analysis in five dimensions on the signal processing algorithm to be converted to obtain an evaluation result, where the five dimensions include complexity, robustness, scalability, scenario requirements, and implementation cost;
[0012] Align the evaluation result with the module functions to be implemented by the FPGA and allocate corresponding hardware resources;
[0013] Divide the signal processing algorithm to be converted into layers according to data dependence and scenario requirements;
[0014] Determine the core steps in the signal processing algorithm to be converted and decompose the core steps into a series of intermediate reasoning steps.
[0015] Furthermore, the complexity is obtained by measuring the time complexity and space complexity of the algorithm;
[0016] The robustness is obtained by judging the anti-interference ability or dynamic channel tracking ability of the signal processing algorithm to be converted;
[0017] The scalability is obtained by various parameters applicable to the signal processing algorithm to be converted, where the parameters at least include the number of antennas, the number of users, and the bandwidth;
[0018] The application scenarios of the signal processing algorithm to be converted at least include satellite communication, the Internet of Things, and 5G communication;
[0019] The implementation cost is obtained by weighing performance improvement and resource consumption and considering hardware cost and time performance.
[0020] Furthermore, in the step of dividing the signal processing algorithm to be converted into layers according to data dependence and scenario requirements, from the perspective of data dependence, if the signal processing algorithm to be converted is decomposed into multiple sequential stages and intermediate results need to be passed between multiple sequential stages, then pipeline operation is selected; if there is no dependence among the elements or tasks within the data block of the signal processing algorithm to be converted, then parallel operation is selected;
[0021] From the perspective of scenario requirements, if the hardware resources are limited and time-division multiplexing of hardware units is required, or if stage delays are acceptable, then pipeline operation is selected; if the hardware resources are sufficient and low latency or real-time performance is required, then parallel operation is selected;
[0022] When a single strategy fails to meet the requirements, pipeline operations and parallel operations are combined.
[0023] Furthermore, the priority of data dependency is higher than the scenario requirements.
[0024] Furthermore, in the step of determining the core steps in the signal processing algorithm to be converted and decomposing the core steps into a series of intermediate inference steps, the core steps in the signal processing algorithm to be converted are determined by means of locating the operations with the greatest impact on performance through mathematical formulas, drawing the data dependency graph of the algorithm, identifying the critical path, and calculating the time complexity and hardware resource requirements of each step.
[0025] The second aspect of the embodiments of the present invention provides a code generation system based on RAG and framework-based COT for implementing the code generation method based on RAG and framework-based COT described in the first aspect. The system includes:
[0026] An input module for inputting the signal processing algorithm to be converted into the framework-based COT large model and building an algorithm framework according to the operation process of the framework-based COT large model, where the signal processing algorithm to be converted is embodied in the form of matlab or C language;
[0027] A generation module for generating basic code according to the algorithm framework;
[0028] An optimization module for focusing on a preset part of the generated basic code and matching professional terms according to user requirements through a retrieval-enhanced generation method to optimize the preset part;
[0029] An output module for outputting the optimized code, where both the basic code and the optimized code are FPGA-executable codes.
[0030] The third aspect of the embodiments of the present invention provides a computer-readable storage medium with a computer program stored thereon, and when the program is executed by a processor, it implements the code generation method based on RAG and framework-based COT provided in the first aspect.
[0031] The fourth aspect of the embodiments of the present invention provides an electronic device including a memory, a processor, and a computer program stored on the memory and running on the processor, and when the processor executes the program, it implements the code generation method based on RAG and framework-based COT provided in the first aspect.
[0032] The beneficial effects of the code generation method and system based on RAG and framework-based COT provided by the present invention are:
[0033] By inputting the signal processing algorithm to be converted into the framework COT large model and building the algorithm framework according to the operation process of the framework COT large model, where the signal processing algorithm to be converted is embodied in the form of Matlab or C language; generating the basic code according to the algorithm framework; according to the user requirements, through the retrieval-augmented generation method, focusing on the preset part of the generated basic code and matching professional terms to optimize the preset part; outputting the optimized code, where both the basic code and the optimized code are FPGA-executable codes. Specifically, in the process of enhancing the potential of HDL development by using LLM, through the application of the framework chain-of-thought prompting technology and retrieval-augmented generation, codes with low latency and high resource utilization are obtained, improving the execution efficiency of FPGA in the hardware circuit from the perspectives of time and space. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 is a flow diagram of the framework COT large model;
[0035] Figure 2 is an implementation flow chart of a code generation method based on RAG and framework COT provided in Embodiment 1 of the present invention;
[0036] Figure 3 is a structural diagram of the FFT basic module;
[0037] Figure 4 is a structural diagram of the FFT optimized by "ping-pong operation";
[0038] Figure 5 is a structural block diagram of a code generation system based on RAG and framework COT provided in Embodiment 3 of the present invention;
[0039] Figure 6 is a structural block diagram of an electronic device. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0041] It should be noted that when an element is referred to as being "fixedly provided on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used in the specification of this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0043] Embodiment 1
[0044] According to an embodiment of the present invention, there is provided an embodiment of a code generation method based on RAG and framework-based COT. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0045] In Embodiment 1 of the present invention, there is provided a code generation method based on RAG and framework-based COT, which can be used in electronic devices, such as computers. For a signal algorithm in the field of wireless communication, it often has characteristics such as directivity, high complexity, and large amount of data. Directivity: Many computing tasks can be divided into multiple subtasks with parallel and priority relationships, and the subtasks can be parallelized only when the underlying priority relationships are not violated. For example, the execution of subtask A depends on the completion of subtask B, then the output of B will determine the input of A. High complexity: Signal processing algorithms involve a variety of mathematical operations, and at the same time, most algorithms also introduce complex numbers and decimals, posing high requirements on the accuracy of calculation results, showing a high degree of complexity of the algorithms. Large amount of data: The processing object of signal processing algorithms is usually large-scale signals, and as the order, time, or dimension increases, the input parameters and the amount of calculation will increase exponentially. If there is a real-time processing requirement, a large amount of data needs to be continuously processed, thereby increasing the calculation burden of the algorithms.
[0046] Due to the characteristics of directivity, high complexity, and large amount of data of signal processing algorithms, the present invention proposes to introduce the framework-based thought chain prompting technology to provide professional knowledge for the large model to generate basic code. The schematic flowchart of the framework-based COT large model is as Figure 1 shown.
[0047] More specifically, please refer to Figure 2 , Figure 2 which shows the implementation flowchart of a code generation method based on RAG and framework-based COT provided in Embodiment 1 of the present invention, specifically including steps S01 to S04.
[0048] Step S01: Input the signal processing algorithm to be converted into the framework-based COT large model, and build an algorithm framework according to the operation process of the framework-based COT large model. Among them, the signal processing algorithm to be converted is embodied in the form of Matlab or C language.
[0049] It should be noted that the signal processing algorithm to be converted is evaluated in two aspects. On the one hand, the mathematical operations involved in the algorithm are analyzed. On the other hand, the algorithm is analyzed from five dimensions to obtain the evaluation results. Among them, the mathematical operations include step operation characteristics and approximation method errors. Specifically, a large number of butterfly operations are used in FFT (Fast Fourier Transform), and multiple multiplication operations in butterfly operations will affect the calculation efficiency of FFT. The MIMO detection algorithm uses series approximation and is more sensitive to approximation errors at high signal-to-noise ratios. The five dimensions include complexity, robustness, scalability, scenario requirements, and implementation cost. Specifically, the complexity is obtained by measuring the time complexity (such as the number of floating-point operations) and space complexity (memory occupancy) of the algorithm. Specifically, the main operations of the algorithm are determined, such as Fourier transform classes (such as FFT), matrix operation classes (such as matrix multiplication, inversion, decomposition), and iterative operation classes (such as gradient descent, Newton's method). The time complexity and space complexity of the algorithm are calculated based on the time complexity and space complexity of the main operations. Taking the complexity of PCA (Principal Component Analysis) as an example, the main operations are covariance matrix calculation and eigen decomposition , for the time complexity, the leading term is , if , it is simplified to , for the space complexity, storing the covariance matrix requires , the data matrix is ;
[0050] The robustness is obtained by judging the anti-interference ability or dynamic channel tracking ability of the signal processing algorithm to be converted. Specifically, the changes in indicators such as bit error rate and signal-to-noise ratio gain are detected in the interference scenario, and the size of the Doppler frequency shift and the length of the channel coherence time are considered in the dynamic scenario to evaluate its dynamic channel tracking ability;
[0051] The scalability is obtained through various parameters applicable to the signal processing algorithm to be converted. Among them, the parameters may include the number of antennas, the number of users, and the bandwidth, etc. Specifically, the parameters are changed to test the changes in algorithm performance and determine the boundary of its scalability;
[0052] The application scenarios of the signal processing algorithm to be converted may include satellite communication, Internet of Things, and 5G communication, etc. Specifically, the application scenarios of 5G and the Internet of Things have different requirements. 5G may focus more on high speed and low latency, while the Internet of Things may focus more on low power consumption and large-scale connection;
[0053] By weighing performance improvement and resource consumption, considering hardware costs and time performance, the implementation cost is obtained. Specifically, use a trade-off curve to find the optimal solution that improves performance without increasing costs or energy consumption;
[0054] Align the evaluation results with the module functions that need to be implemented by the FPGA (Field-Programmable Gate Array), and allocate corresponding hardware resources. Exemplarily, for example, to avoid consuming resources for real-time calculation of rotation factors in FFT, consider using a lookup table (LUT) to pre-store the results; in the Viterbi decoder, to avoid frequent access to external memory, consider using BRAM to store path metrics, etc. Through the alignment operation, the framework-style COT (Chain-of-Thought) large model can obtain the main functions of the algorithm as a whole, generally meet the algorithm performance requirements, set the general framework for the interaction with the FPGA hardware, and accumulate experience for the next step of dividing the algorithm levels;
[0055] Divide the levels of the signal processing algorithm to be converted according to data dependence and scenario requirements. Among them, from the perspective of data dependence, if the signal processing algorithm to be converted is decomposed into multiple sequential stages and intermediate results need to be passed between multiple sequential stages, indicating that the data is dependent, then select pipeline operation; if there is no dependence among the elements or tasks within the data block of the signal processing algorithm to be converted, then select parallel operation;
[0056] From the perspective of scenario requirements, if the hardware resources are limited and time-division multiplexing of hardware units is required, or if stage delays are acceptable, then select pipeline operation; if the hardware resources are sufficient and low latency or real-time performance is required, then select parallel operation;
[0057] When a single strategy cannot meet the requirements, then combine pipeline operation and parallel operation to combine the advantages of both. Exemplarily, such as the MIMO-OFDM system adopts the sub-task parallelization + stage pipelining mode, and the matrix multiplication accelerator adopts the internal pipelining mode of parallel units. Among them, the priority of data dependence is higher than that of scenario requirements, and parallelization can be carried out according to scenario requirements only when the sub-tasks do not violate the underlying priority relationship. This step can obtain the main process and timing requirements of the algorithm implementation. According to user requirements and hardware resources, the framework-style COT large model can flexibly generate code with parallel or pipelined characteristics;
[0058] Identify the core steps in the signal processing algorithm to be converted, and decompose the core steps into a series of intermediate reasoning steps. Specifically, determine the core steps in the signal processing algorithm to be converted by means of locating the operations with the greatest impact on performance (such as bit error rate, capacity) through mathematical formulas, drawing the data dependence graph of the algorithm, identifying the critical path (the path with the greatest impact on latency), calculating the time complexity and hardware resource requirements of each step, etc. The framework-based thought chain prompting technology decomposes the core steps into a series of intermediate reasoning steps to guide the framework-based COT large model to standardize the core steps, so as to improve the flexibility of the code. By formulating a framework for the signal processing algorithm from the overall to the module and then to the details, the framework-based COT large model can produce basic code that reduces storage overhead and has the ability to expand the algorithm scale. On the premise of ensuring the accuracy of the code, it can achieve reasonable resource allocation and improve time efficiency, and can meet the basic needs of users.
[0059] Step S02, generate basic code according to the algorithm framework.
[0060] It can be understood that since the algorithm framework has been built, basic code executable by FPGA can be produced according to this framework.
[0061] Step S03, according to the user's needs, through the retrieval-augmented generation method, focus on a preset part of the generated basic code, match professional terms, and optimize the preset part.
[0062] Exemplarily, in scenarios that require high throughput, low latency, and continuous data processing, ping-pong operations can be introduced to improve the code performance through a double-buffering and synchronization mechanism; in scenarios that require direct control of hardware behavior (such as high-speed interface design), primitives can be used to directly correspond to hardware resources, reducing the logic level and thus reducing latency. RAG (Retrieval-Augmented Generation) not only provides users with more accurate and professional code, but also avoids excessive optimization of the code by the integrated tool due to the hallucination of the large model. While ensuring the functional integrity of the code, it further improves the execution efficiency of FPGA in the hardware circuit. The optimized code output by RAG can be modified again according to the user's needs. After multiple adjustments by the framework-based COT large model, the final output code can be obtained.
[0063] Step S04, output the optimized code, where both the basic code and the optimized code are FPGA-executable codes.
[0064] In summary, for the code generation method based on RAG and framework-based COT in the above embodiments of the present invention, the method inputs the signal processing algorithm to be converted into the framework-based COT large model, and builds an algorithm framework according to the operation process of the framework-based COT large model. Among them, the signal processing algorithm to be converted is embodied in the form of matlab or C language; according to the algorithm framework, basic code is generated; according to user requirements, through the retrieval enhancement generation method, focus on the preset part of the generated basic code, match professional terms, and optimize the preset part; output the optimized code, where both the basic code and the optimized code are FPGA-executable codes. Specifically, in the process of using LLM to enhance the potential of HDL, by applying the framework-based thought chain prompting technology and retrieval enhancement generation, codes with low latency and high resource utilization are obtained, improving the execution efficiency of FPGA in the hardware circuit from the perspectives of time and space.
[0065] Embodiment 2
[0066] To better understand the technical means of a code generation method based on RAG and framework-based COT in Embodiment 1 of the present invention, a specific example of designing an FPGA-executable code for a 64-point FFT is given in Embodiment 2 of the present invention. Specifically, the basic idea of FFT is to use the periodicity, symmetry, particularity of the rotation factor and the interchangeability of the period N to successively decompose the DFT operation of a sequence of length N points into DFT operations of shorter sequences, combine like terms, and greatly reduce the amount of calculation.
[0067] First, using the framework-based COT, it is evaluated that FFT has the characteristics of high time complexity and strong scalability. As Figure 3 shown, it is a schematic structural diagram of the basic module of FFT. To implement FFT in FPGA, the following modules are required: Butterfly operation unit (BPU): including a multiplier and an adder, responsible for performing complex multiplication and addition operations. Each BPU processes two input samples and generates two output samples; dual-port RAM: used to store intermediate results and input / output data; rotation factor unit: responsible for dynamically generating or pre-storing rotation factor values; address transfer module: dynamically generates the correct RAM address according to the current stage and operation phase of FFT; data reordering module: in the final stage of FFT, it is usually necessary to perform bit-reversal sorting on the output data; timing management unit: used to generate and manage the system clock to ensure that each module works synchronously; input / output interface: used to communicate with the external system, receive input data and output the FFT calculation result.
[0068] The design schemes for FFT implementation include sequential processing, cascaded processing, parallel processing, and array processing. Considering the relatively large number of operation points in the 64-point FFT, the parallel processing scheme is adopted. The 64-point FFT needs to be divided into 6 levels of operations, so the 64-point FFT calculation is divided into 6 stages. Each stage uses pipeline registers. At the same time, there are 32 butterfly operations in each level, and 32 BPUs can be designed to work in parallel to complete all butterfly operations at each level simultaneously.
[0069] In the process of FFT implementation, each butterfly operation requires complex multiplication and addition of input data. The rotation factor is a key parameter in complex multiplication. The symmetry and periodicity of the rotation factor can greatly reduce the amount of calculation. Therefore, the calculation method of the normalized rotation factor can directly affect the implementation efficiency of FFT. Taking as an example, the following is the process of gradually converting the normalized rotation factor into a 32-bit binary sequence (with 16-bit imaginary part and 16-bit real part).
[0070] The first step: Represent using Euler's formula , ;
[0071] The second step: Calculate the value of , ;
[0072] The third step: Use 16-bit fixed-point numbers to magnify times,
[0073] Quantize and round the real part, ,
[0074] Quantize and round the imaginary part, ;
[0075] The fourth step: Convert and into binary representations respectively, namely "0111111101100010" and "1111001101110100", and connect the binary representations into a 32-bit binary sequence, using the higher / lower 16 bits as the imaginary / real part.
[0076] Applying the framework-based COT, the LLM improves the execution efficiency of the FPGA in implementing the FFT from the overall algorithm link to the key steps. Under the FFT decomposition strategy, the indexes of the original input sequence are arranged in the natural order, and the single-point indexes after recursive bisection will be arranged in reverse order according to the binary bits, finally forming the bit-reversed order. For the 64-point FFT, the input sequence index values are 0, 1, 2, 3, ..., 62, 63, and the output sequence index values are 0, 32, 16, 48, ..., 31, 63. When the output index value is 1, the corresponding index value of the input sequence should be 4. To correctly calculate and merge the results, the data must be reordered. It is possible to require an improvement in the time efficiency of data reordering. Therefore, retrieval-augmented generation can be used to meet the requirements of maintaining high time efficiency and resource utilization in the data reordering module. Retrieval-augmented generation matches the specific term "ping-pong operation" through requirements. The specific operation is to store the output data in RAM3 at the first clock, in RAM4 at the second clock, in RAM3 at the third clock, in RAM4 at the fourth clock, and so on. The 64-point FFT uses the "ping-pong operation" to store the data in RAM3 and RAM4. The output sequence index values in RAM3 are 0, 16, ..., 31, and the output sequence index values in RAM4 are 32, 48, ..., 63. As Figure 4 shown, it is a schematic diagram of the FFT structure optimized by the "ping-pong operation". By matching professional terms through retrieval-augmented generation, the LLM solves the problem of data continuity in the FFT and improves the time efficiency without consuming a large number of registers.
[0077] Embodiment 3
[0078] Please refer to Figure 5 , Figure 5 which is a structural block diagram of a code generation system based on RAG and framework-based COT provided in Embodiment 3 of the present invention. The code generation system 200 based on RAG and framework-based COT is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used hereinafter, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0079] Specifically, the code generation system 200 based on RAG and framework-based COT includes: an input module 21, a generation module 22, an optimization module 23, and an output module 24, where:
[0080] An input module 21 for inputting a signal processing algorithm to be converted into a framework-based COT large model and building an algorithm framework according to the operation process of the framework-based COT large model, wherein the signal processing algorithm to be converted is embodied in the form of Matlab or C language;
[0081] A generation module 22 for generating basic code according to the algorithm framework;
[0082] An optimization module 23 for focusing on a preset part of the generated basic code and matching professional terms according to user requirements through a retrieval-augmented generation method to optimize the preset part;
[0083] An output module 24 for outputting the optimized code, wherein both the basic code and the optimized code are FPGA-executable codes.
[0084] Further, in some other embodiments of the present invention, the input module 21 includes:
[0085] An evaluation unit for performing mathematical operations and algorithm analysis in five dimensions on the signal processing algorithm to be converted to obtain an evaluation result, wherein the five dimensions include complexity, robustness, scalability, scenario requirements, and implementation cost. Specifically, the complexity is obtained by measuring the time complexity and space complexity of the algorithm;
[0086] The robustness is obtained by judging the anti-interference ability or dynamic channel tracking ability of the signal processing algorithm to be converted;
[0087] The scalability is obtained through various parameters applicable to the signal processing algorithm to be converted, wherein the parameters at least include the number of antennas, the number of users, and the bandwidth;
[0088] The application scenarios of the signal processing algorithm to be converted at least include satellite communication, Internet of Things, and 5G communication;
[0089] The implementation cost is obtained by weighing performance improvement and resource consumption and considering hardware cost and time performance;
[0090] An alignment unit for aligning the evaluation result with the module functions to be implemented by the FPGA and allocating corresponding hardware resources;
[0091] A hierarchical division unit for dividing the signal processing algorithm to be converted into layers according to data dependence and scenario requirements. From the perspective of data dependence, if the signal processing algorithm to be converted is decomposed into multiple sequential stages and intermediate results need to be transmitted between multiple sequential stages, pipeline operation is selected; if there is no dependence between elements or tasks within a data block in the signal processing algorithm to be converted, parallel operation is selected;
[0092] From the perspective of scenario requirements, if the hardware resources are limited, time-division multiplexing of hardware units is required, or if the periodic latency is acceptable, pipeline operation is selected; if the hardware resources are sufficient and low latency or real-time performance is required, parallel operation is selected.
[0093] When a single strategy cannot meet the requirements, pipeline operation and parallel operation are combined.
[0094] In addition, the priority of data dependency is higher than that of scenario requirements.
[0095] A determination unit is configured to determine the core steps in a signal processing algorithm to be converted, and decompose the core steps into a series of intermediate inference steps. Specifically, the core steps in the signal processing algorithm to be converted are determined by means of locating the operation with the greatest impact on performance through mathematical formulas, drawing a data dependency graph of the algorithm, identifying the critical path, and calculating the time complexity and hardware resource requirements of each step.
[0096] Embodiment 4
[0097] Embodiment 4 of the present invention provides an electronic device. Please refer to Figure 6 , which is a structural block diagram of an electronic device, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, the code generation method based on RAG and framework-based COT as described above is implemented.
[0098] Among them, in some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips, and is configured to run the program code stored in the memory 20 or process data, such as executing an access restriction program.
[0099] Among them, the memory 20 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 20 can be an internal storage unit of the electronic device in some embodiments, such as the hard disk of the electronic device. The memory 20 can also be an external storage device of the electronic device in other embodiments, such as a plug-in hard disk equipped on the electronic device, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc. Further, the memory 20 can also include both an internal storage unit and an external storage device of the electronic device. The memory 20 can be used not only to store application software and various types of data of the electronic device, but also to temporarily store the data that has been output or will be output.
[0100] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the code generation method based on RAG and framework-based COT as described above.
[0101] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0102] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or processing in other suitable ways if necessary, and then stored in a computer memory.
[0103] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0104] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0105] The above embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A code generation method based on RAG and framework COT, characterized in that: The method comprises: Input the signal processing algorithm to be converted into the framework COT large model, and build an algorithm framework according to the operation flow of the framework COT large model, wherein the signal processing algorithm to be converted is embodied in the form of matlab or C language; Generate basic code according to the algorithm framework; According to user needs, by searching and enhancing the generation method, focusing on the preset part of the generated basic code, and matching professional terms, the preset part is optimized; Outputting optimized code, wherein the basic code and the optimized code are both FPGA executable codes; According to the operation process of the framework COT large model, the steps of building the algorithm framework include: The signal processing algorithm to be converted is subjected to mathematical operations and algorithm analysis in five dimensions to obtain evaluation results, where the five dimensions include complexity, robustness, scalability, scenario requirements, and implementation cost; Align the evaluation results with the module functions that need to be implemented by the FPGA, and allocate corresponding hardware resources; Divide the signal processing algorithms to be converted into levels according to data dependencies and scenario requirements; The core steps in the signal processing algorithm to be converted are determined and decomposed into a series of intermediate reasoning steps.
2. The code generation method based on RAG and framework COT according to claim 1 is characterized in that: The complexity is obtained by measuring the time complexity and space complexity of the algorithm; The robustness is obtained by judging the anti-interference capability or dynamic channel tracking capability of the signal processing algorithm to be converted; The scalability is obtained by using various parameters applicable to the signal processing algorithm to be converted, wherein the parameters at least include the number of antennas, the number of users, and the bandwidth; The application scenarios of the signal processing algorithm to be converted include at least satellite communication, Internet of Things and 5G communication; The implementation cost is obtained by weighing performance improvement and resource consumption, and considering hardware cost and time performance.
3. The code generation method based on RAG and framework COT according to claim 2 is characterized in that: In the step of dividing the signal processing algorithm to be converted into levels according to data dependency and scenario requirements, from the perspective of data dependency, if the signal processing algorithm to be converted is decomposed into multiple sequential stages, and intermediate results need to be transferred between the multiple sequential stages, then pipeline operation is selected; if the elements or tasks in the data blocks in the signal processing algorithm to be converted have no dependency, then parallel operation is selected; From the perspective of scenario requirements, if hardware resources are limited and hardware units need to be time-division multiplexed, or if periodic delays are acceptable, pipeline operation is selected; If the hardware resources are sufficient and low latency or real-time performance is required, choose parallel operation; When a single strategy cannot meet the needs, pipeline operations and parallel operations are combined.
4. The code generation method based on RAG and framework COT according to claim 3 is characterized in that: Data dependencies take precedence over scenario requirements.
5. The code generation method based on RAG and framework COT according to claim 4 is characterized in that: The method of determining the core steps in the signal processing algorithm to be converted and decomposing the core steps into a series of intermediate reasoning steps locates the operations that have the greatest impact on performance through mathematical formulas, draws the data dependency graph of the algorithm, identifies the critical path, and calculates the time complexity and hardware resource requirements of each step to determine the core steps in the signal processing algorithm to be converted.
6. A code generation system based on RAG and framework COT, characterized in that: For implementing the code generation method based on RAG and framework COT as described in any one of claims 1 to 5, the system comprises: An input module is used to input the signal processing algorithm to be converted into the framework COT large model, and build an algorithm framework according to the operation process of the framework COT large model, wherein the signal processing algorithm to be converted is embodied in the form of matlab or C language; A generation module, used to generate basic code according to the algorithm framework; An optimization module, for optimizing the preset part by searching and enhancing the generation method according to user needs, focusing on the preset part of the generated basic code, matching professional terms, and optimizing the preset part; The output module is used to output the optimized code, wherein the basic code and the optimized code are both FPGA executable codes.
7. A computer-readable storage medium, characterized in that: include: The readable storage medium stores one or more programs, which, when executed by a processor, implement the code generation method based on RAG and framework COT as described in any one of claims 1 to 5.
8. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein: The memory is used to store computer programs; When the processor is used to execute the computer program stored in the memory, the code generation method based on RAG and framework COT as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
LLM-based OpenCL program performance optimization method and system
CN118672590A
Multi-round reasoning and feedback optimization method and system for industrial software codes
CN119440539A