Source code level RVV conversion method for RISC-V vector database
Patent Information
- Application Number
- CN202611294650.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-25
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明的主要目的在于提供了一种面向RISC-V向量数据库的源码级RVV转换方法,通过构建专用知识库与微调大语言模型,自动识别向量化障碍并生成通过编译、语义及性能验证的RVV源码,以克服编译器自动向量化能力不足和手工优化成本高昂的缺陷
[0061]与现有技术相比,本发明提供一种面向RISC-V向量数据库的源码级RVV转换方法。首先,本发明构建RVV内置函数知识库与向量化优化先验知识库,并对大语言模型进行领域微调,建立了从标量C函数到高质量RVV源码的自动化转换通道,该技术方案不仅填补了编译器在复杂循环结构下的自动向量化覆盖的盲区,更通过精准匹配的向量化改写策略与底层指令原语,使原本被迫退化为标量执行的热点函数得以稳定生成符合RVV规范的向量化实现,从而在保持原函数接口和计算语义不变的前提下,提升计算密集型函数的执行效率。
Smart Images

Figure CN122837853A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of next-generation information technology industry, and in particular to a source code-level RVV conversion method for RISC-V vector databases. Background Technology
[0002] Vector databases, as data management systems designed for high-dimensional vector storage and approximate nearest neighbor search, have been widely applied in artificial intelligence scenarios such as recommender systems, semantic search, and computer vision. Their retrieval performance directly determines the response speed of downstream tasks. With the rapid development of large language models and generative artificial intelligence technologies, the demand for vectorized representation and retrieval of unstructured data has increased dramatically, placing higher demands on the computational efficiency of database systems.
[0003] RISC-V vector extensions employ a length-independent design philosophy, supporting parallel computation of datasets and providing a hardware foundation for accelerating vector databases on open-source hardware platforms. However, existing compilers have limited automatic vectorization capabilities for RVV. For complex functions involving loops carrying state dependencies, indirect memory access, or conditional early termination, compilers often conservatively forgo vectorization. Furthermore, the internal function hierarchy of RVV encodes multi-dimensional configurations such as SEW, LMUL, masking strategies, and tail element handling, making its complexity far exceed that of traditional SIMD architectures like x86 AVX2. Manually writing RVV code not only involves complex configuration options, but its correctness verification itself constitutes an independent technical challenge, resulting in high optimization costs.
[0004] While existing research has explored complex code generation and vectorization optimization for large language models, current methods are mostly geared towards translating instructions with general loop structures or single structures. They lack systematic modeling of RVV-specific programming knowledge and mechanisms for effectively communicating compiler diagnostic information, loop structure characteristics, and vectorization rewriting strategies. This results in generated code often containing compilation errors, semantic deviations, or performance degradation, making it difficult to directly apply to vector database systems in generation environments. In particular, in practical systems like pgvector, the vectorization rewriting of the distance computation kernel requires not only identifying the parallel main loop but also properly handling boundary conditions such as tail logic, reduction initial values, floating-point precision, and half-precision widening. These details are crucial for semantic correctness and performance stability.
[0005] In summary, to address the aforementioned problems, this invention provides a source-level RVV conversion method for RISC-V vector databases. Summary of the Invention
[0006] The main objective of this invention is to provide a source-level RVV conversion method for RISC-V vector databases. By constructing a dedicated knowledge base and fine-tuning a large language model, it automatically identifies vectorization obstacles and generates RVV source code that passes compilation, semantic and performance verification, thereby overcoming the shortcomings of insufficient automatic vectorization capabilities of compilers and high costs of manual optimization.
[0007] Based on the first main aspect of the present invention, a source-level RVV conversion method for RISC-V vector databases is provided, comprising the following steps:
[0008] An RVV built-in function knowledge base and a vectorization optimization prior knowledge base are constructed. The RVV built-in function knowledge base stores the prototypes, operation types, and instruction identification of RVV built-in functions, which is used to provide candidate primitive range and type constraints. The vectorization optimization prior knowledge base stores the mapping relationship between vectorization obstacle types and optimization strategies, which is used to provide high-level rewriting guidance.
[0009] Based on the code big language model, the RVV-Coder domain fine-tuning model is constructed by using a function-level hybrid supervised dataset for supervised fine-tuning. The function-level hybrid supervised dataset includes function rewriting supervised samples and API constraint supervised samples.
[0010] Receive the original scalar C function to be optimized, compiler diagnostic information, AST static features and semantic constraints, identify the vectorization obstacle type, retrieve the vectorization optimization prior knowledge base according to the vectorization obstacle type to obtain the rewriting strategy, and generate a structured analysis context;
[0011] The structured analysis context and the original scalar C function to be optimized are input into the RVV-Coder domain fine-tuning model to generate candidate RVV source code;
[0012] The candidate RVV source code is compiled, verified, and differentially tested. Performance testing and acceleration results are then performed on RISC-V hardware. If the verification passes, the candidate RVV source code is output. If the verification fails, the failure information is fed back to the structured analysis context, triggering a new round of analysis and optimization until the preset convergence condition is met.
[0013] The above technical solution constructs a complete technical framework for knowledge constraint, model generation, and closed-loop verification. RVV-specific programming knowledge is explicitly stored in the form of dual knowledge bases, and domain fine-tuning enables the large language model to generate RVV code within the constraint boundaries. Finally, a three-level serial verification mechanism ensures the quality of the output code. This technical solution changes the coverage bottleneck of traditional compiler automatic vectorization under complex loop structures, while overcoming the shortcomings of high development cost and difficulty in guaranteeing correctness when manually writing RVV code. It deeply integrates the large language model, domain knowledge, and compiler verification link to form an automated solution for source code-level optimization of RISC-V databases, providing an engineering-deployable conversion path for computationally intensive functions that cannot be automatically vectorized by the compiler.
[0014] As a further preferred embodiment, in the aforementioned method, the RVV built-in function knowledge base storage is organized in a dual-track manner according to operation semantics and type constraints;
[0015] The steps for constructing the vectorized optimization prior knowledge base are as follows:
[0016] Error messages from compilers that fail to automatically vectorize are categorized, and a set of fine-grained vectorization obstacle categories is established. The obstacle categories include at least: dependency and recursion, boundary and control flow, memory access and layout, call alias side effects, type and numerical semantics, and benefit and cost.
[0017] For each type of obstacle, a corresponding optimization strategy is configured. The optimization strategy includes at least stage separation, block acceptance, temporary buffering, loop swapping, predicate conversion, branch lifting, step instruction, layout reconstruction, index conversion, function inlining, restrict declaration, operation separation, and adaptive selection.
[0018] A mapping rule base is established, which maps compiler error keywords to obstacle categories and then from obstacle categories to optimization strategies. This mapping rule base is used to match optimization strategies based on compiler diagnostic information during the analysis phase.
[0019] By constructing a fine-grained mapping rule base from compiler error keywords to obstacle categories and then to optimization strategies, the reasons for compiler automatic vectorization failures are pre-classified and corresponding optimization strategies are configured. This invention enables the precise matching of rewriting methods based on compiler diagnostic information during the analysis phase. It systematizes scattered, experience-based knowledge of vectorization obstacle handling into a searchable and reusable structured rule base, solving the pain point that vectorization obstacle identification relies on human experience and is difficult to automate. It provides clear high-level strategy guidance for subsequent code generation and improves the strategy hit rate of the first round of generation.
[0020] As a further preferred embodiment, in the aforementioned method, the execution steps for constructing the RVV-Coder domain fine-tuning model are as follows:
[0021] Function rewriting supervision samples are constructed. The input of each sample is a scalar C function and task constraints, and the output is the corresponding RVV source code implementation. The function rewriting supervision samples are loop fragments with vectorization potential extracted from open source code libraries. They are then added to the library after being filtered by compilation verification and differential testing.
[0022] Construct API constraint supervision samples, where each sample takes a natural language query or error prototype as input and outputs an RVV built-in function name or function prototype as output.
[0023] The LoRA parameter efficient fine-tuning method is used to train the code large language model. The LoRA weight update is expressed as:
[0024] ;
[0025] in, Indicates LoRA weights; This represents the pre-trained weight matrix; Represents the lower projection matrix of the low-rank decomposition; Represents the upper projection matrix of the low-rank decomposition; Indicates the scaling factor; Indicates rank;
[0026] Freeze the pre-trained weight matrix during training. Only update the lower projection matrix of the low-rank decomposition. The upper projection matrix of the low-rank decomposition .
[0027] By rewriting supervised samples through functions, the model learns the mapping pattern from scalar loops to RVV vectorized skeletons. By constraining samples through APIs, the model's illusion space in RVV built-in function naming and parameter configuration is compressed from the source. This solves the problem of frequent compilation errors and prototype deviation caused by the lack of domain knowledge in general-purpose large language models in RVV-specific programming scenarios, enabling the 7B-level model to reach a stable level that can be applied in engineering in RVV code generation tasks.
[0028] As a further preferred embodiment, in the aforementioned method, the steps for generating the structured analysis context are as follows:
[0029] The original scalar C function to be optimized is subjected to AST parsing and memory access pattern analysis to extract loop structure features and identify memory access pattern features.
[0030] The loop structure features and memory access pattern features are matched with the obstacle types in the vectorized optimization prior knowledge base; if the match is successful, the corresponding optimization strategy is output; if the match fails, a general vectorized evaluation conclusion is generated.
[0031] The loop structure features, memory access pattern features, and optimization strategies are assembled into a structured analysis context.
[0032] In the context of generating structured analysis, the vector obstacle conditions that were originally implicit in the source code are made explicit and structured. By using rule matching, the static features at the abstract syntax tree level are transformed into executable optimization strategies. This overcomes two key problems in the analysis stage: how to identify vectorization obstacles from the source code and how to effectively pass the identification results to the generative model. This enables the subsequent generation stage to generate code based on precise contextual constraints rather than generalized natural language descriptions.
[0033] As a further preferred embodiment, in the aforementioned method, the execution steps for generating candidate RVV source code are as follows:
[0034] Based on the optimization strategy in the context of structured analysis, the vectorized main loop skeleton topology is determined, which includes unit step mode, step memory access mode or index memory access mode.
[0035] Retrieve a set of candidate built-in functions from the RVV built-in function knowledge base that match the topology of the vectorized main loop skeleton. The set of candidate built-in functions includes at least vector loading functions, vector arithmetic functions, and vector reduction functions.
[0036] The SEW parameter is determined based on the data type and vector elements, the LMUL parameter is determined based on the number of available vector registers, the vsetvl instruction is called to dynamically configure the vector length, and a vector loop body is generated.
[0037] For the end of the loop, reuse the same vsetvl configuration and vector operations as the main loop, and dynamically adjust the actual number of elements processed through vsetvl so that the end is uniformly closed by the same vector loop.
[0038] The call to the reduction instruction reduces the intermediate results in the vector register to scalar values, and returns after completing the distance calculation or inner product calculation.
[0039] The above technical solution transforms the complex decision-making processes in RVV programming, such as vector length configuration, mask processing, register grouping, and reduction paths, from manual coding to automatic retrieval and assembly based on structured constraints. This solves the code adaptation difficulties caused by the VLA design of RVV and the error-prone problems caused by the explosion of data types and LMUL combinations, ensuring that the generated code can run correctly on RISC-V processors with different vector widths.
[0040] As a further preferred embodiment, in the aforementioned method, the execution steps of the differential test are as follows:
[0041] A unified test driver is generated for the original scalar C function to be optimized and the candidate RVV source code. The test driver supports random input generation and boundary input generation. The boundary input includes: vector dimension less than VLEN, vector dimension not an integer multiple of VLEN, vector dimension zero, and vector elements containing extreme values.
[0042] Execute the original scalar C function to be optimized and the candidate RVV source code on the same set of test input distributions, and collect and compare the output results;
[0043] For integer functions, each value must be completely consistent; for floating-point functions, both absolute and relative errors must be controlled within preset thresholds.
[0044] If all test cases meet the error conditions, the semantic equivalence verification is considered successful; otherwise, the difference value of the first failed test case set is recorded and fed back to the optimization phase.
[0045] This technical solution establishes a differential testing mechanism for dedicated RVV code verification, focusing on error-prone boundaries in RVV scenarios such as vector dimension non-alignment, half-precision widening, and sparse index out-of-bounds. It solves the problem that computational semantic deviations caused by details such as tail processing, reduction initialization, and type widening in RVV code are difficult to detect by conventional testing.
[0046] As a further preferred embodiment, in the aforementioned method, the performance detection is performed in the following steps:
[0047] The candidate RVV source code that passed the differential test was compiled and executed on the hardware platform. With parallel construction disabled, a unified performance test framework was used to measure the execution time speedup of the candidate RVV source code compared to the original scalar C function. The calculation formula is as follows:
[0048] ;
[0049] in, Indicates the speedup ratio of vectorization; Indicates the execution time of a scalar; Indicates that vectorization optimizes execution time;
[0050] When vectorization speedup When the value is greater than 1, the performance test passes; when the vectorization speedup is greater than 1, the performance test passes. If the value is less than or equal to 1, the performance test fails.
[0051] This technical solution establishes a performance verification closed loop for RVV optimized code. By quantifying the speedup ratio, it ensures that each code transformation brings actual hardware performance benefits rather than merely maintaining semantic correctness. This avoids the risk of optimization versions potentially degrading due to the lack of real hardware performance in existing large language model-assisted code generation methods, and gives the optimization results deterministic engineering value.
[0052] Based on a second key aspect of the present invention, a source-level RVV conversion system for implementing the aforementioned method for a RISC-V vector database is provided, comprising:
[0053] A dual knowledge base construction module is used to build and store the RVV built-in function knowledge base and the vectorized optimization prior knowledge base;
[0054] The model fine-tuning module is used to perform LoRA-supervised fine-tuning based on the Code Big Language Model and a function-level hybrid supervised dataset to generate the RVV-Coder domain fine-tuning model.
[0055] The analysis module receives the original scalar C function to be optimized, extracts AST static features, compiler diagnostic information and loop structure features, identifies vectorization obstacle types, retrieves optimization strategies from the vectorization optimization prior knowledge base based on the vectorization obstacle type, and generates a structured analysis context.
[0056] The optimization module is used to input the structured analysis context and the original scalar C function into the RVV-Coder domain fine-tuning model to generate candidate RVV source code;
[0057] The verification and iteration module is used to perform dual verification of the candidate RVV source code through differential testing and performance testing. If the verification passes, the candidate RVV source code is output. If the verification fails, the failure information is fed back to the structured analysis context, triggering a new round of analysis and optimization until the preset convergence condition is met.
[0058] Based on a third key aspect of the present invention, an electronic device is provided, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0059] The memory stores a computer program, which, when executed by the processor, causes the processor to perform the aforementioned source-level RVV conversion method for RISC-V vector databases.
[0060] Based on the fourth principal aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed, implements the aforementioned source-level RVV conversion method for a RISC-V vector database.
[0061] Compared with existing technologies, this invention provides a source-level RVV conversion method for RISC-V vector databases. First, this invention constructs an RVV built-in function knowledge base and a vectorization optimization prior knowledge base, and performs domain-specific fine-tuning on large language models, establishing an automated conversion channel from scalar C functions to high-quality RVV source code. This technical solution not only fills the blind spot of compilers' automatic vectorization coverage under complex loop structures, but also, through a precise matching vectorization rewriting strategy and underlying instruction primitives, enables hot-spot functions that were originally forced to degenerate into scalar execution to stably generate vectorized implementations conforming to the RVV specification. This improves the execution efficiency of computationally intensive functions while maintaining the original function interface and computational semantics.
[0062] Secondly, this invention takes a closed-loop iterative mechanism of analysis, generation, verification, and optimization as its core. After each generation, it enforces semantic consistency verification and real hardware performance verification checks, and feeds back the reasons for verification failures to the next round of generation constraints. This ensures the compilability, semantic equivalence, and positive performance benefits of the output code, fundamentally solving the prominent problems of frequent compilation errors, semantic deviations, and performance degradation in existing large language model code generation methods in RVV scenarios, and improving the engineering usability and result stability of source code-level automatic vectorization conversion.
[0063] Finally, based on this, the present invention transforms function-level optimization results into system performance improvements for vector databases, particularly applicable to index structures where execution hotspots are highly concentrated in the distance calculation kernel. This source code-level RVV conversion method possesses general optimization capabilities at the function granularity level and can also extend the acceleration benefits of local computation kernels to the complete database retrieval chain, alleviating the contradiction between development efficiency and optimization quality in traditional compiler optimization and manual coding. It provides a reliable technical solution for the efficient deployment and performance tuning of vector databases on the RISC-V platform. Attached Figure Description
[0064] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, obtaining other drawings based on these drawings without creative effort still falls within the scope of the present invention.
[0065] Figure 1 The following is an execution flowchart of a source-level RVV conversion method for a RISC-V vector database according to an embodiment of the present invention;
[0066] Figure 2The diagram illustrates a system framework diagram of a source-level RVV conversion method for a RISC-V vector database according to one embodiment of the present invention. Detailed Implementation
[0067] The preferred embodiments of the present invention will be described in detail below to provide a clearer understanding of the purpose, features, and advantages of the invention. It should be understood that the following embodiments are not intended to limit the scope of the invention, but are merely illustrative of the essential spirit of the technical solution of the invention.
[0068] In the following description, certain specific details are set forth for the purpose of illustrating various disclosed embodiments in order to provide a thorough understanding of the various disclosed embodiments. However, those skilled in the art will recognize that embodiments may be practiced without one or more of these specific details. In other instances, well-known techniques associated with the invention may not have been shown or described in detail to avoid unnecessarily obscuring the description of the embodiments.
[0069] Throughout this specification, references to "an embodiment" or "an embodiment" indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Therefore, the appearance of "in an embodiment" or "an embodiment" in various places throughout the specification does not necessarily refer to the same embodiment. Furthermore, a particular feature, structure, or characteristic may be combined in any manner in one or more embodiments.
[0070] The specific meanings of the technical terms or English abbreviations that may be used in this invention are explained as follows:
[0071] RISC-V: An open-source instruction set architecture designed based on the principles of reduced instruction set computing. It employs a modular design concept and consists of a basic integer instruction set and a series of optional standard extensions. This invention uses RISC-V as the target hardware platform.
[0072] RVV: A standard extension in the RISC-V instruction set architecture specifically designed for parallel computing of data sets. It adopts a vector length-independent design philosophy, which enables the same vectorized code to run on processors with different vector widths without recompiling.
[0073] VLEN: Vector register bit width. The hardware bit width of the vector register in the RVV architecture is determined by the base processor implementation.
[0074] SEW: A parameter configured via the vsetvl directive in RVV, used to specify the position of each element in vector operations.
[0075] LMUL: Vector Register Set Multiple, a parameter configured via vtype in RVV, used to combine multiple physical vector registers into a single logical register set.
[0076] vsetvl: Vector length instruction, the core configuration instruction of RVV, dynamically sets vtype parameters such as SEW and LMUL at runtime, and returns the actual vector length supported by the hardware.
[0077] GCC: GNU compiler suite, one of the mainstream compilation frameworks for the RISC-V platform, uses two layers of intermediate representation and register transfer language to complete loop analysis and vectorization transformation decisions.
[0078] RVV intrinsic: RISC-V vector built-in functions are RVV instruction-level programming interfaces provided by the compiler for developers that can be directly called in C / C++ source code. Each built-in function corresponds directly to one or more RVV assembly instructions and has a fixed function name, parameter type, and return value type.
[0079] RVV-Coder: The RVV-specific code generation model constructed through supervised fine-tuning in this invention.
[0080] Combination Figure 1 As shown, this invention provides a source-level RVV conversion method for RISC-V vector databases, comprising the following steps:
[0081] S100, construct an RVV built-in function knowledge base and a vectorized optimization prior knowledge base. The RVV built-in function knowledge base stores the prototypes, operation types, and instruction identification of RVV built-in functions, which is used to provide candidate primitive range and type constraints. The vectorized optimization prior knowledge base stores the mapping relationship between vectorization obstacle types and optimization strategies, which is used to provide high-level rewriting guidance.
[0082] S200, based on the code big language model, uses a function-level hybrid supervised dataset for supervised fine-tuning to construct the RVV-Coder domain fine-tuning model; the function-level hybrid supervised dataset includes function rewriting supervised samples and API constraint supervised samples.
[0083] S300 receives the original scalar C function to be optimized, compiler diagnostic information, AST static features and semantic constraints, identifies the vectorization obstacle type, retrieves the vectorization optimization prior knowledge base according to the vectorization obstacle type to obtain the rewriting strategy, and generates a structured analysis context.
[0084] S400, input the structured analysis context and the original scalar C function to be optimized into the RVV-Coder domain fine-tuning model to generate candidate RVV source code;
[0085] S500: Perform differential testing on the candidate RVV source code to verify semantic equivalence, and perform performance testing on RISC-V hardware to verify the acceleration results; if the verification passes, output the candidate RVV source code; if the verification fails, feed back the failure information to the structured analysis context to trigger a new round of analysis and optimization until the preset convergence condition is met.
[0086] The core objective of this invention is to achieve semantic equivalence mapping of function-level code. The system receives the original scalar C function, compiler feedback, and static features of the abstract analytic hierarchy process (AST) to generate source code conforming to the RVV specification. If the generated code does not meet the correctness or performance standards, the system will use verification errors and hotspot information for closed-loop feedback to drive iterative model correction until the requirements are met or the iteration limit is reached.
[0087] Combination Figure 2 As shown, this invention proposes a source-code-level RVV conversion method for RISC-V vector databases. This method does not combine knowledge injection, model capabilities, and function-level rewriting into a single step; instead, it first completes knowledge and model construction, and then performs constrained source code conversion for specific functions.
[0088] The entire process is clearly divided into two parts: "method preparation" and "function optimization." The former involves accumulating knowledge and training the model; the latter focuses on the target function and performs two steps: "analysis and optimization." The system eliminates the single-step mode of directly generating RVV code, breaking it down into two independent stages: analysis and optimization. This aims to thoroughly extract the key features that determine the success or failure of vectorization from the source code and force the model's search space to converge to a constrained and verifiable boundary.
[0089] The method preparation phase proceeds along two parallel lines. On one side, a knowledge foundation is built: the RVV built-in function knowledge base provides primitive prototypes and usage boundaries, and the prior knowledge base is optimized to establish a "compilation barrier—rewrite strategy" mapping. On the other side, generation capabilities are refined: the base model is fine-tuned using a hybrid supervised dataset to produce the RVV-Coder. These two preliminary steps converge before the generation task, defining factual boundaries for the model and endowing it with the ability to manage these boundaries to generate code.
[0090] The function optimization phase focuses on "analysis" and "optimization." The analysis phase restrains the impulse to generate functions blindly. The system ingests the original function, compiler diagnostics, AST static features, and semantic constraints to identify vectorization obstacles, loop characteristics, memory access patterns, and parallel conditions, and assembles them into a structured context.
[0091] This triggers a search, extracting high-level rewriting strategies and low-level primitive constraints, significantly reducing the search space generated subsequently.
[0092] The optimization phase is a closed loop of generation and verification. The RVV-Coder digests the structured context and the original source code, producing candidate RVV code. Then, dual verification is initiated: differential testing compares semantic consistency, and hardware testing evaluates acceleration benefits. Any verification failure is converted into error messages and fed back into the context, triggering a new round of analysis and optimization. This closed loop is not simply a matter of refreshing prompts, but rather a dynamic reorganization of constraints based on the specific failure causes.
[0093] The knowledge base, fine-tuning, and two-stage execution are not simply a loose collection of components. They follow a strict temporal logic: infrastructure is built up before tasks are processed. This architecture clarifies the source of the model's capabilities and gives the rewriting process of each function analyzable, constrainable, and verifiable engineering properties.
[0094] Step S100 is described in one of the following feasible embodiments.
[0095] Step S100 includes sub-step S101: constructing the RVV built-in function knowledge base and sub-step S102: constructing the vectorized optimization prior knowledge base.
[0096] Model fine-tuning alone is insufficient to reliably handle the structural obstacles and RVV details of real-world engineering code. While models can mimic the surface syntax of RVV code, they struggle to determine whether loops are suitable for vectorization and how to rewrite them under complex semantic constraints. Even when the direction is correct, models often exhibit illusions regarding intrinsic selection, parameter configuration, and toolchain compatibility. To address this, this invention constructs a dual knowledge base outside the model, respectively managing high-level strategies and low-level primitive constraints.
[0097] The dual knowledge base approach doesn't simply concatenate large blocks of standard text into the input prompt; instead, it breaks it down into two independently invoked modules. The RVV built-in function knowledge base provides candidate built-in functions and available boundaries, addressing "which primitive to use"; the vectorized optimization prior knowledge base identifies obstacles from compilation failure information and matches strategies, addressing "how to rewrite." Both intervene in the execution process in stages: high-level strategies come first, followed by primitive constraints.
[0098] Sub-step S101: Provides a detailed explanation of how to build the RVV built-in function knowledge base.
[0099] The RVV built-in function knowledge base contains approximately 15,000 built-in function entries, providing precise implementation constraints. Instead of storing code templates, the library stores low-level primitives that can be searched and composed. Each entry includes its prototype, operation category, and instruction semantics.
[0100] The prototype directly defines the return type, parameter format, and vector state; the operation category supports quick retrieval by functional dimensions such as loading, arithmetic, and reduction.
[0101] This knowledge base is organized along two tracks: operational semantics and type constraints, and uses a "prototype string + category index + pattern filtering" retrieval method. This design closely resembles the actual representation of header files and compilers, making it easy to use directly for compilation checks and repairs.
[0102] Specifically, this library undertakes three essential responsibilities in this invention: defining the range of legitimate primitive candidates to prevent the model from fabricating instructions out of thin air; providing accurate prototype references to correct deviations in parameters and types; and combining the available symbol tables of real toolchains to ensure that the selected primitives are indeed compileable. In short, this library does not interfere with vectorization decisions, but only provides a feasible code boundary for the given rewrite direction.
[0103] For example, when the upper-level strategy decides to rewrite using the method of "continuously reading and reducing float32 vectors", the knowledge base not only recalls the continuous loading primitives, but also synchronously matches the corresponding multiplication and reduction instructions.
[0104] Therefore, the library does not output complete code, but rather a combination of primitives that are safe and usable for a specific data type and operation target.
[0105] Sub-step S102: Explain the construction of the vectorized optimization prior knowledge base.
[0106] The vectorized optimization prior knowledge base focuses on obstacle identification and strategy induction. It consists of problem mapping rules and strategy description units: the former transforms compilation failure information into abstract obstacle labels; the latter outputs "features, objectives, techniques, boundaries, and key points" for each type of problem. The analysis module can directly call these structured strategy units, rather than relying on general experience guidance.
[0107] This invention categorizes 11 types of fine-grained problems (covering issues such as unsupported data types, cross-iteration dependencies, and complex control flow), as shown in Table 1, and refines them into 7 major categories.
[0108] Table 1. Vectorized Optimization Prior Knowledge Classification System
[0109] The classification logic does not rely on the structural divisions of programming linguistics, but directly anchors to "how obstacles block RVV rewriting and how to resolve them". For example, loops carrying state dependencies and arithmetic recursion both disrupt cross-iteration parallelism, so they are classified into the same category; AoS / SoA layout transformation and aggregation / scattering both belong to the problem of disrupting memory access continuity, so they are also treated together.
[0110] In actual operation, the mapping rules are responsible for capturing error clues in the compiler report. If a loop carrying a dependency or a write-after-read dependency is detected, it is marked as "inter-iteration data dependency"; if an early exit or a switch branch statement is detected, it is classified as "conditional branch is not speculative".
[0111] Subsequently, the strategy description unit outputs the corresponding rewriting approach based on the tags, such as loop splitting, local snapshotting, or alias resolution. The core mechanism of this library is: first, identify the problem's attribution, and then define the strategy boundaries.
[0112] Step S200 will be described in one of the following feasible embodiments.
[0113] The base model chosen is Qwen2.5-Coder-7B, balancing engineering feasibility and experimental rigor. The 7B size keeps the training and inference costs for edge deployment within a reasonable range, ensuring the reproducibility of the research, and the model's native code capabilities are sufficient to support domain-specific fine-tuning. More importantly, the moderate parameter size eliminates the interference of the parameter advantages of giant models.
[0114] The base model has general code understanding and generation capabilities in advance. Based on this, RVV-specific programming knowledge is injected into the model parameters through supervised fine-tuning, transforming it from a general code model into an RVV-specific code generation model.
[0115] To meet the fine-tuning requirements of RVV-Coder, this invention constructs a function-level hybrid supervision dataset containing 8375 samples. Among them, 7336 function rewrite supervision samples are responsible for establishing the mapping relationship from scalar C code to RVV implementation. An additional 1039 API constraint supervision samples are specifically used to anchor the names, parameters, and return types of built-in functions.
[0116] The function rewriting supervision samples are sourced from open-source code libraries such as Project CodeNet and TheAlgorithms. While these two open-source libraries do not directly cover the inverted index, graph structure, or sparse matching scenarios specific to vector databases, they contain a vast amount of loop, branch, reduction, step size access, and boundary handling patterns, sufficient to support the model in mastering RVV syntax and conventional code vectorization skeletons. This invention positions them as a basic capability data source, rather than a substitute for domain-specific knowledge.
[0117] The sample construction follows a progressive logic from selection to encapsulation. This invention avoids the crude approach of directly injecting project source code, instead precisely extracting loop segments with stable vectorization potential from the source code. Subsequently, parameters, boundaries, and return paths are added to these segments, and they are encapsulated into minimal C functions that can be compiled independently, forcing the task to converge to a pure function-level rewrite.
[0118] Finally, based on this minimized function, a corresponding RVV implementation was written, and the input-output pairing was finalized. Data import has extremely stringent requirements. All candidate samples must pass both compilation verification and differential testing; any code that fails to compile or has semantic deviations will be discarded.
[0119] In the differential testing phase, integer functions require zero error in value-by-value comparisons, while floating-point functions must control both absolute and relative errors within preset thresholds. Test cases are mandated to cover edge scenarios such as random values, zero values, boundary lengths, and lengths not multiples of the variable value (VL). Even after passing machine testing, samples undergo manual or automated quality review to confirm the rationality of instruction selection, memory access strategies, tail-end processing, and reduction paths. Only when all three conditions—compileability, semantic correctness, and efficient vectorization—are met simultaneously are code pairs approved for inclusion in the training set.
[0120] The function is rewritten to uniformly structure the supervised samples into a single-round question-and-answer session. The input side loads the task constraints and scalar code, and the output side only returns the complete RVV code.
[0121] API constraint-supervised samples come from the RVV intrinsic specification and GCC-style header prototypes. Each sample input is a natural language query or error prototype. The model is not required to rewrite the entire function, but rather to output the exact name or prototype of the specified built-in function.
[0122] API constraint supervision samples are divided into two categories. One category is name validity and error correction tasks, which are used to correct near-misses and fake names that the model is prone to produce; the other category is prototype accuracy generation tasks, which are used to strengthen the model's memory of parameter types, return types, and vl positions.
[0123] In terms of data distribution strategy, all API constraint supervision samples are allocated to the training set, while the validation and test sets retain only function rewrite supervision samples. This division directly serves the evaluation endpoint of this invention: assessing the model's hard power in function-level source code conversion.
[0124] API-constrained supervised samples play a supporting role at the underlying level. They do not participate in independent benchmarking, but they can significantly improve the compilation pass rate and prototype hit rate of the model in real-world scenarios from the source.
[0125] The fine-tuning sequence employs the LoRA parameter-efficient fine-tuning method, freezing all weights of the base model and updating only the low-rank decomposition matrix.
[0126] LoRA weight updates are represented as follows:
[0127] ;
[0128] in, Indicates LoRA weights; This represents the pre-trained weight matrix; Represents the lower projection matrix of the low-rank decomposition; Represents the upper projection matrix of the low-rank decomposition; Indicates the scaling factor; Indicates rank;
[0129] The training hyperparameters are configured as follows: LoRA rank is set to 64, scaling factor is set to 128, target module covers the query, key, value, and output projection matrix of the attention layer, as well as the gating, up projection, and down projection matrices of the feedforward layer, training accuracy is bf16, learning rate is set to 1e-4 and cosine annealing scheduling is used in conjunction with the prediction strategy, batch size per card is 2, gradient accumulation steps are 8 to achieve an equivalent total batch size of 32, and maximum sequence length is set to 6144 to cover the complete input of the long function body. A total of 4 training rounds are conducted.
[0130] During training, this invention employs a hybrid sample training strategy. Function rewriting supervision samples serve as the main force, forcing the model to skillfully schedule RVV loops, reduction, and tail-end processing logic while maintaining the ironclad rule of semantic consistency.
[0131] API-constrained supervised samples act as a flanking aid, powerfully compressing the illusion space of function naming and parameter ordering in the model. The combined force of these two signals firmly locks the model's generation trajectory within the "compileable and verifiable" RVV solution space.
[0132] Step S300 is described in one of the following feasible embodiments.
[0133] First, the system receives the original scalar C function to be optimized, compiler diagnostic information, AST static features, and semantic constraints as joint inputs, thereby identifying vector obstacle types, loop characteristics, memory access patterns, and parallel conditions.
[0134] In this process, compiler diagnostic information is included in the failure clues output during the vectorization attempt phase, AST static features are extracted through code parsing to identify loop structures, memory access patterns, and data dependencies, and semantic constraints limit the interface signature and parameter passing conventions of the original scalar C function to be optimized.
[0135] After completing the vectorized obstacle type identification, the effectiveness of knowledge constraints during knowledge retrieval and constraint injection depends on the retrieval and injection methods. This invention abandons open-ended text recall and applies differentiated retrieval strategies based on the execution stage and task objectives.
[0136] The analysis phase primarily utilizes a priori knowledge base. Input information is not the full source code, but rather compilation failure reports, AST structural features, and lightweight semantic constraints. The analysis module cleans the failure information, extracts keywords and compiler features, and calculates the most matching obstacle type based on loop hierarchy, memory access methods, etc., thereby extracting high-level rewrite suggestions. The query results are transformed into structured "knowledge guidance" and injected into the context. This process is essentially rule matching and problem classification, emphasizing the controllability of intervention.
[0137] A lightweight approach combining "compile diagnostic rule matching + lightweight source code semantic extraction" is adopted. High-level strategies are distributed from a priori knowledge base, while immutable function signatures, recursive chains, or broadcast semantics in the source code are handled by a supplementary constraint module. Together, they define the high-level boundaries during the optimization period.
[0138] After the above multi-source information is integrated, the output discards natural language summaries and directly generates structured context for downstream consumption.
[0139] The process has entered the optimization and compilation repair phase, shifting to the built-in function knowledge base and toolchain cross-constraints. The query target has changed from "diagnosing obstacles" to "validating primitives." The `__riscv_` keyword in the candidate code is compared. Use symbols and header file symbol tables to check for unknown instructions. If compilation fails, extract prototype information to correct the model output.
[0140] At this stage, the query process moves beyond knowledge completion and becomes code validation under strong constraints. The two types of knowledge ultimately converge in the input prompts, each fulfilling its specific function.
[0141] The prior knowledge base explicitly indicates the rewrite direction, addressing the "thinking" aspect; the built-in knowledge base defines the implementation boundaries, addressing the "implementation" aspect. This layered scheduling transforms domain knowledge into a specification that the model can strictly execute, meeting the stringent accuracy requirements of source code conversion.
[0142] Step S400 is described in one of the following feasible embodiments.
[0143] The RVV-Coder domain fine-tuning model receives the structured context output from the analysis phase and the original scalar C function to be optimized as joint inputs. Under the premise that the knowledge boundary is clear, it performs constrained directed rewriting rather than open generation.
[0144] The model solved the basic representation problem through fine-tuning, enabling it to stably master the basic syntax of RVV, common vectorized skeletons and built-in function call rules, so that the retrieved knowledge can be transformed into stable code generation capabilities.
[0145] Source code transformation is a strictly constrained function-level mapping, rather than open-ended code completion. Given a scalar C function and constraints, the model must output an RVV implementation with a consistent signature and semantic equivalence. The input to the training data integrates the original code and rule hints, and the output corresponds to the accurate RVV code or built-in function declaration.
[0146] During the generation process, the two types of knowledge ultimately converge in the input prompts, each fulfilling its specific function. The prior knowledge base explicitly indicates the rewriting direction, addressing the "how to think" approach; the built-in knowledge base defines the implementation boundaries, addressing the "how to implement" problem. This hierarchical scheduling transforms domain knowledge into specifications that the model can strictly adhere to, meeting the stringent accuracy requirements of source code conversion.
[0147] Specifically, the structured analysis context accepted by the RVV-Coder domain fine-tuning model has pre-defined vectorized obstacle identification and rewriting strategies, which include clear obstacle types, optimization strategies, prohibited operation hints, loop structure characteristics, and memory access mode features.
[0148] The RVV-Coder domain fine-tuning model must adhere to three constraints during its generation process: First, the optimization strategy within the structured analysis context defines the topology of the main loop skeleton. Second, RVV's built-in function knowledge provides a set of legal candidate primitives and their prototype definitions through retrieval, anchoring instruction boundaries. Third, the mapping relationship between compiler error patterns and fixes stored in the prior knowledge base further narrows the search space, preventing deviations in parameter configuration and policy boundaries.
[0149] In terms of specific generation logic, the RVV-Coder domain fine-tuning model first determines the vectorized skeleton of the main loop based on the loop regularity, memory access continuity and reduction semantic features of the analysis context, including unit step mode, step memory access mode or index memory access mode.
[0150] Subsequently, the model determines the SEW parameters based on the data type, determines the LMUL configuration based on the number of available vector registers, and calls the vsetvl instruction to dynamically configure the actual vector processing length during runtime.
[0151] For half-precision floating-point types, the generated path also needs to explicitly insert a widening instruction to indicate the half type to fp32 precision for cumulative calculation, and then write back the result.
[0152] During the construction of the vectorized loop body, the RVV-Coder domain fine-tuning model recalls vector loading instructions, vector arithmetic instructions, and related reduction instructions that match the current data type and operation target from the built-in function knowledge base.
[0153] The RVV-Coder domain fine-tuning model forces the output of a complete function-level RVV implementation rather than a local code snippet. The generated candidate code must maintain the same function signature, parameter passing convention, register context saving and restoration logic, and final write-back location as the original function.
[0154] For the end of the loop, the generated code reuses the same vsetvl dynamic configuration mechanism as the main loop. The number of elements actually processed is adjusted at runtime and unified by the same vector loop, rather than being split into independent scalar fallback branches, in order to avoid semantic deviations or performance degradation caused by improper handling of the end.
[0155] Step S500 is described in one of the following feasible embodiments.
[0156] In this step, the candidate RVV source code generated in the previous step is input, and the output is a deployable RVV implementation that has passed all verifications or a new round of optimization feedback instructions. The entire verification process follows the serial order of semantic verification and performance verification. If the previous level of verification fails, it will not proceed to the subsequent stage to avoid wasting ineffective test resources.
[0157] In semantic verification, a differential testing method is adopted to compare the compiled candidate code with the original scalar baseline input by input. At the same time, a unified test driver is generated for the original function and the candidate RVV source code. The test input covers both random vectors and boundary inputs. Boundary cases include vector dimensions less than VLEN, vector dimensions not being an integer multiple of VLEN, vector dimensions being zero, and vector elements containing extreme values. The focus is on verifying the tail processing logic, the initial reduction value, the numerical consistency of the widening path, and the coefficient index boundaries.
[0158] For integer type functions, each value must be completely consistent. For floating-point type functions, both absolute and relative error tolerances must be met, within their respective thresholds.
[0159] If all test cases pass, semantic equivalence is determined; otherwise, the difference value and context information of the first failed test case are recorded and fed back to the optimization phase.
[0160] Performance verification is performed after semantic verification passes, and is conducted on a real RISC-V hardware platform, avoiding the use of simulation environments such as QEMU, in order to eliminate the systematic bias caused by the simulator's element-by-element simulation overhead of the RVV memory access path.
[0161] Performance verification employs a unified testing framework. With parallel builds disabled, the execution time of the candidate RVV source code is compared to the original scalar baseline, and the speedup is calculated.
[0162] ;
[0163] in, Indicates the speedup ratio of vectorization; Indicates the execution time of a scalar; Indicates that vectorization optimizes execution time;
[0164] When vectorization speedup When the value is greater than 1, the performance test passes; when the vectorization speedup is greater than 1... If the value is less than or equal to 1, the performance test fails. The distribution of hot instructions and the location of performance bottlenecks are recorded. In the next iteration, the LMUL configuration, loop unrolling factor, or memory access organization method will be adjusted.
[0165] When any level of verification fails, the failure message is parsed and transformed into a hard constraint for the next iteration, rather than a simple retry. Level 3 verification outputs the final RVV assembly or replacement file.
[0166] In one feasible embodiment, experiments were conducted to verify the effectiveness of the present invention. The following is a detailed description of the experiments.
[0167] In this embodiment, the evaluation is carried out at three levels to verify the effectiveness of the invention in different scenarios.
[0168] Experimental metrics include compilation success rate, vectorization success rate, semantic consistency rate, first-time success rate, and geometric mean speedup.
[0169] Compilation success rate:
[0170] ;
[0171] in, Indicates the compilation success rate; This indicates the number of samples that passed the compilation. This indicates the number of evaluation samples.
[0172] Vectorization success rate:
[0173] ;
[0174] in, Indicates the success rate of vectorization; This represents the number of samples that simultaneously meet all three conditions: the candidate implementation can be compiled, passes semantic verification, and actually generates a compliant RVV implementation; This indicates the number of evaluation samples.
[0175] Semantic consistency rate:
[0176] ;
[0177] in, Indicates semantic consistency rate; This represents the number of samples that passed semantic verification. This indicates the number of evaluation samples.
[0178] First-time success rate:
[0179] ;
[0180] in, Indicates the first-time success rate; This represents the number of samples generated in the first round that passed the target validation; This indicates the number of evaluation samples.
[0181] Geometric mean speedup:
[0182] ;
[0183] in, Indicates the first Geometric mean speedup of a sample; Indicates the first Baseline execution time for each sample; Indicates the first The optimized execution time for each sample.
[0184] All experiments were conducted on the Milk-V M1 development board.
[0185] In the first experimental stage, 52 failed GCC automatic vectorization test cases from TSVC were selected as test objects to detect coverage under compiler failure scenarios. Experimental data show that the present invention generates 41 compilable implementations from the 52 failed samples, of which 27 meet the vectorization success criteria, achieving a success rate of 51.92% and a geometric speedup of 3.1620×.
[0186] As shown in Table 2, the test set was further divided into five categories according to the obstacle type. The results showed that the success rate of the dependency recursion type reached 76.92%, and the success rate of other structural constraint types reached 61.54%. The above results confirm that the present invention can fix a certain number of compiler automatic vectorization failure cases, and is particularly good at handling dependency separation and complex loop skeleton reorganization.
[0187] Table 2 Results for each category of TSVC52
[0188] In the second experimental level, 53 basic vector calculation functions selected from the pgvector vector calculation module were used to test the correctness and performance of the methods on real source code. These functions cover four data types: Vector, Halfvec, Sparsevec, and Bit.
[0189] Experimental data show that 45 out of the 53 functions generated by the method provided in this invention pass semantic verification, with a compilation success rate of 90.57%, a semantic consistency rate of 84.91%, and a geometric mean speedup of 2.2966×.
[0190] When broken down by data type, Vector, Halfvec, Sparsevec, and Bit have average speedups of 3.30×, 5.58×, 1.50×, and 1.58×, respectively, all of which are superior to the Clang-O3 compiler baseline.
[0191] Of the 45 valid functions, 39 samples achieved speedup exceeding 1.2 × the threshold, proving that as long as the original loop is regular, memory access is continuous, and the reduction pattern is clear, source code-level RVV rewriting can accurately deliver performance benefits.
[0192] In evaluating the model fine-tuning effect, the experiment merged TSVC52 with 53 pgvector basic functions into 105 hold-out samples to evaluate the native improvement of the base model by fine-tuning in a single generation.
[0193] Experimental data shows that the core value of fine-tuning lies in establishing the stability of the model's generation of valid RVV implementations. Compared to the original base Qwen2.5-Coder-7B-Instruct, the compilation success rate of RVV-Coder on 105 full samples jumped from 13.33% to 59.05%, and the success rate of strict RVV increased from 5.71% to 37.14%. Independent evaluations of TSVC52 and pgvector subsets simultaneously confirm this trend.
[0194] This indicates that the gains from fine-tuning are not a random hit under a specific data distribution, but a real improvement in cross-task RVV programming capabilities.
[0195] The ablation experiment set up four control groups based on the complete system. Vectorization analysis was removed, knowledge base was removed, validation-driven iteration was canceled and only single generation was retained, and domain fine-tuning was removed and replaced with the base model Qwen2.5-Coder-7B-Instruct.
[0196] In the TSVC52 test, removing the knowledge base caused the first-time success rate to drop from 36.54% to 25.00%, proving that the knowledge base directly determines the quality of the first generated code. In the actual pgvector function library, the decrease in semantic consistency rate caused by removing vectorization analysis and the knowledge base was relatively small, remaining above 77%.
[0197] This indicates that when dealing with well-structured real-world workloads, pre-analysis and domain knowledge are primarily used to smooth out fluctuations and improve output stability.
[0198] The end-to-end experiment replaced vector_l2_squared_distance with RVV and examined the performance gains in the query and construction phases on both IVF_FLAT and HNSW index structures.
[0199] IVF_FLAT improves speed by 31%, 19%, 8%, and 8% on the GIST, NYTimes, SIFT, and COCO datasets, respectively; the improvement in speed for each HNSW dataset also falls between 5% and 10%.
[0200] During the query phase, IVF_FLAT throughput increased by 1% to 9%, while HNSW query performance remained at the baseline.
[0201] The above comparison clearly defines the utility boundary of source code-level vectorization. When performing computationally intensive tasks where the hotspots are highly convergent to the rules, single-function modifications can truly realize end-to-end benefits. However, once the performance bottleneck generalizes to the entire chain, such as graph deviation and hash maintenance, the marginal utility of local instruction optimization diminishes sharply.
[0202] The technical terms, principles, or means related to the technical solutions of the present invention mentioned in the above embodiments, which are not described in detail above, are all well-known technologies or common practices that are known to those skilled in the art.
[0203] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A source-level RVV conversion method for RISC-V vector databases, characterized in that, Includes the following steps: Construct an RVV built-in function knowledge base and a vectorized optimization prior knowledge base; Based on the code big language model, supervised fine-tuning is performed using a function-level hybrid supervised dataset to construct the RVV-Coder domain fine-tuning model; Receive the original scalar C function to be optimized, compiler diagnostic information, AST static features and semantic constraints, identify the vectorization obstacle type, retrieve the vectorization optimization prior knowledge base according to the vectorization obstacle type to obtain the rewriting strategy, and generate a structured analysis context; The structured analysis context and the original scalar C function to be optimized are input into the RVV-Coder domain fine-tuning model to generate candidate RVV source code; Differential testing is performed on the candidate RVV source code, and performance testing and acceleration results are verified on RISC-V hardware. If the verification passes, the candidate RVV source code is output. If the verification fails, the failure information is fed back to the structured analysis context, triggering a new round of analysis and optimization until the preset convergence condition is met.
2. The source-level RVV conversion method for RISC-V vector databases according to claim 1, characterized in that, The RVV built-in function knowledge base is organized in a dual-track manner based on operation semantics and type constraints. The steps for constructing the vectorized optimization prior knowledge base are as follows: Classify error messages from compiler-initiated vectorization failures and establish a fine-grained set of vectorization obstacle categories; Configure corresponding optimization strategies for each type of obstacle; A mapping rule base is established, which maps compiler error keywords to obstacle categories and then from obstacle categories to optimization strategies. This mapping rule base is used to match optimization strategies based on compiler diagnostic information during the analysis phase.
3. The source-level RVV conversion method for RISC-V vector databases according to claim 1, characterized in that, The execution steps for constructing the RVV-Coder domain fine-tuning model are as follows: Function rewriting supervision samples are constructed. The input of each sample is a scalar C function and task constraints, and the output is the corresponding RVV source code implementation. The function rewriting supervision samples are loop fragments with vectorization potential extracted from open source code libraries. They are then added to the library after being filtered by compilation verification and differential testing. Construct API constraint supervision samples, where each sample takes a natural language query or error prototype as input and outputs an RVV built-in function name or function prototype as output. The LoRA parameter efficient fine-tuning method is used to train the code large language model; Freeze the pre-trained weight matrix during training. Only update the lower projection matrix of the low-rank decomposition. The upper projection matrix of the low-rank decomposition .
4. The source-level RVV conversion method for RISC-V vector databases according to claim 1, characterized in that, The steps for generating the structured analysis context are as follows: The original scalar C function to be optimized is subjected to AST parsing and memory access pattern analysis to extract loop structure features and identify memory access pattern features. The loop structure features and memory access pattern features are matched with the obstacle types in the vectorized optimization prior knowledge base; if the match is successful, the corresponding optimization strategy is output; if the match fails, a general vectorized evaluation conclusion is generated. The loop structure features, memory access pattern features, and optimization strategies are assembled into a structured analysis context.
5. The source-level RVV conversion method for RISC-V vector databases according to claim 1, characterized in that, The execution steps for generating candidate RVV source code are as follows: Based on the optimization strategy in the context of structured analysis, determine the vectorized main loop skeleton topology; Retrieve a set of candidate built-in functions from the RVV built-in function knowledge base that topologically match the vectorized main loop skeleton; The SEW parameter is determined based on the data type and vector elements, the LMUL parameter is determined based on the number of available vector registers, the vsetvl instruction is called to dynamically configure the vector length, and a vector loop body is generated. For the end of the loop, reuse the same vsetvl configuration and vector operations as the main loop, and dynamically adjust the actual number of elements processed through vsetvl so that the end is uniformly closed by the same vector loop. The call to the reduction instruction reduces the intermediate results in the vector register to scalar values, and returns after completing the distance calculation or inner product calculation.
6. The source-level RVV conversion method for RISC-V vector databases according to claim 1, characterized in that, The execution steps of the differential test are as follows: A unified test driver is generated for the original scalar C function to be optimized and the candidate RVV source code. The test driver supports random input generation and boundary input generation. Execute the original scalar C function to be optimized and the candidate RVV source code on the same set of test input distributions, and collect and compare the output results; For integer functions, each value must be completely consistent; for floating-point functions, both absolute and relative errors must be controlled within preset thresholds. If all test cases meet the error conditions, the semantic equivalence verification is considered successful; otherwise, the difference value of the first failed test case set is recorded and fed back to the optimization phase.
7. The source-level RVV conversion method for RISC-V vector databases according to claim 1, characterized in that, The performance testing steps are as follows: The candidate RVV source code that passed the differential test was compiled and executed on the hardware platform. Under the condition of disabling parallel construction, a unified performance test framework was used to measure the execution time speedup of the candidate RVV source code compared with the original scalar C function. When vectorization speedup When the value is greater than 1, the performance test passes; if the vectorization speedup ratio is greater than 1, the performance test passes. If the value is less than or equal to 1, the performance test fails.
8. A source-level RVV conversion system for implementing the method of any one of claims 1-7 for a RISC-V vector database, characterized in that, include: A dual knowledge base construction module is used to build and store the RVV built-in function knowledge base and the vectorized optimization prior knowledge base; The model fine-tuning module is used to perform LoRA-supervised fine-tuning based on the Code Big Language Model and a function-level hybrid supervised dataset to generate the RVV-Coder domain fine-tuning model. The analysis module receives the original scalar C function to be optimized, extracts AST static features, compiler diagnostic information and loop structure features, identifies vectorization obstacle types, retrieves optimization strategies from the vectorization optimization prior knowledge base based on the vectorization obstacle type, and generates a structured analysis context. The optimization module is used to input the structured analysis context and the original scalar C function into the RVV-Coder domain fine-tuning model to generate candidate RVV source code; The verification and iteration module is used to perform dual verification of the candidate RVV source code through differential testing and performance testing. If the verification passes, the candidate RVV source code is output. If the verification fails, the failure information is fed back to the structured analysis context, triggering a new round of analysis and optimization until the preset convergence condition is met.
9. An electronic device, comprising: The processor, communication interface, memory, and communication bus are connected, with the processor, communication interface, and memory communicating with each other via the communication bus. The feature is that the memory stores a computer program, which, when executed by the processor, causes the processor to perform the source-level RVV conversion method for RISC-V vector databases as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed, it implements the source-level RVV conversion method for RISC-V vector databases as described in any one of claims 1 to 7.