Vectorization processing method, device, equipment, medium and product
By accurately matching the shared library loader implementation class when a vectorized computation request is received, the problem of Gluten's native loading interface being unable to load local libraries and call vectorized functions on domestic operating systems is solved, achieving high efficiency, stability, and compatibility of vectorized computation and improving computational performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-10
Smart Images

Figure CN121833091A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data computing technology, and in particular to a vectorization processing method, apparatus, device, medium, and product. Background Technology
[0002] In the current fields of big data and high-performance computing, vectorized execution engines, as a key technology for improving query and computation efficiency, have been widely integrated into mainstream open-source ecosystems. The Gluten project, as an adaptation layer connecting Apache Spark and high-performance execution engines (such as Velox), supports the use of vectorized functions for mainstream non-domestic operating systems (such as standard Linux distributions and macOS) through an abstract shared library loading mechanism. However, with the deepening of the national information technology application innovation strategy, the deployment of domestic operating systems in critical infrastructure is becoming increasingly widespread, placing higher demands on compatibility and independent controllability.
[0003] In existing technologies, the shared library loading mechanism provided by the Gluten community uses a unified loader interface. On mainstream international operating systems such as Ubuntu and CentOS, it can load native libraries containing vectorized functions and provide them for use by upper-layer Spark tasks via JNI calls. However, some domestic operating systems often employ customized file system layouts, security enhancement mechanisms, and dedicated software installation directories. This means that the Gluten native loading interface cannot support the loading of native libraries and the calling of vectorized functions on domestic operating systems, thus limiting the computational performance of these systems. Summary of the Invention
[0004] This invention provides a vectorization processing method, apparatus, device, medium, and product to enable the loading of the vectorized execution engine functions on which the domestic operating system relies, thereby improving the compatibility and qualitative aspects of vectorized computation in the domestic operating system and enhancing its computational performance.
[0005] According to one aspect of the present invention, a vectorization processing method is provided, which is applied to a heterogeneous operating system hybrid cluster, wherein different operating systems run in the heterogeneous operating system hybrid cluster, and the operating systems include at least two first operating systems; the method includes:
[0006] Upon receiving a vectorized computation request from a first operating system, a shared library loader implementation class corresponding to the first operating system is determined; wherein, the shared library loader implementation class implements a generic shared library loader interface defined by the vectorized execution engine; the generic shared library loader interface is used to implement the function of loading native libraries and calling the vectorized execution engine on at least one second operating system; the second operating system is different from the first operating system;
[0007] Based on the shared library loader implementation class, at least one function of the vectorized execution engine is called;
[0008] Based on the shared library loader implementation class and the called function, data processing is performed according to the vectorized computation request.
[0009] According to another aspect of the present invention, a vectorization processing apparatus is provided, which is deployed in a heterogeneous operating system hybrid cluster, wherein different operating systems run in the heterogeneous operating system hybrid cluster, and the operating systems include at least two first operating systems; the apparatus includes:
[0010] The loader implementation class determination module is used to determine the shared library loader implementation class corresponding to the first operating system when a vectorized computation request is received from the first operating system; wherein, the shared library loader implementation class implements the generic shared library loader interface defined by the vectorized execution engine; the generic shared library loader interface is used to implement the function of loading native libraries and calling the function of the vectorized execution engine on at least one second operating system; the second operating system is different from the first operating system;
[0011] The function determination module is used to call at least one function of the vectorized execution engine based on the shared library loader implementation class;
[0012] The vectorization processing module is used to perform data processing based on the shared library loader implementation class and the called function, according to the vectorization calculation request.
[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0014] At least one processor; and a memory communicatively connected to said at least one processor; wherein,
[0015] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the vectorization processing method described in any embodiment of the present invention.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the vectorized processing method described in any embodiment of the present invention.
[0017] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the vectorization processing method as described in any embodiment of the present invention.
[0018] The technical solution of this invention, upon receiving a vectorized computation request from a first operating system, determines a shared library loader implementation class corresponding to the first operating system. This shared library loader implementation class implements a generic shared library loader interface defined by the vectorized execution engine. The generic shared library loader interface is used to load native libraries and call functions of the vectorized execution engine on at least one second operating system. The second operating system is distinct from the first operating system. Based on the shared library loader implementation class, at least one function of the vectorized execution engine is called. Based on the shared library loader implementation class and the called function, data processing is performed according to the vectorized computation request, thus solving the problem of native Gluten implementations in the prior art. The loading interface was not adapted to domestic operating systems, failing to support the loading of local libraries and the invocation of vectorized functions, which affected the computing performance of domestic operating systems. To address this, a solution was implemented that, upon receiving a vectorized computing request, accurately matches and activates a shared library loader implementation class compatible with the current operating system. This ensures compatibility while efficiently loading local internal libraries, correctly invoking the function calls of the vectorized execution engine, and processing data according to the vectorized computing request. This not only ensures the availability and stability of vectorized computing capabilities in domestic operating system environments but also improves computing efficiency, reduces CPU utilization, disk I / O overhead, and network transmission pressure, thus minimizing hardware resource waste.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a vectorization processing method provided according to an embodiment of the present invention;
[0022] Figure 2 This is a flowchart of a vectorization processing method provided according to an embodiment of the present invention;
[0023] Figure 3 This is a flowchart of the vectorization processing method provided according to an embodiment of the present invention;
[0024] Figure 4This is a schematic diagram of the structure of a vectorization processing device provided according to an embodiment of the present invention;
[0025] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the vectorization processing method of the present invention. Detailed Implementation
[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0027] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution disclosed herein all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to maintain user personal information security and network security. It should also be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution disclosed herein are all conducted with the user's knowledge and consent, and comply with relevant privacy protection regulations.
[0029] Before introducing this technical solution, we can first introduce the application scenario. The technical solution provided in this embodiment can be applied to any scenario where any domestic operating system needs to call the function in the vectorized execution engine to perform vectorized calculation.
[0030] Currently, the shared library loading interface (i.e., the general shared library loader interface) provided by the Gluten community source code for a general operating system (i.e., a non-domestic operating system) has a drawback: the default shared library loading interface does not support shared libraries for domestic operating systems, which means that domestic operating systems cannot use the functions in the vectorized execution engine provided by the Gluten community source code.
[0031] To solve this problem, the technical solution provided in this embodiment can be used to load relevant local internal libraries in various domestic operating system environments, and to call the function in the vectorized execution engine for vectorized calculation.
[0032] Figure 1 This is a flowchart of a vectorization processing method provided by an embodiment of the present invention. This embodiment is applicable to situations where any domestic operating system needs to call function calls in a vectorization execution engine for vectorized computation. This method is applied to heterogeneous operating system hybrid clusters, where different operating systems run, and the operating systems include at least two first operating systems. This method can be executed by a vectorization processing device, which can be implemented in hardware and / or software, and can be configured in a computing device. Figure 1 As shown, the method includes:
[0033] S110. Upon receiving a vectorized computation request from the first operating system, determine the shared library loader implementation class corresponding to the first operating system.
[0034] The term "first operating system" refers to a domestically developed or independently researched operating system. Its kernel can be based on Linux or other independent architectures, and it is applied in various fields such as government affairs, finance, and energy within the context of information technology innovation. For example, a first operating system might be developed based on the open-source operating system Linux / Unix, and it runs native internal libraries. These libraries exist in a format agreed upon by the first operating system, containing low-level functionalities written in languages such as C / C++, which can be called by programming languages through interfaces. The native internal libraries differ between different first operating systems.
[0035] Vectorized computation requests refer to programs or code that require vectorized computation using a native vectorized execution engine (such as Velox). Vectorized computation requests can be data processing tasks (such as SQL queries, filtering, projection, aggregation, and other SQL operator tasks) initiated by big data computing frameworks (such as Apache Spark) and requiring acceleration by a vectorized execution engine; or they can be operation instructions issued by applications that utilize the CPU's SIMD (Single Instruction Multiple Data) instruction set to perform parallel vectorized processing of multiple data elements, including information such as input data, expected computation type (such as dot product, normalization, matrix operations, SIMD acceleration routines, etc.), and output buffers. For example, in AI applications such as intelligent customer service, image recognition, or speech-to-text, a vectorized computation request could be: "Perform fully connected layer matrix multiplication (MatMul) on a batch of 1024 speech feature vectors (each 512-dimensional float32), with a weight matrix of 512×256." This request would trigger the vectorized execution engine to call the highly optimized sgemm (single-precision general matrix multiplication) function (i.e., the function function), utilizing the CPU's AVX-512 or ARM SVE instructions for parallel computation. In scenarios such as financial risk control and risk behavior analysis, a vectorized computation request could be: "Perform vectorized summation on the 'amount' column (float64) of 100 million transaction records and filter out records with amounts greater than 10,000." This request would apply vectorized addition (vector_add) and comparison (vector_compare_gt) operations (i.e., the function function), avoiding line-by-line interpretation and execution. In e-commerce or content platforms, a vectorized computation request can be: "Perform vectorized multiplication (multiply by the original price column) on the 'discount rate' column (float32) of the 'order table' to generate the 'actual payment amount' column." This request will use the vector_fma (a function that combines multiplication and addition) instruction to process 8 to 16 floating-point numbers at a time.
[0036] A shared library loader implementation class refers to a class in a specific programming language (such as C++ or Java) tailored to the first operating system. Internally, it encapsulates the paths, dependency order, and loading logic of native internal libraries (such as .so or .dll files) under that first operating system. The shared library loader implementation class implements the generic shared library loader interface defined by the vectorized execution engine. The vectorized execution engine is a runtime module used for efficient execution of vectorized computations (such as SIMD instruction parallel processing), relying on optimized underlying native internal libraries to achieve high-performance computing. The generic shared library loader interface is a unified abstract interface (such as a pure virtual class in C++) defined by the vectorized execution engine, specifying standard methods for loading native shared libraries, obtaining function pointers, and releasing resources on the second operating system. The generic shared library loader interface is used to implement loading native libraries (i.e., native internal libraries) and calling the function calls of the vectorized execution engine on at least one second operating system. The second operating system is distinct from the first operating system. The second operating system refers to a non-domestic operating system. The distribution of the vectorized execution engine (such as Spark Gluten) does not include an implementation class for the first operating system. This means the generic shared library loader interface cannot load native libraries and invoke the vectorized execution engine's functionalities on the first operating system. It should be noted that the original community version of the vectorized execution engine (i.e., the distribution Gluten) did not provide a shared library loader implementation class adapted to the first operating system, causing the generic shared library loader interface to fail to complete effective native library loading and invocation in the first operating system environment.
[0037] In this embodiment, upon receiving a vectorized computation request, the operating system identifier carried in the request can be parsed to determine whether the current operating system is the first operating system. If it is determined to be the first operating system, the shared library loader implementation class pre-selected for that first operating system is invoked. This shared library loader implementation class encapsulates the loading logic for the internal library path policy, permission model, or security mechanism (such as national cryptographic algorithm verification, trusted execution environment adaptation) of the first operating system, and follows the contract of the general shared library loader interface to ensure that the vectorized execution engine can call the function without being aware of the underlying differences.
[0038] Multiple shared library loader implementation classes can be pre-defined, each corresponding to a different operating system type (including several primary operating systems). During the deployment phase, the name of the loader implementation class to be used is specified in the configuration file according to the target runtime environment of the primary operating system. When a vectorized computation request is received, the configuration file is read, and the shared library loader implementation class corresponding to that primary operating system is dynamically reflected or registered. This approach, by externalizing the implementation class of the primary operating system, facilitates support for newly emerging primary operating system versions without modifying the code, while maintaining interface consistency with loader implementations under non-domestic systems.
[0039] In this embodiment, determining the shared library loader implementation class that matches the first operating system includes: selecting a matching first loading strategy from the preset local library loading strategies of different operating systems based on the system identifier of the first operating system; and retrieving the shared library loader implementation class that matches the first operating system based on the system identifier according to the first loading strategy.
[0040] The system identifier is used to characterize the uniqueness of the operating system; it can be a field used to characterize the uniqueness of the operating system. The native library loading strategy refers to the algorithm or program logic used to load the operating system's native internal libraries. The first loading strategy refers to the native library loading strategy selected during the matching process that corresponds to the current first operating system.
[0041] In one implementation, a mapping table can be loaded during the initialization phase. This table records the mapping relationship between the system identifier of each first operating system and its corresponding native library loading strategy in key-value pairs. When a vectorized computation request is received, the system identifier of the current first operating system can be read, and a precise match can be performed in the mapping table to determine the first loading strategy. This first loading strategy directly contains or points to the shared library loader implementation class corresponding to the current first operating system. Based on this shared library loader implementation class, the native library loading of the current first operating system and the function calls of the vectorized execution engine are implemented.
[0042] In another implementation, a configuration file (such as in JSON or properties format) can be pre-maintained. This file declares the system identifiers of all supported first operating systems and their associated native library loading strategy metadata, including the fully qualified class name of the shared library loader implementation class. When a vectorized computation request is received, the matching first loading strategy can be retrieved from the configuration file based on the current system identifier, and the corresponding shared library loader implementation class can be dynamically created using reflection. This approach extends support for new first operating systems without requiring code recompilation, improving deployment flexibility and ease of maintenance.
[0043] The technical solution provided in this embodiment accurately selects a matching first loading strategy from a preset set of local library loading strategies based on the system identifier of the first operating system, and calls the corresponding shared library loader implementation class accordingly. This effectively solves the problem of local library loading for vectorized execution engines in heterogeneous domestic environments, ensuring that functional functions can be correctly loaded and used according to the characteristics of the first operating system. This not only guarantees the stability and compatibility of vectorized acceleration functions in mixed clusters of multiple types of first operating systems, avoiding the failure of vectorized computing tasks due to library path or dependency mismatch; at the same time, by making the local library loading strategy of the first operating system configurable, it ensures the efficient access of newly added first operating systems.
[0044] In this embodiment, the local library loading logic of each first operating system can also be encapsulated into a local library loading strategy based on the policy pattern.
[0045] The strategy pattern encapsulates native library loading strategies as independent objects, allowing them to be dynamically selected and replaced at runtime based on the context. Native library loading logic refers to the complete operational process required to dynamically load local internal libraries (such as .so files) in the first operating system environment, including determining the library path, handling permissions, calling system loading interfaces (such as dlopen), and resolving symbol addresses. A native library loading strategy can be understood as a concrete strategy implementation class within the strategy pattern, which encapsulates the complete native library loading logic specific to the first operating system.
[0046] In one specific implementation, the primary operating system used by the current runtime environment can be explicitly specified through a configuration file. The policy manager reads this configuration file and selects the local library loading logic that matches the primary operating system from a pre-registered set of local library loading logics. The policy pattern then instantiates the local library loading logic into a local library loading policy. Alternatively, during the initialization phase, the primary operating system can be determined by reading the operating system identifier. The policy pattern then instantiates the local library loading logic of the identified primary operating system into the corresponding local library loading policy.
[0047] In another specific implementation, the native library loading strategy for each first operating system can be compiled into an independent plugin module (such as a JAR package). At runtime, the main program identifies the available strategy plugin modules through the loading strategy naming convention or metadata description, and automatically loads the corresponding native library loading strategy according to the current first operating system, which is then used as the first loading strategy.
[0048] The above approach encapsulates the local library loading logic of each first operating system into an independent local library loading strategy through the strategy pattern. This achieves the adaptive dynamic switching loading strategy of Gluten's vectorized computing tasks under multiple hybrid operating systems. It not only achieves high cohesion and low coupling of loading behavior, but also ensures that the vectorized execution engine obtains a consistent calling experience and reliable library loading guarantee on different first operating systems.
[0049] S120. Based on the shared library loader implementation class, call at least one function of the vectorized execution engine.
[0050] In this context, a function refers to a native method provided by the vectorized execution engine for invocation by the upper Java layer. For example, a function might be a function for initializing the execution context, registering expressions, performing vectorized computations, or releasing resources, and can be invoked through the JNI (Java Native Interface) mechanism.
[0051] In this embodiment, the vectorized computation request can be parsed to determine which function calls (such as "vector dot product" or "batch normalization") are needed. Furthermore, the shared library loader implementation class retrieves at least one function call from the vectorized execution engine. This allows for the coordinated implementation of vectorized computation based on the function calls and the local internal libraries loaded by the shared library loader implementation class, with the libraries being either released or retained for later reuse after execution. Alternatively, the shared library loader implementation class can simply ensure that the local internal libraries are correctly loaded into the JVM runtime without actively calling any function calls. Only when an operator (such as filtering or projection) that can be unloaded into the vectorized execution engine first appears in the Spark execution plan does the relevant execution module call the vectorized execution engine's function call via JNI. Since the local internal libraries have already been pre-loaded by the loader, the function calls can be parsed and executed successfully. This approach avoids unnecessary initialization overhead and improves system startup efficiency.
[0052] For example, after the shared library loader implementation class completes the registration of all native library paths and submits the loading operation, it can automatically trigger calls to the vectorized execution engine's function calls, such as initVeloxContext() or registerNativeExpressions(). This provides ready functions for subsequent vectorized computation.
[0053] In this embodiment, calling at least one function of the vectorized execution engine based on the shared library loader implementation class includes: calling at least one function of the vectorized execution engine from a pre-packaged local program package based on the shared library loader implementation class.
[0054] In this context, pre-packaged native packages refer to binary collections that embed various functionalities that the primary operating system might use into JAR resource files. Vectorized execution engines are written native computing components (such as Velox) that provide columnar processing, SIMD acceleration, and efficient expression evaluation capabilities. Functionalities refer to the interface methods exposed by the vectorized execution engine to the Java layer through JNI.
[0055] Specifically, at least one function of the vectorized execution engine can be pre-obtained and packaged into a local application package. The shared library loader implementation class implements the generic shared library loader interface defined by the vectorized execution engine. At this point, the function of the vectorized execution engine can be called from the local application package corresponding to the current first operating system, based on the shared library loader implementation class.
[0056] For example, a shared library loader implementation class can trigger the loading of a native package containing functional functions. Once the shared library loader implementation class successfully loads the corresponding native library by calling the native library loading method, the JVM can register the JNI-compliant functional functions from the native package into the runtime environment of the current operating system, thereby establishing a static binding relationship with the pre-declared functional functions of the vectorized execution engine. Subsequently, when processing vectorized computation tasks, the upper-layer business logic only needs to directly call these declared functions to transparently execute the high-performance functional functions implemented in the underlying C++, thus acquiring and using the capabilities of the vectorized execution engine.
[0057] The advantage of this setup is that the vectorized execution engine's functions can be called from a pre-packaged local program package by implementing classes based on the shared library loader. This enables the reliable application of vectorized computing capabilities in the first operating system environment, allowing big data analysis tasks to achieve vectorized computing performance in a completely autonomous and controllable environment.
[0058] S130. Implement classes and callable functions based on the shared library loader, and process data according to vectorized computation requests.
[0059] In this embodiment, a shared library loader can be used to load local internal libraries and cache corresponding function calls. Vectorized computation requests are sent to the function calls, which then process the data based on the local internal libraries. Alternatively, vectorized computation requests can be placed in a task queue, where worker threads retrieve them and use the shared library loader's implementation class and its loaded function calls to process the requests and obtain the results. Multiple vectorized computation requests can also be executed concurrently in different threads, efficiently sharing local internal libraries (such as memory pools and expression compilation results), reducing redundant overhead, and improving throughput.
[0060] For example, upon receiving a vectorized computation request, the shared library loader implementation class ensures that the local internal library of the corresponding first operating system has been successfully loaded; it calls the function of the vectorized execution engine, passes the batch of data to be processed (organized in columnar format) to the function, completes the SIMD accelerated computation, and returns the result synchronously.
[0061] In this embodiment, data processing is performed based on the shared library loader implementation class and the called function, according to the vectorized computation request, including: determining and loading the local internal library of the first operating system based on the shared library loader implementation class according to the first loading strategy; and processing the vectorized computation request based on the local internal library and the called function.
[0062] The local internal library refers to the native libraries (such as .so files) compiled and optimized for the first operating system and its supporting hardware platform. The local internal library of the first operating system may include, but is not limited to, high-performance basic computing functions optimized for hardware architecture, such as matrix multiplication, convolution operation, vector addition, vector dot product, batch normalization, activation functions (such as ReLU, Sigmoid), tensor transpose, data type conversion, memory alignment operations, SIMD instruction encapsulation, mathematical library functions (such as trigonometric functions, exponential and logarithmic functions), encryption auxiliary functions, thread scheduling interface, memory allocation and release tools, error code definitions, version information metadata, and adaptation modules integrated with operating system security mechanisms (such as national cryptographic algorithms and trusted execution environments).
[0063] Based on the system identifier of the first operating system, and by selecting a matching first loading strategy from the preset native library loading strategies of different operating systems, the required native internal libraries for the current first operating system can be located from the shared library loader implementation class based on the first loading strategy. After the native internal libraries are loaded, the function uses the native internal libraries to process the vectorized computation request.
[0064] It should be noted that after the local internal library is loaded, all subsequent vectorized calculation requests can bypass the library loading process and directly call the basic functions in the local internal library for data processing of the functional functions.
[0065] The technical solution provided in this embodiment guides the shared library loader through a first loading strategy to achieve precise loading of local internal libraries that are highly compatible with the current first operating system. Combined with the called function, it performs efficient data processing on vectorized computation requests. This not only solves the problem of traditional vectorization engines failing to run in a domestic environment due to missing or incompatible libraries, but also improves the execution efficiency, stability, and security of vectorized computation in domestic software and hardware systems.
[0066] In this embodiment, the process of determining and loading the local internal libraries of the first operating system based on the shared library loader implementation class according to the first loading strategy includes: determining the local library path compatible with the first operating system based on the library loading method of the shared library loader implementation class according to the first loading strategy; and loading the local internal libraries under the local library path.
[0067] The library loading method is a method in the shared library loader implementation class used to register and prepare the operating system's local internal library path, and dynamically load the local internal library at runtime. For example, the library loading method may include steps such as opening the library file, resolving symbols, and establishing function mappings. The local library path refers to the specific directory path in the first operating system where the local internal libraries are stored. In the first operating system, this local library path may differ from the standard Linux path (such as / usr / lib) due to security policies, package management mechanisms, or distribution specifications; for example, it may be located in a proprietary directory path such as / opt / xxx / lib or / usr / local / xxx / lib.
[0068] Specifically, based on the first loading strategy, the shared library loader implementation class can be parsed to determine the library loading methods carried within it. From these methods, native library paths compatible with the first operating system can be extracted. Furthermore, the native internal libraries of the first operating system can be loaded from these local library paths.
[0069] The advantage of this setup is that the first loading strategy guides the shared library loader implementation class to accurately determine the local library path compatible with the first operating system in its library loading method, and successfully load the local internal library under that path. This ensures the reliable location and loading of the local internal library, thereby ensuring the correctness and stability of the function execution in the subsequent vectorized execution engine.
[0070] Based on the above technical solutions, we can also define shared library loader implementation classes for each first operating system based on the general shared library loader interface provided by the vectorized execution engine; and register the local library path of the first operating system in the library loading method of the shared library loader implementation class.
[0071] In one implementation, a corresponding shared library loader implementation class can be defined for each supported first operating system, and each shared library loader implementation class implements a generic shared library loader interface. Within the library loading method of each shared library loader implementation class, the installation path of the vectorized native library under that first operating system, i.e., the native library path, is specified through hard coding or predefined constants. For example, the native library path ` / usr / lib / u / vectorlib` is registered in the loading method of the shared library loader implementation class corresponding to the first operating system U; and the native library path ` / opt / k / vector / lib` is registered in the loading method of the shared library loader implementation class corresponding to the first operating system k.
[0072] In another implementation, all shared library loader implementations of the first operating system can share a single loading logic framework. However, during initialization or loading, the local library path corresponding to the current first operating system is read from a configuration file (such as JSON or INI format). The configuration file can be pre-filled according to the specific local library path of the first operating system. The read local library path of the current first operating system is then loaded into the library loading method of the shared library loader implementation class.
[0073] In another implementation, the shared library loader implementation class of each first operating system does not hardcode the local library path in its library loading method. Instead, it dynamically derives the expected installation path of the local internal libraries by analyzing the system identifier of the current first operating system. For example, if the system identifier is detected as 'k', the path " / usr / lib / k / " is automatically combined; if it is identified as 'U', the path " / opt / u / vector / " is combined. The advantage of this approach is that it embeds the path decision logic within the shared library loader implementation class, preserving the binding relationship between the implementation class and the operating system, and enhancing the efficiency of determining the local library path for different versions of the same operating system.
[0074] For example, the shared library loading implementation interface (i.e., shared library loader implementation class) of various domestic operating systems can be customized and extended. Meanwhile, in the shared library loading implementation interface implemented by each domestic operating system, the loadLib method (i.e., the library loading method) sequentially loads and creates links (i.e., local library paths) of multiple local internal libraries through JniLibLoader (JNI library loader or Java Native Interface Library loader), and finally calls the commit method to complete the submission of these library loading methods, which is used to load the relevant local internal libraries in each domestic operating environment.
[0075] The advantage of the above approach is that by explicitly defining a dedicated shared library loader implementation class for each first operating system and accurately registering the local library path adapted to that first operating system in its library loading method, the local internal libraries of the corresponding first operating system can be loaded accurately and reliably. This ensures that at least one function of the loaded vectorized execution engine can be correctly applied. This not only ensures that the native components required by the vectorized execution engine can be loaded correctly according to the operating system characteristics, but also avoids loading failures caused by incorrect paths, missing symbols, dependency conflicts, or incompatible permissions. This improves the deployment efficiency and operational stability of vectorized functions in a domestic environment.
[0076] The technical solution provided in this embodiment, upon receiving a vectorized computation request from a first operating system, determines a shared library loader implementation class corresponding to the first operating system. This shared library loader implementation class implements a generic shared library loader interface defined by the vectorized execution engine. The generic shared library loader interface is used to load native libraries and call functions of the vectorized execution engine on at least one second operating system. The second operating system is distinct from the first operating system. Based on the shared library loader implementation class, at least one function of the vectorized execution engine is called. Based on the shared library loader implementation class and the called function, data processing is performed according to the vectorized computation request, thus solving the problem of Gluten's original... The original loading interface was not adapted to domestic operating systems, failing to support the loading of local libraries and the invocation of vectorized functions, which affected the computing performance of domestic operating systems. To address this, a solution was implemented that, upon receiving a vectorized computing request, accurately matches and activates a shared library loader implementation class compatible with the current operating system. This ensures compatibility while efficiently loading local internal libraries, correctly invoking the function calls of the vectorized execution engine, and processing data according to the vectorized computing request. This not only ensures the availability and stability of vectorized computing capabilities in domestic operating system environments but also improves computing efficiency, reduces CPU utilization, disk I / O overhead, and network transmission pressure, thus minimizing hardware resource waste.
[0077] Figure 2This is a flowchart of a vectorization processing method provided by an embodiment of the present invention. Based on the foregoing embodiments, different operating systems may include a host operating system and at least one working operating system. Correspondingly, when a distributed computing task is received, at least one vectorized computing subtask in the distributed computing task can be distributed to at least one working operating system based on the host operating system. For each working operating system, when receiving a vectorized computing subtask, it selects a matching second loading strategy from the local library loading strategies of different operating systems based on its system identifier. Based on the second loading strategy, it calls the shared library loader implementation class corresponding to the working operating system based on the system identifier. Based on the working operating system, it calls at least one function of the vectorized execution engine from the pre-packaged local program package based on the shared library loader implementation class. Based on the working operating system, it executes the vectorized computing subtask based on the shared library loader implementation class and the called function. Specific implementation methods can be found in the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.
[0078] like Figure 2 As shown, the method specifically includes the following steps:
[0079] S210. Upon receiving a distributed computing task, the host operating system distributes at least one vectorized computing subtask from the distributed computing task to at least one working operating system.
[0080] Distributed computing tasks can be big data processing jobs initiated by the master operating system or upper-layer business services and executed in parallel on multiple working operating systems. The master operating system refers to the operating system environment running the master node (Driver or Master), possessing task scheduling and resource coordination capabilities. The working operating systems refer to the operating systems running on each worker node (Executor or Worker) in a heterogeneous operating system cluster. The working operating system can be the primary operating system. The master operating system can be either the primary or secondary operating system. Distributed computing tasks can be decomposed into multiple vectorized computing subtasks. Vectorized computing subtasks refer to subtasks within a distributed computing task that require acceleration using a high-performance native vectorized execution engine (such as Velox), relying on the local internal libraries of the primary operating system to complete columnar SIMD computation.
[0081] After receiving a distributed computing task, the master operating system can decompose the task into several vectorized computing subtasks according to preset partitioning rules (such as data block size, computational complexity, or hardware capability tags). Based on the registration information of the working operating systems (such as IP address, CPU architecture, and whether it is a domestically developed operating system), one or more vectorized computing subtasks are distributed to one or more matching working operating systems. For example, subtasks optimized for operating system A architecture are preferentially distributed to working nodes running operating system A.
[0082] Alternatively, before distributing vectorized computation subtasks, the master operating system can collect dynamic metrics in real time, such as the load status (e.g., CPU utilization, memory availability, network bandwidth) and computing capabilities (e.g., whether it supports specific vectorized instruction sets). Based on these dynamic metrics, a load balancing strategy (e.g., minimum load priority or capability matching priority) is used to distribute the vectorized computation subtasks to the most suitable working operating system. For example, when a domestic working system A is idle and supports AVX2, the master operating system prioritizes allocating high-throughput vectorized subtasks to domestic working system A. This approach can effectively improve the overall resource utilization and task response speed of the cluster.
[0083] Alternatively, the master operating system can pre-maintain a capability database for each working operating system, recording information such as system type, hardware platform, installed local internal library versions, and historical task execution performance. When a distributed computing task is received, the vectorized computing subtasks within it are first subjected to feature analysis (such as required function type, data size, and precision requirements). Then, a matching algorithm (such as a rule engine or lightweight model) is used to match the features of the vectorized computing subtasks with the capability features of each working operating system in the capability database, accurately distributing the vectorized computing subtasks to the working operating system with the most suitable capabilities. For example, subtasks that need to call the national cryptographic acceleration library are only distributed to the security-enhanced version nodes that have deployed the corresponding local internal library.
[0084] S220. For each operating system, when the operating system receives the vectorized computation subtask, it selects a matching second loading strategy from the local library loading strategies of different operating systems based on its system identifier.
[0085] In this embodiment, the second loading strategy refers to the local library loading strategy that the working operating system matches from the local library loading strategy set based on its own system identifier when the vectorized computing subtask is executed. The method for determining the second loading strategy is similar to the method for determining the first loading strategy, and will not be described in detail here.
[0086] S230. Based on the second loading strategy, the shared library loader implementation class corresponding to the working operating system is called according to the system identifier.
[0087] In this embodiment, the method of calling the shared library loader implementation class corresponding to the working operating system is similar to the method of calling the shared library loader implementation class matching the first operating system, and will not be described in detail here.
[0088] S240. Based on the operating system and the shared library loader implementation class, call at least one function of the vectorized execution engine from the pre-packaged local program package.
[0089] It should be noted that the local program package can be sent to each working operating system in advance based on the main control operating system; alternatively, the local program package can be sent to each working operating system when the distributed computing task is received for the first time.
[0090] The local packages corresponding to each working operating system can be the same or different. If they are different, the master operating system can pre-determine the system identifier (or system type) and corresponding local packages of each working operating system through cluster registration information or configuration manifests. Based on the system identifier of each working operating system, the master operating system selects a matching local package from its local resource library and distributes it as an auxiliary resource along with the task description (i.e., the vectorized computation subtask) to the corresponding working operating system during the task scheduling phase. Before executing the task, the working operating system loads the function calls from the local package to ensure that vectorized computation can proceed normally.
[0091] Alternatively, a separate local package service center can be pre-deployed, where all pre-compiled local packages are organized and stored according to operating system identifier (or system type). When the working operating system receives a vectorized computation subtask, it can proactively request the service center to retrieve the matching local package. This approach reduces the load on the main operating system and improves scalability.
[0092] S250, based on the operating system and the shared library loader, implements classes and calls functional functions to execute vectorized computation subtasks.
[0093] When the operating system receives the vectorized computation subtask, it determines and loads the operating system's native libraries based on the shared library loader implementation class using the second loading strategy. Furthermore, it processes the vectorized computation subtask by combining the native libraries with the called function calls.
[0094] The technical solution provided in this embodiment achieves efficient and accurate execution of vectorized computing subtasks in heterogeneous domestic operating system environments by introducing an operating system-aware task distribution and local library loading mechanism into the distributed computing architecture. Specifically, after the master operating system distributes the vectorized subtasks to the adapted working nodes, each working operating system dynamically selects a matching second loading strategy based on its own system identifier and calls the corresponding shared library loader implementation class to load the vectorized execution engine function optimized for its platform from the pre-integrated local program package, thereby completing the subtask computation. This not only ensures that different domestic operating systems can correctly load their exclusive local internal libraries and call high-performance vectorized functions, avoiding execution failures caused by path incompatibility or missing libraries, but also improves the overall throughput and resource utilization of distributed tasks. At the same time, since the local program package has pre-installed multi-platform adaptation libraries during the deployment phase, there is no need to download or compile at runtime, further reducing startup latency and network overhead, and enhancing the stability of data processing.
[0095] As an optional embodiment of the above embodiments, the technical solution provided in this embodiment can be applied to clusters based on a hybrid of domestic operating systems. In Spark batch processing, it enables multiple domestic operating systems to undergo Velox (i.e., function vectorization in a vector execution engine) vectorization, thereby optimizing the computing performance of the domestic operating systems, reducing CPU utilization, disk I / O overhead, and network transmission pressure, and minimizing hardware resource waste. Simultaneously, it meets the batch processing and high concurrency requirements of Spark on domestic operating systems without modifying the business code.
[0096] See Figure 3 The specific implementation methods of vectorization processing can include:
[0097] In the first phase, the Gluten ShareLibraryLoader interface (i.e., the generic shared library loader interface) can be extended to register dependent shared libraries (i.e., native internal libraries) from various domestic operating systems, resulting in shared library loader implementation classes. This allows the relevant native internal libraries to be loaded based on the shared library loader implementation classes in various domestic operating system environments during subsequent vectorized instruction set calls. For example, the shared library loader implementation classes for four different domestic operating systems are: Shared Library Loader Implementation Class 1, Shared Library Loader Implementation Class 2, Shared Library Loader Implementation Class 3, and Shared Library Loader Implementation Class 4.
[0098] In the second phase, an abstraction layer called VeloxListenerApi (Velox listener interface) can be created. The Strategy Pattern can be used to encapsulate various domestic operating systems into independent strategy classes (i.e., native library loading strategies). The registration and matching of native library loading strategies are managed through OSLibraryStrategyFactory (operating system native library loading strategy factory), resolving the differences and compatibility issues of multiple domestically produced shared libraries, and achieving adaptive dynamic switching of Gluten's execution plan across various mixed operating systems. For example, the native library loading strategies for four different domestic operating systems are: Native Library Loading Strategy 1, Native Library Loading Strategy 2, Native Library Loading Strategy 3, and Native Library Loading Strategy 4.
[0099] In the third stage, the operating system ID and version number (i.e., system identifier, such as system identifier 1, system identifier 2, system identifier 3, and system identifier 4) can be extracted from the / etc / os-release file by extending the compilation script. A separate compilation processing branch is designed for the domestic operating system. The compilation parameters and dependency installation logic are modified according to the compilation toolchain and environment limitations of the domestic operating system. The required third-party libraries and Gluten components (including function functions in the vectorized execution engine) are packaged to generate a local program package adapted to the domestic operating system, ensuring that it can be used directly in the domestic environment.
[0100] The technical solution provided by this invention, based on the Adaptive Query Execution (AQE) mechanism introduced in native Spark, integrates the Gluten vectorization extension plugin. During the execution of the DAG Stage (Directed Acyclic Graph Stage), it can dynamically adapt and push down the supported computation operators to the Velox execution library (i.e., the vectorized execution engine) for vectorized computation according to the system identifier of different domestic operating systems in the computing cluster.
[0101] To enable those skilled in the art to further understand the technical solutions of the embodiments of the present invention, specific application scenario examples are provided. For details, please refer to the following specific content.
[0102] In the first phase, the ShareLibraryLoader interface capabilities of Gluten can be extended.
[0103] Specifically, separate implementation classes for the shared library loader of the domestic operating system to be adapted can be defined, implementing (or inheriting) the ShareLibraryLoader interface. The loadLib method in the shared library loader implementation class is overridden. The loadLib method receives a JniLibLoader instance (a JNI library loading tool wrapped by Gluten) and completes a series of library loading operations through chained calls. The commit method is overridden to commit the loading transaction, triggering JniLibLoader to execute the actual library loading operation (loading the local internal libraries into the JVM process space) and complete the creation of symbolic links. At this point, all specified local internal libraries are in a usable state.
[0104] In the second phase, the VeloxListenerApi can be modified based on the strategy pattern to support multiple domestic operating systems.
[0105] Specifically, we can abstract the shared library loading behavior corresponding to the operating system (i.e., the native library loading logic) and define a strategy interface OSLibraryStrategy (i.e., the operating system's native library loading strategy). We implement specific native library loading strategies for each operating system (including domestically developed operating systems). We create a strategy factory class (i.e., the strategy pattern) to manage all native library loading strategies and match the corresponding loading strategy based on the operating system information. We obtain the loading strategy through the strategy pattern and execute the loading logic.
[0106] In the third stage, the modifiable compilation scripts use macro definitions to differentiate between domestic OS (i.e., the first operating system) and mainstream Linux (i.e., the second operating system) to achieve packaging generation.
[0107] Specifically, the script extracts the operating system ID (e.g., openEuler, anolis, bclinux) and version number (i.e., system identifier) from the ` / etc / os-release` file (the operating system distribution information file). In `get_velox.sh` (the Velox acquisition script), the `setup_linux` function (a Linux installation or configuration script) distinguishes between domestic operating systems and calls the corresponding configuration functions (e.g., `process_setup_openeuler2203`, `process_setup_anolis86`, etc.), designing separate processing branches for domestic operating systems. In the `package` (local package), `process_setup_*` macros (process setup or process initialization macros) for each domestic operating system are defined. Adapted local libraries are copied according to the local library paths of the domestic operating system (` / usr / lib64`, ` / usr / local / lib`), inapplicable compilation options are disabled, and dependency installation commands are adjusted (e.g., commenting out `epel-release` and `conda` related configurations to avoid conflicts with system-provided tools). This ensures that the compiled local package can correctly link to the local libraries, guaranteeing direct usability in a domestic environment.
[0108] Through the above steps, targeted adaptations to mainstream Linux and domestic OS have been achieved, which are reflected in source code modification and extension, dependency library selection, compilation parameter adjustment, and system feature compatibility, ensuring that Gluten components can be correctly compiled, run, and packaged in a domestic environment.
[0109] The technical solution provided in this embodiment abstracts the shared library loader implementation class for operating system adaptation and encapsulates the library loading logic and dependency management rules of domestic operating systems into independent policy classes. Combined with the policy pattern, it achieves fully automated adaptation of the entire process, including system identification, policy matching, and automatic execution. This allows for adaptation without modifying the Spark Gluten code on the business side when adding a new domestic operating system; only the corresponding policy class (i.e., the local library loading policy) needs to be added. It also supports seamless switching between mainstream Linux distributions and domestic operating systems, forming a reusable cross-operating system adaptation architecture applicable to components requiring multi-OS compatibility, such as those used in big data applications. Furthermore, the shared library loader implementation class masks system library version differences. By parsing ` / etc / os-release` to extract domestic operating system identifiers and combining this with a predefined dependency list, it automatically detects and completes missing libraries, resolving library compatibility issues between domestic operating systems and mainstream Linux distributions. This achieves zero-manual intervention in dependency management, lowering the deployment threshold for components like Gluten in domestic environments, ensuring stable operation of components in hybrid architectures with multiple domestic operating systems, and improving deployment efficiency through automatic dependency completion.
[0110] Figure 4 This is a schematic diagram of the structure of a vectorization processing device provided according to an embodiment of the present invention. Figure 4 As shown, the device is deployed in a heterogeneous operating system hybrid cluster, in which different operating systems run, and the operating systems include at least two first operating systems; the device includes: a loader implementation class determination module 310, a function determination module 320, and a vectorization processing module 330.
[0111] The loader implementation class determination module 310 is used to determine the shared library loader implementation class corresponding to the first operating system when a vectorized computation request is received from the first operating system. The shared library loader implementation class implements a generic shared library loader interface defined by the vectorized execution engine. The generic shared library loader interface is used to implement loading native libraries and calling functions of the vectorized execution engine on at least one second operating system. The second operating system is different from the first operating system. The function determination module 320 is used to call at least one function of the vectorized execution engine based on the shared library loader implementation class. The vectorization processing module 330 is used to perform data processing according to the vectorized computation request based on the shared library loader implementation class and the called function.
[0112] The technical solution of this embodiment, upon receiving a vectorized computation request from a first operating system, determines a shared library loader implementation class corresponding to the first operating system. This shared library loader implementation class implements a generic shared library loader interface defined by the vectorized execution engine. The generic shared library loader interface is used to load native libraries and call functions of the vectorized execution engine on at least one second operating system. The second operating system is distinct from the first operating system. Based on the shared library loader implementation class, at least one function of the vectorized execution engine is called. Based on the shared library loader implementation class and the called function, data processing is performed according to the vectorized computation request, thus solving the problem of native Gluten loading in the prior art. The interface was not adapted to domestic operating systems, failing to support the loading of local libraries and the invocation of vectorized functions, which affected the computing performance of domestic operating systems. To address this, a solution was implemented that, upon receiving a vectorized computing request, accurately matches and activates a shared library loader implementation class compatible with the current operating system. This ensures compatibility while efficiently loading local internal libraries, correctly invoking the function calls of the vectorized execution engine, and processing data according to the vectorized computing request. This not only ensures the availability and stability of vectorized computing capabilities in domestic operating system environments but also improves computing efficiency, reduces CPU utilization, disk I / O overhead, and network transmission pressure, thus minimizing hardware resource waste.
[0113] Based on the above-described device, optionally, the loader implementation class determination module 310 includes:
[0114] The first loading strategy determination unit is used to select a matching first loading strategy from the preset local library loading strategies of different operating systems based on the system identifier of the first operating system.
[0115] The shared library loader implementation class determination unit is used to retrieve the shared library loader implementation class that matches the first operating system based on the system identifier according to the first loading strategy.
[0116] Based on the above-mentioned device, optionally, the function determination module 320 is used to call at least one function of the vectorized execution engine from a pre-packaged local program package based on the shared library loader implementation class.
[0117] Based on the above-mentioned device, optionally, the vectorization processing module 330 includes:
[0118] An internal library loading unit is used to determine and load the local internal libraries of the first operating system based on the shared library loader implementation class according to the first loading strategy.
[0119] The data processing unit is used to process the vectorized computation request based on the local internal library and the called function.
[0120] Based on the above-mentioned device, optionally, an internal library loading unit is used to determine a local library path compatible with the first operating system based on the library loading method of the shared library loader according to the first loading strategy; and load the local internal library under the local library path.
[0121] Optionally, based on the above-described apparatus, the apparatus may further include:
[0122] The definition unit is used to define a shared library loader implementation class for each type of the first operating system based on the general shared library loader interface provided by the vectorized execution engine.
[0123] A registration unit is used to register the local library path of the first operating system in the library loading method of the shared library loader implementation class.
[0124] Optionally, based on the above-described apparatus, the apparatus may further include:
[0125] The encapsulation unit is used to encapsulate the local library loading logic of each of the first operating systems into a local library loading strategy based on the strategy pattern.
[0126] Based on the above-described device, optionally, the different operating systems include a main control operating system and at least one working operating system; the device further includes:
[0127] The subtask distribution unit is used to distribute at least one vectorized computing subtask in the distributed computing task to at least one working operating system based on the main control operating system when a distributed computing task is received.
[0128] The matching unit is used to select a matching second loading strategy from the local library loading strategies of different operating systems based on the system identifier of each operating system when the operating system receives the vectorized computing subtask.
[0129] The first invocation unit is used to invoke the shared library loader implementation class corresponding to the working operating system based on the second loading strategy and according to the system identifier.
[0130] The second calling unit is used to call at least one function of the vectorized execution engine from a pre-packaged local program package based on the shared library loader implementation class of the operating system.
[0131] An execution unit is used to execute the vectorized computation subtask based on the operating system, the shared library loader implementation class, and the called function.
[0132] The vectorization processing apparatus provided in the embodiments of the present invention can execute the vectorization processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0133] Figure 5 This is a schematic diagram of the structure of an electronic device implementing the vectorization processing method of embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0134] like Figure 5As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory 12 or a random access memory 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory 12 or a computer program loaded from storage unit 18 into the random access memory 13. The random access memory 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, read-only memory 12, and random access memory 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0135] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0136] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as vectorized processing methods.
[0137] In some embodiments, the vectorization processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory 12 and / or communication unit 19. When the computer program is loaded into random access memory 13 and executed by processor 11, one or more steps of the vectorization processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the vectorization processing method by any other suitable means (e.g., by means of firmware).
[0138] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0139] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0140] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0141] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0142] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0143] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0144] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from read-only memory 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.
[0145] This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the vectorization processing method provided in any embodiment of this invention.
[0146] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0147] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0148] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A vectorization processing method, characterized in that, The method is applied to a heterogeneous operating system hybrid cluster, wherein different operating systems run in the heterogeneous operating system hybrid cluster, and the operating systems include at least two first operating systems; the method includes: Upon receiving a vectorized computation request from a first operating system, a shared library loader implementation class corresponding to the first operating system is determined; wherein, the shared library loader implementation class implements a generic shared library loader interface defined by the vectorized execution engine; the generic shared library loader interface is used to implement the function of loading native libraries and calling the vectorized execution engine on at least one second operating system; the second operating system is different from the first operating system; Based on the shared library loader implementation class, at least one function of the vectorized execution engine is called; Based on the shared library loader implementation class and the called function, data processing is performed according to the vectorized computation request.
2. The method according to claim 1, characterized in that, The step of determining the shared library loader implementation class that matches the first operating system includes: Based on the system identifier of the first operating system, a matching first loading strategy is selected from the preset local library loading strategies of different operating systems; Based on the first loading strategy and the system identifier, the shared library loader implementation class that matches the first operating system is retrieved.
3. The method according to claim 1, characterized in that, The method of calling at least one function of the vectorized execution engine based on the shared library loader implementation class includes: Based on the shared library loader implementation class, at least one function of the vectorized execution engine is called from the pre-packaged local program package.
4. The method according to claim 1, characterized in that, The implementation class based on the shared library loader and the called function, performing data processing according to the vectorized computation request, includes: Based on the first loading strategy and the implementation class of the shared library loader, the local internal libraries of the first operating system are determined and loaded; Based on the local internal library and the called function, the vectorized computation request is processed.
5. The method according to claim 4, characterized in that, The step of determining and loading the native internal libraries of the first operating system based on the shared library loader implementation class according to the first loading strategy includes: Based on the first loading strategy, the library loading method of the class implemented by the shared library loader is used to determine the local library path compatible with the first operating system; Load the local internal library under the specified local library path.
6. The method according to claim 1, characterized in that, The method further includes: Based on the generic shared library loader interface provided by the vectorized execution engine, a shared library loader implementation class corresponding to each of the first operating systems is defined; Register the local library path of the first operating system in the library loading method of the shared library loader implementation class.
7. The method according to claim 1, characterized in that, The method further includes: Based on the strategy pattern, the local library loading logic of each of the first operating systems is encapsulated into a local library loading strategy.
8. The method according to claim 1, characterized in that, The different operating systems include a host operating system and at least one working operating system; the method further includes: Upon receiving a distributed computing task, the host operating system distributes at least one vectorized computing subtask of the distributed computing task to at least one of the working operating systems. For each of the operating systems, when the operating system receives the vectorized computing subtask, it selects a matching second loading strategy from the local library loading strategies of different operating systems based on its system identifier. Based on the second loading strategy, the shared library loader implementation class corresponding to the working operating system is called according to the system identifier; Based on the operating system and the shared library loader implementation class, at least one function of the vectorized execution engine is called from the pre-packaged local program package; Based on the operating system, the vectorized computation subtask is executed according to the shared library loader implementation class and the called function.
9. A vectorization processing device, characterized in that, Deployed in a heterogeneous operating system hybrid cluster, wherein different operating systems run in the heterogeneous operating system hybrid cluster, and the operating systems include at least two first operating systems; the device includes: The loader implementation class determination module is used to determine the shared library loader implementation class corresponding to the first operating system when a vectorized computation request is received from the first operating system; wherein, the shared library loader implementation class implements the generic shared library loader interface defined by the vectorized execution engine; the generic shared library loader interface is used to implement the function of loading native libraries and calling the function of the vectorized execution engine on at least one second operating system; the second operating system is different from the first operating system; The function determination module is used to call at least one function of the vectorized execution engine based on the shared library loader implementation class; The vectorization processing module is used to perform data processing based on the shared library loader implementation class and the called function, according to the vectorization calculation request.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to said at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the vectorization processing method according to any one of claims 1-8.