Data processing method and device, equipment and storage medium
By determining the corresponding computing engine type of the client in the data processing system and establishing a session connection, the problem that a single computing scenario cannot be applied to complex computing scenarios is solved, and efficient data processing of diversified services is achieved.
Patent Information
- Application Number
- CN202510193646.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-13
AI Technical Summary
Computing engines in a single computing scenario cannot be suitable for complex computing scenarios with diversified services, affecting the data processing efficiency of massive data.
By obtaining the client's data query request, determining the computing engine type corresponding to the client, and when the service instance corresponding to the computing engine type has established a session connection, the data query request is processed through the computing engine corresponding to the service instance to obtain the data query result.
It realizes the integration of different computing engine types, supports the needs of hybrid computing scenarios, is suitable for complex computing scenarios with diversified services, and improves the data processing efficiency of massive data storage, processing, analysis, mining, etc.
Smart Images

Figure CN120144848A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of data processing, and particularly relates to a data processing method, apparatus, device, and storage medium. Background Art
[0002] The diversity of computing scenarios in the big data field reflects the complexity and flexibility of big data computing. Different computing scenarios will adopt different computing engines to handle data processing processes such as storage, processing, analysis, and mining of massive data. For example, batch computing scenarios use Spark or Hive computing engines, streaming computing scenarios use Spark Streaming or Flink computing engines, and interactive data analysis scenarios use Trino or StarRocks computing engines. However, the computing engines in a single computing scenario cannot be applied to the complex computing scenarios of diversified services, which affects the data processing efficiency of massive data in complex computing scenarios. Summary of the Invention
[0003] Embodiments of this application provide a data processing method, apparatus, device, and storage medium, which can solve the problem that the computing engines in a single computing scenario in the related art cannot be applied to the complex computing scenarios of diversified services.
[0004] In a first aspect, embodiments of this application provide a data processing method, which may include:
[0005] Obtain a data query request from a client;
[0006] Determine a computing engine type corresponding to the client according to the data query request;
[0007] When a session connection has been established with a service instance corresponding to the computing engine type, process the data query request through a first computing engine corresponding to the service instance to obtain a data query result;
[0008] Send the data query result to the client.
[0009] In a second aspect, embodiments of this application provide a data processing system, including a front-end server, a session manager, and a back-end server; wherein,
[0010] The front-end server is configured to obtain a query request from a client;
[0011] The session manager is configured to determine a computing engine type corresponding to the client according to the data query request;
[0012] The back-end server is configured to, when a session connection has been established with a service instance corresponding to the computing engine type, process the data query request through a first computing engine corresponding to the service instance to obtain a data query result.
[0013] The front-end server is also used to send the data query result to the client.
[0014] In a third aspect, an embodiment of the present application provides a data processing device, including:
[0015] An acquisition module, configured to acquire a data query request of a client;
[0016] A determination module, configured to determine a computing engine type corresponding to the client according to the data query request;
[0017] A processing module, configured to process the data query request through a first computing engine corresponding to the service instance to obtain a data query result when a session connection has been established with the service instance corresponding to the computing engine type;
[0018] A sending module, configured to send the data query result to the client.
[0019] In a fourth aspect, an embodiment of the present application provides a computer device, which includes: a processor and a memory storing computer program instructions;
[0020] When the processor executes the computer program instructions, the data processing method as shown in the first aspect is implemented.
[0021] In a fifth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the data processing method as shown in the first aspect is implemented.
[0022] In a sixth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the data processing method as shown in the first aspect.
[0023] In a seventh aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the data processing method as shown in the first aspect.
[0024] The data processing method, apparatus, device, and storage medium according to the embodiments of the present application determine the type of computing engine corresponding to the client through the data query request of the client, and when a session connection has been established with the service instance corresponding to the computing engine type, process the data query request through the first computing engine corresponding to the service instance to obtain a data query result, and send the data query result to the client. In this way, there is no need to configure the query engines of different clients one by one in advance, reducing the resource amount for configuring the computing engines for each client one by one. Moreover, when a session connection has been established with the service instance corresponding to the computing engine type, through the service instance corresponding to this engine type, the first computing engine corresponding to the service instance is called to process the data query request, realizing the integration of computing engines of different computing engine types, supporting the requirements of hybrid computing scenarios, being applicable to complex computing scenarios of diversified services, and improving the data processing efficiency such as storage, processing, analysis, and mining of massive data in complex computing scenarios of diversified services. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the accompanying drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.
[0026] Figure 1 FIG. is a schematic structural diagram of a data processing system provided by an embodiment of the present application;
[0027] Figure 2 FIG. is a flowchart of a data processing method provided by an embodiment of the present application;
[0028] Figure 3 FIG. is a schematic structural diagram of a data processing system involving resource isolation steps provided by an embodiment of the present application;
[0029] Figure 4 FIG. is a schematic structural diagram of a data processing system involving authentication steps provided by an embodiment of the present application;
[0030] Figure 5 FIG. is a schematic structural diagram of a data processing apparatus provided by an embodiment of the present application;
[0031] Figure 6 FIG. is a schematic structural diagram of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0032] The features and exemplary embodiments of various aspects of the present application will be described in detail below. To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than limiting the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only intended to provide a better understanding of the present application by showing examples of the present application.
[0033] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0034] In the technical solution of the present application, the acquisition, storage, use, processing, etc. of data (including but not limited to the features, information, etc. in the text) all comply with the relevant provisions of national laws and regulations.
[0035] In the related art, a single computing engine is used to solve specific technical scenarios, and each computing engine has its own unique calling method. Taking the submission of a batch computing component job of the spark computing engine as an example, scala / java / python / R application program code can be first written based on the Application Programming Interface (API) of spark and compiled and packaged into a jar package. Then, the running environment of the spark-client is deployed and configured. Then, the jar package of the submitted job is submitted using spark-submit of the spark-client to wait for the job to finish running. It can be seen that a single computing engine can only serve its respective fixed fields and cannot be applied to complex computing scenarios with diversified services, affecting the data processing efficiency of massive data in complex computing scenarios.
[0036] To solve the problems occurring in the related art, the embodiments of the present application provide a data processing method, system, device, equipment, and storage medium, which can provide a unified query engine to achieve the integration of computing engines of different computing engine types, support the requirements of hybrid computing scenarios, be applicable to complex computing scenarios of diversified services, and improve the data processing efficiency of storing, processing, analyzing, and mining massive data in complex computing scenarios of diversified services.
[0037] Based on this, the data processing system, method, device, computer equipment, and storage medium of the embodiments of the present application will be described in detail below with reference to the accompanying Figures 1 to 6 drawings. It should be noted that these embodiments are not used to limit the scope of the disclosure of the present application.
[0038] Figure 1 FIG. is a schematic structural diagram of a data processing system provided by an embodiment of the present application.
[0039] As Figure 1 shown, a data processing system provided by an embodiment of the present application can be implemented by an entity service gateway or a virtual service gateway, and the form of the service gateway is not limited herein. Specifically, the data processing system in the embodiments of the present application can be configured as a Structured Query Language (SQL) service gateway to implement an SQL interface and support hybrid computing of computing engines such as Spark, Hive, and Flink, so that the client can access different computing engines such as Spark, Hive, Flink, and Trino through a simple connection method, without having to configure the query engines of different clients one by one in advance, nor having to configure and switch different calling devices for different computing engines, supporting the requirements of hybrid computing scenarios, being applicable to complex computing scenarios of diversified services, and improving the data processing efficiency of storing, processing, analyzing, and mining massive data in complex computing scenarios of diversified services.
[0040] In some embodiments, the data processing system in the embodiments of the present application may include a Frontend Server 101, a Session Manager 102, and a Backend Server 103.
[0041] Among them, the Frontend Server 101 is used to provide a standard Java Database Connectivity (JDBC) and SQL interface to the client, and transmit the SQL data query request or operation request submitted by the client to the Backend Server for subsequent processing.
[0042] Session Manager 102 is used to manage the life cycle of sessions, including session creation, destruction, authentication, resource allocation, and status tracking, etc. Among them, the session can be a session between the client and the data processing system, or specifically a session between the client and the service instance.
[0043] Backend Server 103 is used to respond to the SQL data query request or operation request transmitted by the Frontend Server, and based on the service instance corresponding to the type of computing engine requested by different requests, call different computing engines corresponding to the service instance to process the SQL data query request, and return the processed data query result to the Frontend Server to be fed back to the client through the Frontend Server.
[0044] Based on each part in the above-mentioned data processing system, the following will be combined with Figure 1 to illustrate the process of executing the data processing method flow of the above-mentioned data processing system.
[0045] The Frontend Server receives the data query request from the client. The data query request can include an SQL query request, or can include an SQL query request and a connection request of the JDBC client. Among them, the connection request can be a request to establish a session connection with the client once, or can be a request to establish multiple session connections for different SQL query requests of the client. Here, an example will be given for illustration with the data query request including an SQL query request and a connection request.
[0046] For each connection request, the Frontend Server will authenticate the client based on the connection request, and when the authentication result indicates that the client passes the authentication, it is determined that the client has the access right to access the Backend Server. Among them, if the client passes the authentication, the Frontend Server can create a service instance for the client, determine the type of computing engine used to process the SQL query request according to the SQL query request, and then, according to the type of computing engine, establish a session connection between the client and the service instance, and assign a unique Session identifier (Identity document, ID) to it, and its service instance can correspond to the type of computing engine.
[0047] The Frontend Server forwards the received SQL query request to the Session Manager so that the Session Manager can allocate the SQL query request to the corresponding Backend Service for execution.
[0048] After the Session Manager receives the SQL query request transmitted by the Frontend Server, it selects the corresponding Backend Server according to the session configuration to execute the request.
[0049] The Backend Server is used to submit the SQL query requests corresponding to different calculation engine types of SQL query requests to different SQL execution engines to execute specific SQL queries. Specifically, the Backend Server receives the SQL query request transmitted by the Session Manager, and according to the calculation engine type corresponding to the session connection between the foregoing and the service instance, submits its SQL query request to at least one calculation engine corresponding to the calculation engine type, so as to process the SQL query request through at least one calculation engine corresponding to the calculation engine type to obtain a data query result. For example, if the calculation engine type is Spark, a Spark Session corresponding to the calculation engine Spark1 is created and the SQL query request is submitted to Spark1 through its Spark Session for execution; or, if the calculation engine type is Hive, a Hive Session corresponding to the calculation engine Hive1 is created and the SQL query request is submitted to Hive1 through its Hive Session for execution.
[0050] The Backend Server will return its data query result to the Backend Server in the form of a result set (ResultSet). The Backend Server then returns the result to the Session Manager, and finally the result is returned to the Frontend Server through the Session Manager, so as to send the data query result to the client through the Frontend Server.
[0051] Thus, the data processing system in the embodiments of the present application can provide a unified query engine for a large number of clients through a unified service gateway, so as to support data query requests of different calculation engine types in mixed scenarios such as batch processing, stream processing, and interactive query at the same time, such as SQL query requests of the Hive type, SQL query requests of the Spark type, and SQL query requests of the Flink type, thereby realizing unified support for mixed scenarios such as batch processing, stream processing, and interactive query.
[0052] In addition, in some embodiments of the present application, in order to achieve a balance between resource utilization efficiency and resource isolation, the Session Manager in the embodiments of the present application can also be used to provide a resource isolation function, so as to improve the resource utilization efficiency while also improving the resource isolation level. The client can specify the isolation level of the computing engine through a configuration parameter class. The isolation levels include an exclusive isolation level and a shared isolation level, and the security level of the former is greater than that of the latter. Among them, in the exclusive isolation level, one session corresponds to one computing engine; in the shared isolation level, one client corresponds to one computing engine, that is, multiple sessions under the same client share one computing engine. The Session Manager manages the life cycle of the computing engine according to the isolation parameters of the isolation level, such as exclusive isolation parameters and shared isolation parameters, and the connection status of the session.
[0053] Specifically, the Session Manager can establish a session connection with the service instance according to the computing engine configuration parameter class corresponding to the client. For example, if the computing engine configuration parameter class includes exclusive isolation parameters, a new computing engine corresponding to the computing engine type will be created. That is, when there is a new session connection and the exclusive isolation parameters are used, the SessionManager will create a new computing engine. For another example, if the computing engine configuration parameter class includes shared isolation parameters, according to the shared isolation parameters, it is determined whether the computing engine corresponding to the service instance includes a shared computing engine corresponding to the client; in the case where it is determined that the computing engine corresponding to the service instance includes a shared computing engine corresponding to the client, a session connection corresponding to the shared computing engine is established, that is, the Sesion Manager will assign the session connection to the existing computing engine.
[0054] In addition, in order to save the computing resources of the data processing system, when a session is disconnected and no other session shares the computing engine, the Session Manager will destroy the computing engine.
[0055] In this way, in the exclusive isolation level with high security, each session has an independent computing engine, which can avoid mutual conflicts and interferences of resources. In the shared isolation level with low security, multiple sessions of one client will share one computing engine, which can improve resource utilization. In addition, through the unified query engine, that is, the service network, by supporting both shared and independent isolation levels at the same time, it is possible to improve resource utilization efficiency while ensuring the relative isolation of resources.
[0056] In addition, in some embodiments of the present application, to avoid the problem of resource leakage caused by the current unauthorized control mode of computing engines such as Spark and Flink, the data processing system in the embodiments of the present application may further include a permission verification module (Ranger). This permission verification module can be independently configured within the data processing system, or can be set within the Session Manager, or can be set within the corresponding permission verification plug-ins RangerPlugin of computing engines such as Hive, Spark, and Flink. Ranger can be used to achieve unified configuration and verification of permission policies for computing engines such as Hive, Spark, and Flink.
[0057] In this way, the permission configuration and management of the service network are entrusted to Ranger, thereby achieving more flexible and secure permission management.
[0058] The following details its permission verification logic.
[0059] When the client sends an SQL query request to the unified service gateway, the unified service gateway will call RangerPlugin to intercept the SQL query request. RangerPlugin extracts the verification information in the SQL query request, such as information about the client's identity, operation type, computing engine type, etc., and encapsulates this information into a permission verification request, and passes it through a permission verification module such as Ranger Admin Server. Ranger Admin Server performs an identity verification result on the SQL query request according to the configured policy to determine whether the client has the permission to execute the operation of accessing the computing engine. RangerAdmin Server returns the identity verification result to RangerPlugin, and RangerPlugin decides whether to allow the SQL query request to continue to execute according to the result returned by Ranger AdminServer. If allowed, the SQL query request is forwarded to the corresponding computing engine; if not allowed, an error message indicating insufficient permissions is returned to the client.
[0060] In addition, the permission verification module can also be used to provide fine-grained permission control, which can be accurate to tables, columns, or even rows; and can also be used for unified permission management, that is, the permission management and configuration for Hive, Spark, and Flink are unified into Ranger, thereby achieving unified permission management; and can also be used for flexible permission policies, that is, based on Ranger, not only can the permission policies of libraries, tables, and fields be configured, but also access control can be defined based on roles, and the allow, deny, and exception access control policies for the control policy can be customized.
[0061] Thus, through this permission verification module, unified fine-grained permission management can be achieved, thereby overcoming the deficiencies in permission control of computing engines such as Spark and Flink themselves. In summary, the data processing system provided by the embodiments of the present application can provide services for the query engine through a unified service gateway, which can not only achieve unified SQL query request access to the computing engine and permission control, but also reduce the cost of manually configuring multiple query engines and save the migration cost between query engines for different computing engines, further improving data productivity.
[0062] Based on the above data processing system, the embodiments of the present application provide a data processing method applied to the data processing system. The following will be described in detail in conjunction with Figure 2 this.
[0063] Figure 2 FIG. is a flowchart of a data processing method provided by an embodiment of the present application.
[0064] As Figure 2 shown, this data processing method can be applied to a data processing system, and the method can specifically include the following steps:
[0065] Step 210, obtain a data query request from the client; Step 220, determine the type of computing engine corresponding to the client according to the data query request; Step 230, when a session connection has been established with the service instance corresponding to the computing engine type, process the data query request through the first computing engine corresponding to the service instance to obtain a data query result; Step 240, send the data query result to the client.
[0066] Exemplarily, if the data query request is an SQL query request, the Frontend Server can receive the SQL query request sent by the client through the SQL interface. Determine the type of computing engine that the client wants to access for this SQL query request, such as the Spark type, and create a service instance corresponding to the computing engine type for the client based on the SQL query request and the computing engine type to establish a session connection between the computing engine and the client. Then, the Session Manager establishes a session connection with the service instance of the Spark type, and when a session connection has been established with the service instance of the Spark type, forwards the SQL query request to the Backend Server. The Backend Server processes the SQL query request according to the first computing engine corresponding to the service instance, such as Spark1, to obtain a data query result, and sends its data query result to the Frontend Server so that the Frontend Server can feedback its data query result to the client.
[0067] In this way, there is no need to configure the query engines of different clients one by one in advance, reducing the resource amount for configuring the computing engines for each client one by one. Moreover, when a session connection has been established with a service instance corresponding to the computing engine type, the first computing engine corresponding to the service instance is called through the service instance corresponding to the engine type to process the data query request, realizing the integration of computing engines of different computing engine types, supporting the requirements of hybrid computing scenarios, being applicable to complex computing scenarios of diversified services, and improving the data processing efficiency such as storage, processing, analysis, and mining of massive data in complex computing scenarios of diversified services.
[0068] The above steps will be described in detail as follows.
[0069] Regarding step 210, in some embodiments of the present application, a data query request sent by a client can be received through a JDBC interface or an SQL interface. The data query request can include an SQL query request; or the data query request can include an SQL query request and a connection request of a JDBC client. Among them, the connection request is used to create a session connection between the client and the service instance.
[0070] Regarding step 220, in some embodiments of the present application, the data query request carries the computing engine type for the client to process the data query request. Based on this, step 220 can specifically include: parsing the data query request and extracting the computing engine type from the parsed data query request.
[0071] In some other embodiments of the present application, the data query request carries the identity identifier of the client. Based on this, step 220 can specifically include: obtaining the pre-set computing engine type corresponding to the identity identifier of the client according to the association relationship between the pre-set identity identifier and the pre-set computing engine type; and determining the pre-set computing engine type corresponding to the identity identifier of the client as the computing engine type corresponding to the client.
[0072] In still some other embodiments of the present application, step 220 can specifically include: according to the association relationship between the request type of the pre-set data query request and the pre-set computing engine type, determining the pre-set computing engine type corresponding to the request type of the data query request; and determining the pre-set computing engine type corresponding to the request type of the data query request as the computing engine type corresponding to the client.
[0073] It should be noted that the above steps for determining the computing engine type corresponding to the client can be used independently or in combination, and there is no limitation on using at least one of the above methods to determine the computing engine type corresponding to the client.
[0074] Regarding step 230, in some embodiments of the present application, in order to achieve a balance between resource utilization efficiency and resource isolation, embodiments of the present application also provide steps for resource isolation to improve resource utilization efficiency while also increasing the level of resource isolation. Based on this, before step 230, the data processing method may further include step 3101 and step 3102.
[0075] Step 3101, determine a service instance corresponding to the computing engine type from a preset service instance list.
[0076] Further, this step 3101 may specifically include:
[0077] According to preset filtering conditions, determine a service instance corresponding to the computing engine type from a preset service instance list; where the preset filtering conditions include at least one of the following: load balancing, the status of the server to which the service instance belongs, and the geographical location information of the server and the client.
[0078] Exemplarily, the service instance list in the distributed application coordination service (zookeeper) can be queried, and one or more instances of the service instance can be deployed. Zookeeper will return information about all currently registered service instances, including information such as addresses and ports. According to the instance list obtained from zookeepr, a service instance of a service module is selected according to strategies such as load balancing, the status of the server, and geographical location information.
[0079] Thus, it adopts a high availability (HA) design, effectively preventing single point of failure, and achieving load balancing through a multi-node deployment method, thereby ensuring the continuity of the data processing system and application services related to data query requests.
[0080] Step 3102, establish a session connection with the service instance according to the computing engine configuration parameter class corresponding to the client, where the computing engine configuration parameter class is used to indicate the isolation level of the first computing engine.
[0081] Exemplarily, the client can specify the isolation level of the computing engine through the configuration parameter class. The isolation levels include an exclusive isolation level and a shared isolation level, and the security level of the former is higher than that of the latter. Among them, for the exclusive isolation level, one session corresponds to one computing engine; for the shared isolation level, one client corresponds to one computing engine, that is, multiple sessions under the same client share one computing engine. Session Manager manages the life cycle of the computing engine according to the isolation parameters of the isolation level, such as exclusive isolation parameters and shared isolation parameters, and the connection status of the session.
[0082] Thus, the client can specify the isolation level of the computing engine through the configuration parameter class, achieving a balance between resource utilization efficiency and resource isolation. While improving the resource utilization efficiency, the resource isolation level can also be enhanced.
[0083] Based on this, the isolation levels in the embodiments of the present application may include an exclusive isolation level and a shared isolation level. The computing engine configuration class may include independent isolation parameters corresponding to the exclusive isolation level and shared isolation parameters corresponding to the shared isolation level. Below, the process of establishing a session connection with a service instance will be described in detail in combination with different types of isolation parameters.
[0084] In some embodiments, the computing engine configuration parameter class includes exclusive isolation parameters. The exclusive isolation parameters are used to indicate a one-to-one correspondence between the computing engine and the session connection. The session connection with the service instance includes a session connection with the first computing engine. Based on this, step 3102 above may specifically include:
[0085] According to the exclusive isolation parameters, create a computing engine corresponding to the computing engine type through the service instance;
[0086] Determine the created computing engine as the first computing engine;
[0087] Establish a session connection with the first computing engine.
[0088] Exemplarily, if the computing engine configuration parameter class includes exclusive isolation parameters, then create a new computing engine corresponding to the computing engine type, and use the newly created computing engine as the first computing engine. That is, when there is a new session connection and it is an exclusive isolation parameter, a new computing engine can be created as the first computing engine. At this time, this session connection is the session connection between the client and the first computing engine.
[0089] In some other embodiments, the computing engine configuration parameter class includes shared isolation parameters. The shared isolation parameters are used to indicate the correspondence between the computing engine and the client. The session connection with the service instance includes a session connection with the shared computing engine. Based on this, step 3102 above may specifically include:
[0090] According to the shared isolation parameters, determine whether the computing engine corresponding to the service instance includes a shared computing engine corresponding to the client;
[0091] In the case where it is determined that the computing engine corresponding to the service instance includes a shared computing engine corresponding to the client, establish a session connection corresponding to the shared computing engine. The first computing engine includes the shared computing engine.
[0092] Exemplarily, the computing engine configuration parameter class includes shared isolation parameters. According to the shared isolation parameters, it is determined whether the computing engine corresponding to the service instance includes a shared computing engine corresponding to the client; in the case where it is determined that the computing engine corresponding to the service instance includes a shared computing engine corresponding to the client, a session connection corresponding to the shared computing engine is established, that is, the session connection will be assigned to the existing computing engine corresponding to the client.
[0093] To better understand the above steps of establishing a session connection with a service instance in combination with different types of isolation parameters, the following will be described in detail in combination with Figure 3 this.
[0094] As Figure 3 shown, the data processing system may include a routing module, a service module, and a computing engine.
[0095] Among them, the routing module is used to provide routing services for the client to the service module and the computing engine. The routing module can be zookeeper or a service routing component with similar functions. After the service instance of the service module starts the process, it will register the instance in the service module in zookeeper. When the corresponding executor in the computing engine is called, it will also register the service instance in zookeeper. The registration information includes information such as the address, port, status, and type of the service instance. Among them, the computing engine in the embodiment of the present application can be configured as an executor, such as a Presto executor, a Hive executor, a Spark executor, a Trino executor, a StartRocks executor, a Flink executor.
[0096] Based on this, when the client creates a data query request through the JDBC client, it can proceed according to the following steps and processes.
[0097] First, the client will first query the service instance list in zookeeper. One or more service instances can be deployed. Zookeeper will return the information of all currently registered service instances, including information such as the address and port of the service module. Then, according to the instance list obtained from zookeepr, a service module instance is selected according to strategies such as load balancing, server status, and geographical location information. Then, the client establishes a session with the selected service module instance. During the session establishment process, the service module can create a session with the computing engine according to information such as the engine type and isolation level in the client JDBC connection string (Uniform Resource Locator, url).
[0098] The service module and the computing engine create a session according to the following policy. If the isolation level is the connection level, i.e., the exclusive isolation level, and each session does not share the service instance corresponding to the computing engine, a new computing engine is directly created according to the type of the computing engine, and the service instance session corresponding to the computing engine is registered in ZooKeeper. If the isolation level is the server level, i.e., the shared isolation level, it is queried in ZooKeeper whether there is an available computing engine in the current service instance. If an available computing engine is found, it is directly used.
[0099] It should be noted that, in order to save the computing resources of the data processing system, corresponding computing engines can be destroyed based on the actual situation. Based on this, the data processing method provided by the embodiments of the present application may further include step 3201 and step 3202.
[0100] Step 3201, when the session connection is disconnected, determine whether there is a shared computing engine corresponding to the client.
[0101] Step 3202, when it is determined that there is no shared computing engine corresponding to the client, destroy the first computing engine.
[0102] Thus, the data processing method provided by the embodiments of the present application can support the multi-tenant mode, allow multiple applications or application programs to share the same server, and at the same time ensure resource isolation and high concurrency processing capabilities, thereby significantly improving the resource utilization effect, improving the query efficiency and avoiding resource waste.
[0103] In other embodiments of the present application, in order to avoid the problem of resource leakage caused by the current lack of permission control mode for computing engines such as Spark and Flink, the embodiments of the present application also provide steps for authentication to determine access permissions. Based on this, before step 230, the data processing method may further include step 4101 and step 4102.
[0104] Step 4101, authenticate the client to obtain an authentication result.
[0105] Step 4102, when the authentication result indicates that the client passes the authentication, establish a session connection with the service instance.
[0106] Further, in some embodiments, step 4101 may specifically include:
[0107] Intercept and parse the data query request through the service instance to obtain the verification information of the data query request;
[0108] When the verification information meets the verification conditions, determine the authentication result indicating that the client passes the authentication.
[0109] It should be noted that the verification conditions in the embodiments of the present application include at least one of the following: rejection exception condition, rejection condition, permission exception condition, and permission condition. The priorities of the verification conditions are, from high to low, rejection exception condition, rejection condition, permission exception condition, and permission condition.
[0110] Based on this, before the above step 4102, the data processing method may further include step 4103 and step 4104.
[0111] Step 4103, when the verification conditions include at least two verification conditions, verify the verification information according to the priority order of the verification conditions.
[0112] Step 4104, when the verification information meets each verification condition in at least two verification conditions, determine an authentication result indicating that the client passes the authentication.
[0113] Exemplarily, the verification process is carried out according to the priorities and conditions of the policies, including permission condition (Allow), permission exception condition (Exclude from Allow), rejection condition (Deny), and rejection exception condition (Exclude from Deny). The priorities of different conditions are, from high to low, rejection exception condition > rejection condition > permission exception condition > permission condition. The specific verification logic is as follows: First, verify whether the rejection condition is hit. If the rejection condition is hit, continue to verify whether the rejection exception condition is hit: If the rejection exception condition is not hit, directly return the permission rejection information; if the rejection exception condition is hit, continue with the subsequent verification process. If the rejection condition is not hit, continue with the subsequent verification process. Second, verify whether the permission condition is hit. If the permission condition is hit, continue to verify whether the permission exception condition is hit: If the permission exception condition is not hit, directly return the permission approval information; if the permission exception condition is hit, continue with the subsequent verification process. If the permission condition is not hit, also continue with the subsequent verification process. Finally, if there is no feedback information of approval or rejection in the previous two steps, return the default permission policy. The default permission control policy can be defined as approval or rejection. In this way, after the permission verification of querying data is passed, the administrator account of the service instance can be used to proxy the user to execute SQL.
[0114] To better understand the above steps of authentication to determine access permissions, the following is combined with Figure 4 for detailed description.
[0115] As Figure 4As shown, Ranger is an open-source permission control component used for permission control at the levels of libraries, tables, fields, and files; HMS is an open-source Hive metadata manager used to define and manage metadata information such as libraries, tables, and fields; YARN is an open-source resource manager in the Hadoop ecosystem used to control resource scheduling in big data clusters; HDFS is an open-source distributed storage in the Hadoop ecosystem used to store data such as HDFS files or Hive tables. The service instance intercepts and parses data query requests and executes them according to the following steps and processes. Here, the Spark computing engine is used as an example for illustration. The principles of other computing engines are similar and will not be elaborated here.
[0116] First, deploy the Ranger client plugin (RangerPlugin) for the Spark component in the computing engine executor, and configure the spark.sql.extension parameter for Spark to load and call RangerPlugin during operation. Then, after establishing a connection with the client for the service instance, the service instance starts the executor of the Spark computing engine. After the computing engine executor starts, it applies for computing resources from YARN and generates an Application Master and Executors in YARN. Here, the Application Master is the driver program of the computing engine executor, and the Executor is the execution program of the computing engine. After starting, the Application Master fetches the permission policy information from the Ranger server through the Ranger Plugin for subsequent permission verification. The permission policy information includes users, resources, and corresponding permission information. Then, the Application Master receives the user SQL sent by the service instance and verifies it with the user's permission policy information using the information such as Hive tables, fields, and data storage locations stored in the storage component (Metastore). The verification process is based on the priority and conditions of the policy, including Allow, Exclude from Allow, Deny, and Exclude from Deny. The priorities of different conditions from high to low are Exclude from Deny > Deny > Exclude from Allow > Allow. The specific verification logic is as follows: First, check if the Deny condition is met. If the Deny condition is met, continue to check if the Exclude from Deny condition is met: If the Exclude from Deny condition is not met, directly return the permission denial information; if the Exclude from Deny condition is met, continue with the subsequent verification process. If the Deny condition is not met, continue with the subsequent verification process. Second, check if the Allow condition is met. If the Allow condition is met, continue to check if the Exclude from Allow condition is met: If the Exclude from Allow condition is not met, directly return the permission approval information; if the Exclude from Allow condition is met, continue with the subsequent verification process. If the Allow condition is not met, also continue with the subsequent verification process. Finally, if there is no feedback information of approval or denial in the previous two steps, return the default permission policy. The default permission control policy can be defined as Allow or Deny. In this way, after the permission verification for data query is passed, the administrator account of the service instance can be used to proxy the user to execute the SQL.
[0117] Therefore, the data processing method provided by the embodiments of the present application can protect the network security between the client and the server through a centralized authentication model, ensure that different tenants have different security authentications and permission controls for resource acquisition and data access, thereby reducing the risk of data and resource leakage.
[0118] Regarding step 240, in some embodiments of the present application, the service instance submits an SQL query request to the big data cluster, obtains the return result of the query request, and returns the result to the client. After the computing engine completes the calculation, it feeds back the result and status of the query request to the service instance, and the service instance returns the result and status of the final query request to the user side through the communication protocol of JDBC. In this way, a unified SQL access interface and permission control module can be implemented through a unified service gateway, realizing the integration of computing engines of different computing engine types, supporting the requirements of hybrid computing scenarios, being applicable to complex computing scenarios of diversified services, improving the data processing efficiency of storing, processing, analyzing, and mining massive data in complex computing scenarios of diversified services, and further enabling it to easily integrate into the existing Hadoop ecosystem and reducing the user learning and migration costs.
[0119] It should be noted that the data processing method provided by the embodiments of the present application can provide a unified query engine to support hybrid application scenarios such as batch processing, stream processing, and interactive query at the same time.
[0120] The present application also provides a data processing device, which will be specifically described in conjunction with Figure 5 for detailed description.
[0121] Figure 5 is a schematic structural diagram of a data processing device provided by an embodiment of the present application.
[0122] In some embodiments of the present application, Figure 5 the data processing device shown can be set in the computer device provided by the embodiments of the present application.
[0123] As Figure 5 shown, the data processing device 50 may specifically include:
[0124] An obtaining module 501, configured to obtain a data query request of the client;
[0125] A determining module 502, configured to determine a computing engine type corresponding to the client according to the data query request;
[0126] A processing module 504, configured to process the data query request through a first computing engine corresponding to the service instance to obtain a data query result when a session connection has been established with the service instance corresponding to the computing engine type;
[0127] A sending module 504, configured to send a data query result to a client.
[0128] In this way, in the data processing device according to the embodiment of the present application, there is no need to configure the query engines of different clients one by one in advance, reducing the resource amount for configuring the computing engines for each client one by one. Moreover, when a session connection has been established with a service instance corresponding to the computing engine type, through the service instance corresponding to this engine type, the first computing engine corresponding to the service instance is called to process the data query request, realizing the integration of computing engines of different computing engine types, supporting the requirements of hybrid computing scenarios, being applicable to complex computing scenarios of diversified services, and improving the data processing efficiency such as storage, processing, analysis, and mining of massive data in complex computing scenarios of diversified services.
[0129] The data processing device 50 in the embodiment of the present application will be described in detail below.
[0130] In some embodiments of the present application, the determining module 502 may further be configured to determine a service instance corresponding to the computing engine type from a preset service instance list;
[0131] The data processing device 50 in the embodiment of the present application may further include an establishing module, configured to establish a session connection with the service instance according to the computing engine configuration parameter class corresponding to the client, where the computing engine configuration parameter class is used to indicate the isolation level of the first computing engine.
[0132] In some embodiments of the present application, the data processing device 50 in the embodiment of the present application may further include a creating module, configured to, when the computing engine configuration parameter class includes an exclusive isolation parameter, where the exclusive isolation parameter is used to indicate that the computing engine corresponds one-to-one to the session connection, and the session connection with the service instance includes a session connection with the first computing engine, create a computing engine corresponding to the computing engine type through the service instance according to the exclusive isolation parameter;
[0133] The determining module 502 may further be configured to determine the created computing engine as the first computing engine;
[0134] The data processing device 50 in the embodiment of the present application may further include an establishing module, configured to establish a session connection with the first computing engine.
[0135] In some embodiments of the present application, the determining module 502 may further be configured to, when the computing engine configuration parameter class includes a shared isolation parameter, where the shared isolation parameter is used to indicate that the computing engine corresponds to the client, and the session connection with the service instance includes a session connection with the shared computing engine, determine whether the computing engine corresponding to the service instance includes a shared computing engine corresponding to the client according to the shared isolation parameter;
[0136] The data processing device 50 in the embodiments of the present application may further include an establishment module, configured to establish a session connection corresponding to the shared computing engine when it is determined that the computing engine corresponding to the service instance includes a shared computing engine corresponding to the client, and the first computing engine includes the shared computing engine.
[0137] In some embodiments of the present application, the determination module 502 may further be configured to determine whether there is a shared computing engine corresponding to the client when the session connection is disconnected;
[0138] The data processing device 50 in the embodiments of the present application may further include a destruction module, configured to destroy the first computing engine when it is determined that there is no shared computing engine corresponding to the client.
[0139] In some embodiments of the present application, the determination module 502 may specifically be configured to determine, according to a preset screening condition, a service instance corresponding to the computing engine type from a preset service instance list; wherein, the preset screening condition includes at least one of the following: load balancing, the status of the server to which the service instance belongs, and the geographical location information of the server and the client.
[0140] In some embodiments of the present application, the data processing device 50 in the embodiments of the present application may further include an authentication module, configured to perform authentication on the client to obtain an authentication result;
[0141] The data processing device 50 in the embodiments of the present application may further include an establishment module, configured to establish a session connection with the service instance when the authentication result indicates that the client has passed the authentication.
[0142] In some embodiments of the present application, the acquisition module 501 may further be configured to intercept and parse a data query request through the service instance to obtain verification information of the data query request;
[0143] The determination module 502 may further be configured to determine an authentication result indicating that the client has passed the authentication when the verification information meets the verification condition.
[0144] In some embodiments of the present application, the data processing device 50 in the embodiments of the present application may further include an authentication module, configured to, when the verification condition includes at least one of the following: rejection exception condition, rejection condition, permission exception condition, permission condition, the priority of the verification conditions is in descending order as rejection exception condition, rejection condition, permission exception condition, permission condition, and when the verification condition includes at least two verification conditions, verify the verification information according to the priority order of the verification conditions;
[0145] The determination module 502 can also be used to determine an authentication result indicating that the client has passed authentication when the verification information meets each of at least two verification conditions.
[0146] This application also provides a computer device. Specifically, it will be described in detail in conjunction with Figure 5 for elaboration.
[0147] Figure 5 is a schematic structural diagram of a computer device provided by an embodiment of this application.
[0148] As Figure 5 shown, the computer device may include at least one of the following involved in the embodiments of this application: an electronic device, a service gateway as Figure 1 shown. Among them, the computer device may include a processor 501 and a memory 502 storing computer program instructions.
[0149] Specifically, the above-mentioned processor 501 may include a central processing unit (CPU), or an application specific integrated circuit (ASTC), or may be configured to implement one or more integrated circuits of the embodiments of this application.
[0150] The memory 502 may include a mass storage for data or instructions. By way of example and not limitation, the memory 502 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 502 may include removable or non-removable (or fixed) media. In a suitable case, the memory 502 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 502 is a non-volatile solid-state memory. In a specific embodiment, the memory 502 includes solid-state storage (ROM). In a suitable case, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.
[0151] The processor 501 reads and executes the computer program instructions stored in the memory 502 to implement any one of the data processing methods in the above embodiments.
[0152] In one example, the computer device may further include a communication interface 503 and a bus 510. Among them, as Figure 5As shown, a processor 501, a memory 502, and a communication interface 503 are connected via a bus 510 to complete communication with each other.
[0153] The communication interface 503 is mainly used to implement communication between various modules, devices, units, and / or apparatuses in the embodiments of the present application.
[0154] The bus 510 includes hardware, software, or both, and couples components of a flow control device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front-Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 510 may include one or more buses. Although the embodiments of the present application describe and illustrate a specific bus, the present application contemplates any suitable bus or interconnect.
[0155] The detection device for face recognition attacks can execute the data processing method in the embodiments of the present application, thereby implementing the combination of Figures 1 to 4 the described data processing method and apparatus.
[0156] In addition, in combination with the data processing method in the above embodiments, the embodiments of the present application can be implemented by providing a computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the data processing methods in the above embodiments is implemented.
[0157] It should be clear that the present application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.
[0158] The functional blocks shown in the above structural block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0159] It should also be noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.
[0160] The above is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, modules, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present application.
Claims
1. A data processing method, comprising: Get the client's data query request; Determining a computing engine type corresponding to the client according to the data query request; In the case where a session connection has been established with the service instance corresponding to the computing engine type, the data query request is processed by a first computing engine corresponding to the service instance to obtain a data query result; The data query result is sent to the client.
2. The method according to claim 1, wherein: After determining the computing engine type corresponding to the client according to the data query request, the method further includes: Determine a service instance corresponding to the computing engine type from a preset service instance list; A session connection with the service instance is established according to the computing engine configuration parameter class corresponding to the client, wherein the computing engine configuration parameter class is used to indicate an isolation level of the first computing engine.
3. The method according to claim 2, wherein: The computing engine configuration parameter class includes an exclusive isolation parameter, where the exclusive isolation parameter is used to indicate that computing engines correspond to session connections one by one, and the session connection with the service instance includes a session connection with the first computing engine; The establishing a session connection with the service instance according to the computing engine configuration parameter class corresponding to the client includes: Creating a computing engine corresponding to the computing engine type through the service instance according to the exclusive isolation parameter; Determine the created computing engine as the first computing engine; A session connection is established with the first computing engine.
4. The method according to claim 2, wherein: The computing engine configuration parameter class includes a shared isolation parameter, and the shared isolation parameter is used to indicate that the computing engine corresponds to the client, and the session connection with the service instance includes a session connection with the shared computing engine; The establishing a session connection corresponding to the service instance according to the computing engine configuration parameter class corresponding to the client includes: Determining, according to the shared isolation parameter, whether the computing engine corresponding to the service instance includes a shared computing engine corresponding to the client; In the case of determining that the computing engines corresponding to the service instance include a shared computing engine corresponding to the client, a session connection corresponding to the shared computing engine is established, and the first computing engine includes the shared computing engine.
5. The method according to any one of claims 2 to 4, wherein: The method further comprises: In the case where the session connection is disconnected, determining whether there is a shared computing engine corresponding to the client; If it is determined that there is no shared computing engine corresponding to the client, the first computing engine is destroyed.
6. The method according to claim 2, characterized in that The determining the service instance corresponding to the computing engine type from the preset service instance list includes: According to preset filtering conditions, determine the service instance corresponding to the computing engine type from the preset service instance list; wherein the preset filtering conditions include at least one of the following: load balancing, the status of the server to which the service instance belongs, and the geographic location information of the server and the client.
7. The method according to claim 1, wherein: Before the data query request is processed by the first computing engine corresponding to the service instance to obtain the data query result, the method further includes: Authentication of the client is performed to obtain an authentication result; When the identity authentication result indicates that the client has passed the identity authentication, a session connection with the service instance is established.
8. The method according to claim 7, wherein: The performing identity authentication on the client to obtain an identity authentication result includes: intercepting and parsing the data query request through the service instance to obtain verification information of the data query request; In a case where the verification information satisfies the verification condition, an identity verification result indicating that the client has passed the identity verification is determined.
9. The method according to claim 8, wherein: The verification condition includes at least one of the following: a reject exception condition, a reject condition, an allow exception condition, and an allow condition, and the priority of the verification condition from high to low is the reject exception condition, the reject condition, the allow exception condition, and the allow condition; After obtaining the identity authentication information of the data query request, the method further includes: In a case where the verification condition includes at least two verification conditions, verifying the verification information according to the priority order of the verification conditions; In a case where the verification information satisfies each of the at least two verification conditions, an identity verification result indicating that the client has passed the identity verification is determined.
10. A data processing system comprising a front-end server, a session manager and a back-end server; wherein: The front-end server is used to obtain the query request from the client; The session manager is used to determine the computing engine type corresponding to the client according to the data query request; The backend server is used to process the data query request through the first computing engine corresponding to the service instance to obtain the data query result when a session connection has been established with the service instance corresponding to the computing engine type; The front-end server is also used to send the data query result to the client.
11. A data processing device, comprising: The acquisition module is used to obtain the data query request from the client; A determination module, configured to determine a computing engine type corresponding to the client according to the data query request; a processing module, configured to, when a session connection has been established with a service instance corresponding to the computing engine type, process the data query request through a first computing engine corresponding to the service instance to obtain a data query result; A sending module is used to send the data query result to the client.
12. An electronic device, comprising: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the steps of the data processing method according to any one of claims 1 to 9 are implemented.
13. A storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the steps of the data processing method according to any one of claims 1 to 9.
14. A computer program product, characterized in that The program product is stored in a storage medium, and the program product is executed by at least one processor to implement the steps of the data processing method according to any one of claims 1 to 9.