A Method for Realizing Remote Debugging of Big Data Development Based on K8S Technology

Remote debugging of big data development through K8S technology solves the cumbersome debugging process caused by the need for security tool jumps in existing technology, improves efficiency and security, supports the use of a variety of development tools, and achieves efficient remote debugging and environmental isolation.

CN114741280BActive Publication Date: 2025-07-25KEDADUOCHUANG CLOUD NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210301491.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-24
Publication Date
2025-07-25
Estimated Expiration
2042-03-24

AI Technical Summary

Technical Problem

During the big data development process, developers' code needs to be redirected through security tools such as VPN and bastion machines before it can be deployed on the big data platform, resulting in cumbersome debugging process and inefficient efficiency.

Method used

Remote debugging is realized through K8S technology, using the remote debugging service management platform and Kerberos authentication, combining the K8s cluster and big data platform, achieving seamless connection between local development tools and big data platforms, using remote debugging tools for code development and debugging, and using Kerberos tickets for user authentication and resource isolation.

Benefits of technology

It improves the debugging efficiency of developers, reduces operating costs, improves business response speed, and ensures security through environmental isolation, supporting the compatibility and expansion of multiple development tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114741280B_ABST
    Figure CN114741280B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for realizing remote debugging of big data development based on K8S technology, belonging to the technical field of big data development, including: S1: Deployment and configuration; S2: Service application; S3: Service startup; S4: Service connection; S5: Development and debugging; S6: Iterative optimization. On the premise of ensuring the security of the big data platform, the present invention integrates the complex debugging process with the development process through cloud native technology, improves the debugging efficiency of developers, reduces the operation cost, and improves the business response speed; uses the remote debugging tool of the development tool, combines the actual situation of k8s and the big data environment, implements the use of the remote debugging tool in the actual scenario, promotes the use scenario of the tool, and improves the labor productivity at the same time; remotely debugs and connects to the big data platform environment, but connects through the intermediate environment client, and each remote debugging pod has its own unique authentication ticket, realizing the isolation of user data and resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of big data development, and particularly relates to a method for realizing remote debugging of big data development based on K8S technology. Background Art

[0002] The enterprise big data platform aggregates important assets of the enterprise. It is in the core data area and belongs to the isolation protection level on the network, and direct connection to the public network is not allowed. During the big data development process, the code developed by developers needs to be packaged and can only be deployed on the big data platform after jumping through security tools such as vpn and bastion hosts. The process of repeatedly modifying the code and uploading it to the big data platform for debugging is cumbersome and the debugging efficiency is low. For this reason, a method for realizing remote debugging of big data development based on K8S technology is proposed. Summary of the Invention

[0003] The technical problem to be solved by the present invention is: how to solve the problem that in the existing big data development process, the code developed by developers needs to be packaged and can only be deployed on the big data platform after jumping through security tools such as vpn and bastion hosts, and the process of repeatedly modifying the code and uploading it to the big data platform for debugging is cumbersome and the debugging efficiency is low. A method for realizing remote debugging of big data development based on K8S technology is provided. This method organically connects the local development tool and the big data platform, realizing the convenience of local development while eliminating the pain points of remote debugging of the big data platform.

[0004] The present invention solves the above technical problem through the following technical solutions. The present invention includes the following steps:

[0005] S1: Deployment and Configuration

[0006] Deploy the remote debugging service management platform and the k8s cluster, and configure the development tool on the user's computer;

[0007] S2: Service Application

[0008] The user submits an application for the remote debugging service to the remote debugging service management platform, and the remote debugging service management platform performs user information creation, resource allocation, and ticket generation;

[0009] S3: Service Startup

[0010] After the user submits the service application, the remote debugging service management platform starts the service, performs big data development client environment configuration, user ticket loading, and deploys Pod instances;

[0011] S4: Service Connection

[0012] User remote debugging service connection: through the service address information, local development tools can remotely connect to the big data development and debugging server;

[0013] S5: Development and debugging

[0014] Users develop and debug codes in local development tools, including code development, task submission, and log viewing.

[0015] S6: Iterative Optimization

[0016] Users perform iterative code optimization, analyze based on exceptions, modify the code logic and re-run it until the correct running results are obtained.

[0017] Furthermore, in step S1, the specific process of deploying the remote debugging service management platform and the k8s cluster includes the following steps:

[0018] S101: Build a k8s cluster and connect the network environment of the k8s cluster and the big data platform;

[0019] S102: Create a basic image, which includes client configuration that can connect to the big data platform environment, including HDFS, YARN, and Spark client environment parameter configurations, so that the client can remotely connect to the big data platform;

[0020] S103: A remote connection server tool is configured in the image so that the local development tool can be connected to the running container environment.

[0021] Furthermore, in step S1, the specific process of configuring the user computer development tool includes the following steps:

[0022] S201: Configure local development tools, install the Spark development environment, and enable the ability to write Spark programs;

[0023] S202: Install the ssh remote connection plug-in to enable the ability to connect to the remote environment of the k8s container to achieve remote debugging;

[0024] S203: Configure the IP, port, user name, and password of the remote SSH environment, save them locally, and implement the ability to connect to the remote environment under the default configuration.

[0025] Furthermore, in the step S2, the process of creating user information is as follows: when a user registers, the remote debugging service management platform generates user information, uses the kerberors user system, and generates corresponding user accounts on each host in the big data platform and the k8s cluster.

[0026] Furthermore, in the step S2, the resource allocation includes user storage directory generation, user computing resource allocation, and data permission allocation.

[0027] Furthermore, in the step S2, when generating the ticket, the remote debugging service management platform generates a kerberos ticket for use when the user logs in to the remote debugging service management platform, serving as a credential for user authentication.

[0028] Furthermore, in the step S3, the specific process of configuring the big data development client environment includes the following steps:

[0029] S301: Download the base image;

[0030] S302: Download the deployment file corresponding to the user;

[0031] S303: After the pod instance starts, load the user ticket and configure the environment variables;

[0032] S304: Complete the environment preparation for the big data development remote debugging client that only serves this user.

[0033] Furthermore, in the step S3, the specific process of loading the user ticket includes the following steps:

[0034] S311: According to the user identification, download the ticket file corresponding to the user in the Pod instance;

[0035] S312: Execute the user login command to implement the authentication operation of the user to the big data platform.

[0036] Furthermore, in the step S3, the specific process of deploying the Pod instance includes the following steps:

[0037] S321: Generate different deployment configuration files according to different users;

[0038] S322: Submit the deployment file to the k8s container management platform;

[0039] S323: Deploy and start the pod instance according to the deployment file.

[0040] Furthermore, in the step S322, the k8s container management platform is the basic platform for downloading the base image, configuring the operation configuration required for pod instance initialization, and starting, stopping, or managing the pod instance. The remote debugging service management platform manages the remote debugging service through this platform.

[0041] The present invention has the following advantages compared with the prior art: on the premise of ensuring the security of the big data platform, the complex debugging process and the development process are integrated through cloud-native technology, improving the debugging efficiency of developers, reducing the operation cost, and enhancing the business response speed; by using the remote debugging tool of the development tool and combining the actual situation of the k8s and big data environments, the use of the remote debugging tool is implemented in the actual scenario, promoting the use scenario of the tool, and at the same time, due to the convenience of the tool, the labor productivity is enhanced; the remote debugging does not directly connect to the big data platform environment, but is connected through the intermediate environment client, and each remote debugging pod has its own unique authentication ticket, realizing the isolation of user data and resources and effectively ensuring the production security; the development tool is not limited to VSCode, and more tools can be encapsulated in the remote debugging service, so that more scenarios can be realized, and it can be well compatible and extended with other development tools; only one set of k8s resources and application software is required to complete the management of the remote debugging service for big data development. The solution is flexible and non-invasive, and can be seamlessly docked with the current situation. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 FIG. 6 is a schematic diagram of the system architecture for the k8s cluster to implement the management of the remote debugging service in Embodiment 1 of the present invention;

[0043] Figure 2 FIG. 10 is a schematic diagram of the implementation process of the remote debugging method for big data development in Embodiment 2 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The embodiments of the present invention will be described in detail below. The following embodiments are implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given. However, the protection scope of the present invention is not limited to the following embodiments.

[0045] As Figure 1 described above, this embodiment mainly realizes the management of the remote debugging service through a k8s cluster. The user customizes the big data remote debugging service and provides the user with the remote debugging for big data development. The main modules include a client group (k8s cluster) and a remote debugging service management platform.

[0046] The specific process of remote debugging for big data development is as follows:

[0047] (1) When the user registers, the remote debugging service management platform generates user information, uses the Kerberos user system, and at the same time generates corresponding user accounts on each host of the big data platform and the client group.

[0048] Kerberos is a network authentication protocol, whose design goal is to provide a strong authentication service for client / server applications through a key system. The implementation of this authentication process does not depend on the authentication of the host operating system, does not require trust based on the host address, does not require the physical security of all hosts on the network, and assumes that data packets transmitted on the network can be read, modified, and data inserted arbitrarily. In the above situations, Kerberos, as a trusted third-party authentication service, performs the authentication service through traditional cryptographic techniques (such as: shared keys).

[0049] (2) Use the remote debugging service management platform to generate Kerberos tickets for use when the user logs in to the remote debugging service management platform, serving as the credential for user authentication.

[0050] The Kerberos server provides an API that can manage users. The remote debugging service management platform connects to the Kerberos management terminal through the management account to generate user Kerberos tickets.

[0051] (3) The remote debugging service management platform allocates resources to users, including generating user storage directories, allocating user computing resources, and allocating data permissions.

[0052] (4) Generate a big data platform client base image on the remote debugging service management platform, including installing clients of common tool components such as hdfs, hbase, hive, and spark, and at the same time install the necessary software for the VSCode server (ssh remote debugging server program); start the image, create a user account, set a password, create a user directory, load the user Kerberos ticket, start the sshd service, and generate a pod to run.

[0053] (5) Automatically proxy the ssh service to the load balancer and expose it to the outside world.

[0054] The interface design of the big data development remote debugging service is as follows:

[0055] Table 1 Request Parameter Table

[0056] Serial number Encoding Name Field type Remarks 1 name Task name string 2 controllerType Service type int 3 image Base image name string 4 vpu Number of CPU cores int 5 memory Memory size int 6 port Port int 7 teplicas Number of replicas int 8 tesetid Namespace string 9 user User name string 10 passwd Password string

[0057] Table 2 Response Parameter Table

[0058]

[0059]

[0060] (6) The user obtains the IP and port, username, and password of the big data development remote debugging service.

[0061] (7) The user opens the VSCode development tool on the local machine, installs the Remote-SSH extension, and enters the IP, port, username, and password provided by the big data development remote debugging service.

[0062] (8) The user connects to the big data development remote debugging service pod and imports the Spark project.

[0063] (9) Use a script to automatically compile, package, and run the project, and submit tasks using the yarn-cluster mode.

[0064] (10) The log of the task running can be viewed on the console, and the code is iteratively modified based on the log.

[0065] It should be noted that in this embodiment, pod refers to the running instance in the k8s cluster; spark, an engine for big data computing; yarn-cluster, a submission mode for big data computing tasks; ssh, a protocol service for remotely connecting to a host for interactive operations.

[0066] In this embodiment, the big data platform is a big data storage and computing cluster composed of different numbers of computers for different users to use for data processing. At the same time, kerberors account management is provided to achieve user security authentication, different permissions for different users, and resource isolation.

[0067] Embodiment 2

[0068] As Figure 2 shown, the implementation process of the present invention is as follows:

[0069] (1) The remote debugging service management platform and the k8s cluster are deployed, and the development tool on the user's computer is configured.

[0070] (2) The user applies for the service, submits an application for the remote debugging service to the remote debugging service management platform, and the remote debugging service management platform creates user information, allocates resources, generates tickets, etc.

[0071] (3) The user service is started. After the user submits the service application, the remote debugging service management platform starts the service, including configuring the big data development client environment, loading the user ticket, and generating the Pod instance.

[0072] (4) The user connects to the remote debugging service. Through the service address information, the local development tool remotely connects to the big data development debugging server.

[0073] (5) The user develops and debugs the code, including code development, task submission, log viewing, etc.

[0074] (6) Iteratively optimize the user code, analyze it based on exceptions, modify the code logic and re-run it until the correct running result is obtained.

[0075] In summary, the method for realizing remote debugging of big data development based on the K8S technology in the above embodiments first realizes the what-you-see-is-what-you-get coding method. The code developed locally is saved in the big data platform in real time, and the code can be updated and iterated in real time. Secondly, code testing can be performed on the local development tool client, and the running logs can be output, and problems can be located according to the logs. In addition, different users are completely isolated from the environment, and data security is effectively guaranteed. This method efficiently supports the big data development process while ensuring security, greatly improves the efficiency of developers, and enables agile business response, and is worthy of being promoted and used.

[0076] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for realizing remote debugging of big data development based on K8S technology, characterized in that, It includes: S1: Deployment and Configuration Deploy the remote debugging service management platform and the k8s cluster, and configure the development tools on the user's computer; In step S1, the specific process of deploying the remote debugging service management platform and the k8s cluster includes the following steps: S101: Set up the k8s cluster and connect the network environment of the k8s cluster to the big data platform; S102: Create a base image, which includes the client configuration for connecting to the big data platform environment, specifically the configuration of hdfs, yarn, and spark client environment parameters; S103: Configure the remote connection server tool in the base image; S2: Service Application The user submits an application for the remote debugging service to the remote debugging service management platform, and the remote debugging service management platform performs user information creation, resource allocation, and ticket generation; In step S2, the process of creating user information is as follows: when the user registers, the remote debugging service management platform generates user information, uses the kerberors user system, and at the same time generates corresponding user accounts on each host in the big data platform and the k8s cluster; In step S2, resource allocation includes user storage directory generation, user computing resource allocation, and data permission allocation; In step S2, when generating the ticket, the remote debugging service management platform generates the kerberors ticket for the user to use when logging in to the remote debugging service management platform, as a user authentication credential; S3: Service Startup After the user submits the service application, the remote debugging service management platform starts the service, performs big data development client environment configuration, user ticket loading, and deploys Pod instances; In step S3, the specific process of big data development client environment configuration includes the following steps: S301: Download the base image; S302: Download the deployment file corresponding to the user; S303: After the pod instance starts, load the user ticket and configure the environment variables; S304: Complete the environment preparation for the big data development remote debugging client that only serves this user; In step S3, the specific process of user ticket loading includes the following steps: S311: Download the ticket file corresponding to the user in the Pod instance according to the user identifier; S312: Execute the user login command to implement the user authentication operation to the big data platform; In step S3, the specific process of deploying Pod instances includes the following steps: S321: Generate different deployment configuration files according to different users; S322: Submit the deployment file to the k8s container management platform; S323: Deploy and start the pod instance according to the deployment file; S4: Service Connection The user connects to the remote debugging service and remotely connects the local development tool to the big data development debugging server through the service address information; S5: Development and Debugging The user performs code development and debugging in the local development tool, including code development, task submission, and log viewing; S6: Iterative Optimization The user performs code iterative optimization, analyzes based on exceptions, modifies the code logic and runs it again until the correct running result is obtained.

2. The method for realizing remote debugging of big data development based on K8S technology according to claim 1, wherein: In step S1, the specific process of configuring the user computer development tool includes the following steps: S201: Configure the local development tool and install the spark development environment; S202: Install the ssh remote connection plugin to enable the ability to connect to the k8s container remote environment; S203: Configure the ip, port, username, and password of the remote ssh environment and save them locally.

3. The method for remotely debugging big data development based on K8S technology according to claim 2, characterized in that: In the step S322, the k8s container management platform is the basic platform for downloading the basic image, configuring the operation configuration required for the initialization of the pod instance, and starting, stopping, or managing the pod instance. The remote debugging service management platform manages the remote debugging service through this platform.