Big data job starting method and device, electronic equipment and storage medium

By pre-downloading dependent packages from object storage to the network file system in a cloud native big data environment and starting these dependent packages when job submission is solved, the problems of high bandwidth throughput and slow job startup speed caused by dependent package pulling are achieved, and more efficient job startup is achieved.

CN120144209APending Publication Date: 2025-06-13BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510286569.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In a cloud-native big data environment, relying on package pulling leads to high bandwidth throughput of object storage, which in turn leads to the problem of slow job startup speed.

Method used

By obtaining the configuration information of the target job, parsing the specified directory of the dependency package, downloading the dependency package from the object store and storing it in the network file system. When submitting the job, start the dependency package in the network file system.

Benefits of technology

Effectively reduces the overall bandwidth throughput and access to QPS caused by dependent packet pulling, and improves job startup speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144209A_ABST
    Figure CN120144209A_ABST
Patent Text Reader

Abstract

The invention provides a big data job starting method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining target job configuration information of a target job; under the condition that the target job configuration information is analyzed to obtain a dependency package specified directory, downloading a target dependency package pointed by the dependency package specified directory from the object storage according to the dependency package specified directory, and storing the target dependency package in a network file system; and when the target job is submitted, starting the target dependent package in the network file system. According to the method and the device, the target dependent packet can be directly pulled from the network file system as long as the target file with the job configuration information as the target job configuration information is submitted, so that the technical effect of effectively reducing the overall bandwidth throughput and related access QPS caused by pulling of the dependent packet is achieved; therefore, the problem of low operation starting speed caused by high bandwidth throughput of object storage due to dependent packet pulling in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data technology, and in particular to a method and device for starting a big data job, an electronic device, and a storage medium. Background Art

[0002] At present, in the past few years of the development of cloud-native big data architecture, more and more companies have migrated big data clusters to cloud-native environments to achieve the effect of reducing costs and increasing efficiency. In the past, big data jobs (Spark, Flink, etc.) were run in the Hadoop environment (an open source distributed computing framework designed to handle the storage and processing of large-scale data sets (big data)). Now, big data jobs are run in cloud-native K8s.

[0003] During the submission and running of a big data job, there are often some dependent packages related to the job itself. The main function of these dependent packages is to assist the execution of the internal logic of the job and to avoid problems with the job due to the lack of certain system dependencies in a relatively clean operating system.

[0004] In a Hadoop big data cluster, dependent files are usually placed in the distributed file storage HDFS, so that when a job is submitted, it can be read uniformly from HDFS. However, in cloud-native big data, if a distributed HDFS system is deployed separately, it will bring great difficulties and unnecessary complexity to the operation and maintenance of the entire cluster. In addition, in the cloud-native design architecture, if you have to rely on a single independent system to supplement capabilities, it is actually not in line with the original intention of cloud-native design.

[0005] Therefore, in the related art, the dependent files are directly stored in the object storage, and then when the job is running, the address of the object storage is specified, so that the relevant dependent packages can be pulled from the object storage during the task operation. However, using this method will result in high-concurrency jobs, such as a Spark job starting thousands of concurrent tasks, which will cause thousands of concurrent tasks to access the object storage system at the same time, resulting in a relatively high overall bandwidth throughput of the object storage, and the related access QPS (Queries Per Second) will also be very high. Furthermore, the storage of dependent packages in the object storage will lead to an increase in the object storage QPS, and the bandwidth throughput of the object storage is high, which will lead to the problem of slower job startup.

[0006] This shows that there is a technical problem in the related technology that the bandwidth throughput of object storage is high due to reliance on package pulling, resulting in slow job startup. Summary of the invention

[0007] This application provides a method and apparatus for starting a big data job, an electronic device, and a storage medium, so as to at least solve the problem in the related art that the job startup speed is slow due to high bandwidth throughput of object storage caused by dependency package pulling.

[0008] According to one aspect of the embodiments of the present application, a method for starting a big data job is provided, including:

[0009] Obtain the target job configuration information of the target job;

[0010] When parsing the target job configuration information to obtain the specified directory of the dependency package, download the target dependency package pointed to by the specified directory of the dependency package from the object storage according to the specified directory of the dependency package, and store the target dependency package in the network file system;

[0011] When submitting the target job, start the target dependency package in the network file system.

[0012] Optionally, in the method as described above, the parsing of the target job configuration information to obtain the specified directory of the dependency package includes:

[0013] Parse the target job configuration information, and when a target field is parsed, determine the field value of the target field;

[0014] Determine the field value as the specified directory of the dependency package.

[0015] Optionally, in the method as described above, the starting of the target dependency package in the network file system when submitting the target job includes:

[0016] When submitting the target job, modify the specified directory of the dependency package in the target job configuration information of the target job to the local directory of the target dependency package in the network file system;

[0017] Start the target dependency package in the network file system according to the local directory.

[0018] Optionally, in the method as described above, the downloading of the target dependency package pointed to by the specified directory of the dependency package from the object storage according to the specified directory of the dependency package includes:

[0019] Obtain whether the first data feature value of the target dependency package in the object storage is consistent with the second data feature value of the existing dependency package in the network file system, where the existing dependency package is the dependency package downloaded from the specified directory of the dependency package before;

[0020] If the first data feature value is inconsistent with the second data feature value, download the target dependency package from the object storage and delete the existing dependency package.

[0021] Optionally, as in the foregoing method, the method further includes:

[0022] If the first data feature value is consistent with the second data feature value, do not download the target dependency package from the object storage.

[0023] Optionally, as in the foregoing method, storing the target dependency package in the network file system includes:

[0024] Determine the working space of the target job;

[0025] Generate a local directory in the network file system according to the working space and the dependency package specified directory;

[0026] Store the target dependency package in the local directory.

[0027] Optionally, as in the foregoing method, the method further includes:

[0028] Obtain the lifecycle of each historical dependency package in the network file system, where the lifecycle is used to indicate the duration from the most recent use of the corresponding historical dependency package to the current moment;

[0029] If it is determined that there is a target historical dependency package with a lifecycle longer than a preset duration among all historical dependency packages, delete the target historical dependency package.

[0030] According to another aspect of the embodiments of the present application, there is also provided a big data job startup device, including:

[0031] An acquisition module, configured to acquire target job configuration information of a target job;

[0032] A download module, configured to, when parsing the target job configuration information to obtain a dependency package specified directory, download a target dependency package pointed to by the dependency package specified directory from an object storage according to the dependency package specified directory, and store the target dependency package in a network file system;

[0033] A startup module, configured to start the target dependency package in the network file system when submitting the target job.

[0034] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. Among them, the memory is used to store a computer program. The processor is used to execute the method steps in any of the above embodiments by running the computer program stored on the memory.

[0035] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. A computer program is stored in the storage medium. Among them, the computer program is set to execute the method steps in any of the above embodiments when running.

[0036] In the embodiments of the present application, a method is adopted in which the dependent package is pre-downloaded from the object storage to the network file system. By obtaining the target job configuration information of the target job; when parsing the target job configuration information to obtain the specified directory of the dependent package, according to the specified directory of the dependent package, the target dependent package pointed to by the specified directory of the dependent package is downloaded from the object storage and stored in the network file system; when submitting the target job, the target dependent package in the network file system is started. Since the files in the network file system can be shared, it can be realized that as long as the target file with the job configuration information being the target job configuration information is submitted, the target dependent package can be directly pulled from the network file system without repeatedly downloading the target dependent package from the object storage, achieving the technical effect of effectively reducing the overall bandwidth throughput and related access QPS caused by pulling the dependent package, and further solving the problem in the related technology that the bandwidth throughput of the object storage is high due to pulling the dependent package, resulting in a slow job startup speed. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application and used together with the description to explain the principles of the present application.

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0039] Figure 1 is a schematic diagram of the hardware environment of an optional big data job startup method according to the embodiments of the present application;

[0040] Figure 2 is a schematic flowchart of an optional big data job startup method according to the embodiments of the present application;

[0041] Figure 3 It is a schematic flowchart of another alternative method for starting a big data job according to an embodiment of the present application;

[0042] Figure 4 It is a schematic flowchart of another alternative method for starting a big data job according to an embodiment of the present application;

[0043] Figure 5 It is a schematic flowchart of another alternative method for starting a big data job according to an embodiment of the present application;

[0044] Figure 6 It is a block diagram of the system architecture for implementing an alternative method for starting a big data job according to an embodiment of the present application;

[0045] Figure 7 It is a block diagram of an alternative big data job startup device according to an embodiment of the present application;

[0046] Figure 8 It is a block diagram of an alternative electronic device according to an embodiment of the present application. Detailed implementation manners

[0047] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0049] First, some nouns or terms that appear in the process of describing the embodiments of the present application are applicable to the following explanations:

[0050] 1. Dependency Package: Refers to the external libraries, frameworks, or toolkits that are necessary for an application to run. These dependency packages provide additional functional support, enabling developers to focus on the core business logic without having to implement all functions from scratch.

[0051] 2. A job refers to an independent unit of work submitted by a user to a cluster, which contains a series of tasks or operations to be executed. It is the most basic execution unit in a distributed computing environment, ensuring the consistency and reliability of data processing.

[0052] According to one aspect of the embodiments of the present application, a method for starting a big data job is provided. As an alternative embodiment, in this embodiment, the above-mentioned method for starting a big data job can be applied to a hardware environment composed of a terminal 1402 and a server 1404 as shown in Figure 1 the figure. As shown in Figure 1 the figure, the server 1404 is connected to the terminal 1402 through a network and can be used to provide services (such as game services, application services, etc.) for the terminal or the client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 1404.

[0053] The above network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal is not limited to a PC, a mobile phone, a tablet computer, etc.

[0054] The method for starting a big data job in the embodiments of the present application can be executed by the server, or by the terminal, or jointly by the server and the terminal. Among them, when the terminal executes the method for starting a big data job in the embodiments of the present application, it can also be executed by the client installed on it.

[0055] Taking the execution of the method for starting a big data job in this embodiment by the server as an example, Figure 2 a method for starting a big data job provided by the embodiments of the present application includes the following steps:

[0056] Step S202, obtaining the target job configuration information of the target job.

[0057] The method for starting a big data job in this embodiment can be applied to scenarios where a big data job is started under a cloud-native big data architecture, such as scenarios of data analysis or business intelligence jobs, scenarios of machine learning and artificial intelligence jobs, etc., or other job scenarios, which are not listed one by one here.

[0058] Specifically, the target job can be a type of job corresponding to a specific workspace and with job configuration information being the target job configuration information.

[0059] The target job configuration information can be obtained based on the job configuration information of the target job obtained for the first time.

[0060] Optionally, it can be implemented by the system architecture as shown in Figure 6 The big data job startup method in this embodiment can be implemented. The spark-tools (i.e., a spark submission program) can be customized in advance to upload the user's spark script (i.e., containing the target job configuration information) and dependency packages to the application server, that is, the spark-service, through the spark-tools.

[0061] Step S204, when parsing the target job configuration information to obtain the directory specified by the dependency package, download the target dependency package pointed to by the directory specified by the dependency package from the object storage according to the directory specified by the dependency package, and store the target dependency package in the network file system.

[0062] Specifically, in this embodiment, the spark-service can be used to parse the submitted target job configuration information. Generally, the directory corresponding to the dependency package corresponds to a specific configuration item. Therefore, after parsing the corresponding configuration item, the directory specified by the dependency package can be obtained.

[0063] The directory specified by the dependency package is the directory of the dependency package retrieved by the target job in the object storage. Furthermore, the target dependency package pointed to by the directory specified by the dependency package can be downloaded from the object storage according to this directory specified by the dependency package and stored in the network file system (i.e., the NFS service). Then, the target dependency package in this network file system can be submitted to the spark deployment environment for running.

[0064] Step S206, when submitting the target job, start the target dependency package in the network file system.

[0065] When it is determined to submit the target job, there is no need to pull the target dependency package from the object service anymore. Just directly start the target dependency package from the network file system. Furthermore, even if multiple target jobs are started, it is only necessary to download the target dependency package from the object storage once, without multiple downloads.

[0066] In this embodiment, the dependent package is pre-downloaded from the object storage to the network file system by obtaining the target job configuration information of the target job; when the target job configuration information is parsed to obtain the specified directory of the dependent package, the target dependent package pointed to by the specified directory of the dependent package is downloaded from the object storage according to the specified directory of the dependent package, and the target dependent package is stored in the network file system; when the target job is submitted, the target dependent package in the network file system is started. Since the files in the network file system can be shared, it is possible to achieve the purpose that as long as the target file with the job configuration information being the target job configuration information is submitted, the target dependent package can be directly pulled from the network file system without repeatedly downloading the target dependent package from the object storage, achieving the technical effect of effectively reducing the overall bandwidth throughput and related access QPS caused by the dependent package pulling, and further solving the problem in the related technology that the job startup speed is slow due to the high bandwidth throughput of the object storage caused by the dependent package pulling.

[0067] As Figure 3 shown, as an optional embodiment, for the method as described above, the above-mentioned step S202 of parsing the target job configuration information to obtain the specified directory of the dependent package is implemented through the following steps S302 and S304:

[0068] Step S302: Parse the target job configuration information, and when the target field is parsed, determine the field value of the target field.

[0069] That is to say, after obtaining the target configuration information, by parsing the target job configuration information, the target field needs to be parsed. The target field is pre-agreed and is used to store the field of the directory of the dependent package. Therefore, after the target field is parsed, the field value of the target field can be determined.

[0070] Step S304: Determine the field value as the specified directory of the dependent package.

[0071] After the field value is determined, since it is pre-agreed that the field value of the target field is the directory of the dependent package to be pulled by the target job, the field value can be determined as the specified directory of the dependent package.

[0072] For example, after spark-service parses --archives (i.e., an optional configuration item, that is, the target field) in the target configuration information, it will automatically parse out the value of the --archives configuration as the specified target of the dependent package, so as to facilitate the later download of the target dependent package from the object storage according to the specified directory of the dependent package and save it to the network file service.

[0073] Through the method of this embodiment, the specified directory of the dependent package can be automatically parsed, facilitating the accurate pulling of the target dependent package later.

[0074] As Figure 4 shown, as an alternative embodiment, for the method as described above, the foregoing step S206 of starting the target dependent package in the network file system when submitting the target job can be implemented through the following steps S402 and S404:

[0075] Step S402, when submitting the target job, modify the specified directory of the dependent package in the target job configuration information of the target job to the local directory of the target dependent package in the network file system.

[0076] That is to say, when submitting the target job with the job configuration information being the target job configuration information, the target job configuration information in the target job will be modified, and the originally recorded specified directory of the dependent package will be modified to the local directory of the target dependent package in the network file system.

[0077] Step S404, start the target dependent package in the network file system according to the local directory.

[0078] After determining the local directory, the target dependent package in the network file system can be started according to this local directory.

[0079] As Figure 5 shown, as an alternative embodiment, for the method as described above, the foregoing step S204 of downloading the target dependent package pointed to by the specified directory of the dependent package from the object storage according to the specified directory of the dependent package can be implemented through the following steps S502 to S506:

[0080] Step S502, obtain whether the first data feature value of the target dependent package in the object storage is consistent with the second data feature value of the existing dependent package in the network file system, where the existing dependent package is the dependent package downloaded from the specified directory of the dependent package before.

[0081] That is to say, during the process of synchronizing the dependent package in the object storage to the local, it is necessary to determine whether the dependent package is pre-stored locally. The first data feature value of the target dependent package can be determined by calculating the feature value of the target dependent package, and the feature value of the existing dependent package in the network file system is also calculated in the same way, and the second data feature value is obtained.

[0082] Step S504, if the first data feature value is inconsistent with the second data feature value, download the target dependent package from the object storage and delete the existing dependent package.

[0083] After obtaining the first data eigenvalue and the second data eigenvalue, it is possible to determine whether the target dependent package is consistent with the existing dependent package by judging whether the first data eigenvalue is consistent with the second data eigenvalue. If the first data eigenvalue is not consistent with the second data eigenvalue, the target dependent package is not consistent with the existing dependent package. Furthermore, the target dependent package can be downloaded from the object storage, and the existing dependent package on the local side can be deleted.

[0084] Step S506, if the first data eigenvalue is consistent with the second data eigenvalue, the target dependent package is not downloaded from the object storage.

[0085] If the first data eigenvalue is consistent with the second data eigenvalue, the target dependent package is consistent with the existing dependent package. Furthermore, it is not necessary to download the target dependent package from the object storage.

[0086] For example, in dependent package synchronization, the MD5 value of the target dependent package and the existing dependent package can be calculated to obtain the first MD5 value (i.e., the first data eigenvalue) and the second MD5 value (i.e., the second data eigenvalue) corresponding to the target dependent package and the existing dependent package respectively, and based on the first MD5 value and the second MD5 value, it is judged whether the target dependent package is consistent with the existing dependent package. If they are consistent, the pull operation is no longer executed, avoiding repeated pulling every time a job is submitted and affecting the job submission performance.

[0087] As an optional embodiment, for the method as described above, the target dependent package can be stored in the network file system through the following steps: determine the working space of the target job; generate a local directory in the network file system according to the working space and the dependent package specified directory; store the target dependent package in the local directory. Specifically, the working space of the target job refers to an independent environment or directory for the target job, which is used to store the resources, configuration files, temporary data, and output results required by the job. Therefore, an identifier uniquely corresponding to the working space can be obtained, and then a local directory in the network file system can be generated according to the working space and the dependent package specified directory, so that the local directory can not only indicate the address of the dependent package in the object storage, but also be distinguished locally through the identifier corresponding to the working space, so that the directory of the target dependent package in the object storage is consistent, and at the same time, there will be no file information with the same name conflict. This judgment logic is used to avoid the same name conflict, and when saving in the local directory, it is also strictly distinguished according to the namespace + the address of the object storage directory.

[0088] As an alternative embodiment, like the foregoing method, the method further includes the following steps: obtaining the lifecycle of each historical dependency package in the network file system, where the lifecycle is used to indicate the duration from the most recent use of the corresponding historical dependency package to the current moment; and deleting the target historical dependency package when it is determined that there is a target historical dependency package with a lifecycle longer than a preset duration among all the historical dependency packages. That is to say, the historical dependency packages stored in the network file system may correspond to dependency package information, which is used to maintain the lifecycle of a corresponding historical dependency package, and regularly clean up related dependency packages to avoid excessive dependency package files occupying the space of the local disk for a long time. For example, when there is a target historical dependency package with a lifecycle longer than a preset duration (e.g., the default is 7 days) among all the historical dependency packages, the target historical dependency package is cleaned up.

[0089] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of this application.

[0091] According to another aspect of the embodiments of this application, there is also provided a big data job startup device for implementing the above big data job startup method. Figure 7 is a structural block diagram of an alternative big data job startup device according to the embodiments of this application, as Figure 7 shown. The device may include:

[0092] An obtaining module 71, configured to obtain the target job configuration information of the target job;

[0093] A download module 72, configured to, when parsing the target job configuration information to obtain the specified directory of the dependency package, download from the object storage the target dependency package pointed to by the specified directory of the dependency package according to the specified directory of the dependency package, and store the target dependency package in the network file system;

[0094] A start module 73, configured to start the target dependency package in the network file system when submitting the target job.

[0095] It should be noted that the obtaining module 71 in this embodiment may be used to execute the above step S202, the download module 72 in this embodiment may be used to execute the above step S204, and the start module 73 in this embodiment may be used to execute the above step S206.

[0096] Through the above modules, by adopting the method of pre-downloading the dependency package from the object storage to the network file system, the target job configuration information of the target job is obtained; when parsing the target job configuration information to obtain the specified directory of the dependency package, according to the specified directory of the dependency package, download from the object storage the target dependency package pointed to by the specified directory of the dependency package, and store the target dependency package in the network file system; when submitting the target job, start the target dependency package in the network file system. Since the files in the network file system can be shared, it can be realized that as long as the target file with the job configuration information being the target job configuration information is submitted, the target dependency package can be directly pulled from the network file system without repeatedly downloading the target dependency package from the object storage, achieving the technical effect of effectively reducing the overall bandwidth throughput and related access QPS caused by the pulling of the dependency package, and further solving the problem in the related technology that the slow job startup speed is caused by the high bandwidth throughput of the object storage due to the pulling of the dependency package.

[0097] The device in this embodiment, in addition to including the above modules, may further include modules for executing any method in any of the foregoing embodiments of the big data job startup method.

[0098] It should be noted here that the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run in a hardware environment as shown in Figure 1 and can be implemented by software or by hardware, where the hardware environment includes a network environment.

[0099] According to another aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above big data job startup method, and the electronic device may be a server, a terminal, or a combination thereof.

[0100] According to another embodiment of the present application, an electronic device is further provided, including: as Figure 8 shown, the electronic device may include: a processor 1501, a communication interface 1502, a memory 1503, and a communication bus 1504. Among them, the processor 1501, the communication interface 1502, and the memory 1503 complete mutual communication through the communication bus 1504.

[0101] The memory 1503 is used to store a computer program;

[0102] When the processor 1501 executes the program stored on the memory 1503, the following steps are implemented:

[0103] Step S202, obtaining the target job configuration information of the target job.

[0104] Step S204, when parsing the target job configuration information to obtain the specified directory of the dependent package, downloading the target dependent package pointed to by the specified directory of the dependent package from the object storage according to the specified directory of the dependent package, and storing the target dependent package in the network file system;

[0105] Step S206, when submitting the target job, starting the target dependent package in the network file system.

[0106] Optionally, in this embodiment, the above-mentioned communication bus may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the above-mentioned electronic device and other devices.

[0107] The memory may include a random access memory (RAM), and may also include a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0108] As an example, the above-mentioned memory 1503 may but is not limited to include the acquisition module 61, the download module 62, and the start module 63 in the above-mentioned big data job start device. In addition, it may also include but is not limited to other module units in the above-mentioned big data job start device, which will not be elaborated in this example.

[0109] The above-mentioned processor can be a general-purpose processor, including but not limited to: CPU (Central Processing Unit, central processing unit), NP (Network Processor, network processor), etc.; it can also be a DSP (Digital Signal Processor, digital signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA (Field-Programmable Gate Array, field-programmable gate array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0110] The embodiment of the present application also provides a computer-readable storage medium. The storage medium includes a stored program. When the program runs, it executes the method steps of the above-mentioned method embodiment.

[0111] Optionally, in this embodiment, the above-mentioned storage medium can include but not limited to: various media that can store program codes such as USB flash drives, ROMs, RAMs, mobile hard disks, magnetic disks, or optical discs.

[0112] The serial numbers of the above-mentioned embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0113] If the integrated unit in the above-mentioned embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above-mentioned computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in the storage medium and includes several instructions for causing one or more computer devices (which can be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application.

[0114] In the above-mentioned embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0115] In several embodiments provided by this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in electrical or other forms.

[0116] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution provided in this embodiment.

[0117] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0118] The above is only the preferred embodiment of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A method for starting a big data job, characterized in that: include: Get the target job configuration information of the target job; When the target job configuration information is parsed to obtain a specified directory of a dependency package, the target dependency package pointed to by the specified directory of the dependency package is downloaded from the object storage according to the specified directory of the dependency package, and the target dependency package is stored in a network file system; When the target job is submitted, the target dependent package in the network file system is started.

2. The method according to claim 1, characterized in that The target job configuration information is parsed to obtain a specified directory of a dependent package, including: Parsing the target job configuration information, and determining a field value of the target field when a target field is obtained through the parsing; The field value is determined as the specified directory of the dependent package.

3. The method according to claim 1, characterized in that When submitting the target job, starting the target dependency package in the network file system includes: When submitting the target job, modifying the dependency package specified directory in the target job configuration information of the target job to a local directory of the target dependency package in the network file system; According to the local directory, the target dependent package in the network file system is started.

4. The method according to claim 1, characterized in that: The step of downloading the target dependency package pointed to by the specified directory of the dependency package from the object storage according to the specified directory of the dependency package includes: Obtaining whether a first data characteristic value of a target dependency package in the object storage is consistent with a second data characteristic value of an existing dependency package in the network file system, wherein the existing dependency package is a dependency package previously downloaded from a specified directory of the dependency package; If the first data characteristic value is inconsistent with the second data characteristic value, the target dependency package is downloaded from the object storage, and the existing dependency package is deleted.

5. The method according to claim 4, characterized in that The method further comprises: If the first data characteristic value is consistent with the second data characteristic value, the target dependency package is not downloaded from the object storage.

6. The method according to claim 1, characterized in that The storing the target dependent package in a network file system includes: Determining a workspace for the target operation; Generate a local directory in the network file system according to the workspace and the specified directory of the dependent package; The target dependent package is stored in the local directory.

7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: Obtaining the life cycle of each historical dependency package in the network file system, wherein the life cycle is used to indicate the length of time from the last use of the corresponding historical dependency package to the current moment; When it is determined that there is a target historical dependency package whose life cycle is longer than the preset time period among all historical dependency packages, the target historical dependency package is deleted.

8. A big data operation startup device, characterized in that: include: An acquisition module, used to acquire target job configuration information of a target job; A download module, for parsing the target job configuration information to obtain a specified directory of a dependency package, downloading a target dependency package pointed to by the specified directory of the dependency package from the object storage according to the specified directory of the dependency package, and storing the target dependency package in a network file system; A startup module is used to start the target dependent package in the network file system when the target job is submitted.

9. An electronic device comprising a processor, a communication interface, a memory and a communication bus, wherein: The processor, the communication interface and the memory communicate with each other via the communication bus, wherein: The memory is used to store computer programs; The processor is configured to execute the method according to any one of claims 1 to 7 by running the computer program stored in the memory.

10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when executed.