A Big Data Offline Processing Method and System Based on Multiple Computing Engines

The method and system leverage Akka's distributed computing framework to enhance the efficiency and reliability of big data offline processing by creating independent computing units, addressing computational efficiency challenges in existing frameworks.

CN117076559BActive Publication Date: 2025-07-15XIAMEN MEIYA PICO INFORMATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310875339.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-17
Publication Date
2025-07-15
Estimated Expiration
2043-07-17

AI Technical Summary

Technical Problem

The computing efficiency of existing big data offline computing is not high and it is difficult to meet the practical application needs.

Method used

Using a multi-computing engine-based method, an offline computing system is built using the Akka distributed computing framework. Through the engine actuator control layer interface, job submission actuator, offline computing Actor and other modules, the job processing process is optimized to achieve efficient concurrent and reliable data processing.

Benefits of technology

It improves the efficiency and reliability of offline computing of big data, reduces system maintenance costs, and enhances system flexibility and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117076559B_ABST
    Figure CN117076559B_ABST
Patent Text Reader

Abstract

The present invention proposes a big data offline processing method and system based on multiple computing engines. The method includes the following steps: In response to a job request initiated by a third-party application through the Web side, the job request calls the engine executor control layer interface EECI; the engine job submission executor EJSR polls the independent engine job queue and retrieves the job information JCI and submits it to an instance WSActor of the offline computing engine WebActor; an instance SJEActor of the offline computing engine Actor is started, and a start message is sent to an instance WSActor of the offline computing engine WebActor; at the same time, the job executor JER is started. The job executor JER continuously polls the job queue to check whether there is job information JCI in the job queue. When the job information JCI is polled, the job information JCI is created and submitted to the job executor JERA. The job executor JERA calls a specific engine for processing. By calling the interface of the big data offline processing system based on multiple computing engines during the actual modeling process, the business scenario of big data offline computing can be solved efficiently and reliably.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of big data, and particularly relates to a big data offline processing method and system based on multiple computing engines. Background Art

[0002] With the rapid development and popularization of Internet technology, all walks of life are actively exploring new data processing and analysis methods. The arrival of the big data era has pushed this demand to a brand new height. In the big data application scenario, how to efficiently process and analyze massive data has become a focus issue that companies and institutions are competing to study. As a traditional data processing method, offline computing has also been widely used in the big data application scenario. The distributed computing framework is one of the core technologies of big data offline computing. Currently, the most popular open-source distributed computing frameworks include Hadoop, Spark, etc. These frameworks can effectively improve the efficiency and reliability of big data processing. Although a series of achievements have been made in big data offline computing, there are still some technical challenges in practical applications. One of the more prominent difficulties is the computing efficiency problem. Although there are currently multiple distributed computing frameworks to choose from, how to further improve the computing efficiency of big data offline computing is still an important issue.

[0003] Akka is an excellent distributed computing framework with multiple advantages such as high scalability, high concurrency, high reliability, distributed computing, and reactive programming. Through the design based on the Actor model, Akka can divide the application program into independent computing units. These units can asynchronously receive and send messages to each other and create more Actors when necessary, thus easily expanding to multiple nodes to achieve highly scalable application programs. At the same time, Akka can also achieve efficient concurrent programming, improve the performance and throughput of the application program; ensure the high reliability of the application program through the monitoring and recovery mechanism to avoid problems such as system crashes or data loss; support deploying Actors on different nodes to build complex distributed application programs; and support the reactive programming model to achieve non-blocking IO operations, improve the reaction speed and scalability of the application program. Therefore, adopting Akka to build an offline computing architecture system can improve the data processing efficiency, computing reliability, flexibility and scalability of the system, and at the same time is conducive to reducing the system maintenance cost.

[0004] Akka itself is not a framework specifically for offline computing, but some offline computing functions can be implemented through Akka. In Akka, an Actor can be regarded as an independent computing unit that can perform operations such as processing tasks, saving states, and sending messages. Therefore, Akka distributed Actors can be used to implement the offline computing of big data.

[0005] In view of this, it is very meaningful to propose a big data offline processing method and system based on multiple computing engines. Summary of the Invention

[0006] In order to solve the problem of low computing efficiency in existing big data offline computing, the present invention provides a big data offline processing method and system based on multiple computing engines to solve the above-mentioned existing technical defect problems.

[0007] In a first aspect, the present invention proposes a big data offline processing method based on multiple computing engines, and the method includes the following steps:

[0008] In response to a job request initiated by a third-party application through the Web side, the job request invokes the engine executor control layer interface EECI;

[0009] The engine job submission executor EJSR will poll the independent engine job queue and take out the job information JCI and submit it to an instance WSActor of the offline computing engine WebActor;

[0010] Start an instance SJEActor of the offline computing engine Actor, and send startup information to an instance WSActor of the offline computing engine WebActor;

[0011] At the same time, start the job executor JER. The job executor JER continuously polls the job queue to check if there is job information JCI in the job queue. When the job information JCI is polled, create and submit the job information JCI to the job executor JERA, and the job executor JERA invokes a specific engine for processing.

[0012] Preferably, it further includes:

[0013] After the SJEActor instance is started, it will send startup information to the WSActor instance and create a first timer. After the WSActor instance receives the startup information of the SJEActor instance, it creates a second timer;

[0014] Then send the job information JCI to the SJEActor instance. After the SJEActor instance receives the job information JCI, it closes the second timer, and at the same time stores the job information JCI in the job queue and sends a message confirming the receipt of JCI to the WSActor instance. After the WSActor instance receives the confirmation message, it closes the first timer.

[0015] Further preferably, it further includes:

[0016] After receiving a job request, the Engine Executor Control Layer Interface (EECI) will call an Engine Executor Service Class Interface (EESI).

[0017] It will perform parameter verification on the passed parameters. If the job ID is empty, an exception will be thrown. If the job ID is not empty, the job information (JCI) will be submitted to the Web-side independent engine job queue according to the running mode of the engine.

[0018] Further preferably, it also includes:

[0019] Obtain the corresponding job ID according to the request parameters, and start an instance (WSActor) of the offline computing engine WebActor named this job ID, and then pass in the job information (JCI).

[0020] After receiving the job information, the instance (WSActor) of the offline computing engine WebActor will call the engine management service to construct a start command for an instance (SJEActor) of the offline computing engine Actor.

[0021] Then start the job program command with the IP address of the WSActor instance and the job ID. This program establishes a connection channel with the WSActor instance through the IP address and the job ID.

[0022] Further preferably, it also includes:

[0023] If the first timer or the second timer waits for too long, an exception will be thrown and it will be uniformly persisted to the distributed storage database. The third-party application will regularly pull the corresponding persisted table in the distributed storage database to obtain the relevant request result information.

[0024] Further preferably, when the job executor (JERA) calls a specific engine for processing, it specifically includes:

[0025] The job executor (JERA) will call the initialization engine method, and then call the preprocessing method before job execution once before executing the job.

[0026] After completion, it will call the execution method for starting the job. After the job ends, it will call the cleanup operation method and persist the result to the distributed storage database.

[0027] If an error occurs during job execution, the corresponding exception handling method will be called, and the exception information will be persisted to the distributed storage database.

[0028] Finally, the third-party application will regularly pull the corresponding persisted table in the distributed storage database to obtain the relevant request result information.

[0029] Second aspect, embodiments of the present invention further provide a big data offline processing system based on multiple computing engines, and the system includes:

[0030] A job request module, configured to respond to a job request initiated by a third-party application through the Web side, and the job request invokes an engine executor control layer interface EECI;

[0031] An engine job submission executor EJSR module, configured to poll an independent engine job queue and retrieve job information JCI to submit to an instance WSActor of an offline computing engine WebActor;

[0032] An offline computing engine WebActor module, configured to start an instance SJEActor of an offline computing engine Actor and send startup information to an instance WSActor of the offline computing engine WebActor;

[0033] A job executor JER module, configured to start a job executor JER, and the job executor JER continuously polls the job queue to check whether there is job information JCI in the job queue. When the job information JCI is polled, the job information JCI is created and submitted to a job executor JERA, and the job executor JERA invokes a specific engine for processing.

[0034] Further preferably, it further includes:

[0035] An offline computing engine Actor module, configured to start a job program command with the IP address and job ID of the WSActor instance, and the program establishes a connection channel with the WSActor instance through the IP address and job ID;

[0036] A timer module, configured to receive confirmation information;

[0037] A parameter passing verification module, configured to perform parameter verification on the passed parameters.

[0038] Third aspect, embodiments of the present invention provide an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner of the first aspect.

[0039] Fourth aspect, embodiments of the present invention provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0040] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0041] The present invention can efficiently and reliably solve the business scenarios of big data offline computing by calling the interface of the big data offline processing system based on multiple computing engines during the actual modeling process; through reasonable design and optimization, this method makes full use of the distributed computing power of Akka to construct an engine system specifically for offline computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated in and constitute a part of this specification. The drawings illustrate the embodiments and, together with the description, are used to explain the principles of the invention. Other embodiments and many of the intended advantages of the embodiments will be readily appreciated as they become better understood by reference to the following detailed description. The elements of the drawings are not necessarily to scale relative to each other. Like reference numerals refer to corresponding like parts.

[0043] Figure 1 is a schematic flow diagram of the big data offline processing method based on multiple computing engines according to an embodiment of the present invention;

[0044] Figure 2 is a schematic overall architecture diagram of the big data offline processing method based on multiple computing engines according to an embodiment of the present invention;

[0045] Figure 3 is a schematic diagram of the big data offline processing system based on multiple computing engines according to an embodiment of the present invention;

[0046] Figure 4 is a schematic structural diagram of a computer device of an electronic device suitable for implementing the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] In the following detailed description, reference is made to the accompanying drawings which form a part hereof, and in which are shown by way of illustration specific illustrative embodiments in which the invention may be practiced. In this regard, directional terms such as “top,” “bottom,” “left,” “right,” “up,” “down,” etc. are used with reference to the orientation of the described figures. Since the components of the embodiments may be positioned in several different orientations, for purposes of illustration the directional terms are used and are in no way limiting. It should be understood that other embodiments may be utilized or logical changes may be made without departing from the scope of the present invention. Accordingly, the following detailed description is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.

[0048] Figure 1 An embodiment of the present invention discloses a big data offline processing method based on multiple computing engines, as Figure 1 shown, the method includes the following steps:

[0049] S1, responding to a job request initiated by a third-party application through a Web terminal, the job request calls an engine executor control layer interface EECI;

[0050] S2, the engine job submission executor EJSR polls the independent engine job queue, and takes out the job information JCI and submits it to the instance WSActor of the offline computing engine WebActor;

[0051] S3, start the instance SJEActor of the offline computing engine Actor, and send the startup information to the instance WSActor of the offline computing engine WebActor;

[0052] S4. At the same time, the job executor JER is started. The job executor JER continuously polls the job queue to check whether there is job information JCI in the job queue. When the job information JCI is polled, it creates and submits the job information JCI to the job executor JERA. The job executor JERA calls a specific engine for processing.

[0053] Specifically, Figure 2 The overall framework schematic diagram of the present invention is shown as follows: Figure 2 As shown. Figure 2 The architecture diagram of the present invention is mainly divided into two stages, which include core interaction and job queue modules. The specific method is described as follows:

[0054] Step 1: A third-party application initiates a job request through the Web, which calls the Engine Executor Control Layer Interface (EECI).

[0055] Step 2: After receiving the job request, the Engine Executor Control Layer Interface (EECI) interface calls an Engine Executor Service Class Interface (EESI);

[0056] Step 3: Verify the parameters of the passed parameters. If the job ID is empty, an exception will be thrown.

[0057] Step 4: If the job ID is not empty, the job information (JCI) will be submitted to the Web-side independent engine job queue according to the engine's operating mode;

[0058] In step 5, the engine job submission executor (EJSR) polls the independent engine job queue and takes out the job information (JCI) and submits it to the offline computing engine WebActor (WSActor);

[0059] Step 6: Obtain the corresponding job ID according to the request parameters, and start an instance of the offline computing engine WebActor (WSActor) named after this job ID. Then, pass in the job information (JCI). After receiving the job information, the offline computing engine WebActor (WSActor) calls the engine management service to construct a start command for the offline computing engine Actor (SJEActor), and then starts the job program command with the IP address of the WSActor instance and the job ID. This program establishes a connection channel with the WSActor instance through the IP address and the job ID;

[0060] Step 7: After the SJEActor instance starts, it sends a start message to the WSActor instance (such as the start message sent in 7.1 in Figure 2 ), and creates a timer (such as the one created in 7.2 in Figure 2 ). After receiving the start message from the SJEActor instance, the WSActor instance creates a timer (such as the one created in 7.3 in Figure 2 ), and then sends the job information (JCI) to the SJEActor instance (such as the JCI sent in 7.4 in Figure 2 ). After receiving the job information (JCI), the SJEActor instance closes the timer (such as the one closed in 7.5 in Figure 2 ), and at the same time stores the job information (JCI) in the job queue (such as the JCI sent in 7.6 in Figure 2 ), and sends a message confirming the receipt of the JCI to the WSActor instance (such as the message confirming the receipt of the JCI sent in 7.7 in Figure 2 ). After receiving the confirmation message, the WSActor instance closes the timer (such as the one closed in 7.8 in Figure 2 ).

[0061] During this period, exceptions caused by the timer waiting for timeout will be thrown and uniformly persisted to the distributed storage database. Third-party applications will periodically pull the corresponding persisted table in the distributed storage database to obtain relevant request result information.

[0062] Step 8: When the web service starts, it starts the job executor (JER). After the job executor (JER) starts, it continuously polls the job queue to check if there is job information (JCI) in the job queue. Once it polls the job information (JCI), it creates and submits the job information (JCI) to the job executor (JERA);

[0063] In Step 9, the job executor (JERA) will call the initialization engine method, and then call the preprocessing method before job execution before executing the job; after completion, it will call the execution method for starting the job; after the job ends, it will call the cleanup operation method and persist the results to the distributed storage database; if an error occurs during job execution, it will call the corresponding exception handling method and persist the exception information to the distributed storage database. Finally, the third-party application will pull the corresponding persisted table in the distributed storage database at regular intervals to obtain the relevant request result information.

[0064] This method is implemented in the Tianyuan application of the Qiankun big data operating system and has been verified many times in the actual combat modeling work at the ministry and provincial levels in recent years for smart cities, public health event prevention and control, anti-fraud, etc. In the actual modeling process, by calling the interface of the big data offline processing method based on multiple computing engines, the business scenarios of big data offline computing can be solved efficiently and reliably.

[0065] Secondly, the embodiment of the present invention also discloses a big data offline processing system based on multiple computing engines, as Figure 3 shown, the system includes: a job request module 31, an engine job submission executor EJSR module 32, an offline computing engine WebActor module 33, a job executor JER module 34, an offline computing engine Actor module 35, a timer module 36, and a parameter passing verification module 37.

[0066] In a specific embodiment, the job request module 31 is configured to respond to a job request initiated by a third-party application through the Web side, and the job request calls the engine executor control layer interface EECI; the engine job submission executor EJSR module 32 is configured to poll the independent engine job queue and take out the job information JCI and submit it to an instance WSActor of the offline computing engine WebActor; the offline computing engine WebActor module 33 is configured to start an instance SJEActor of the offline computing engine Actor and send a start message to the instance WSActor of the offline computing engine WebActor.

[0067] The Job Executor JER module 34 is configured to start the Job Executor JER. The Job Executor JER continuously polls the job queue to check if there is job information JCI in the job queue. When job information JCI is polled, it creates and submits the job information JCI to the Job Executor JERA, and the Job Executor JERA calls a specific engine for processing; the Offline Computing Engine Actor module 35 is configured to start a job program command with the IP address and job ID of the WSActor instance. This program establishes a connection channel with the WSActor instance through the IP address and job ID; the Timer module 36 is configured to receive confirmation information; the Parameter Verification module 37 is configured to perform parameter verification on the passed parameters.

[0068] Reference is now made to Figure 4 , which shows a schematic structural diagram of a computer device 600 suitable for use in implementing an embodiment of the present invention (such as Figure 1 the server or terminal device shown). Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.

[0069] As Figure 4 shown, the computer device 600 includes a central processing unit (CPU) 601 and a graphics processing unit (GPU) 602, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 603 or programs loaded from a storage section 609 into a random access memory (RAM) 606. In the RAM 604, various programs and data required for the operation of the device 600 are also stored. The CPU 601, GPU 602, ROM 603, and RAM 604 are connected to each other via a bus 605. An input / output (I / O) interface 606 is also connected to the bus 605.

[0070] The following components are connected to the I / O interface 606: an input section 607 including a keyboard, a mouse, etc.; an output section 608 including, for example, a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 609 including a hard disk, etc.; and a communication section 610 including a network interface card such as a LAN card, a modem, etc. The communication section 610 performs communication processing via a network such as the Internet. A drive 611 can also be connected to the I / O interface 606 as needed. A removable medium 612, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 611 as needed so that a computer program read from it can be installed into the storage section 609 as needed.

[0071] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 610, and / or installed from the removable medium 612. When the computer program is executed by the central processing unit (CPU) 601 and the graphics processing unit (GPU) 602, the above-mentioned functions defined in the method of the present invention are executed.

[0072] It should be noted that the computer-readable medium described in the present invention can be a computer-readable signal medium or a computer-readable medium or any combination of the two. The computer-readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor device, apparatus, or component, or any combination of the above. More specific examples of the computer-readable medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution device, apparatus, or component. And in the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution device, apparatus, or component. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0073] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0074] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based device that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0075] The modules described in the embodiments of the present invention may be implemented in software or in hardware. The described modules may also be provided in a processor.

[0076] As another aspect, the present invention also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or may exist alone without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to: in response to a job request initiated by a third-party application through the Web side, the job request calls the engine executor control layer interface EECI; the engine job submission executor EJSR polls the independent engine job queue and retrieves the job information JCI and submits it to an instance WSActor of the offline computing engine WebActor; starts an instance SJEActor of the offline computing engine Actor and sends a start message to the instance WSActor of the offline computing engine WebActor; at the same time, starts a job executor JER, and the job executor JER continuously polls the job queue to check whether there is job information JCI in the job queue. When the job information JCI is polled, the job information JCI is created and submitted to the job executor JERA, and the job executor JERA calls a specific engine for processing.

[0077] The above description is only a preferred embodiment of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present invention is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, a technical solution formed by mutually replacing the above features with technical features (but not limited to) having similar functions disclosed in the present invention.

Claims

1. A big data offline processing method based on multiple computing engines, characterized in that The method includes the following steps: In response to a job request initiated by a third-party application through the Web side, the job request calls the engine executor control layer interface EECI; The engine job submission executor EJSR polls the independent engine job queue and retrieves the job information JCI to submit to an instance WSActor of the offline computing engine WebActor; Start an instance SJEActor of the offline computing engine Actor and send startup information to an instance WSActor of the offline computing engine WebActor; At the same time, start the job executor JER. The job executor JER continuously polls the job queue to check if there is job information JCI in the job queue. When the job information JCI is polled, create and submit the job information JCI to the job executor JERA, and the job executor JERA calls a specific engine for processing.

2. The big data offline processing method based on multiple computing engines according to claim 1, characterized in that It also includes: After the SJEActor instance is started, it sends startup information to the WSActor instance and creates a first timer. After receiving the startup information of the SJEActor instance, the WSActor instance creates a second timer; Then send the job information JCI to the SJEActor instance. After receiving the job information JCI, the SJEActor instance closes the second timer, stores the job information JCI in the job queue, and sends a message confirming the receipt of JCI to the WSActor instance. After receiving the confirmation message, the WSActor instance closes the first timer.

3. The big data offline processing method based on multiple computing engines according to claim 2, wherein It also includes: After receiving the job request, the engine executor control layer interface EECI interface calls an engine executor service class interface EESI; Perform parameter verification on the passed parameters. If the job ID is empty, an exception will be thrown. If the job ID is not empty, the job information JCI will be submitted to the Web-side independent engine job queue according to the running mode of the engine.

4. The big data offline processing method based on multiple computing engines according to claim 3, wherein It also includes: Obtain the corresponding job ID according to the request passed parameters, start an instance WSActor of the offline computing engine WebActor named with this job ID, and then pass in the job information JCI; After receiving the job information, an instance WSActor of the offline computing engine WebActor calls the engine management service to construct a startup command for an instance SJEActor of the offline computing engine Actor; Then start the job program command with the IP address of the WSActor instance and the job ID. This program establishes a connection channel with the WSActor instance through the IP address and the job ID.

5. The big data offline processing method based on multiple computing engines according to claim 4, wherein It also includes: If the first timer or the second timer waits for a timeout, an exception will be thrown and uniformly persisted to the distributed storage database. The third-party application will periodically pull the corresponding persisted table in the distributed storage database to obtain relevant request result information.

6. The big data offline processing method based on multiple computing engines according to claim 5, characterized in that The job executor JERA calling a specific engine for processing specifically includes: The job executor JERA will call the initialization engine method, and then call a preprocessing method before job execution before executing the job; After completion, call the execution method of starting the job. After the job ends, call the cleanup operation method and persist the results to the distributed storage database; If an error occurs during the job execution, call the corresponding exception handling method and persist the exception information to the distributed storage database; Finally, the third-party application will pull the corresponding persisted table in the distributed storage database at regular intervals to obtain the relevant request result information.

7. A big data offline processing system based on multiple computing engines, characterized in that, The system includes: A job request module configured to respond to a job request initiated by a third-party application through the Web side. The job request calls the engine executor control layer interface EECI; An engine job submission executor EJSR module configured to poll the independent engine job queue and retrieve the job information JCI and submit it to an instance WSActor of the offline computing engine WebActor; An offline computing engine WebActor module configured to start an instance SJEActor of the offline computing engine Actor and send startup information to an instance WSActor of the offline computing engine WebActor; A job executor JER module configured to start the job executor JER. The job executor JER continuously polls the job queue to check if there is job information JCI in the job queue. When the job information JCI is polled, create and submit the job information JCI to the job executor JERA, and the job executor JERA calls a specific engine for processing.

8. The big data offline processing system based on multiple computing engines according to claim 7, characterized in that It further includes: An offline computing engine Actor module configured to start a job program command with the IP address of the WSActor instance and the job ID. This program establishes a connection channel with the WSActor instance through the IP address and the job ID; A timer module configured to receive confirmation information; A parameter verification module configured to verify the passed parameters.

9. An electronic device, including: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multitask concurrent executive system and method for hybrid network service

    CN101741850A

  • System and method supporting single software code base using actor / director model separation

    US20180052712A1