Calculation Method, Device and Electronic Device of Database
By obtaining a copy of the pending data of the bottleneck data node in the database and allocating it to the idle data server for processing, the problem of data analysis tasks that cannot be completed caused by the bottleneck data node is solved, and the completion of data analysis tasks in the bottleneck situation is achieved.
Patent Information
- Application Number
- CN202110875959.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-30
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-07-30
AI Technical Summary
In a relational database with a large-scale parallel processing architecture, when bottleneck data nodes appear, the existing technology cannot complete the data analysis task and needs to wait for the bottleneck node to complete the processing.
By obtaining a copy of the pending data file of the bottleneck data server in the control server and assigning it to the target data server in the idle state, the data in the copy is processed through the target data server, and finally the processing results are fed back to the control server.
When bottleneck data nodes occur, processing tasks can be shared through idle data servers to ensure that data analysis tasks can be completed, thereby improving the processing performance of the database.
Smart Images

Figure CN113590590B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of distributed data storage, and particularly to a calculation method, apparatus, and electronic device for a database. Background Art
[0002] In a relational database with a Massive Parallel Processing (MPP) architecture, the data of each table is distributed to data nodes (DNs) in the cluster according to rules (such as hash modulo). Data analysis tasks need to be coordinated by a control node (Coordinator Node, CN) that does not actually store data to complete the data processing work in parallel on each data node. During the actual operation of the database, when due to a special reason (such as data skew), the data processing progress of a data node is much slower than that of other data nodes, the task needs to wait for this data node to complete the data processing work before the entire data analysis task can be summarized and completed. That is, the execution performance of the data analysis task is limited by the slowest-running data node, which is also called the bottleneck data node.
[0003] When a bottleneck node appears, the prior art often waits for the data processing work of the bottleneck data node to complete, or feedbacks the status of task execution failure when the waiting duration exceeds a predetermined duration.
[0004] It can be seen that when a bottleneck data node appears, the database in the prior art cannot complete the data analysis task. Summary of the Invention
[0005] The purpose of the embodiments of this application is to provide a calculation method, apparatus, and electronic device for a database, so as to still be able to complete the data analysis task when a bottleneck node appears in the database.
[0006] To solve the above technical problems, the embodiments of this specification provide a calculation method for a database, including: when it is determined that a first data server has not feedback a calculation result within a predetermined time period, obtaining a copy of the data file to be processed by the first data server; allocating the copy to a target data server, where the target data server is a data server in an idle state; processing the data in the copy by the target data server to obtain a processing result; and feedbacking the processing result as the calculation result of the first data server.
[0007] The embodiments of this specification also provide a calculation method for a database. The method includes: when the control server determines that the first data server has not fed back the calculation result within a predetermined time period, obtaining a copy of the data file to be processed by the first data server and sending it to the management and control server, and using the data server that is currently in an idle state as the target data server, and sending the server identifier of the target data server to the management and control server; the management and control server queries the link information of the target data server according to the server identifier, and sends the link information and the received copy to the data middleware; the data middleware allocates virtual disks for each target data server, mounts the virtual disks on the corresponding target data servers according to the link information, and distributes the data in the copy to each virtual disk; the control server triggers the computing resources of each target data server to process the data in the mounted virtual disks, receives and aggregates the calculation results fed back by each target data server, and uses the aggregated result as the calculation result of the first data server.
[0008] The calculation method, device, and electronic device for a database provided by the embodiments of this specification can enable the data server in an idle state to share the task to be processed by the bottleneck data server, so that the data analysis task can be completed even when there is a bottleneck data server. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0010] Figure 1 Shows a schematic structural diagram of distributed data;
[0011] Figure 2 Shows a schematic structural diagram of a database provided by the embodiments of this specification;
[0012] Figure 3 Shows a flowchart of a calculation method for a database according to the embodiments of this specification;
[0013] Figure 4 Shows a flowchart of another calculation method for a database according to the embodiments of this specification;
[0014] Figure 5A Shows a flowchart of yet another calculation method for a database according to the embodiments of this specification;
[0015] Figure 5B The figure shows a specific method flow chart for processing data in a copy through a target data server to obtain a processing result;
[0016] Figure 6 The figure shows a schematic block diagram of a computing device of a database according to an embodiment of the present specification;
[0017] Figure 7 The figure shows a schematic block diagram of an electronic device according to an embodiment of the present specification. Detailed implementation manners
[0018] In order to enable those skilled in the art of the present technology to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this application.
[0019] Figure 1 The figure shows a schematic structural diagram of distributed data. In a distributed database architecture, it includes a control node and multiple data nodes, and the control node is communicatively connected to each data node. The control node receives an operation instruction of the database, such as an SQL statement, parses the operation instruction, and allocates computing tasks to multiple data nodes according to the parsing result. After the data nodes complete the computing tasks, they will feedback the computing results to the control node. The control node aggregates the computing results fed back by each data node, and uses the aggregated computing result as the execution result of the operation instruction and feeds it back to the user. Each database node includes a control module, a computing module, and a storage module. The storage module is used to store the data in the database table, the computing module is responsible for the computing tasks of the data in the storage module, and the control module is used for data interaction with the control node (such as receiving tasks assigned by the control, sending computing results to the control node, etc.), controlling the internal processing flow of the control node, etc.
[0020] It can be seen from this that the data nodes of the database are independent of each other, that is, the data nodes do not communicate with each other, and there will be no situation where a data node directly processes the data stored in a bottleneck data node.
[0021] On the other hand, in the actual usage process, the database only has a control node server (hereinafter referred to as the control server) and a data node server (hereinafter referred to as the data server), and there are no other servers. The database architecture based on which the database is built has defined that the control node and the data node each implement their own functions according to the above content, and the database architecture does not define what operations the control node or the data node should perform when a bottleneck data node appears.
[0022] Therefore, the database structure itself does not have the conditions to handle bottleneck data nodes.
[0023] For this reason, in the embodiments of this specification, a management and control server and a data middleware are introduced on the basis of the original database architecture. Figure 2 Fig. shows a schematic structural diagram of a database provided by the embodiments of this specification. Comparing Figure 1 It can be seen that the management and control server is communicatively connected to the control server, and the data middleware is communicatively connected to the management and control server. This database structure still follows the original database architecture.
[0024] Based on this, the embodiments of this specification provide a calculation method for a database to solve the calculation problem of bottleneck data nodes. As Figure 3 shown, the method includes the following steps:
[0025] S310: When the control server determines that the first data server has not fed back the calculation result within a predetermined time period, it obtains a copy of the data file to be processed by the first data server and sends it to the management and control server; and uses the data server that is currently in an idle state as the target data server, and sends the server identifier of the target data server to the management and control server.
[0026] Since the control server needs to aggregate the calculation results of multiple data servers, it starts waiting for the fed-back calculation results after assigning calculation tasks to the data servers. When it is detected that a data server has not fed back the calculation result within a predetermined time period, it means that there is a bottleneck data node.
[0027] The copy and the server identifier are used for the management and control server to control the data middleware to allocate virtual disks for each target data server, mount the virtual disks on the corresponding target data servers according to the link information, and allocate the data in the copy to each virtual disk.
[0028] When each data server processes data, it only obtains data from its own storage unit for processing, and the control node or the outside world does not provide the data to be processed. Therefore, the control server can only obtain a copy of the file to be processed from the data node.
[0029] The control server will record the working status of each data server in real time (the working status, such as the idle status, the working status). Therefore, the control server can directly find out the identifier of the data server that is currently in the idle status from the working status data it records.
[0030] Specifically, as Figure 4 shown, the control server can perform the following operations in step S310.
[0031] S311: In the case of determining that the first data server has not fed back the calculation result within a predetermined time period, send a rollback instruction to the first data server to control the first data server to roll back to the data processing task that has been executed and completed.
[0032] Rollback corresponds to the rollback operation in database operations.
[0033] The data server usually has problems when processing a database transaction. A database transaction is a sequence of database operations that access and may operate on various data items. These operations are either all executed or all not executed, and it is an indivisible unit of work. A transaction consists of all database operation instructions executed between the start and end of the transaction. The rollback operation enables the first data server to resume to the database operation instructions that have been executed and completed.
[0034] For example, the database transaction includes 120 database operation instructions. When the first data server has not returned the data processing result to the control server all the time when executing the 25th database operation instruction, after the control server controls the first data server to roll back to the 24th database operation instruction, the 25th to 120th database operation instructions are used as the data tasks to be processed, and the data files to be processed correspond to the data tasks to be processed.
[0035] S312: Control the first data server to send a copy of the data file to be processed after rollback to the management and control server.
[0036] S313: Obtain a copy of the data file to be processed by the first data server.
[0037] S314: Query the data servers that are currently in the idle status, and use the data servers that are currently in the idle status as the target servers.
[0038] S315: Query the server identifier of the target server.
[0039] S316: Send the copy and the server identifier to the management and control server.
[0040] This step is used to control the server control data middleware to allocate virtual disks for each target data server, obtain the link information of the target server according to the server identifier, mount the virtual disks on the corresponding target data servers according to the link information, and allocate the data in the replicas to each virtual disk.
[0041] S320: The control server queries the link information of the target data server according to the server identifier, and sends the link information and the received replica to the data middleware.
[0042] Since the data nodes under the database architecture can only transmit data with the control node, the control server can only obtain the replicas of the data files to be processed by the bottleneck data node server through the control server.
[0043] After receiving the server identifier of the target data server, the control server can query the link information of the target data server corresponding to the server identifier according to the pre-stored correspondence list of server identifiers and link information, or query other devices communicatively connected to it.
[0044] The link information can be, for example, network address, network port and other information, and can also include the directory structure of the file system. In some embodiments, if the control server and multiple data servers are connected through the Internet, the network address in the link information can be an IP address. In some embodiments, if the control server and multiple data servers are connected through a local area network, the network address in the link information can be a local area network address.
[0045] In the database architecture, the control node does not store and does not have the ability to query this link information.
[0046] Specifically, as Figure 4 shown, the control server in step S320 can perform the following operations.
[0047] S321: Receive the replica of the data file to be processed by the first data server sent by the control server, and the server identifier of the target data server.
[0048] S322: Query the link information of the target data server according to the server identifier.
[0049] S323: Send the link information and the replica to the data middleware.
[0050] S324: Send a data allocation instruction to the data middleware.
[0051] The data allocation instruction is used to instruct the data middleware to allocate virtual disks for each target data server, mount the virtual disks on the corresponding target data servers according to the link information, and allocate the data in the replicas to each virtual disk.
[0052] In some embodiments, as Figure 4 shown, before step S324, the following steps are further included.
[0053] S325: Obtain the number of target data servers and the file size of the replicas.
[0054] The number of target data servers can be sent by the control server or obtained by counting the number of identifiers of the target servers sent by the control server.
[0055] The file size of the replicas, such as 2 TB, 200 MB, etc.
[0056] S326: Generate an allocation plan for the replica files for the data middleware, where the file size allocated to each target data server is equal.
[0057] For example, if the size of the replica file is 90 MB and there are 3 target data servers, the allocation plan generated by the management and control server is to allocate 30 MB of data volume to each target data server. When the data middleware allocates virtual disks, it can allocate a virtual disk of more than 30 MB to each target server. For example, it can allocate virtual disks of 40 MB or 50 MB.
[0058] S327: Send the allocation plan of the replica files to the data middleware.
[0059] Correspondingly, when the data middleware allocates virtual disks to each target data server, it allocates based on the allocation plan of the replica files.
[0060] S330: The data middleware allocates virtual disks to each target data server, mounts the virtual disks on the corresponding target data servers according to the link information, and distributes the data in the replicas to each virtual disk.
[0061] The virtual disk can be mounted under a fixed file directory of the target data server, or the directory structure of the file system includes the file directories where the disk can be mounted.
[0062] The data files to be processed by the bottleneck data server usually may include many data records (each record may include multiple fields). "Distributing the data in the replicas to each virtual disk" is to distribute these data records to each virtual disk.
[0063] The data in the replica can be distributed to each virtual disk according to a preset distribution strategy. In some embodiments, the data in the replica can be distributed in a way that equally divides the data file sizes, that is, the sizes of the data files allocated to each virtual disk are equal. For example, the allocated data files are all 20 megabytes. In this case, the capacities of the virtual disks are usually the same.
[0064] In some embodiments, the computing resources of the data server can be divided into multiple parts. Each time a task assigned by the control server is executed, only some of the computing resources are mobilized, and some computing resources are in an idle state. In this case, the preset distribution strategy can also be based on the proportion of the idle computing resources. For example, if the proportion of the idle computing resources of target data servers 1 and 2 is 1:2, then the sizes of the data files allocated to these two target data servers can also be in the ratio of 1:2.
[0065] Specifically, as Figure 4 shown, the data middleware in step S330 can perform the following operations.
[0066] S331: Obtain the link information of the target data server sent by the management and control server and the replica of the data file to be processed by the first data server.
[0067] S332: Allocate virtual disks for each target data server.
[0068] S333: Mount the virtual disks on the corresponding target data servers according to the link information.
[0069] S334: Distribute the data in the replica to each virtual disk.
[0070] This step is used for the control server to control the computing resources of the target data server to process the data in the mounted virtual disks, summarize the calculation results fed back by the target data servers, and use the summarized calculation results as the calculation results of the first data server.
[0071] In some embodiments, as Figure 4 shown, step S334 further includes the following steps.
[0072] S3341: Obtain the distribution plan of the replica files, and the difference in the file sizes allocated to each target data server is within a predetermined range.
[0073] S3342: According to the distribution plan, divide and distribute each data record in the replica file to the virtual disks corresponding to each target data server.
[0074] The distribution plan of the replica files specified by the management and control server is only a distribution plan based on file size, which can be used to determine the size of the data allocated to each virtual disk. The distribution performed by the data middleware in step S3342 is to intercept multiple data records from the replica files for storage in each virtual disk, and the total data volume of the intercepted multiple data records is the size of the data allocated to each virtual disk in the distribution plan.
[0075] S340: The control server triggers the computing resources of each target data server to process the data in the mounted virtual disk, receives and aggregates the calculation results fed back by each target data server, and uses the aggregated result as the calculation result of the first data server.
[0076] In the database architecture, it is defined that the control node can trigger the data node to process the data in its predetermined file directory and feed back the calculation result to the control node. The method provided in this specification enables the control server to use the aggregated result as the calculation result of the first data server, that is, the calculation result of the bottleneck data server.
[0077] As Figure 4 shown, in some embodiments, after step S340, the control server sends an end notification to the management and control server to notify the management and control server that the calculation task of the first data server has been completed. After receiving the end notification, the management and control server sends a cancellation instruction to the data middleware to cancel the virtual disk.
[0078] In the above database calculation method, on the basis of the control server and data server defined in the original database architecture, a management and control server and a data middleware are set up. In the case of a bottleneck data node, the control server and the management and control server are used to send the replica of the data file to be processed by the bottleneck data server and the link information of the idle data server to the data middleware in sequence. The data middleware allocates virtual disks for each idle data server and mounts the virtual disks to the corresponding idle data servers according to the link information, so that the control server can control the computing resources of the idle data servers to process the data in the mounted virtual disks, that is, process the data files to be processed by the bottleneck data server, aggregate the fed-back calculation results, and use the aggregated result as the calculation result of the bottleneck data server.
[0079] On the other hand, the above database calculation method sends the replica of the data file to be processed by the bottleneck data server to the data middleware, enabling the idle data servers to process the replica files on the data middleware instead of directly controlling the computing resources of the idle data servers to obtain data from the storage module of the bottleneck data server, thus not damaging the original data stored in the bottleneck data server and ensuring the security and stability of the data stored in the database.
[0080] The embodiments of this specification also provide a calculation method for a database to solve the calculation problem of bottleneck data nodes. As Figure 5A shown, the method includes the following steps:
[0081] S510: When it is determined that the first data server has not fed back the calculation result within a predetermined period, obtain a copy of the data file to be processed by the first data server.
[0082] S520: Allocate the copy to the target data server, where the target data server is a data server in an idle state.
[0083] S530: Process the data in the copy through the target data server to obtain a processing result.
[0084] S540: Feed back the processing result as the calculation result of the first data server.
[0085] The above steps S510 to S540 can all be executed by one server, or as Figure 3 in the corresponding embodiment, introduce a control server and a data middleware to cooperate in execution. This embodiment does not limit the execution entity of these steps.
[0086] For the relevant descriptions and beneficial effects of the above steps S510 to S540, reference can be made to Figure 3 the corresponding embodiment for understanding, and details will not be repeated here.
[0087] In some embodiments, as Figure 5B shown, step S530 includes the following steps:
[0088] S531: Allocate a virtual disk to the target data server.
[0089] S532: Mount the allocated virtual disk to the target data server.
[0090] S533: Allocate the data in the copy to the virtual disk according to a preset allocation strategy.
[0091] S534: Process the allocated data in the mounted virtual disk through the target data server to obtain a first result.
[0092] S535: Aggregate and process the first results of each target data server to obtain a processing result.
[0093] In some embodiments, step S532 includes the following steps: Obtain the link information of the target data server; Mount the allocated virtual disk to the target data server through the link information.
[0094] The above steps S531 to S535 can all be executed by one server, or as in the Figure 3 or Figure 4 corresponding embodiments, introduce a control server and a data middleware to cooperate in execution. This embodiment does not limit the execution entity of these steps.
[0095] For the relevant descriptions and beneficial effects of the above steps S531 to S535, reference can be made to Figure 3 or Figure 4 the corresponding embodiments for understanding, and will not be elaborated here.
[0096] In some embodiments, after determining that the first data server has not fed back the calculation result within a predetermined time period, it further includes: sending a rollback instruction to the first data server to control the first data server to roll back to the data processing tasks that have been executed; controlling the first data server to return the data that has not been processed after the rollback as a copy. For the relevant descriptions and beneficial effects, reference can be made to Figure 3 or Figure 4 the corresponding embodiments for understanding, and will not be elaborated here.
[0097] The embodiments of this specification provide a database that can be used to implement Figure 3 the calculation method of the database shown. As Figure 2 shown, the database includes a control server, multiple data servers, a control server, and a data middleware.
[0098] The control server is used to, when determining that the first data server has not fed back the calculation result within a predetermined time period, obtain a copy of the data file to be processed by the first data server and send it to the control server, and use the data server that is currently in an idle state as the target data server, and send the server identifier of the target data server to the control server.
[0099] The control server is used to query the link information of the target data server according to the server identifier, and send the link information and the received copy to the data middleware.
[0100] The data middleware is used to allocate virtual disks for each target data server, mount the virtual disks on the corresponding target data servers according to the link information, and allocate the data in the copy to the virtual disks.
[0101] The control server is further used to trigger the computing resources of each target data server to process the data in the mounted virtual disks, receive and aggregate the calculation results fed back by each target data server, and use the aggregated result as the calculation result of the first data server.
[0102] In some embodiments, the data middleware may include a data transfer mapper and a disk array. The data transfer mapper may be used to receive the link information and replicas sent by the management and control server, allocate virtual disks for each target data server, mount the virtual disks on the corresponding target data servers according to the link information, and allocate the data in the replicas to the virtual disks. The disk array is used to store data and is allocated by the data transfer mapper to form each virtual disk. That is, each virtual disk is actually a different partition in the disk array.
[0103] The embodiments of this specification provide a computing device for a database, which can be used to implement Figure 3 the computing method of the database shown. As Figure 6 shown, the device includes an acquisition module 10, an allocation module 20, a processing module 30, and a feedback module 40.
[0104] The acquisition module 10 is used to obtain a copy of the data file to be processed by the first data server when it is determined that the first data server has not fed back the calculation result within a predetermined time period.
[0105] The allocation module 20 is used to allocate the replica to the target data server, where the target data server is a data server in an idle state.
[0106] The processing module 30 is used to process the data in the replica through the target data server to obtain a processing result.
[0107] The feedback module 40 is used to feed back the processing result as the calculation result of the first data server.
[0108] In some embodiments, the processing module 30 includes a first allocation sub-module, a mounting sub-module, a second allocation sub-module, a processing sub-module, and a summarization sub-module.
[0109] The first allocation sub-module is used to allocate virtual disks for the target data server. The mounting sub-module is used to mount the allocated virtual disks on the target data server. The second allocation sub-module is used to allocate the data in the replica to each virtual disk. The processing sub-module is used to process the allocated data in the mounted virtual disk through the target data server to obtain a first result. The summarization sub-module is used to perform a summarization process on the first results of each target data server to obtain a processing result.
[0110] In some embodiments, the mounting sub-module first obtains the link information of the target data server, and then mounts the allocated virtual disk on the target data server through the link information.
[0111] In some embodiments, as Figure 6 shown, the computing device for this database further includes a sending module 50 and a control return module 60.
[0112] The sending module 50 is configured to send a rollback instruction to the first data server to control the first data server to roll back to the data processing task that has been executed. The control return module 60 is configured to control the first data server to return the unprocessed data after rollback as a copy.
[0113] The related description and effects of the computing device of the above database can be understood by referring to the computing method of the database, which will not be elaborated here.
[0114] An embodiment of the present invention further provides an electronic device, as Figure 7 shown. The electronic device may include a processor 71 and a memory 72, where the processor 71 and the memory 72 may be connected through a bus or other means, Figure 7 taking the connection through the bus as an example.
[0115] The processor 71 may be a central processing unit (CPU). The processor 71 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. chips, or a combination of the above types of chips.
[0116] The memory 72, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the computing method of the database in the embodiment of the present invention (for example, Figure 6 the obtaining module 10, the allocation module 20, the processing module 30, and the feedback module 40 shown). The processor 71 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 72, that is, implements the computing method of the database in the above method embodiment.
[0117] The memory 72 may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data created by the processor 71 and the like. In addition, the memory 72 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 72 may optionally include a memory remotely disposed relative to the processor 71, and these remote memories may be connected to the processor 71 through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0118] The one or more modules are stored in the memory 72 and, when executed by the processor 71, perform the calculation method of the database in the embodiments as Figure 3 shown.
[0119] Specific details of the above electronic device may be understood by referring to the corresponding relevant descriptions and effects in the embodiments of Figure 3 and will not be elaborated herein.
[0120] Those skilled in the art can understand that to implement all or part of the processes in the above method embodiments, it can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. Among them, the storage medium may be a magnetic disk, an optical disc, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD), etc.; the storage medium may also include a combination of the above types of memories.
[0121] In the 1950s, it was obvious to distinguish whether an improvement in a technology was a hardware improvement (e.g., improvement in circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement in method processes). However, with the development of technology, many improvements in method processes today can be regarded as direct improvements in hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structures by programming the improved method processes into the hardware circuits. Therefore, it cannot be said that an improvement in a method process cannot be implemented with hardware entity modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can program by themselves to "integrate" a digital system on a piece of PLD without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip 2. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called Hardware Description Language (HDL), and there is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog 2. Those skilled in the art should also be clear that only by slightly logically programming the method process with the above-mentioned several hardware description languages and programming it into the integrated circuit can the hardware circuit implementing the logical method process be easily obtained.
[0122] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.
[0123] The systems, devices, modules or units described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions.
[0124] For the convenience of description, when describing the above devices, they are described as various units according to functions. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.
[0125] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of certain parts of each embodiment of the present application.
[0126] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0127] The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0128] Although the present application is depicted through embodiments, those of ordinary skill in the art know that the present application has many variations and changes without departing from the spirit of the present application. It is hoped that the appended claims will cover these variations and changes without departing from the spirit of the present application.
Claims
1. A calculation method for a database, characterized in that, the database includes a control server and multiple data servers, and the control server is communicatively connected to each data server; the method includes introducing a management and control server and a data middleware on the basis of the database architecture, and: When the control server determines that the first data server has not fed back the calculation result within a predetermined time period, it obtains a copy of the data file to be processed by the first data server and sends it to the management and control server, and uses the data server that is currently in an idle state as the target data server, and sends the server identifier of the target data server to the management and control server; The management and control server queries the link information of the target data server according to the server identifier, and sends the link information and the received copy to the data middleware; The data middleware allocates virtual disks for each target data server, mounts the virtual disks on the corresponding target data servers according to the link information, and distributes the data in the copy to each virtual disk; The control server triggers the computing resources of each target data server to process the data in the mounted virtual disks, receives and aggregates the calculation results fed back by each target data server, and uses the aggregated result as the calculation result of the first data server.
2. The method according to claim 1, characterized in that, the management and control server queries the link information of the target data server according to the server identifier, including: Querying the link information of the target data server corresponding to the server identifier according to the pre-stored correspondence list of the server identifier and the link information, or querying other devices communicatively connected to the server corresponding to the server identifier.
3. The method according to claim 1, characterized in that, the control server does not store and does not have the ability to query link information.
4. The method according to claim 1, characterized in that, the link information at least includes: the local area network address and the directory structure of the file system.
5. The method according to claim 1, characterized in that, after the control server determines that the first data server has not fed back the calculation result within a predetermined time period and before obtaining a copy of the data file to be processed by the first data server and sending it to the management and control server, it further includes: Sending a rollback instruction to the first data server to control the first data server to roll back to the data processing tasks that have been completed; Controlling the first data server to return the unprocessed data after rollback as a copy.
6. A calculation device for a database, characterized in that, the database includes a control server and multiple data servers, and the control server is communicatively connected to each data server; the device introduces a management and control server and a data middleware on the basis of the database architecture, and the device includes: An acquisition module, configured to, when it is determined that the first data server has not fed back a calculation result within a predetermined time period, acquire a copy of the data file to be processed by the first data server, send it to the management and control server, use the data server currently in an idle state as the target data server, and send the server identifier of the target data server to the management and control server; An allocation module, configured to enable the management and control server to query the link information of the target data server according to the server identifier, and send the link information and the received copy to the data middleware; The data middleware allocates virtual disks for each target data server, mounts the virtual disks on the corresponding target data servers according to the link information, and allocates the data in the copy to each virtual disk, so as to allocate the copy to the target data servers, where the target data servers are data servers in an idle state; A processing module, configured to enable the control server to trigger the computing resources of each target data server to process the data in the mounted virtual disk, so as to process the data in the copy through the target data server to obtain a processing result; A feedback module, configured to feedback the aggregated processing result as the calculation result of the first data server.
7. A database, characterized in that, it includes: A control server, multiple data servers, a management and control server, and a data middleware, where the control server is communicatively connected to each data server, and: The control server is configured to, when it is determined that the first data server has not fed back a calculation result within a predetermined time period, acquire a copy of the data file to be processed by the first data server, send it to the management and control server, use the data server currently in an idle state as the target data server, and send the server identifier of the target data server to the management and control server; The management and control server is configured to query the link information of the target data server according to the server identifier, and send the link information and the received copy to the data middleware; The data middleware is configured to allocate virtual disks for each target data server, mount the virtual disks on the corresponding target data servers according to the link information, and allocate the data in the copy to the virtual disks according to a preset allocation strategy; The control server is further configured to trigger the computing resources of each target data server to process the data in the mounted virtual disk, receive and aggregate the calculation results fed back by each target data server, and use the aggregated result as the calculation result of the first data server.
8. An electronic device, characterized in that, it includes: A memory and a processor, where the processor and the memory are communicatively connected to each other, the memory stores computer instructions, and the processor realizes the steps of the method according to any one of claims 1 to 5 by executing the computer instructions.
9. A computer storage medium, characterized in that, the computer storage medium stores computer program instructions, and when the computer program instructions are executed, the steps of the method according to any one of claims 1 to 5 are realized.
Citation Information
Patent Citations
Task scheduling method, device, electronic equipment and medium
CN111641678A