Distributed computing system and method based on functions as a service
Patent Information
- Application Number
- TW113135142
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-09-16
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2044-09-15
AI Technical Summary
Existing distributed computing frameworks for big data processing require users to build and configure infrastructure, which is not user-friendly for those lacking technical skills, and processing large files within time constraints is impossible.
A distributed computing system based on Function as a Service (FaaS) that dynamically segments file data into smaller clusters and performs layered processing using cloud functions to manage data within time constraints.
Enables efficient completion of data processing within limited time frames by reducing the data load per function, overcoming the limitations of single large files.
Smart Images

Figure TWG2TB001908536_001 
Figure TWG2TB001908536_002 
Figure TWG2TB001908536_003
Abstract
Description
[Technical Field]
[0001] This invention relates to a distributed computing method and system, and more particularly to a distributed computing system and method based on Function as a Service. [Previous Technology]
[0002] Distributed computing for big data is a technology that processes massive datasets by distributing data and tasks across multiple computers. This approach can significantly improve data processing speed and efficiency by performing computations simultaneously on multiple nodes, reducing the risk of single points of failure and improving system scalability.
[0003] Commonly used distributed computing frameworks include Apache Hadoop, Apache Spark, and Apache Flink. These frameworks support a wide range of applications, from batch data processing to real-time data stream processing. However, these distributed computing frameworks require users to build and configure their own infrastructure using Infrastructure as a Service (IaaS). This is not user-friendly for those lacking the relevant technical skills.
[0004] Therefore, there is a need for a distributed system that can perform big data computing without having to build and configure the infrastructure itself. [Summary of the Invention]
[0005] This invention provides a distributed computing system based on Function as a Service, comprising: a first storage device for storing file data uploaded by a local device; a hierarchical control module coupled to the first storage device for planning a first processing layer and a second processing layer according to a processing flow of the file data, and for calling corresponding cloud functions to process the file data according to the processing programs of the first processing layer and the second processing layer; a detection and judgment module coupled to the first storage device for determining the category of the file data, and for generating a clustering rule according to the category of the file data, wherein the hierarchical control module executes the processing program of the first processing layer and calls a first cloud function to segment the file data into multiple clusters of file data according to the clustering rule; and a second storage device for storing the multiple clusters of file data, and the hierarchical control module executes the processing program of the second processing layer and calls multiple second cloud functions to synchronously process the multiple clusters of file data.
[0006] In some embodiments, the detection and judgment module generates the grouping rule based on the processing capability of each of the plurality of second cloud functions for the file data category within a limited time.
[0007] In some embodiments, the processing capacity is the amount of archive data that can be processed within the time limit.
[0008] In some embodiments, a third storage device is further included for storing the plurality of group file data after the plurality of second cloud functions have been synchronously processed.
[0009] In some embodiments, the hierarchical control module further plans a third processing layer according to the processing flow of the archive data.
[0010] In some embodiments, the hierarchical control module executes the processing program of the third processing layer, calling multiple third cloud functions to synchronously process the multiple group file data after the multiple second cloud functions have synchronously processed.
[0011] In some embodiments, a fourth storage device is further included for storing the plurality of group file data after the plurality of third cloud functions have been synchronously processed.
[0012] In some embodiments, a difference analysis element is further included, coupled to the ground element, for comparing file data before and after uploading to generate a difference file.
[0013] This invention provides a distributed computing method based on Function as a Service, comprising: using a hierarchical control module to plan a first processing layer and a second processing layer according to a processing flow of a file data; using the hierarchical control module to determine whether a third processing layer needs to be set according to the processing flow of the file data; when the hierarchical control module determines that a third processing layer does not need to be set, using the hierarchical control module to execute the first processing layer to call a first cloud function to cut the file data into multiple clusters of file data according to a clustering rule; and using the hierarchical control module to execute the second processing layer to call multiple second cloud functions to synchronously process the cut multiple clusters of file data.
[0014] In some embodiments, the grouping rule is generated using a detection and judgment module based on the processing capability of each of the plurality of second cloud functions for the archive data within a limited time.
[0015] This invention provides a distributed computing method and system based on Function as a Service (SaaS). By dynamically segmenting file data to reduce its size and performing layered processing, the amount of file data processed by a single function is reduced, thereby controlling the completion of data computation within time constraints. Therefore, it can solve the problem of data processing being impossible to complete under limited time constraints when a single file is too large.
Implementation Method
[0017] The spirit of this case will be clearly explained below with diagrams and detailed description. Anyone with ordinary knowledge in the relevant technical field can make changes and modifications based on the technology taught in this case after understanding the embodiments of this case, without departing from the spirit and scope of this case.
[0018] The terminology used herein is for the purpose of describing specific embodiments only and is not intended to limit the scope of this application. Singular forms such as “a,” “this,” “this,” “this,” and “the” as used herein also include plural forms.
[0019] The terms "coupled" or "connected" as used herein can refer to two or more components or devices making direct physical contact with each other, or making indirect physical contact with each other, or to two or more components or devices operating or acting on each other.
[0020] The terms 'include', 'including', 'have', 'contain', etc. used in this document are all open-ended terms, meaning that they include but are not limited to.
[0021] The term "and / or" as used herein includes any or all of the things mentioned.
[0022] The terms used herein, unless otherwise specified, generally have their ordinary meaning in the context of the art, the subject matter, and the specific content. Certain terms used to describe the subject matter will be discussed below or elsewhere in this specification to provide additional guidance to those skilled in the art in describing the subject matter.
[0023] Because data computation via Infrastructure as a Service (IaaS) requires the construction and configuration of the infrastructure itself, it is not user-friendly for those lacking the relevant technical skills. Therefore, Function-as-a-Service (FaaS) is often used for data computation, as it does not require the construction and maintenance of related infrastructure. However, if FaaS is used for data computation, the problem of time constraints for single function processing arises. If the file data is too large, data processing cannot be completed within the time limit. Therefore, this invention provides a distributed computing method and system based on FaaS, which reduces the file data size by dynamically splitting the file data and performing layered processing, thereby reducing the amount of file data processed by a single function and controlling the completion of data computation within the time limit.
[0024] Figure 1 is a schematic diagram of a function-as-a-service (MAA) based distributed computing system according to a preferred embodiment of the present invention. The MAA-based distributed computing system 100 includes a first storage device 110, a second storage device 120, a detection and judgment module 130, and a hierarchical control module 140. In some embodiments, the first storage device 110 and the second storage device 120 are cloud storage devices, such as any type of random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), or similar elements or combinations thereof.
[0025] In some embodiments, the layered control module 140 plans the number of processing layers for the file data based on the cloud function operation process to be performed on the file data uploaded by the ground element 150 through the network 200. In some embodiments, the ground element 150 uploads file data conforming to the Portable Document Format (PDF) through the network 200, so as to convert the portable document format file data into file data conforming to the Electronic Publication Standard (EPUB) through cloud function operation. In this embodiment, the layered control module 140 can plan the number of processing layers for this file data based on the cloud function operation process to be performed on this file data, and call the corresponding cloud function to process the file data according to the number of processing layers to be executed. In some embodiments, since this invention is mainly used for file data that cannot be processed under time constraints, the first processing layer of the layered control module 140 is to perform file data segmentation. Accordingly, the layered control module 140 converts portable file format archive data into archive data conforming to e-book standards. The planned processing layers include a first processing layer that segments the portable file format archive data, a second processing layer that converts the portable file format archive data into image format data, a third processing layer that converts the image format data into archive data conforming to e-book standards, and a fourth processing layer that encrypts the archive data conforming to e-book standards. The fourth processing layer is an optional processing layer; that is, it can be omitted if encryption is not necessary. However, it is worth noting that the above-described number of planned layers is only one embodiment and is not intended to limit the scope of this invention. In other embodiments, additional intermediate processing layers can be changed or added depending on the required archive format. Furthermore, the layered control module 140 is not limited to performing the above-described archive conversion plan; in other embodiments, different plans can be implemented based on the archive data uploaded by the ground element 150. Furthermore, the layered control module 140 can pre-plan the corresponding number of processing layers for different types of file data, so that when the ground terminal element 150 uploads file data, it can immediately call the corresponding processing layer for processing based on the uploaded file data.
[0026] In some embodiments, the first storage device 110 is coupled to the hierarchical control module 140 for storing a file of data uploaded by the ground element 150 via the network 200.
[0027] In some embodiments, the detection and judgment module 130 is coupled to the first storage device 110 and the hierarchical control module 140, and segments the file data according to the first processing layer planned by the hierarchical control module 140. In some embodiments, the detection and judgment module 130 determines the file data size and file data category, and generates a clustering rule based on the processing capability of the corresponding cloud function for the file data category under time constraints in the number of processing layers planned by the hierarchical control module 140. Accordingly, the hierarchical control module 140 executes the first processing layer call to the corresponding cloud function to segment the file data according to this clustering rule. In some embodiments, the file data uploaded by the ground element 150 through the network 200 includes 100,000 files conforming to the portable file format. Under time constraints, the processing capacity of the corresponding cloud function in the processing layer to convert the file data into file data conforming to e-book standards is 10,000 records. Therefore, the detection and judgment module 130 generates a grouping rule, and the hierarchical control module 140 executes the planned first processing layer, calling the corresponding cloud function to divide the file data into 10 groups of 10,000 records each. Then, the hierarchical control module 140 executes the planned subsequent processing layers, calling the corresponding cloud functions to perform parallel conversion processing on these 10 groups of file data. Since the data processed by each cloud function is reduced from the original 100,000 records to 10,000 records per group that can be processed within the time constraint, it can be ensured that the file data uploaded by the ground element 150 can be processed normally. In some embodiments, the detection and judgment module 130 can be implemented by a processor executing a judgment program stored in memory, or by firmware.
[0028] In some embodiments, the second storage device 120 is coupled to the detection and judgment module 130 and the first storage device 110, and is used to store multi-group file data generated by segmenting file data according to the grouping rules. In some embodiments, according to the grouping rules of the detection and judgment module 130, the file data is segmented into 10 groups of 10,000 records each, and the hierarchical control module 140 executes the first processing layer to call the corresponding cloud function to segment the file data into 10 groups. Therefore, the second storage device 120 is used to store these 10 groups of file data.
[0029] Next, the layered control module 140 executes the second processing layer to call the corresponding cloud functions to perform parallel processing on the 10 groups of file data. In some embodiments, the second processing layer executed by the layered control module 140 is to convert the portable file format file data into image format data. Accordingly, the layered control module 140 will call 10 cloud functions to synchronously convert the 10 groups of file data, converting the portable file format into image format. Next, the layered control module 140 executes the third processing layer to call the corresponding cloud functions, converting the 10 groups of file data into image format, and then converting them into file data conforming to the e-book standard. Finally, the layered control module 140 executes the fourth processing layer to call the corresponding cloud functions, encrypting the converted e-book standard file data before outputting it.
[0030] It is worth noting that only the second storage device 120 is shown in the first figure of this case. However, in other embodiments, a third storage device may also be included to store 10 groups of file data converted into image format. A fourth storage device is used to store 10 groups of file data converted into electronic book standard. A fifth storage device is used to store 10 groups of file data encrypted and conforming to electronic book standard. Through the third, fourth, and fifth storage devices, the ground element 150 can retrieve the corresponding format files for subsequent processing as needed via the network 200.
[0031] In another preferred embodiment, the ground element 150 is further coupled to a difference analysis element 160. Since some of the file data uploaded by the ground element 150 does not need to be uploaded for conversion, the difference analysis element 160 can be used to compare the differences between the files before and after the upload to generate a difference file.
[0032] It is worth noting that the distributed computing system 100 based on function as a service is not limited to performing the above-mentioned file conversion. In other embodiments, other types of cloud functions can also be executed to process file data. As long as the uploaded file data cannot be processed within the time limit, it can be processed through the distributed computing system 100 based on function as a service in this case.
[0033] Figure 2 shows a schematic diagram of a distributed computing process based on Function as a Service according to a preferred embodiment of this invention. Please refer to Figures 1 and 2. The distributed computing process 205 based on Function as a Service includes step 210, setting at least two processing layers for hierarchical control. In some embodiments, the hierarchical control module 140 plans the number of processing layers for file data to set hierarchical control based on the cloud function operation process to be performed on the file data uploaded by the ground element 150 through the network 200, thereby calling the corresponding cloud function to process the file data according to the number of processing layers executed. In some embodiments, the hierarchical control module 140 plans a first processing layer and a second processing layer according to the file data processing process.
[0034] Step 220: Determine whether to set additional processing layers. In some embodiments, this invention is mainly used for file data that cannot be processed under time constraints. Therefore, the first processing layer of the layered control module 140 is used to cut the file data. The second processing layer is the layer that performs corresponding cloud function processing on the cut file data. Therefore, in step 220, the layered control module 140 determines whether to set additional processing layers, that is, the layered control module 140 determines whether to set a third processing layer based on the processing flow of the file data. If no additional processing layers are required, then step 230 is executed, and the layered control module 140 executes the first processing layer to call a first cloud function to cut the file data into multiple clusters of file data according to a clustering rule. If additional processing layers are required, then step 240 is executed, and the layered control module 140 plans additional processing layers, such as a third layer, based on the file data, performs layered processing on the file data, and executes the first processing layer to call a first cloud function to cut the file data into multiple clusters of file data according to a clustering rule. Finally, in step 250, subsequent processing layers sequentially call the corresponding cloud functions to synchronously process the segmented complex group file data. In some embodiments, the hierarchical control module 140 calls the corresponding cloud functions to process the file data according to the number of processing layers executed.
[0035] In summary, this invention provides a distributed computing method and system based on Function as a Service (SaaS). By dynamically segmenting file data to reduce its size and performing layered processing, the amount of file data processed by a single function is reduced, thereby enabling data computation to be completed within time constraints. Therefore, it solves the problem of data processing being impossible to complete under limited time constraints due to excessively large single file data.
[0036] Although the present invention has been disclosed above by way of embodiments, it is not intended to limit the present invention. Anyone with ordinary knowledge in the art may make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims. [Simplified Explanation of the Diagram]
[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the specification, serve to explain the technical solutions of the embodiments of the present invention. Figure 1 is a schematic diagram of a distributed computing system based on Function as a Service according to a preferred embodiment of the present invention. Figure 2 is a schematic diagram of a distributed computing process based on Function as a Service according to a preferred embodiment of the present invention.
Claims
1. A distributed computing system based on Function as a Service, comprising: A first storage device for storing file data uploaded by a ground terminal element; A hierarchical control module, coupled to the first storage device, is used to plan a first processing layer and a second processing layer according to a processing flow of the file data, and to call corresponding cloud functions to process the file data according to the processing programs of the first processing layer and the second processing layer; a detection and judgment module, coupled to the first storage device, is used to determine the category of the file data, and to generate a clustering rule according to the category of the file data, wherein the hierarchical control module executes the processing program of the first processing layer and calls a first cloud function to segment the file data into multiple clusters of file data according to the clustering rule; a second storage device is used to store the multiple clusters of file data, and the hierarchical control module executes the processing program of the second processing layer and calls multiple second cloud functions to process the multiple clusters of file data synchronously; And a third storage device for storing the complex group file data after the complex second cloud function has been synchronously processed.
2. The distributed computing system based on Function as a Service as described in Request 1, wherein the detection and judgment module generates the clustering rule based on the processing capability of each of the plurality of second cloud functions for the file data category within a limited time.
3. A distributed computing system based on Function as a Service as described in claim 2, wherein the processing capacity is the amount of archive data that can be processed within the time limit.
4. The distributed computing system based on function as a service as described in claim 1, wherein the hierarchical control module further plans a third processing layer according to the processing flow of the file data.
5. The distributed computing system based on Function as a Service as described in claim 4, wherein the hierarchical control module executes the processing program of the third processing layer, and calls a plurality of third cloud functions to synchronously process the plurality of group file data after the plurality of second cloud functions have been synchronously processed.
6. The distributed computing system based on function as a service as described in claim 5, further comprising a fourth storage device for storing the complex group file data after being synchronously processed by the complex third cloud function.
7. The distributed computing system based on Function as a Service as described in claim 1, further comprising a difference analysis element coupled to the local element for comparing file data before and after uploading to generate a difference file.
8. A distributed computing method based on Function as a Service, comprising: A hierarchical control module is used to plan a first processing layer and a second processing layer according to a processing flow of a file data. The hierarchical control module determines whether a third processing layer needs to be set based on the processing flow of the file data. When the hierarchical control module determines that a third processing layer does not need to be set, the hierarchical control module executes the first processing layer to call a first cloud function to segment the file data into multiple clusters of file data according to a clustering rule. The hierarchical control module executes the second processing layer to call multiple second cloud functions to synchronously process the multiple clusters of file data. The hierarchical control module also uses a storage device to store the multiple clusters of file data synchronously processed by the multiple second cloud functions.
9. The distributed computing method based on Function as a Service as described in Request 8 further includes using a detection and judgment module to generate the clustering rule based on the processing capability of each of the multiple second cloud functions for the file data within a limited time.
Citation Information
Patent Citations
Archive classification general library realized based on cloud archive integrated platform
CN114201447A
Big data file analysis processing method and system in cloud computing environment
CN118535577A
Dispersing-type algorithm system applicable to image monitoring platform
TW201303753A
Job decomposition processing method for distributed computing
US20230350652A1