Data Parallel Processing Method, Apparatus, Computer Device, and Readable Storage Medium

By generating unique identification information in the real-time data processing engine, splitting data in parallel processing, combined with the results of blood relationship metadata merging, the problems of data processing timing and computing unit blocking are solved, and fast and orderly parallel data processing is achieved.

CN114691356BActive Publication Date: 2025-07-08ROOTCLOUD TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210223346.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-09
Publication Date
2025-07-08
Estimated Expiration
2042-03-09

AI Technical Summary

Technical Problem

When the existing real-time data processing engine is processed in parallel, it is difficult to ensure the timing requirements of the same type of data, and large-scale data calculations are likely to cause the calculation unit to block, affecting the processing progress.

Method used

By generating unique identification information split data into multiple sub-data, it is allocated to different data processors in parallel according to preset packet rules, and the blood metadata merge processing results are used to merge the processing results, and the ordered dictionary merge processing results are used to ensure timing and calculation efficiency.

Benefits of technology

It realizes the rapid and orderly processing of data in large-scale real-time data processing, avoids blocking of computing units, and ensures the integrity and accuracy of data processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114691356B_ABST
    Figure CN114691356B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a method, apparatus, computer device, and readable storage medium for data parallel processing. The method includes: obtaining real-time data to be processed; generating unique identification information corresponding to the data to be processed; splitting the data to be processed according to a pre-designed calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed, where each piece of sub-data is associated with lineage metadata corresponding to the data to be processed, and each piece of the lineage metadata includes the unique identification information; allocating all the sub-data to different data processors for parallel calculation according to a preset grouping rule; and combining the processing results fed back by each data processor according to a preset ordered dictionary and the lineage metadata to obtain a target processing result corresponding to the data to be processed. Through associating lineage metadata during data splitting, calculation, and combination, and in the way of grouping and modeling, the present application realizes fast and orderly parallel processing of large-scale real-time data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular, to a data parallel processing method, apparatus, computer device, and readable storage medium. Background Art

[0002] In existing real-time data processing engines, a parallel processing method is usually adopted to process data. Existing real-time data processing engines usually perform numbering processing according to the splitting order of the real-time data, and merge the processing results according to the numbering order.

[0003] There are mainly two major problems with existing real-time data processing engines: one is that there may be strict timing requirements before and after the calculation of the same type of data. For example, the user's operation timing is login-browsing-adding to the shopping cart - placing an order - logging out, corresponding to five operation logs with a fixed timing. However, when the current real-time data processing engine is executed in parallel in different computing units, the timing of the data processing results cannot be guaranteed. If all data of the same type is handed over to a single computing unit for execution during grouping, it often blocks the processing progress of the computing unit.

[0004] The second is that existing real-time data processing engines often store the fields of a large amount of data in a single computing unit for calculation when dealing with large-scale parallel computing problems, resulting in a long calculation time for a single piece of data and blocking the processing of subsequent data.

[0005] Therefore, there is an urgent need for a data parallel processing method that can quickly process real-time data to solve the problems in the actual operation of real-time data processing engines. Summary of the Invention

[0006] To solve the above technical problems, embodiments of the present application provide a data parallel processing method, apparatus, computer device, and readable storage medium. The specific solutions are as follows:

[0007] In a first aspect, embodiments of the present application provide a data parallel processing method, the method including:

[0008] Obtain real-time data to be processed;

[0009] Generate unique identification information corresponding to the data to be processed;

[0010] Split the data to be processed according to a pre-designed calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed, where each piece of sub-data is associated with lineage metadata corresponding to the data to be processed, and each piece of the lineage metadata includes the unique identification information;

[0011] According to a preset grouping rule, allocate all sub-data to different data processors for parallel calculation;

[0012] Merge the processing results fed back by each data processor according to the preset ordered dictionary and the genealogy metadata, so as to obtain the target processing result corresponding to the data to be processed.

[0013] According to a specific implementation manner of an embodiment of the present application, the step of splitting the data to be processed according to a preset calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed includes:

[0014] Obtain the task to be processed associated with the data to be processed, where the task to be processed includes a data input / output mode and a mapping calculation logic;

[0015] Generate a parallel calculation model of multiple data processors with grouping serial numbers according to the data input / output mode and the mapping calculation logic;

[0016] Divide the data to be processed according to the parallel calculation model to obtain the number of attribute grouping groups of the data to be processed and multiple pieces of sub-data with attribute grouping serial numbers.

[0017] According to a specific implementation manner of an embodiment of the present application, the step of allocating all sub-data to different data processors for parallel calculation according to a preset grouping rule includes:

[0018] Allocate the sub-data with interdependent calculation logics to the same sub-data group according to the parallel calculation model;

[0019] Allocate all sub-data to different data processors according to the calculation logic corresponding to each sub-data group and the grouping serial numbers of each data processor, so that each data processor performs parallel calculation.

[0020] According to a specific implementation manner of an embodiment of the present application, the step of splitting the data to be processed according to a preset calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed further includes:

[0021] Associate corresponding genealogy metadata with each piece of sub-data, where the genealogy metadata includes the unique identifier of the data to be processed to which it belongs, the attribute grouping serial number corresponding to the sub-data, and the number of grouping groups of the data to be processed to which it belongs. According to a specific implementation manner of an embodiment of the present application, the step of merging the processing results fed back by each data processor according to the preset ordered dictionary and the genealogy metadata to obtain the target processing result corresponding to the data to be processed includes:

[0022] Collect the processing results fed back by each data processor into a preset ordered dictionary, and the ordered dictionary is a storage space with a preset storage capacity;

[0023] Merge the processing results in the ordered dictionary according to the arrangement order of the grouping serial numbers of the attributes of the data to be processed and the number of groups, so as to obtain the target processing result.

[0024] According to a specific implementation manner of the embodiment of the present application, the step of merging the processing results fed back by each data processor according to the preset ordered dictionary and the blood relationship metadata to obtain the target processing result corresponding to the data to be processed further includes:

[0025] If the storage space occupied by the processing result in the ordered dictionary exceeds the preset capacity threshold, merge the processing result with the earliest order in the ordered dictionary.

[0026] In a second aspect, the embodiment of the present application further provides a data parallel processing device, and the device includes:

[0027] An acquisition module, configured to acquire real-time data to be processed;

[0028] A generation module, configured to generate unique identification information corresponding to the data to be processed;

[0029] A data splitting module, configured to split the data to be processed according to a preset calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed, where each piece of sub-data is associated with the blood relationship metadata corresponding to the data to be processed, and each piece of the blood relationship metadata includes the unique identification information;

[0030] A parallel computing module, configured to allocate all sub-data to different data processors for parallel computing according to a preset grouping rule;

[0031] A data merging module, configured to merge the processing results fed back by each data processor according to the preset ordered dictionary and the blood relationship metadata to obtain the target processing result corresponding to the data to be processed.

[0032] According to a specific implementation manner of the embodiment of the present application, the data splitting module is specifically configured to acquire a task to be processed associated with the data to be processed, and the task to be processed includes a data input / output mode and a mapping calculation logic;

[0033] Generate a parallel computing model of multiple data processors with grouping serial numbers according to the data input / output mode and the mapping calculation logic;

[0034] Divide the data to be processed according to the parallel computing model to obtain the number of attribute groups of the data to be processed and multiple pieces of sub-data with attribute grouping serial numbers.

[0035] In a third aspect, an embodiment of the present application further provides a computer device, including a processor and a memory. The memory stores a computer program, and when the computer program runs on the processor, it executes the data parallel processing method described in the first aspect.

[0036] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium, and when the computer program runs on a processor, it executes the data parallel processing method described in the first aspect.

[0037] An embodiment of the present application provides a data parallel processing method, apparatus, computer device, and readable storage medium. The method includes: obtaining real-time data to be processed; generating unique identification information corresponding to the data to be processed; splitting the data to be processed according to a pre-designed calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed, where each piece of sub-data is associated with lineage metadata corresponding to the data to be processed, and each piece of lineage metadata includes the unique identification information; allocating all the sub-data to different data processors for parallel calculation according to a preset grouping rule; and merging the processing results fed back by each data processor according to a preset ordered dictionary and the lineage metadata to obtain a target processing result corresponding to the data to be processed. Through the data parallel processing method of the present application, large-scale real-time data can be processed quickly and orderly in parallel. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the protection scope of the present invention. In each drawing, similar components are numbered similarly.

[0039] Figure 1 shows a schematic flowchart of a data parallel processing method provided by an embodiment of the present application;

[0040] Figure 2 shows a schematic diagram of device modules of a data parallel processing apparatus provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0042] The components of the embodiments of the present invention that are typically depicted and shown in the accompanying drawings herein may be arranged and designed in a variety of different configurations. Accordingly, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0043] Hereinafter, the terms "including", "having" and their cognates that may be used in various embodiments of the present invention are only intended to denote specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be construed as precluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or as adding the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.

[0044] In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0045] Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the various embodiments of the present invention belong. The terms (such as those defined in a commonly used dictionary) will be construed to have the same meaning as the contextual meaning in the relevant technical field and will not be construed to have an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present invention.

[0046] Reference Figure 1 , is a schematic flowchart of a data parallel processing method provided for an embodiment of the present application. The data parallel processing method provided for the embodiment of the present application, as Figure 1 shown, the data parallel processing method includes:

[0047] Step S101, obtaining real-time data to be processed;

[0048] In a specific embodiment, the data parallel processing method can be applied to various real-time data processing engines, and the real-time data processing engines can simultaneously perform data processing in different computing units to maximize the utilization rate of computing resources.

[0049] A general real-time data processing engine can, according to user-defined real-time data processing logic, implement reading data in a data source, parsing and converting the type of the original data according to the mode of the input data, executing mapping calculation logic to obtain output data, and writing the output data into a data sink according to a preset mode.

[0050] The real-time data processor engine of this embodiment can also read the real-time data to be processed from the data source. Specifically, the data to be processed is associated with corresponding tasks to be processed.

[0051] The data source can be any database or database server in the prior art.

[0052] Specifically, the real-time data processing engine of this embodiment can also obtain the corresponding real-time data to be processed from the data source according to the type of the task to be processed.

[0053] The real-time data to be processed can be a single piece of data or a large-scale data set. The scale of the real-time data to be processed is not specifically limited here.

[0054] Step S102, generate unique identification information corresponding to the data to be processed;

[0055] In a specific embodiment, after obtaining the real-time data to be processed, the data splitter in the real-time data processing engine of this embodiment will automatically generate corresponding unique identification information for the data to be processed.

[0056] Specifically, the unique identification information can be a data field with a fixed number of bytes. When splitting the data to be processed, each piece of sub-data split from the data to be processed will be associated with the unique identification information as the attribution identification information of the sub-data, so that other data processing structures in the data processing engine can classify and merge the sub-data according to the attribution identification information.

[0057] Step S103, split the data to be processed according to a pre-designed calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed, where each piece of sub-data is associated with lineage metadata corresponding to the data to be processed, and each piece of the lineage metadata includes the unique identification information;

[0058] In a specific embodiment, after the data splitter in the real-time data processing engine of this embodiment generates the unique identification information of the data to be processed, it starts to perform a preset splitting process on the data to be processed.

[0059] Before splitting the data to be processed, the data splitter first needs to group the input-output models and mapping calculation logics of each data processor in the real-time data processing engine to ensure that the sub-data with computational logic dependencies can be placed in the same data processor.

[0060] The data splitter splits the data to be processed according to the input-output models and mapping calculation logics of different groups, and distributes the multiple sub-data after splitting to different data processors for parallel execution. Since the attributes between the input-output models and mapping calculation logics of different groups are independent of each other, after splitting the data to be processed into multiple sub-data, each data processor can process the received sub-data in parallel.

[0061] According to a specific implementation manner of an embodiment of the present application, the step of splitting the data to be processed according to a pre-designed calculation model to obtain multiple sub-data corresponding to the data to be processed includes:

[0062] Obtain the task to be processed associated with the data to be processed, where the task to be processed includes a data input-output mode and a mapping calculation logic;

[0063] Generate a parallel calculation model of multiple data processors with group numbers according to the data input-output mode and the mapping calculation logic;

[0064] Divide the data to be processed according to the parallel calculation model to obtain the number of attribute grouping groups of the data to be processed and multiple sub-data with attribute group numbers.

[0065] In a specific embodiment, the real-time data engine of this embodiment constructs a calculation model according to the input mode, output mode, and the mapping relationship between the input mode and the output mode.

[0066] Specifically, the input-output mode, that is, the input schema and the output schema, includes schema objects in the mode, and the schema objects can be tables, columns, data types, views, stored procedures, relationships, primary keys, and foreign keys, etc. The components of the input schema and the output schema include element declarations, attribute declarations, simple and complex data types, model groups and attribute groups, attribute usage, and element particles.

[0067] In this embodiment, before processing the real-time data to be processed by the real-time data engine, it is necessary to preset the input schema and output schema of the real-time data engine, and perform pre-grouping according to the calculation mapping relationship between the input schema and the output schema. Thus, it can be ensured that after the real-time data engine obtains the task to be processed of the data to be processed, it can group and split the calculation logic of the task to be processed according to the preset input schema, output schema, and the mapping calculation logic between the two.

[0068] Specifically, during the real-time data processing, the real-time data processing engine can also generate a parallel computing model with multiple data processors having grouping serial numbers according to the input-output mode and mapping calculation logic of the real-time task to be processed.

[0069] For example, the definitions of the input-output mode and mapping calculation logic can be shown in the following table:

[0070] Table 1

[0071]

[0072] Table 2

[0073]

[0074]

[0075] In a specific embodiment, as shown in Table 1 and Table 2, the input mode of the task to be processed can be split according to the attribute declaration and attribute type, and the input mode is divided into the reference type Number corresponding to the numerical value, the reference type Boolean corresponding to the Boolean value, and the object wrapper type String of the string.

[0076] According to the attribute declaration and attribute type of the output mode, it is split, and the output mode is divided into the reference type Number corresponding to the numerical value and the object wrapper type String of the string.

[0077] According to the calculation logic of the task to be processed, it is grouped and split. In Table 2, the complete calculation logic of the task to be processed is split into two calculation rules, and corresponding serial numbers are associated with the calculation rules.

[0078] When grouping each attribute in the input mode, the attributes with interrelated calculation logics are grouped into one group. As shown in Table 1 and Table 2, the calculation rule corresponding to the output attribute x is interrelated with the input attributes a, b, and c, so the serial number 1 is assigned to all the input attributes a, b, and c. The calculation rule corresponding to the output attribute y is interrelated with the input attributes b, c, d, and e, so the serial number 2 is assigned to all the input attributes b, c, d, and e.

[0079] In this embodiment, the parallel computing model includes the above input-output mode definition and mapping calculation logic, and the parallel computing model will generate data processors corresponding to the number of serial numbers. By processing the real-time data to be processed through the parallel computing model, the data to be processed can be grouped and divided according to the attribute type and grouping serial number of the input model, so as to obtain the number of grouped sets of the data to be processed and multiple sub-data with attribute grouping serial numbers.

[0080] According to a specific implementation manner of an embodiment of the present application, the step of splitting the data to be processed according to a preset calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed further includes:

[0081] Associate corresponding lineage metadata with each piece of sub-data, where the lineage metadata includes the unique identifier of the data to be processed to which it belongs, the attribute grouping serial number corresponding to the sub-data, and the number of grouping sets of the data to be processed to which it belongs.

[0082] In a specific embodiment, after the data splitter splits the data to be processed, it will also associate corresponding lineage metadata with each piece of sub-data. After associating corresponding lineage metadata with each piece of sub-data, lineage analysis can be performed based on each piece of sub-data.

[0083] The lineage analysis is used for comprehensive tracking of the data processing process, so as to find all related metadata objects starting from a certain data object and the relationships between these metadata objects. The relationships between metadata objects specifically refer to the data flow input-output relationships representing these metadata objects.

[0084] In this embodiment, the lineage metadata includes the unique identifier information of the data to be processed to which the sub-data belongs, the attribute grouping serial number corresponding to the sub-data, and the number of grouping sets of the data to be processed to which the sub-data belongs. Thus, when each data processing finishes processing the sub-data, data merging processing will be performed on all output results according to the associated lineage metadata in the processing results.

[0085] It should be noted that when the data splitter splits the data to be processed, null-valued sub-data may appear in the split sub-data. For the split null-valued sub-data, the data splitter will also associate lineage metadata with it. After the null-valued sub-data enters the data processor following a preset grouping, the output result of the data processor is also associated with lineage metadata, so as to ensure that no part of the data to be processed is missed during data merging, and to ensure the integrity and accuracy of the parallel processing output result.

[0086] Step S104, according to a preset grouping rule, allocate all sub-data to different data processors for parallel calculation;

[0087] In a specific embodiment, after the data splitter splits the data to be processed according to the calculation model to obtain multiple pieces of sub-data associated with lineage metadata, it will group the data processing units according to the data ownership identifier and data grouping serial number of the sub-data.

[0088] After the grouping process of the sub-data is completed, the sub-data groups of different groups are dispersed to different parallel degrees for calculation, that is, they are allocated to different data processors for parallel calculation. Since the scale of the sub-data group is smaller than the scale of the data to be processed, the data processing parallelism will not occupy the same computing resources as the data to be processed.

[0089] According to a specific implementation manner of the embodiment of the present application, the step of allocating all sub-data to different data processors for parallel calculation according to a preset grouping rule includes:

[0090] Allocate the sub-data whose calculation logics are interdependent to the same sub-data group according to the parallel calculation model;

[0091] Allocate all sub-data to different data processors according to the calculation logics corresponding to each sub-data group and the grouping serial numbers of each data processor, so that each data processor performs parallel calculation.

[0092] In a specific embodiment, the specific structure of the parallel calculation model may refer to the description in the above embodiment and will not be elaborated here.

[0093] According to the parallel calculation model, multiple sub-data whose calculation logics are interdependent can be allocated to the same data group to obtain multiple sub-data groups corresponding to the grouping serial numbers of the data processors.

[0094] After the grouping process of the sub-data is completed, the data splitter will allocate each sub-data group to different data processors for each data processor to perform parallel processing.

[0095] By allocating the sub-data whose calculation logics are interdependent to the same data processor and numbering each data processor, the data processing speed can be increased while ensuring that the overall data processing logic will not be chaotic.

[0096] Step S105, merge the processing results fed back by each data processor according to the preset ordered dictionary and the lineage metadata to obtain the target processing result corresponding to the data to be processed.

[0097] In a specific embodiment, after the data processing unit in the real-time data processing engine of this embodiment calculates the output results of each group of sub-data groups, it will send each output result to the data merger to perform the data merging operation.

[0098] Specifically, the data merger will pre - establish an ordered dictionary as the storage space of the data merger. The keys of the ordered dictionary are the unique identifiers of the data, and the values are the combined data structures composed of the output attribute list set and the expected number of groups. When the data merger receives an output data, it will go to the corresponding value structure according to the ordered dictionary, merge the output data into the attribute list set, and then check whether the data grouping serial number and the expected number of groups included in the attribute list are the same as the grouping number of the lineage metadata. If the data grouping serial number and the expected number of groups included in the attribute list are the same as the grouping number of the lineage metadata, the data merger outputs the data merger result.

[0099] According to a specific implementation manner of an embodiment of the present application, the step of merging the processing results fed back by each data processor according to a preset ordered dictionary and the lineage metadata to obtain a target processing result corresponding to the data to be processed includes:

[0100] Collect the processing results fed back by each data processor into a preset ordered dictionary, where the ordered dictionary is a storage space with a preset storage capacity;

[0101] Merge the processing results in the ordered dictionary according to the arrangement order of the attribute grouping serial numbers of the data to be processed and the number of groups to obtain the target processing result.

[0102] In a specific embodiment, after receiving the output result sent by the data processor, the data merger will send the data output result to the ordered dictionary according to the lineage metadata of the data output result, and the ordered dictionary can assign a corresponding timestamp to each output result according to the time sequence of receiving the output result.

[0103] When the data output results in the ordered dictionary meet the corresponding output conditions, that is, when the number of data output results in the ordered dictionary is the same as the number of groups of the data to be processed, the data merger merges all the data output results with the same lineage metadata in the ordered dictionary to obtain the processing result of the data to be processed.

[0104] According to a specific implementation manner of an embodiment of the present application, the step of merging the processing results fed back by each data processor according to a preset ordered dictionary and the lineage metadata to obtain a target processing result corresponding to the data to be processed further includes:

[0105] If the storage space occupied by the processing results in the ordered dictionary exceeds the preset capacity threshold, merge the processing results at the earliest order in the ordered dictionary.

[0106] In a specific embodiment, considering that the output results processed by different data processors may arrive at the data combiner in batches, without a certain protection mechanism for the status of the data combiner, the un-settled status data that needs to be temporarily stored may surge in a short period of time, eventually exhausting the storage space of the ordered dictionary in the data combiner.

[0107] In this embodiment, an emergency processing rule is set for the data combiner of the real-time data processing engine. If the storage space occupied by the output results in the ordered dictionary exceeds the preset capacity threshold, that is, when the calculation results of one of the sub-data groups of the data to be processed have not been output to the data combiner, the data combiner will merge the data in the ordered dictionary in advance.

[0108] Specifically, the ordered dictionary will automatically call the processing result with the earliest order in the storage space, and merge and settle the processing result with the earliest order in advance. Among them, the processing result with the earliest order is the data output result that enters the ordered dictionary earliest according to the time stamp. For the part of the output result that has not been obtained, it is processed with a null output result.

[0109] Thus, through the emergency processing rule of the data combiner, it can be ensured that the ordered dictionary space of the data processor is in a limited and stable range, and the calculation of the entire real-time data processing engine will not be blocked due to slow processing of one piece of data.

[0110] The data parallel processing method in this application can correctly and effectively process large-scale real-time data through a preset calculation model splitting method, and can ensure that there is no data skew problem in the computing resources. By combining the merging method of lineage metadata, parallel processing can be carried out correctly and quickly, optimizing the traditional data parallel processing process.

[0111] Reference Figure 2 , is a schematic diagram of device module 200 of a data parallel processing device provided by an embodiment of this application. The data parallel processing device 200 provided by the embodiment of this application, as Figure 2 shown, the data parallel processing device 200 includes:

[0112] An acquisition module 201, configured to acquire real-time data to be processed;

[0113] A generation module 202, configured to generate unique identification information corresponding to the data to be processed;

[0114] A data splitting module 203, configured to split the data to be processed according to a preset calculation model to obtain multiple sub-data corresponding to the data to be processed, where each sub-data is associated with the lineage metadata corresponding to the data to be processed, and each piece of the lineage metadata includes the unique identification information;

[0115] A parallel computing module 204, configured to allocate all sub-data to different data processors for parallel computing according to a preset grouping rule;

[0116] A data merging module 205, configured to merge the processing results fed back by each data processor according to a preset ordered dictionary and the lineage metadata, so as to obtain a target processing result corresponding to the data to be processed.

[0117] According to a specific implementation manner of an embodiment of the present application, the data splitting module 203 is specifically configured to obtain a to-be-processed task associated with the to-be-processed data, where the to-be-processed task includes a data input / output mode and a mapping calculation logic;

[0118] Generate a parallel computing model of multiple data processors with grouping serial numbers according to the data input / output mode and the mapping calculation logic;

[0119] Divide the to-be-processed data according to the parallel computing model, so as to obtain the number of attribute grouping groups of the to-be-processed data and multiple pieces of sub-data with attribute grouping serial numbers.

[0120] In addition, an embodiment of the present application further provides a computer device, including a processor and a memory, where the memory stores a computer program, and when the computer program runs on the processor, it executes the data parallel processing method in the above embodiment.

[0121] An embodiment of the present application further provides a computer-readable storage medium, where a computer program is stored in the computer-readable storage medium, and when the computer program runs on a processor, it executes the data parallel processing method in the above embodiment.

[0122] In summary, an embodiment of the present application provides a data parallel processing method, apparatus, computer device, and readable storage medium, which can process large-scale data sets according to preset splitting and merging methods, and perform serial number arrangement according to the attribute types of the data, and can ensure that the data output result can retain the original time sequence; by associating the lineage metadata, it can ensure that the real-time data processing engine can correctly process the situation of partial attribute data; and on the premise of ensuring the correctness of the splitting and merging logic, it can effectively control the stable non-overflow of the storage space of the real-time data processing engine. In addition, for the specific implementation manners of the data parallel processing apparatus, computer device, and computer-readable storage medium provided in the embodiments of the present application, reference may be made to the specific implementation manners of the above method embodiments, which will not be elaborated herein one by one.

[0123] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the structure diagram and / or flowchart, as well as the combination of blocks in the structure diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0124] In addition, in each embodiment of the present invention, the various functional modules or units can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.

[0125] If the above functions are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0126] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. A data parallel processing method, characterized in that, The method includes: obtaining real-time data to be processed; generating unique identification information corresponding to the data to be processed; splitting the data to be processed according to a pre-designed calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed, where each piece of sub-data is associated with lineage metadata corresponding to the data to be processed, and each piece of the lineage metadata includes the unique identification information; allocating all the sub-data to different data processors for parallel calculation according to a preset grouping rule; merging the processing results fed back by each data processor according to a preset ordered dictionary and the lineage metadata to obtain a target processing result corresponding to the data to be processed; wherein the step of splitting the data to be processed according to a pre-designed calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed includes: associating corresponding lineage metadata with each piece of sub-data, and the lineage metadata includes the unique identification of the data to be processed to which it belongs, the attribute grouping serial number corresponding to the sub-data, and the number of grouping sets of the data to be processed to which it belongs; wherein associating corresponding lineage metadata with each piece of sub-data is used for performing lineage analysis based on each piece of sub-data, and the lineage analysis is used for comprehensively tracking the data processing process to determine all relevant metadata objects starting from a certain data object and the relationships between the metadata objects; the step of merging the processing results fed back by each data processor according to a preset ordered dictionary and the lineage metadata to obtain a target processing result corresponding to the data to be processed includes: collecting the processing results fed back by each data processor into a preset ordered dictionary, and the ordered dictionary is a storage space with a preset storage capacity. If the storage space occupied by the processing results in the ordered dictionary exceeds the preset capacity threshold, the ordered dictionary will automatically call the processing result at the forefront of the storage space to perform a merge settlement on the processing result at the forefront in advance. The processing result at the forefront is the data output result that enters the ordered dictionary earliest according to the timestamp. For the part of the output result that has not been obtained, it is processed with a null value output result; merging the processing results in the ordered dictionary according to the arrangement order of the attribute grouping serial numbers of the data to be processed and the number of grouping sets to obtain the target processing result.

2. The method according to claim 1, wherein the step of splitting the data to be processed according to a pre-designed calculation model to obtain multiple pieces of sub-data corresponding to the data to be processed includes: obtaining a task to be processed associated with the data to be processed, and the task to be processed includes a data input / output mode and mapping calculation logic; generating a parallel calculation model of multiple data processors with grouping serial numbers according to the data input / output mode and the mapping calculation logic; dividing the data to be processed according to the parallel calculation model to obtain the number of attribute grouping sets of the data to be processed and multiple pieces of sub-data with attribute grouping serial numbers.

3. The method according to claim 2, wherein the step of allocating all the sub-data to different data processors for parallel calculation according to a preset grouping rule includes: allocating the sub-data with interdependent calculation logics to the same sub-data group according to the parallel calculation model; All sub-data are allocated to different data processors according to the calculation logics corresponding to the respective sub-data groups and the grouping serial numbers of the respective data processors, so that the respective data processors perform parallel calculations.

4. A data parallel processing device, characterized in that, The device includes: an acquisition module configured to acquire real-time data to be processed; a generation module configured to generate unique identification information corresponding to the data to be processed; a data splitting module configured to split the data to be processed according to a pre-designed calculation model to obtain multiple sub-data corresponding to the data to be processed, wherein each sub-data is associated with lineage metadata corresponding to the data to be processed, and each of the lineage metadata includes the unique identification information; a parallel calculation module configured to allocate all sub-data to different data processors for parallel calculation according to a preset grouping rule; a data merging module configured to merge the processing results fed back by the respective data processors according to a preset ordered dictionary and the lineage metadata to obtain a target processing result corresponding to the data to be processed; wherein the data splitting module is specifically configured to: associate corresponding lineage metadata with each sub-data, the lineage metadata including the unique identification of the data to be processed to which it belongs, the attribute grouping serial number corresponding to the sub-data, and the number of grouping sets of the data to be processed to which it belongs; wherein associating corresponding lineage metadata with each sub-data is used for performing lineage analysis according to the respective sub-data, and the lineage analysis is used for comprehensively tracking the data processing process to determine all relevant metadata objects starting from a certain data object and the relationships between these metadata objects; The data merging module is specifically configured to: aggregate the processing results fed back by the respective data processors into a preset ordered dictionary, the ordered dictionary being a storage space having a preset storage capacity. If the storage space occupied by the processing results in the ordered dictionary exceeds a preset capacity threshold, the ordered dictionary automatically calls the processing result at the forefront in the storage space and performs a merge settlement on the processing result at the forefront in advance. The processing result at the forefront is the data output result that enters the ordered dictionary earliest according to the timestamp judgment. For the part of the output result that has not been obtained, it is processed with a null value output result; the processing results in the ordered dictionary are merged according to the arrangement order of the attribute grouping serial numbers of the data to be processed and the number of grouping sets to obtain the target processing result.

5. The device according to claim 4, wherein The data splitting module is specifically configured to obtain a task to be processed associated with the data to be processed, the task to be processed including a data input / output mode and a mapping calculation logic; generate a parallel calculation model of multiple data processors with grouping serial numbers according to the data input / output mode and the mapping calculation logic; divide the data to be processed according to the parallel calculation model to obtain the number of attribute grouping sets of the data to be processed and multiple sub-data with attribute grouping serial numbers.

6. A computer device, characterized in that, It includes a processor and a memory, and the memory stores a computer program which, when running on the processor, executes the data parallel processing method according to any one of claims 1-3.

7. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program runs on a processor, it executes the data parallel processing method according to any one of claims 1-3.

Citation Information

Patent Citations

  • Data processing method and device, scheduling server and medium

    CN111459659A