Data processing method and device, electronic equipment, storage medium and product

By dynamically allocating data processing tasks to appropriate computing nodes in a distributed system and utilizing preset processing models, the inefficiency problem of traditional data processing methods is solved, and fast and efficient processing and analysis of large-scale data is achieved.

CN120743519APending Publication Date: 2025-10-03CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510855836.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional data processing methods have problems of delay and low efficiency when processing large-scale data, especially because the manually designed feature extraction rules have poor adaptability, resulting in low efficiency in processing data of different fields and types.

Method used

By obtaining the resource information of the data to be processed and the computing nodes in the distributed system, determining the data type and selecting the target processing model, dynamically allocating tasks to the processing nodes for calculation, and using the preset processing model to improve data processing efficiency.

Benefits of technology

It enables rapid processing of large-scale data, improves data processing efficiency, and supports efficient analysis of different types of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120743519A_ABST
    Figure CN120743519A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method and device, electronic equipment, a storage medium and a product, relates to the technical field of data processing, and can be applied to a medical service scene or a financial service scene. The method comprises the steps of obtaining multiple pieces of to-be-processed data required for executing a to-be-processed task and residual resources of each computing node; determining data types of the multiple pieces of to-be-processed data, wherein the data types are used for representing association relationships among the multiple pieces of to-be-processed data; determining a target processing model from a plurality of preset processing models according to the data type; determining a target computing node in the plurality of computing nodes according to the model identifier of the target processing model, the residual resources of each computing node and the resource usage amount for executing the to-be-processed task; and calculating the plurality of pieces of data to be processed through a target processing model deployed by the target calculation node to obtain an execution result corresponding to the task to be processed. The embodiment of the invention can improve the data processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data processing method, device, electronic device, storage medium and product. Background Art

[0002] Data processing refers to the process of processing and analyzing acquired data, with the goal of transforming disordered and fragmented data into valuable information to support decision-making. However, due to hardware resource limitations, when processing large-scale data (such as medical data or financial data), traditional data processing methods have significant delays and low data processing efficiency. In addition, traditional data processing methods are usually based on manually designed feature extraction rules to extract data features, and then perform subsequent processing and analysis based on the extracted features. Due to the poor universality of manually designed feature extraction rules, when processing data from different fields and types, it is necessary to adaptively adjust the feature extraction rules in advance according to the field and type to which the data belongs, which also leads to the problem of low data processing efficiency in existing data processing methods. Summary of the Invention

[0003] The main purpose of the embodiments of the present application is to propose a data processing method, device, electronic device, storage medium and product, aiming to improve data processing efficiency.

[0004] To achieve the above objectives, a first aspect of an embodiment of the present application provides a data processing method, which is applied to a distributed system including multiple computing nodes. The method includes:

[0005] Acquire a plurality of to-be-processed data required to execute the to-be-processed task, and the remaining resources of each of the computing nodes;

[0006] Determining data types of the plurality of data to be processed, where the data types are used to characterize association relationships between the plurality of data to be processed;

[0007] determining a target processing model from a plurality of preset processing models according to the data type;

[0008] Determining a target computing node among the plurality of computing nodes according to a model identifier of the target processing model, remaining resources of each computing node, and resource usage for executing the task to be processed;

[0009] The target processing model deployed on the target computing node is used to calculate the plurality of data to be processed, and obtain the execution results corresponding to the tasks to be processed.

[0010] In some embodiments, determining a target computing node among the plurality of computing nodes based on the model identifier of the target processing model, the remaining resources of each computing node, and the resource usage for executing the task to be processed includes:

[0011] Determining, according to the model identifier of the target processing model, a plurality of candidate computing nodes from the plurality of computing nodes, the candidate computing nodes being computing nodes on which the target processing model is deployed among the plurality of computing nodes;

[0012] A target computing node among the plurality of computing nodes to be selected is determined according to the remaining resources of each computing node to be selected and the resource usage for executing the task to be processed.

[0013] In some embodiments, the number of the to-be-selected computing nodes is N, where N is an integer greater than 1;

[0014] The step of determining a target computing node from among the plurality of computing nodes to be selected based on the remaining resources of each computing node to be selected and the resource usage for executing the task to be processed comprises:

[0015] If there is a candidate computing node whose remaining resources are greater than the resource usage among the N candidate computing nodes, selecting any one candidate computing node whose remaining resources are greater than the resource usage from the N candidate computing nodes as the target computing node;

[0016] If there is no candidate computing node among the N candidate computing nodes whose remaining resources are greater than the resource usage, the task to be processed is split into multiple subtasks, and based on the resource usage of each subtask and the remaining resources of each candidate computing node, M candidate computing nodes are determined from the N candidate computing nodes as the target computing nodes, where M is an integer greater than 1 and less than or equal to N.

[0017] In some embodiments, determining a target processing model from a plurality of preset processing models according to the data type includes:

[0018] Matching the data type with a plurality of preset types in a preset relationship to obtain a model identifier corresponding to the data type, wherein the preset relationship includes the plurality of preset types and a model identifier corresponding to each preset type;

[0019] The target processing model is determined according to the model identifier corresponding to the data type.

[0020] In some embodiments, obtaining a plurality of to-be-processed data required to execute a to-be-processed task includes:

[0021] Acquire multiple initial data corresponding to the task to be processed;

[0022] Performing preset processing on the plurality of initial data to obtain the plurality of data to be processed, wherein the preset processing includes at least one of the following:

[0023] Determining duplicate data in the plurality of initial data, and deduplicating the duplicate data;

[0024] Identifying a plurality of data missing locations in the initial data, and performing data filling on the data missing locations;

[0025] Performing clustering processing on the plurality of the initial data by using a clustering algorithm, determining outliers in the plurality of the initial data, and deleting the outliers;

[0026] Based on the preset data screening conditions, the plurality of initial data are screened.

[0027] In some embodiments, the data screening condition is obtained according to the following process:

[0028] receiving a configuration operation for the data screening condition;

[0029] In response to the configuration operation, the data screening condition is determined according to the configuration information indicated by the configuration operation, the data screening condition adopts a Boolean logic relationship or a regular expression, and the configuration information is used to fill in the Boolean logic relationship or the regular expression.

[0030] To achieve the above-mentioned object, a second aspect of an embodiment of the present application provides a data processing device, which is applied to a distributed system, wherein the distributed system includes multiple computing nodes, and the device includes:

[0031] An acquisition module, configured to acquire a plurality of to-be-processed data required to execute the to-be-processed task, and the remaining resources of each of the computing nodes;

[0032] A first determining module is used to determine the data types of the plurality of data to be processed, wherein the data types are used to represent the association relationship between the plurality of data to be processed;

[0033] a second determining module, configured to determine a target processing model from a plurality of preset processing models according to the data type;

[0034] A third determining module is configured to determine a target computing node among the plurality of computing nodes according to a model identifier of the target processing model, remaining resources of each computing node, and resource usage for executing the task to be processed;

[0035] The computing module is used to compute the plurality of the to-be-processed data through the target processing model deployed by the target computing node, and obtain the execution results corresponding to the to-be-processed tasks.

[0036] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and the processor implements the data processing method described in the first aspect when executing the computer program.

[0037] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the data processing method described in the first aspect above.

[0038] To achieve the above-mentioned purpose, the fifth aspect of the embodiments of the present application proposes a computer program product. When the instructions in the computer program product are executed by an electronic device, the electronic device executes the data processing method described in the first aspect above.

[0039] The data processing method, device, electronic device, storage medium and product proposed in this application, after obtaining multiple data to be processed required to execute the task to be processed and the remaining resources of each computing node, determine the data type of the multiple data to be processed, and then determine the target processing model from multiple preset processing models based on the data type. Then, based on the model identifier of the target processing model, the remaining resources of each computing node and the resource usage for executing the task to be processed, determine the target computing node among the multiple computing nodes, and finally calculate the multiple data to be processed through the target processing model deployed by the target computing node to obtain the execution result corresponding to the task to be processed. The above steps, by dynamically allocating the task to be processed to the processing computing node, and the target processing model deployed by the processing computing node to calculate the multiple data to be processed, can still achieve rapid processing of the data to be processed for large-scale data to be processed, thereby improving data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is one of the flow charts of the data processing method provided in the embodiment of the present application;

[0041] Figure 2 This is the second flow chart of the data processing method provided in the embodiment of the present application;

[0042] Figure 3 This is the third flow chart of the data processing method provided in the embodiment of the present application;

[0043] Figure 4 This is the fourth flow chart of the data processing method provided in the embodiment of the present application;

[0044] Figure 5 is a structural diagram of a data processing device provided in an embodiment of the present application;

[0045] Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0047] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0049] First, let’s analyze some of the terms used in this application:

[0050] A distributed system is a system architecture consisting of multiple independent computing nodes interconnected through a network, working together to achieve a unified goal. Its core characteristic is the decomposition of tasks among multiple computing nodes, the sharing and coordination of data through communication mechanisms, and the ultimate presentation of the entire system as a single entity.

[0051] Compute node: The basic unit of a distributed system. A compute node is a physical or virtual device with independent computing power, storage resources, and network communication capabilities. These compute nodes are interconnected through a network and collaborate to process distributed tasks, achieving the overall goals of the distributed system.

[0052] Resource usage: In a distributed system, the consumption and occupation of system resources by a single computing node (such as a server, computer, virtual machine, etc.) when performing a specific task, usually reflected in the form of CPU usage, memory usage, and GPU usage.

[0053] Based on this, the embodiments of the present application provide a data processing method, device, electronic device, storage medium and product, aiming to improve data processing efficiency.

[0054] The data processing method, device, electronic device, storage medium and product provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the data processing method in the embodiments of the present application is described.

[0055] The data processing method provided in the embodiment of the present application relates to the field of data processing technology. The data processing method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the data processing method, etc., but is not limited to the above forms.

[0056] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0057] Figure 1 This is one of the flow charts of the data processing method provided in the embodiment of the present application. Please refer to Figure 1 The data processing method provided in the embodiment of the present application can be applied to a distributed system, which includes multiple computing nodes. Figure 1 The method may include but is not limited to steps S101 to S105.

[0058] Step S101, obtaining a plurality of to-be-processed data required to execute the to-be-processed task and the remaining resources of each computing node;

[0059] Step S102, determining data types of the plurality of data to be processed, wherein the data types are used to represent associations between the plurality of data to be processed;

[0060] Step S103, determining a target processing model from a plurality of preset processing models according to the data type;

[0061] Step S104, determining a target computing node among the plurality of computing nodes according to the model identifier of the target processing model, the remaining resources of each computing node, and the resource usage for executing the task to be processed;

[0062] Step S105 , calculating the plurality of data to be processed by the target processing model deployed on the target computing node to obtain the execution results corresponding to the tasks to be processed.

[0063] Among them, the task to be processed can be a data analysis task for financial data, a prediction task for medical data, or other data processing tasks, which are not limited here. The data to be processed is related to the task to be processed. For example, when the task to be processed is a data analysis task for financial data (such as stock data), multiple data to be processed (such as multiple stock data) required to execute the task to be processed can be obtained, and then the data processing method of the embodiment of the present application is used to perform subsequent calculations on the multiple data to be processed, complete the task to be processed, and obtain the execution result corresponding to the task to be processed (such as the changing trend of stock data). Alternatively, when the task to be processed is a prediction task for medical data, multiple data to be processed (such as multiple medical detection data) required to execute the task to be processed can be obtained, and then the data processing method of the embodiment of the present application is used to perform subsequent calculations on the multiple data to be processed, complete the task to be processed, and obtain the execution result corresponding to the task to be processed (such as the predicted disease risk level of the related object). Remaining resources refer to the hardware or system resources of the computing node that are not currently occupied by other tasks and can be used by the task to be processed.

[0064] In the data processing process, first, multiple pending data required for executing pending tasks and the remaining resources of each computing node can be obtained. Here, multiple pending data can be obtained from various data sources such as databases, comma-separated value files or json files, and an asynchronous data loading mechanism is adopted in the acquisition process to improve data reading efficiency. The remaining resources of each computing node can be determined based on the total resources of each computing node and the amount of resources used by each computing node. The amount of resources used by each computing node can be obtained after monitoring through the application programming interface (Application Programming Interface, API) or a third-party library. For example, if the total resources of a computing node are 10 gigabytes and the amount of resources used is 2 gigabytes, then the remaining resources of the computing node are 8 gigabytes.

[0065] Subsequently, a data type can be determined to characterize the association relationship between the multiple data to be processed, such as a time series data type, a sequence data type, a complex sequence relationship data type, or an image data type. Based on the data types of the multiple data to be processed, a target processing model can be determined from multiple preset processing models. The multiple preset processing models can be pre-trained using distributed training techniques. For example, the multiple preset processing models may include a first processing model, a second processing model, and a third processing model. The first processing model may adopt the architecture of a convolutional neural network model, the second processing model may adopt the architecture of a recurrent neural network model or a long short-term memory network model, and the third processing model may adopt the architecture of a transformer model. Here, the first processing model is suitable for processing data of a sequence or image data type, so the first processing model can be used as the processing model corresponding to the sequence or image data type. The second processing model is suitable for processing data of a time series data type, so the second processing model can be used as the processing model corresponding to the time series data type. The third processing model is suitable for processing data of a complex sequence relationship data type, so the third processing model can be used as the processing model corresponding to the complex sequence relationship data type. In this way, a correspondence between data types and processing models can be determined. After obtaining the data types of the plurality of data to be processed, a target processing model may be determined from a plurality of preset processing models based on the correspondence between the data types and the processing models and the data types of the plurality of data to be processed.

[0066] It is understandable that there is the possibility that some computing nodes have not deployed the target processing model. Therefore, after determining the target processing model, the target computing node among the multiple computing nodes can be determined based on the model identification of the target processing model, the remaining resources of each computing node, and the resource usage of executing the task to be processed. The model identification can be in the form of numbers, letters, or other forms that can distinguish different processing models. Exemplarily, the computing node deployed with the target processing model among the multiple computing nodes can be determined based on the model identification of the target processing model, and then the target computing node can be determined based on the remaining resources of the computing node deployed with the target processing model and the resource usage of executing the task to be processed. Alternatively, the computing node with sufficient remaining resources to execute the task to be processed can be determined based on the remaining resources of each computing node and the resource usage of executing the task to be processed. Subsequently, the computing node with sufficient remaining resources to execute the task to be processed can be used as the target computing node according to the model identification of the target processing model. Through the target processing model deployed by the target computing node, multiple data to be processed can be calculated, thereby obtaining the execution result corresponding to the task to be processed. It should be noted that when computing multiple data items, parallel processing is supported, meaning multi-threading and multi-processing techniques can be used to improve task processing efficiency. Furthermore, a feedback mechanism allows updates to the data processing process based on the completion status of pending tasks and the resource utilization of the distributed system.

[0067] In addition to the execution results, the target processing model can also output the confidence level corresponding to the execution results. When the task to be processed is a classification task, the target processing model can also output the probability distribution of multiple categories. In this way, based on the execution results, confidence levels or probability distributions output by the target processing model, corresponding visualization charts and analysis data can be generated. Here, the visualization charts and analysis data can be in a variety of formats (such as hypertext markup language format or portable document format), and the visualization charts and analysis data can be in the same format or in different formats.

[0068] In the steps S101 to S105 shown in the embodiment of the present application, after obtaining the multiple data to be processed required to execute the task to be processed and the remaining resources of each computing node, the data types of the multiple data to be processed are determined, and then according to the data types, the target processing model is determined from multiple preset processing models, and then according to the model identifier of the target processing model, the remaining resources of each computing node and the resource usage for executing the task to be processed, the target computing node among the multiple computing nodes is determined, and finally, the target processing model deployed by the target computing node is used to calculate the multiple data to be processed and obtain the execution results corresponding to the task to be processed. The above steps, by dynamically allocating the task to be processed to the processing nodes, and the target processing model deployed by the processing nodes to calculate the multiple data to be processed, can still achieve rapid processing of the data to be processed for large-scale data to be processed, thereby improving the processing efficiency of the data.

[0069] In some embodiments, obtaining a plurality of to-be-processed data required to execute a to-be-processed task includes:

[0070] Acquire multiple initial data corresponding to the task to be processed;

[0071] Performing preset processing on the plurality of initial data to obtain the plurality of data to be processed, wherein the preset processing includes at least one of the following:

[0072] Determining duplicate data in the plurality of initial data, and deduplicating the duplicate data;

[0073] Identifying a plurality of data missing locations in the initial data, and performing data filling on the data missing locations;

[0074] Performing clustering processing on the plurality of the initial data by using a clustering algorithm, determining outliers in the plurality of the initial data, and deleting the outliers;

[0075] Based on the preset data screening conditions, the plurality of initial data are screened.

[0076] Specifically, multiple initial data corresponding to the task to be processed can be obtained, and the multiple initial data can be preset processed to obtain multiple data to be processed. Here, the preset processing can include at least one of deduplication processing, missing value filling processing, outlier detection processing or screening processing. When the preset processing includes deduplication processing, the duplicate data in the multiple initial data can be determined, and the duplicate data can be deduplicated. For example, the duplicate data in the multiple initial data can be determined by a hash table or a Bloom filter, and the duplicate data can be deduplicated. When the preset processing includes missing value filling processing, the data missing positions in the multiple initial data can be identified, and the data missing positions can be filled with data. For example, interpolation method (such as K nearest neighbor interpolation or mean filling) or machine learning model (such as matrix decomposition) can be used to fill the data missing positions and complete the missing value filling processing. When the preset processing includes outlier detection processing, the multiple initial data are clustered by a clustering algorithm (such as a density-based clustering algorithm or a K-means clustering algorithm), the outliers in the multiple initial data are determined, and the outliers are deleted. In addition, outliers in multiple initial data can also be determined by a standard score algorithm or an interquartile range algorithm. When the preset processing includes a screening process, multiple initial data can be screened based on preset data screening conditions. In addition, a semi-supervised learning method can be used to clean multiple initial data using a small amount of labeled data. It should be noted that when the preset processing includes multiple items of deduplication processing, missing value filling processing, outlier detection processing or screening processing, the multiple processing can be performed simultaneously, or can be performed in a pre-set processing order to obtain multiple data to be processed. By processing multiple initial data, multiple data to be processed with higher data quality can be obtained, so that the calculation accuracy can be improved when the target processing model is used to calculate the multiple data to be processed.

[0077] In some embodiments, after obtaining multiple data to be processed, the multiple data to be processed can be standardized or normalized, and the processed multiple data to be processed can be encoded (for example, using one-hot encoding or target encoding), so as to convert the multiple data to be processed into a form that can be processed by the target processing model.

[0078] In some embodiments, the data screening condition is obtained according to the following process:

[0079] receiving a configuration operation for the data screening condition;

[0080] In response to the configuration operation, the data screening condition is determined according to the configuration information indicated by the configuration operation, the data screening condition adopts a Boolean logic relationship or a regular expression, and the configuration information is used to fill in the Boolean logic relationship or the regular expression.

[0081] The data screening conditions adopt Boolean logic relations or regular expressions and can be configured by relevant objects. Specifically, an interactive module that supports data visualization and interactive analysis is provided. The interactive module includes a human-computer interaction interface, which allows relevant objects to customize data analysis rules or data screening conditions. In the process of the relevant objects defining data screening conditions, configuration operations on the data screening conditions can be received. The configuration operation includes configuration information. Here, the configuration information includes at least one of the characters that should be included in the filtered data, the text format that should be met, or the conditional information that should be met. In this way, the data screening conditions can be obtained by filling the configuration information indicated by the configuration operation into the Boolean logic relations or the regular expression. In this way, when the preset processing includes screening processing, the screening of multiple initial data can be achieved by configuring the data screening conditions.

[0082] Figure 2 This is the second flow chart of the data processing method provided in the embodiment of the present application. Please refer to Figure 2 . Figure 2 The method includes but is not limited to steps S201 to S206, wherein:

[0083] Step S201, obtaining a plurality of to-be-processed data required to execute the to-be-processed task and the remaining resources of each computing node;

[0084] Step S202, determining data types of the plurality of data to be processed, wherein the data types are used to represent associations between the plurality of data to be processed;

[0085] Step S203: Match the data type with a plurality of preset types in a preset relationship to obtain a model identifier corresponding to the data type, wherein the preset relationship includes the plurality of preset types and a model identifier corresponding to each preset type;

[0086] Step S204, determining the target processing model according to the model identifier corresponding to the data type;

[0087] Step S205, determining a target computing node among the plurality of computing nodes according to the model identifier of the target processing model, the remaining resources of each computing node, and the resource usage for executing the task to be processed;

[0088] Step S206 , calculating the plurality of data to be processed by the target processing model deployed by the target computing node to obtain the execution results corresponding to the tasks to be processed.

[0089] The above steps S201 to S202 may refer to steps S101 to S102, and the above steps S205 to S206 may refer to steps S104 to S105, which will not be repeated here.

[0090] According to the data types of multiple data to be processed, the target processing model can be determined from multiple preset processing models. Specifically, a pre-built preset relationship is obtained. Here, the preset relationship includes multiple preset types and a model identifier corresponding to each preset type. Then, the data type can be matched with the multiple preset types in the preset relationship, and the model identifier corresponding to the successfully matched preset type can be used as the model identifier corresponding to the data type. Each model identifier corresponds to a preset processing model. Therefore, according to the model identifier corresponding to the data type, the preset processing model corresponding to the model identifier can be used as the target processing model. In this way, the target processing model can be determined, which facilitates the subsequent determination of the target computing node, so that the multiple data to be processed can be calculated through the target processing model deployed by the target computing node.

[0091] Figure 3 This is the third flow chart of the data processing method provided in the embodiment of the present application. Please refer to Figure 3 . Figure 3 The method includes but is not limited to steps S301 to S306, wherein:

[0092] Step S301, obtaining a plurality of to-be-processed data required to execute the to-be-processed task and the remaining resources of each computing node;

[0093] Step S302: determining data types of the plurality of data to be processed, wherein the data types are used to represent associations between the plurality of data to be processed;

[0094] Step S303, determining a target processing model from a plurality of preset processing models according to the data type;

[0095] Step S304: determining a plurality of candidate computing nodes from the plurality of computing nodes according to the model identifier of the target processing model, wherein the candidate computing nodes are computing nodes on which the target processing model is deployed among the plurality of computing nodes;

[0096] Step S305, determining a target computing node among the plurality of computing nodes to be selected based on the remaining resources of each computing node to be selected and the resource usage for executing the task to be processed;

[0097] Step S306 , calculating the plurality of data to be processed by the target processing model deployed on the target computing node to obtain the execution results corresponding to the tasks to be processed.

[0098] The above steps S301 to S303 may refer to steps S101 to S103, and step S306 may refer to step S105, which will not be repeated here.

[0099] There is a possibility that some computing nodes have not deployed the target processing model. Therefore, after determining the target processing model, the target computing node among the multiple computing nodes can be determined based on the model identifier of the target processing model, the remaining resources of each computing node, and the resource usage of executing the pending task. Exemplarily, the model identifier of the target processing model can be matched with the model identifier stored in each computing node. When the model identifier of the target processing model successfully matches any model identifier stored in a certain computing node, it can be determined that the target processing model is deployed in the computing node, and this computing node is then used as a candidate computing node. In this way, multiple candidate computing nodes can be obtained. Subsequently, the target computing node can be determined based on the remaining resources of the multiple candidate computing nodes and the resource usage of executing the pending task. In this way, the target computing node can be determined from the multiple computing nodes, which facilitates the subsequent processing of multiple pending data by the target processing model deployed on the target computing node, thereby improving data processing efficiency.

[0100] Figure 4 This is the fourth flow chart of the data processing method provided in the embodiment of the present application. Please refer to Figure 4 . Figure 4 The method includes but is not limited to steps S401 to S407, wherein:

[0101] Step S401, obtaining a plurality of to-be-processed data required to execute the to-be-processed task and the remaining resources of each computing node;

[0102] Step S402: determining data types of the plurality of data to be processed, wherein the data types are used to represent associations between the plurality of data to be processed;

[0103] Step S403, determining a target processing model from a plurality of preset processing models according to the data type;

[0104] Step S404: determining a plurality of candidate computing nodes from the plurality of computing nodes according to the model identifier of the target processing model, wherein the candidate computing nodes are computing nodes on which the target processing model is deployed among the plurality of computing nodes;

[0105] Step S405: If there is a candidate computing node whose remaining resources are greater than the resource usage among the N candidate computing nodes, any one candidate computing node whose remaining resources are greater than the resource usage is selected from the N candidate computing nodes as the target computing node;

[0106] Step S406: If there is no candidate computing node with remaining resources greater than the resource usage among the N candidate computing nodes, split the task to be processed into multiple subtasks, and determine M candidate computing nodes from the N candidate computing nodes as the target computing nodes based on the resource usage of each subtask and the remaining resources of each candidate computing node, where M is an integer greater than 1 and less than or equal to N.

[0107] Step S407 , calculating the plurality of data to be processed by the target processing model deployed by the target computing node to obtain the execution results corresponding to the tasks to be processed.

[0108] The above steps S401 to S404 may refer to steps S301 to S304, and step S407 may refer to step S306, which will not be repeated here.

[0109] A target computing node among the N candidate computing nodes can be determined based on the remaining resources of each candidate computing node and the resource usage of executing the pending task. N is an integer greater than 1. Specifically, if there is a candidate computing node among the N candidate computing nodes whose remaining resources are greater than the resource usage, in order to reduce the possibility of resource fragmentation of the distributed system (i.e., the dispersion of the available resources of the distributed system), any candidate computing node whose remaining resources are greater than the resource usage can be selected from the N candidate computing nodes as the target computing node.

[0110] If none of the N candidate compute nodes has remaining resources greater than its resource usage, the task to be processed can be split into multiple subtasks to improve processing efficiency for the multiple data items. Based on the resource usage of each subtask and the remaining resources of each candidate compute node, M candidate compute nodes are determined as target compute nodes from the N candidate compute nodes, with M being an integer greater than 1 and less than or equal to N, in accordance with the aforementioned process for selecting target compute nodes from the N candidate compute nodes.

[0111] In one example, after selecting M candidate computing nodes as target computing nodes, the target computing nodes can be used to compute the to-be-processed data corresponding to each of the multiple subtasks, thereby obtaining the execution results corresponding to each of the multiple subtasks. Here, the number of the multiple subtasks can be M. After merging the execution results corresponding to the multiple subtasks, the execution result corresponding to the to-be-processed task is obtained.

[0112] In this way, a target computing node can be determined from multiple computing nodes, so that multiple data to be processed can be processed through the target processing model deployed on the target computing node, thereby improving data processing efficiency.

[0113] The data processing method of this application, firstly, is based on a distributed system that uses deep learning technology, which can demonstrate higher efficiency in processing large-scale data. Secondly, by learning the complex relationships in the data, it can more accurately extract effective information, thereby accelerating the entire data analysis process. Thirdly, through real-time optimization, it can ensure that the distributed system fully utilizes hardware resources and can more quickly generate the execution results of the tasks to be processed.

[0114] Figure 5 This is a structural diagram of the data processing device provided in the embodiment of the present application. Figure 5 The present application also provides a data processing device 500, which is applied to a distributed system including multiple computing nodes. The data processing device 500 can implement the above-mentioned data processing method, and the device 500 includes:

[0115] An acquisition module 501 is configured to acquire a plurality of to-be-processed data required for executing a to-be-processed task and the remaining resources of each computing node;

[0116] A first determining module 502 is configured to determine a data type of the plurality of data to be processed, wherein the data type is used to represent an association relationship between the plurality of data to be processed;

[0117] A second determining module 503 is configured to determine a target processing model from a plurality of preset processing models according to the data type;

[0118] A third determining module 504 is configured to determine a target computing node among the plurality of computing nodes according to the model identifier of the target processing model, the remaining resources of each computing node, and the resource usage for executing the task to be processed;

[0119] The calculation module 505 is used to calculate the plurality of data to be processed through the target processing model deployed by the target computing node to obtain the execution result corresponding to the task to be processed.

[0120] In some embodiments, the third determining module 504 includes:

[0121] A first determining submodule is configured to determine, based on the model identifier of the target processing model, a plurality of candidate computing nodes from the plurality of computing nodes, the candidate computing nodes being computing nodes on which the target processing model is deployed among the plurality of computing nodes;

[0122] The second determining submodule determines a target computing node among the plurality of computing nodes to be selected according to the remaining resources of each computing node to be selected and the resource usage for executing the task to be processed.

[0123] In some embodiments, the second determining submodule includes:

[0124] A selection unit is configured to select any one of the N candidate computing nodes whose remaining resources are greater than the resource usage as the target computing node if there is a candidate computing node whose remaining resources are greater than the resource usage among the N candidate computing nodes;

[0125] A determination unit is used to split the task to be processed into multiple subtasks if there is no candidate computing node with remaining resources greater than the resource usage among the N candidate computing nodes, and determine M candidate computing nodes from the N candidate computing nodes as the target computing nodes based on the resource usage of each subtask and the remaining resources of each candidate computing node, where M is an integer greater than 1 and less than or equal to N.

[0126] In some embodiments, the second determining module 503 includes:

[0127] a matching submodule, configured to match the data type with a plurality of preset types in a preset relationship to obtain a model identifier corresponding to the data type, wherein the preset relationship includes the plurality of preset types and a model identifier corresponding to each preset type;

[0128] The third determining submodule is configured to determine the target processing model according to the model identifier corresponding to the data type.

[0129] In some embodiments, the acquisition module 501 includes:

[0130] An acquisition submodule, configured to acquire a plurality of initial data corresponding to the task to be processed;

[0131] The processing submodule is configured to perform a preset process on the plurality of initial data to obtain a plurality of the data to be processed, wherein the preset process includes at least one of the following:

[0132] Determining duplicate data in the plurality of initial data, and deduplicating the duplicate data;

[0133] Identifying a plurality of data missing locations in the initial data, and performing data filling on the data missing locations;

[0134] Performing clustering processing on the plurality of the initial data by using a clustering algorithm, determining outliers in the plurality of the initial data, and deleting the outliers;

[0135] Based on the preset data screening conditions, the plurality of initial data are screened.

[0136] In some embodiments, the data screening condition is obtained according to the following process:

[0137] receiving a configuration operation for the data screening condition;

[0138] In response to the configuration operation, the data screening condition is determined according to the configuration information indicated by the configuration operation, the data screening condition adopts a Boolean logic relationship or a regular expression, and the configuration information is used to fill in the Boolean logic relationship or the regular expression.

[0139] The specific implementation of the data processing device 500 can refer to the specific embodiment of the above-mentioned data processing method, which will not be described in detail here.

[0140] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned data processing method when executing the computer program. The electronic device can be any smart terminal including a tablet computer, an in-vehicle computer, or the like.

[0141] Figure 6 This is a hardware structure diagram of the electronic device provided in the embodiment of the present application. Figure 6 . Electronic equipment includes:

[0142] The processor 601 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0143] The memory 602 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 602 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 602 and is called by the processor 601 to execute the data processing method of the embodiments of this application.

[0144] Input / output interface 603, used to implement information input and output;

[0145] Communication interface 604, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0146] Bus 605 , which transmits information between various components of the device (e.g., processor 601 , memory 602 , input / output interface 603 , and communication interface 604 );

[0147] The processor 601 , the memory 602 , the input / output interface 603 and the communication interface 604 are connected to each other in communication within the device via a bus 605 .

[0148] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned data processing method is implemented.

[0149] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0150] The data processing method, device, electronic device, storage medium and product provided by the embodiment of the present application, after obtaining multiple data to be processed required to execute the task to be processed and the remaining resources of each computing node, determine the data type of the multiple data to be processed, and then determine the target processing model from multiple preset processing models based on the data type, and then determine the target computing node among the multiple computing nodes based on the model identifier of the target processing model, the remaining resources of each computing node and the resource usage for executing the task to be processed, and finally calculate the multiple data to be processed by the target processing model deployed by the target computing node to obtain the execution result corresponding to the task to be processed. The above steps, by dynamically allocating the task to be processed to the processing computing node, and the target processing model deployed by the processing computing node to calculate the multiple data to be processed, can still achieve rapid processing of the data to be processed for large-scale data to be processed, thereby improving data processing efficiency.

[0151] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0152] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0153] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0154] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0155] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0156] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0157] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0158] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0159] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0160] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0161] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A data processing method, characterized in that: Applied to a distributed system, the distributed system including a plurality of computing nodes, the method comprising: Acquire a plurality of to-be-processed data required to execute the to-be-processed task, and the remaining resources of each of the computing nodes; Determining data types of the plurality of data to be processed, where the data types are used to characterize association relationships between the plurality of data to be processed; determining a target processing model from a plurality of preset processing models according to the data type; determining a target computing node among the plurality of computing nodes according to a model identifier of the target processing model, remaining resources of each computing node, and resource usage for executing the task to be processed; The target processing model deployed on the target computing node is used to calculate the plurality of data to be processed, and obtain the execution results corresponding to the tasks to be processed.

2. The method according to claim 1, characterized in that The determining of a target computing node among the plurality of computing nodes according to the model identifier of the target processing model, the remaining resources of each computing node, and the resource usage for executing the task to be processed includes: Determining, according to the model identifier of the target processing model, a plurality of candidate computing nodes from the plurality of computing nodes, the candidate computing nodes being computing nodes on which the target processing model is deployed among the plurality of computing nodes; A target computing node among the plurality of computing nodes to be selected is determined according to the remaining resources of each computing node to be selected and the resource usage for executing the task to be processed.

3. The method according to claim 2, characterized in that The number of the computing nodes to be selected is N, where N is an integer greater than 1; The step of determining a target computing node from among the plurality of computing nodes to be selected based on the remaining resources of each computing node to be selected and the resource usage for executing the task to be processed includes: If there is a candidate computing node whose remaining resources are greater than the resource usage among the N candidate computing nodes, selecting any one candidate computing node whose remaining resources are greater than the resource usage from the N candidate computing nodes as the target computing node; If there is no candidate computing node among the N candidate computing nodes whose remaining resources are greater than the resource usage, the task to be processed is split into multiple subtasks, and based on the resource usage of each subtask and the remaining resources of each candidate computing node, M candidate computing nodes are determined from the N candidate computing nodes as the target computing nodes, where M is an integer greater than 1 and less than or equal to N.

4. The method according to claim 1, wherein Determining a target processing model from a plurality of preset processing models according to the data type includes: Matching the data type with a plurality of preset types in a preset relationship to obtain a model identifier corresponding to the data type, wherein the preset relationship includes the plurality of preset types and a model identifier corresponding to each preset type; The target processing model is determined according to the model identifier corresponding to the data type.

5. The method according to claim 1, wherein Get multiple pieces of data to be processed required to execute the tasks to be processed, including: Acquire multiple initial data corresponding to the task to be processed; Performing preset processing on the plurality of initial data to obtain the plurality of data to be processed, wherein the preset processing includes at least one of the following: Determining duplicate data in the plurality of initial data, and deduplicating the duplicate data; Identifying a plurality of data missing locations in the initial data, and performing data filling on the data missing locations; Performing clustering processing on the plurality of the initial data by using a clustering algorithm, determining outliers in the plurality of the initial data, and deleting the outliers; Based on the preset data screening conditions, the plurality of initial data are screened.

6. The method according to claim 5, characterized in that The data screening conditions are obtained according to the following process: receiving a configuration operation for the data screening condition; In response to the configuration operation, the data screening condition is determined according to the configuration information indicated by the configuration operation, the data screening condition adopts a Boolean logic relationship or a regular expression, and the configuration information is used to fill in the Boolean logic relationship or the regular expression.

7. A data processing device, characterized in that: Applied to a distributed system, the distributed system includes multiple computing nodes, and the device includes: An acquisition module, configured to acquire a plurality of to-be-processed data required to execute the to-be-processed task, and the remaining resources of each of the computing nodes; A first determining module is used to determine the data types of the plurality of data to be processed, wherein the data types are used to represent the association relationship between the plurality of data to be processed; a second determining module, configured to determine a target processing model from a plurality of preset processing models according to the data type; A third determining module is configured to determine a target computing node among the plurality of computing nodes according to a model identifier of the target processing model, remaining resources of each computing node, and resource usage for executing the task to be processed; The computing module is used to compute the plurality of the to-be-processed data through the target processing model deployed by the target computing node, and obtain the execution results corresponding to the to-be-processed tasks.

8. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the data processing method according to any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that When the instructions in the computer program product are executed by an electronic device, the electronic device is caused to execute the data processing method according to any one of claims 1 to 6.