Data processing method and processing equipment thereof

By checking and securely connecting multi-source data and dynamically adjusting resource configuration, the problems of low data quality and resource utilization in the existing technology are solved, and an efficient and reliable data processing process is achieved.

CN119938653AInactive Publication Date: 2025-05-06GUANGZHOU LAXUN TECH DEV CO LTD

Patent Information

Application Number
CN202411826109.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology lacks a verification mechanism when processing multi-source data, which leads to the inability to guarantee data source and timeliness, affecting the data quality and reliability of processing results, and the resource configuration is static and cannot be adjusted dynamically, resulting in low resource utilization and limited processing speed.

Method used

By verifying the identification parameters and access timestamps of multi-source data, the authenticity and timeliness of data are ensured; data connection is used for security protocols, processors and memory resources are dynamically allocated, resource allocation strategies are adjusted according to task priorities, and data processing paths and load allocation are optimized.

Benefits of technology

It improves the data quality and reliability of processing results, enhances the security of data during transmission, improves the speed and response time of data processing, optimizes the data processing process, and reduces processing delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938653A_ABST
    Figure CN119938653A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a data processing method and processing equipment thereof.The data processing method comprises the following steps that on the basis of multi-source data to be processed, identification parameters and access timestamps of data sources are verified, access data are extracted through a data interface, and the data are stored in a temporary buffer area after the data type is verified to be matched with the format; and obtaining access verification data. According to the method, the authenticity and the timeliness of the data are ensured by checking the identification parameters and the access timestamps of the multi-source data, the data quality is improved, data connection is performed by using a security protocol, the protection strength of the data in the transmission process is increased, processor and memory resources are dynamically allocated, a resource allocation strategy is adjusted according to the task priority, and the resource allocation efficiency is improved. The data processing speed and response time are improved, the data processing process is optimized by adjusting the data processing parameters, the processing delay is reduced, and the data processing accuracy and speed are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a data processing method and a processing device thereof. Background Art

[0002] The field of data processing technology focuses on the acquisition, verification, storage, protection and processing of data in information systems. This field covers the entire process from data input, verification to processing and transformation so that the data can be further analyzed and utilized. In modern information technology, data processing not only includes basic data format conversion and cleaning, but also involves more advanced operations such as data fusion, intelligent classification and real-time data stream processing. The development of this technology field has promoted the application and advancement of related technologies such as big data technology, cloud computing and machine learning, and provided support for business intelligence, decision support systems and automation systems.

[0003] Among them, data processing methods focus on developing and implementing methodologies to process data efficiently and effectively. Methods include data cleaning, data integration, anomaly detection and repair, and optimized data storage technology. It has a wide range of uses, not only serving the field of data science for data mining and analysis, but also supporting enterprises to improve processing speed and accuracy when processing large-scale data sets, thereby enhancing the ability to make data-driven decisions. Such methods are crucial to ensuring data quality and availability, and are the basis for advanced data analysis and business intelligence analysis.

[0004] Traditional processing methods lack a verification mechanism when processing multi-source data, resulting in the inability to guarantee the source and timeliness of the data, affecting data quality and the reliability of processing results. Traditional methods are usually static in terms of resource allocation and cannot be dynamically adjusted according to the actual needs of the task, which leads to low resource utilization and limited processing speed. In terms of data processing process optimization, they fail to fully consider how to reduce delays and increase processing speed, which is particularly obvious when processing large-scale data sets, causing data processing tasks to be delayed and affecting the timeliness of decision-making. Summary of the invention

[0005] The purpose of the present invention is to solve the shortcomings in the prior art and to propose a data processing method and a processing device thereof.

[0006] In order to achieve the above object, the present invention adopts the following technical solution: a data processing method, comprising the following steps:

[0007] S1: Based on the multi-source data to be processed, the identification parameters and access timestamp of the data source are verified, the access data is extracted through the data interface, and after verifying that the data type and format match, it is stored in a temporary buffer to obtain the access verification data;

[0008] S2: Based on the access verification data, check the data integrity and consistency, exclude the data that does not meet the rules, and re-request the data to obtain aggregated complete data;

[0009] S3: Based on the aggregated complete data, data classification is performed to identify sensitive information and non-sensitive information, random noise is applied to the sensitive information, and data is re-encoded to obtain noisy data;

[0010] S4: Based on the noise-added data, dynamically allocate processor and memory resources according to task priority, configure data processing paths, adjust data flow and processing node loads, and obtain resource configuration information;

[0011] S5: Based on the resource configuration information, according to the data characteristics and processing objectives, a matching processing algorithm is selected to process the data to obtain optimized processing information;

[0012] S6: Based on the optimization processing information, the data results are sorted, the processing results are visualized, and the output format of the data processing results is adjusted to obtain processing output information.

[0013] As a further solution of the present invention, the access verification data includes a data source identifier, data access time, data type and data format; the aggregated complete data includes a data summary that complies with accuracy verification rules, a data consistency index and a data integrity index; the noisy data includes sensitive information processed with random noise, encrypted personal identification information and re-encoded data items; the resource configuration information includes processor allocation information, memory usage and data processing path configuration; the optimized processing information includes adjusted calculation parameters, selected data processing algorithms and optimized data flows; the processing output information includes data analysis charts, data explanatory information and formatted data output.

[0014] As a further solution of the present invention, based on the multi-source data to be processed, the identification parameters and access timestamp of the data source are verified, the access data is extracted through the data interface, and after verifying that the data type matches the format, it is stored in a temporary buffer. The steps of obtaining the access verification data are specifically as follows:

[0015] S101: Based on the multi-source data to be processed, the identification parameters and access timestamp of each data source are checked to verify the accuracy and validity of the data source identification and timestamp, optimize the timeliness of the data and the validity of the source, and obtain a data verification record;

[0016] S102: Based on the data verification record, using a secure transmission protocol, establishing a stable secure connection with the data source, and verifying the encryption level and integrity of the protocol, optimizing the security of the data during transmission, and obtaining a secure connection status;

[0017] S103: Based on the security connection status, data in the data source is extracted through the data interface, and before extraction, the data type and format are checked by a format verification script to see if they meet the requirements. Illegal data or data that does not match the format will be denied access, and legal data will be stored in a temporary buffer to obtain access verification data.

[0018] As a further solution of the present invention, based on the access verification data, checking the data integrity and consistency, excluding the data that does not conform to the rules, and re-requesting the data to obtain the aggregated complete data are specifically as follows:

[0019] S201: Based on the access verification data, perform data integrity check, check each data item, detect whether all data items are complete, have no missing fields or damaged data structures, and obtain data integrity information;

[0020] S202: Based on the data integrity information, screen and exclude data identified as not meeting preset rules, issue a re-request for missing and erroneous data items, optimize the quality and consistency of the data set, and obtain a data re-request status;

[0021] S203: Based on the data re-request status, the integrity index of the data set is evaluated, and the quality and security of the data set are checked. The data that passes the verification will be compiled into the data set to obtain aggregated complete data.

[0022] As a further solution of the present invention, based on the aggregated complete data, data classification is performed to identify sensitive information and non-sensitive information, random noise is applied to the sensitive information, and data is re-encoded to obtain the noisy data in the following steps:

[0023] S301: Based on the aggregated complete data, data classification is performed, sensitive information and non-sensitive information are distinguished by using data labels, and data items are annotated to optimize the accuracy of information classification, thereby obtaining a classified data index;

[0024] S302: adding random noise to the identified sensitive data based on the classified data index, and performing protection processing on the sensitive data by adjusting noise parameters, including intensity and frequency, to avoid direct information leakage, thereby obtaining noise-enhanced sensitive data;

[0025] S303: Based on the noise-enhanced sensitive data, the information involving personal identification is encrypted using asymmetric encryption technology, and the entire data set is re-encoded to optimize the security of the data during network transmission to obtain noisy data.

[0026] As a further solution of the present invention, based on the noise-added data, according to the task priority, dynamically allocate processor and memory resources, configure the data processing path, adjust the data flow and the load of the processing node, and obtain the resource configuration information in the following steps:

[0027] S401: Based on the noisy data, analyze the priority of the task, combine the processing requirements and resource usage of the differentiated tasks, evaluate the resource configuration weight, dynamically allocate processor and memory resources, adjust the resource allocation strategy, optimize the effective use of resources, and obtain basic configuration information;

[0028] S402: Based on the basic configuration information, design a data processing path, adjust the data flow and the load of the processing node, optimize the data processing flow, and improve the efficiency of the processing node to obtain a resource adjustment status;

[0029] S403: Based on the resource adjustment status, continuously monitor the usage of the processor and memory, timely adjust resource allocation according to changes in data processing tasks, optimize each task to obtain sufficient resources on demand, and obtain resource configuration information.

[0030] As a further solution of the present invention, the calculation formula for evaluating resource allocation weights is:

[0031]

[0032] Among them, R is the resource configuration weight, T represents the task priority, w1 is the weight coefficient of the task priority, P represents the complexity of the task processing requirements, w2 is the weight coefficient of the processing requirement complexity, D represents the data volume, w3 is the weight coefficient of the data volume, S represents the total system resources, and w4 is the weight coefficient of the total system resources.

[0033] As a further solution of the present invention, based on the resource configuration information, according to the data characteristics and processing objectives, a matching processing algorithm is selected to process the data, and the steps of obtaining optimized processing information are specifically as follows:

[0034] S501: Based on the resource configuration information, analyze data characteristics, change the data processing flow, adjust the data segmentation size and the number of parallel processing, match the processing target, optimize the processing efficiency, and obtain the processing parameter configuration;

[0035] S502: Based on the processing parameter configuration, select a matching processing algorithm to process the data, and adjust the processing parameters to reduce processing delays, optimize the response speed and accuracy of data processing, and obtain optimized configuration information;

[0036] S503: Based on the optimized configuration information, the load distribution of the processing nodes is adjusted, and the data transmission between the differentiated nodes is changed to optimize the continuity and efficiency of data processing, thereby obtaining optimized processing information.

[0037] As a further solution of the present invention, based on the optimization processing information, the data results are sorted, the processing results are visualized, and the output format of the data processing results is adjusted to obtain the processing output information in the following steps:

[0038] S601: Based on the optimization processing information, sort the data processing results, sort and index the processed data, classify and mark the key processing results, and obtain a data result index;

[0039] S602: Based on the data result index, perform result visualization, select a matching chart type to display the data processing result, adjust the design elements of the chart, including color depth and display ratio, optimize the visual presentation of information, and obtain result visualization information;

[0040] S603: Based on the result visualization information, the output format of the data processing result is adjusted to match the subsequent use target, and the processing output information is obtained.

[0041] A data processing device, the data processing device is used to execute the above data processing method, the processing device comprises:

[0042] The data access module verifies the identification parameters and access timestamp of the data source based on the multi-source data to be processed, connects the data source through a security protocol, extracts the access data through the data interface, verifies that the data type matches the format, and stores it in a temporary buffer to obtain access verification data;

[0043] The data verification module verifies the data based on the access verification data, checks the data integrity and consistency, excludes the data that does not meet the rules, and re-requests the data, evaluates the data integrity index, verifies the data quality and security, and obtains the aggregated complete data;

[0044] The data noise protection module classifies data based on the aggregated complete data, identifies sensitive information and non-sensitive information, applies random noise to sensitive information, uses asymmetric encryption technology to process personal identification information, re-encodes data, optimizes the security of data during transmission, and obtains noise-added data;

[0045] The data processing implementation module dynamically allocates processor and memory resources based on the noise-added data and task priorities, configures the data processing path, adjusts the data flow and the load of the processing node, selects a matching processing algorithm for data processing based on data characteristics and processing objectives, and adjusts processing parameters to obtain optimized processing information;

[0046] The processing result output module organizes the data results based on the optimized processing information, performs processing result visualization, displays the data processing results through charts and graphs, adjusts the output format of the data processing results, optimizes the readability and comprehensibility of the data processing results, and obtains processing output information.

[0047] Compared with the prior art, the advantages and positive effects of the present invention are:

[0048] In the present invention, by verifying the identification parameters and access timestamps of multi-source data, the authenticity and timeliness of the data are ensured, the data quality is improved, the data connection is carried out using a security protocol, the protection of the data during transmission is increased, the risk of data leakage is reduced, the accuracy and integrity of the data set are guaranteed by screening out data that does not meet the standards, sensitive and non-sensitive information is identified during the data classification process, and random noise and asymmetric encryption are applied to sensitive information, thereby enhancing the security of personal information. The processor and memory resources are dynamically allocated, and the resource allocation strategy is adjusted according to the task priority, which improves the speed and response time of data processing. By adjusting the data processing parameters, the data processing process is optimized, the processing delay is reduced, and the accuracy and speed of data processing are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic diagram of the workflow of the present invention;

[0050] Figure 2 This is a detailed flow chart of S1 of the present invention;

[0051] Figure 3 This is a detailed flow chart of S2 of the present invention;

[0052] Figure 4 This is a detailed flow chart of S3 of the present invention;

[0053] Figure 5 This is a detailed flow chart of S4 of the present invention;

[0054] Figure 6 This is a detailed flow chart of S5 of the present invention;

[0055] Figure 7 This is a detailed flow chart of S6 of the present invention;

[0056] Figure 8 It is a flow chart of the processing equipment of the present invention. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0058] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, in the description of the present invention, "multiple" means two or more, unless otherwise clearly and specifically defined.

[0059] See also Figure 1 The present invention provides a technical solution: a data processing method, comprising the following steps:

[0060] S1: Based on the multi-source data to be processed, the identification parameters and access timestamp of the data source are verified, the data source is connected through a security protocol, the access data is extracted through the data interface, and after verifying that the data type and format match, it is stored in a temporary buffer to obtain access verification data;

[0061] S2: Based on the access verification data, verify the data, check the data integrity and consistency, exclude the data that does not meet the rules, re-request the data, evaluate the data integrity indicators, verify the data quality and security, and obtain aggregated complete data;

[0062] S3: Based on the aggregated complete data, data classification is performed to identify sensitive information and non-sensitive information, random noise is applied to sensitive information, personal identification information is processed using asymmetric encryption technology, data is re-encoded, and the security of data during transmission is optimized to obtain noisy data;

[0063] S4: Based on the noisy data and task priorities, dynamically allocate processor and memory resources, configure data processing paths, adjust data flows and processing node loads, and monitor computing resource status in real time, adjust resource allocation, match changes in processing requirements, and obtain resource configuration information;

[0064] S5: Based on the resource configuration information, according to the data characteristics and processing objectives, select a matching processing algorithm for data processing, and adjust the processing parameters to optimize the data processing speed and accuracy, reduce processing delays, and obtain optimized processing information;

[0065] S6: Based on the optimized processing information, the data results are sorted, the processing results are visualized, the data processing results are displayed through charts and graphs, and the output format of the data processing results is adjusted to optimize the readability and comprehensibility of the data processing results to obtain the processing output information.

[0066] Access verification data includes data source identification, data access time, data type and data format. Aggregate complete data includes data summary that complies with accuracy verification rules, data consistency indicators and data integrity indicators. Noise-added data includes sensitive information processed with random noise, personal identification information after encryption and re-encoded data items. Resource configuration information includes processor allocation information, memory usage and data processing path configuration. Optimized processing information includes adjusted calculation parameters, selected data processing algorithms and optimized data flows. Processing output information includes data analysis charts, data explanatory information and formatted data output.

[0067] See also Figure 2 , based on the multi-source data to be processed, the identification parameters and access timestamp of the data source are verified, the data source is connected through a security protocol, the access data is extracted through the data interface, and after verifying that the data type matches the format, it is stored in a temporary buffer. The specific steps for obtaining the access verification data are as follows:

[0068] S101: Based on the multi-source data to be processed, the identification parameters and access timestamp of each data source are checked to verify the accuracy and validity of the data source identification and timestamp, optimize the timeliness of the data and the validity of the source, and obtain a data verification record;

[0069] Based on the multi-source data to be processed, the identity and time of each data source are verified. For the data source to be verified, the retrieval standard of identification information and access timestamp is established, covering the type of data source, expected access time, access frequency and data scale. The identification information and timestamp of each data source are collected through database query or API call, and preliminary screening is performed to identify the data sources that do not meet the preset standards. The timestamp of the data that meets the standards is verified, and the timestamp of each data source is compared with the current time of the system to determine the real-time and validity of the timestamp. The data source with inconsistent or outdated timestamp is marked and processed. At the same time, the process and results of each data verification are recorded to ensure that each step of the operation is recorded and checked. Data verification records are generated, including the identity authentication results of the data source and the validity test results of the timestamp, ensuring the timeliness of the data and the validity of the source.

[0070] S102: Based on the data verification record, using the secure transmission protocol, establishing a stable secure connection with the data source, and verifying the encryption level and integrity of the protocol, optimizing the security of the data during transmission, and obtaining a secure connection status;

[0071] Based on the data verification record, a secure connection is established according to the data source. Standard secure transmission protocols such as SSL or TLS are used. Necessary security certificates and keys are configured on the server side to ensure the privacy of the connection and encrypted transmission of data. A security assessment is performed on the data source, including the strength of the encryption algorithm and the validity of the certificate to prevent the data from being intercepted or tampered with during transmission. Subsequently, two-way authentication is performed to ensure the identity authentication of the data source and the server. During the verification process, the encryption level and integrity of the protocol are checked to ensure that the data of the communicating parties cannot be accessed by a third party. Through these steps, a stable secure connection is established, and the security status of each connection is recorded and stored in the server's security log to obtain the secure connection status.

[0072] S103: Based on the security connection status, extract data from the data source through the data interface, and check whether the data type and format meet the requirements through the format verification script before extraction. Illegal data or data that does not match the format will be denied access, and legal data will be stored in a temporary buffer to obtain access verification data;

[0073] Based on the security connection status, according to the preset data format and type requirements, the format and type of the incoming data are checked, each incoming data packet is scanned, format errors and type mismatches in the data packet are identified, and data packets that do not conform to the format or type are recorded and denied access. Legal data are temporarily stored in a temporary buffer, and the access status and format verification results of each data packet are recorded to ensure that only data that passes the verification can be further processed. The access verification data obtained includes the format and type verification results of each data packet, as well as the final status of whether it is accepted into the system.

[0074] See also Figure 3 , based on the access verification data, verify the data, check the data integrity and consistency, exclude the data that does not meet the rules, and re-request the data, evaluate the data integrity indicators, and verify the data quality and security. The specific steps to obtain the aggregated complete data are:

[0075] S201: Based on the access verification data, perform data integrity check, check each data item, detect whether all data items are complete, have no missing fields or damaged data structures, and obtain data integrity information;

[0076] Based on the access verification data, each field in the data is identified, and the required fields in the database or preset template are compared. The algorithm is used to detect whether there are missing or abnormal fields in the data items. Each data item is independently verified to ensure that no fields are missed. At the same time, the data structure is deeply scanned to identify any structural damage or data format errors. For each problem detected, a detailed log record is generated, including the problem type, location and possible causes. Each inspection action is strictly recorded to ensure traceability and the feasibility of subsequent analysis. Through inspection and recording, data integrity information is obtained, which lists in detail the integrity status of all data and the specific problems that exist.

[0077] S202: based on the data integrity information, screen and exclude the data identified as not meeting the preset rules, issue a re-request for missing and erroneous data items, optimize the quality and consistency of the data set, and obtain a data re-request status;

[0078] Based on the data integrity information, the data cleaning process is performed to identify and filter out data that does not meet the preset rules. Data cleaning tools are used to accurately locate errors and missing data in the data set. Data re-requests are issued for missing data, and requests are automatically sent to the original data source or backup data storage, requiring resubmission of complete and correct data items, including analysis of error types, tracking of missing data, and logging of re-requests. To ensure the consistency and quality of the data set, the status of all re-requests is updated in the master data table, including the requested data items, the time when the request was initiated, and the response status of the data source. The data re-request status is obtained through continuous inspection and adjustment.

[0079] S203: Based on the data re-request status, the integrity index of the data set is evaluated, and the quality and security of the data set are checked. The data that passes the verification will be compiled into the data set to obtain aggregated complete data;

[0080] Based on the data re-request status, the integrity indicators of the data set are evaluated, and data quality and security checks are performed, including the use of automated tools to perform statistical analysis on the data set, evaluate data integrity indicators such as missing value ratio, number of outliers and data consistency level, and perform multi-dimensional security checks on the data, including verification of data access rights and inspection of data encryption status to ensure the security of the data set. After comprehensive quality and security checks, the verified data is compiled into the final data set, and records include the qualified status of the data items, the statistical indicators of the data set and the final data integrity assessment results to obtain aggregated complete data.

[0081] See also Figure 4, based on the aggregation of complete data, data classification, identification of sensitive information and non-sensitive information, random noise is applied to sensitive information, personal identification information is processed using asymmetric encryption technology, data is re-encoded, and the security of data during transmission is optimized. The specific steps for obtaining noisy data are as follows:

[0082] S301: Based on the aggregated complete data, data classification is performed, sensitive information and non-sensitive information are distinguished by data labels, and data items are annotated to optimize the accuracy of information classification and obtain a classified data index;

[0083] Based on the aggregation of complete data, machine learning algorithms are used to automatically distinguish between sensitive and non-sensitive information in the data, identify and mark sensitive categories such as personal information and financial data, and annotate each data item in detail, involving the content, sensitivity and related labels of the data to ensure the speed and accuracy of data classification. After the classification is completed, a classified data index is generated. The index records in detail the location, category and sensitivity level of each type of data. The index is used to guide subsequent data processing and access control.

[0084] S302: adding random noise to the identified sensitive data based on the classified data index, and performing protection processing on the sensitive data by adjusting noise parameters, including intensity and frequency, to avoid direct information leakage, thereby obtaining noise-enhanced sensitive data;

[0085] Based on the classified data index, protection measures are taken for sensitive data, and parameter configurations for noise addition are set, including the intensity and frequency of noise. Parameters are adjusted for data of different sensitivity levels to achieve the best protection effect. Noise data is generated using a random number generation algorithm, which ensures the randomness and unpredictability of noise. The noise data is then mixed with sensitive data, and the algorithm ensures that the noise is evenly distributed in the data to avoid potential information leakage risks. The processed data is marked and stored to ensure that the status of each noise addition is recorded and verified to maintain the integrity and security of the data, and to obtain noise-enhanced sensitive data. While protecting sensitive information, the data can still be used for subsequent analysis and processing.

[0086] S303: Based on noise enhancement of sensitive data, asymmetric encryption technology is used to encrypt information involving personal identification, and the entire data set is re-encoded to optimize the security of data during network transmission to obtain noise-enhanced data;

[0087] Based on noise enhancement of sensitive data, asymmetric encryption technology is applied, and appropriate public and private keys are selected to ensure the security of the encryption process and the confidentiality of the data. Special processing is performed on data items containing personal identification information. During the encryption process, advanced encryption algorithms such as RSA or ECC are used. The algorithm supports high-security data processing and protects data from unauthorized access. After encryption is completed, the entire data set is re-encoded to ensure that each data item meets the security standards for network transmission. Each step in the encryption and encoding process is recorded in detail for auditing and tracking. Through these meticulous operations, noisy data is obtained, and the security of data during network transmission is greatly improved, ensuring the integrity and confidentiality of the data.

[0088] See also Figure 5 ,Based on the noisy data, according to the task priority, dynamically allocate processor and memory resources, configure the data processing path, adjust the data flow and the load of the processing node, and monitor the computing resource status in real time, adjust resource allocation, and match the changes in processing requirements. The specific steps to obtain resource configuration information are:

[0089] S401: Based on the noisy data, analyze the priority of the task, combine the processing requirements and resource usage of the differentiated tasks, evaluate the resource configuration weight, dynamically allocate processor and memory resources, adjust the resource allocation strategy, optimize the effective use of resources, and obtain basic configuration information;

[0090] The calculation formula for evaluating resource allocation weights is:

[0091]

[0092] Among them, R is the resource configuration weight, T represents the task priority, w1 is the weight coefficient of the task priority, P represents the complexity of the task processing requirements, w2 is the weight coefficient of the processing requirement complexity, D represents the data volume, w3 is the weight coefficient of the data volume, S represents the total system resources, and w4 is the weight coefficient of the total system resources.

[0093] formula:

[0094]

[0095] Parameter details and how to obtain them:

[0096] T stands for the priority of the task, which is set based on the urgency and importance of the task. Priority values ​​are usually set manually by the task management system or administrator based on the deadline and importance of the task, ranging from 1 to 10, with 1 being the lowest priority and 10 being the highest priority.

[0097] w1 is the weight coefficient of task priority, which is usually set according to task type and business needs to ensure that high-priority tasks can obtain sufficient resources. The setting of this coefficient needs to be based on historical data analysis and business logic optimization.

[0098] P represents the complexity of the task processing requirements, measured by the number of CPU cores required for the task. This parameter can be obtained from the task queue management system, which calculates the number of CPU cores required based on the type of task and the estimated processing time.

[0099] w2 is the weight coefficient for processing demand complexity, which is used to adjust the impact of complexity on resource allocation. The coefficient is usually determined based on performance test results or empirical data.

[0100] D represents the total amount of data that needs to be processed during the task processing, in GB. This data volume parameter is pre-estimated by the data management system or specified by the user when submitting the task.

[0101] w3 is the weight coefficient of data volume, which emphasizes the importance of data volume in resource allocation. This coefficient is usually set based on the storage and processing capabilities of the system.

[0102] S represents the total resources of the system, that is, the total amount of memory available to the system, also in GB. This parameter is determined by the hardware configuration of the system and is usually determined when the system is deployed.

[0103] w4 is the weight coefficient of the total amount of system resources, which is used to regulate the impact of the total amount of resources on the allocation strategy. Its setting is based on the overall load capacity of the system and the expected operating efficiency.

[0104] Calculation example:

[0105] Set the task priority T to 8, the processing requirement P requires a 4-core CPU, the data volume D is 20GB, and the total system resources S is 100GB. The weight coefficients w1 = 0.5, w2 = 1.0, w3 = 1.5, w4 = 0.8. The resource allocation weight R is calculated as follows:

[0106]

[0107] The calculation result is R≈0.085, which represents the resource allocation weight that a task should obtain based on its priority, processing requirement complexity, data volume, and system resource configuration. A higher weight value means that the task should tend to obtain more resources, which is used in practice to dynamically adjust the resource allocation of tasks to ensure that high-priority and resource-intensive tasks receive sufficient system support.

[0108] S402: Based on the basic configuration information, design a data processing path, adjust the data flow and the load of the processing node, optimize the data processing flow, and improve the efficiency of the processing node to obtain a resource adjustment status;

[0109] Based on the basic configuration information, the data processing path is designed. By identifying the data flow direction, including the source and target nodes of the data, the structure of the data flow is optimized according to the resource adjustment status, and the load distribution of the processing nodes is adjusted. The process uses load balancing technology to ensure balanced transmission of data between the processing nodes and reduce the overload of any single node. By monitoring the data flow dynamics in real time, it can quickly respond to performance changes of any node and adjust the data path and node load in time. This flexible adjustment strategy greatly improves the efficiency of data processing and the operation efficiency of the node, and obtains the resource adjustment status.

[0110] S403: Based on the resource adjustment status, continuously monitor the usage of the processor and memory, timely adjust resource allocation according to changes in data processing tasks, optimize each task to obtain sufficient resources as needed, and obtain resource configuration information;

[0111] Based on the resource adjustment status, the usage of processor and memory is continuously monitored. Resource monitoring tools are used to collect and analyze processor load, memory occupancy and its changing trend in real time. Resource reallocation is automatically performed based on monitoring data. The allocation ratio of processor and memory is adjusted according to the real-time changes of data processing tasks to ensure that each task obtains appropriate resources according to its processing complexity and real-time requirements. During the resource allocation process, a detailed configuration log is generated. The log records the initial configuration, adjustment process and final status of the resources to obtain resource configuration information.

[0112] See also Figure 6 , based on resource configuration information, according to data characteristics and processing objectives, select matching processing algorithms for data processing, and adjust processing parameters to optimize data processing speed and accuracy, reduce processing delays, and obtain optimized processing information. The specific steps are:

[0113] S501: Based on the resource configuration information, analyze data characteristics, change the data processing flow, adjust the data segmentation size and the number of parallel processing, match the processing target, optimize the processing efficiency, and obtain the processing parameter configuration;

[0114] Based on resource configuration information, adjust the data processing flow to improve efficiency. By analyzing data characteristics, determine key parameters in data processing, such as data segmentation size and number of parallel processing. Dynamically adjust parameters according to data type and processing complexity. Through simulation operation, test the impact of different segmentation sizes on processing time, determine the optimal segmentation strategy, and adjust the number of parallel processing threads at the same time to ensure that each processing unit can run efficiently without waiting too long. Adjustments are made based on real-time resource usage data and task priorities. Constantly update and optimize these parameters to match different processing goals. Through these detailed operation rates, obtain processing parameter configuration.

[0115] S502: Based on the processing parameter configuration, a matching processing algorithm is selected to process the data, and the processing parameters are adjusted to reduce processing delays, optimize the response speed and accuracy of data processing, and obtain optimized configuration information;

[0116] Based on the processing parameter configuration, select the processing algorithm that suits the current data characteristics, and make adjustments to the algorithm performance to reduce processing delays, including optimizing the algorithm's input parameters and making subtle adjustments to the algorithm itself, such as improving the efficiency of sorting or searching algorithms, and reducing unnecessary calculation steps. By monitoring the response speed and accuracy of algorithm execution in real time, adjust processing parameters such as thread allocation, memory usage, and CPU time slices based on feedback results. Adjustments ensure maximum efficiency and best results in data processing. Through continuous optimization, the generated optimized configuration information reflects the status and effectiveness of the current algorithm configuration.

[0117] S503: Based on the optimized configuration information, adjust the load distribution of the processing nodes, and change the data transmission between the differentiated nodes to optimize the continuity and efficiency of data processing, and obtain optimized processing information;

[0118] Based on the optimized configuration information, the load distribution of processing nodes and the data transmission strategy between nodes are adjusted. Load analysis and data flow re-planning are carried out according to differentiated node performance and data processing requirements. The data transmission path is adjusted through network optimization technology to reduce network congestion and data delay, ensure that data can be efficiently transmitted between nodes, and continuously monitor the performance and load of nodes. The load distribution and data flow are automatically adjusted according to real-time data. The adjustment makes the data processing process more coherent and efficient. The obtained optimized processing information records the adjustment results of the processing nodes and the optimization measures for data transmission.

[0119] See also Figure 7 , based on the optimization processing information, the data results are sorted, the processing results are visualized, the data processing results are displayed through charts and graphs, and the output format of the data processing results is adjusted to optimize the readability and comprehensibility of the data processing results. The specific steps for obtaining the processing output information are:

[0120] S601: Based on the optimization processing information, the data processing results are sorted, the processed data is sorted and indexed, the key processing results are classified and marked, and a data result index is obtained;

[0121] Based on the optimization processing information, the processed data is organized in order and sorted using a sorting algorithm, which includes comparing the importance and relevance of each data item and determining the order of each data item. Next, the index marking technology is applied to assign a unique index to each data item, which facilitates quick access and retrieval. During the indexing process, the system automatically marks the key results of data processing, such as the highest frequency values, anomalies, and trend changes, to obtain the data result index.

[0122] S602: Based on the data result index, perform result visualization, select a matching chart type to display the data processing result, adjust the design elements of the chart, including color depth and display ratio, optimize the visual presentation of information, and obtain result visualization information;

[0123] Based on the data result index, select a chart type that is suitable for displaying data characteristics, such as a bar chart, line chart, or scatter chart. For each chart, adjust the design elements to optimize the visual effect, including adjusting the color depth to highlight different data levels and adjusting the display ratio to accommodate different visual presentation requirements. The adjustment is based on the content of the data and the best practices of visual presentation to ensure that the information is conveyed clearly and accurately. Through design adjustments, the generated result visualization information clearly displays the processed data results, making the interpretation of the data more intuitive and easy to understand.

[0124] S603: Based on the result visualization information, the output format of the data processing result is adjusted to match the subsequent use target, and the processing output information is obtained;

[0125] Based on the result visualization information, adjust the data output format to match the subsequent use or storage requirements of the data, including selecting a suitable file format such as CSV, JSON or XML, adjusting the file structure and encoding, ensuring the compatibility of the format with the target system or platform, optimizing the accuracy and detail of the data according to the output target, and using automated scripts to test the loading and reading speeds of different formats to ensure the efficiency and accuracy of the output data and obtain processed output information.

[0126] See also Figure 8 , a data processing device, the data processing device is used to execute the above data processing method, the processing device comprises:

[0127] The data access module verifies the identification parameters and access timestamp of the data source based on the multi-source data to be processed, connects the data source through a security protocol, extracts the access data through the data interface, verifies that the data type matches the format, and stores it in a temporary buffer to obtain access verification data;

[0128] The data verification module verifies the data based on the access verification data, checks the data integrity and consistency, excludes the data that does not meet the rules, and re-requests the data, evaluates the data integrity indicators, verifies the data quality and security, and obtains the aggregated complete data;

[0129] The data noise protection module aggregates complete data, classifies data, identifies sensitive information and non-sensitive information, applies random noise to sensitive information, uses asymmetric encryption technology to process personal identification information, re-encodes data, optimizes the security of data during transmission, and obtains noisy data;

[0130] The data processing implementation module dynamically allocates processor and memory resources based on the noisy data and task priority, configures the data processing path, adjusts the data flow and the load of the processing node, monitors the computing resource status in real time, adjusts resource allocation, matches the changes in processing requirements, selects matching processing algorithms for data processing based on data characteristics and processing objectives, and adjusts processing parameters to optimize data processing speed and accuracy, reduce processing delays, and obtain optimized processing information;

[0131] The processing result output module is based on optimizing processing information, organizing data results, performing processing result visualization, displaying data processing results through charts and graphs, and adjusting the output format of data processing results to optimize the readability and comprehensibility of data processing results to obtain processing output information.

[0132] The above are only preferred embodiments of the present invention and are not intended to limit the present invention in other forms. Any technician familiar with the profession may use the technical contents disclosed above to change or modify them into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still falls within the protection scope of the technical solution of the present invention.

Claims

1. A data processing method, characterized in that: The following steps are involved: Based on the multi-source data to be processed, the identification parameters and access timestamp of the data source are verified, the access data is extracted through the data interface, and after verifying that the data type and format match, it is stored in a temporary buffer to obtain the access verification data; Based on the access verification data, check the data integrity and consistency, exclude the data that does not meet the rules, and re-request the data to obtain aggregated complete data; Based on the aggregated complete data, data classification is performed to identify sensitive information and non-sensitive information, random noise is applied to the sensitive information, and data is re-encoded to obtain noisy data; Based on the noise-added data, according to the task priority, dynamically allocate processor and memory resources, configure the data processing path, adjust the data flow and the load of the processing node, and obtain resource configuration information; Based on the resource configuration information, according to data characteristics and processing objectives, a matching processing algorithm is selected to process the data to obtain optimized processing information; Based on the optimized processing information, the data results are sorted, the processing results are visualized, and the output format of the data processing results is adjusted to obtain processing output information.

2. The data processing method according to claim 1, characterized in that: The access verification data includes data source identification, data access time, data type and data format; the aggregated complete data includes data summary, data consistency index and data integrity index that comply with accuracy verification rules; the noisy data includes sensitive information processed with random noise, personal identification information after encryption and re-encoded data items; the resource configuration information includes processor allocation information, memory usage and data processing path configuration; the optimized processing information includes adjusted calculation parameters, selected data processing algorithm and optimized data flow; the processing output information includes data analysis charts, data explanatory information and formatted data output.

3. The data processing method according to claim 1, characterized in that: Based on the multi-source data to be processed, the identification parameters and access timestamp of the data source are verified, the access data is extracted through the data interface, and after verifying that the data type matches the format, it is stored in the temporary buffer. The specific steps for obtaining the access verification data are as follows: Based on the multi-source data to be processed, the identification parameters and access timestamps of each data source are checked to verify the accuracy and validity of the data source identification and timestamp, optimize the timeliness of the data and the validity of the source, and obtain data verification records; Based on the data verification record, a stable secure connection is established with the data source using a secure transmission protocol, and the encryption level and integrity of the protocol are verified to optimize the security of the data during transmission and obtain a secure connection status; Based on the secure connection status, data in the data source is extracted through the data interface, and before extraction, a format verification script is used to check whether the data type and format meet the requirements. Illegal or format-mismatched data will be denied access, and legal data will be stored in a temporary buffer to obtain access verification data.

4. The data processing method according to claim 1, characterized in that: Based on the access verification data, the steps of checking data integrity and consistency, excluding data that does not conform to the rules, and re-requesting data to obtain aggregated complete data are as follows: Based on the access verification data, a data integrity check is performed to check each data item to detect whether all data items are complete, without missing fields or damaged data structures, and obtain data integrity information; Based on the data integrity information, data identified as not meeting preset rules are screened and excluded, and a re-request is issued for missing and erroneous data items to optimize the quality and consistency of the data set and obtain a data re-request status; Based on the data re-request status, the integrity index of the data set is evaluated, and the quality and security of the data set are checked. The data that passes the verification will be incorporated into the data set to obtain aggregated complete data.

5. The data processing method according to claim 1, characterized in that: Based on the aggregated complete data, data classification is performed, sensitive information and non-sensitive information are identified, random noise is applied to the sensitive information, and data is re-encoded to obtain the noisy data in the following steps: Based on the aggregated complete data, data classification is performed, sensitive information and non-sensitive information are distinguished by using data labels, and data items are annotated to optimize the accuracy of information classification and obtain a classified data index; Based on the classified data index, random noise is added to the identified sensitive data, and the sensitive data is protected by adjusting noise parameters, including intensity and frequency, to avoid direct information leakage, thereby obtaining noise-enhanced sensitive data; Based on the noise-enhanced sensitive data, asymmetric encryption technology is used to encrypt information related to personal identification, and the entire data set is re-encoded to optimize the security of the data during network transmission to obtain noisy data.

6. The data processing method according to claim 1, characterized in that: Based on the noise-added data, according to the task priority, dynamically allocate processor and memory resources, configure the data processing path, adjust the data flow and the load of the processing node, and obtain the resource configuration information in the following steps: Based on the noisy data, analyze the priority of the task, combine the processing requirements and resource usage of the differentiated tasks, evaluate the resource configuration weight, dynamically allocate processor and memory resources, adjust the resource allocation strategy, optimize the effective use of resources, and obtain basic configuration information; Based on the basic configuration information, a data processing path is designed, the data flow direction and the load of the processing node are adjusted, the data processing flow is optimized, and the efficiency of the processing node is improved to obtain a resource adjustment status; Based on the resource adjustment status, the use of the processor and memory is continuously monitored, resource allocation is adjusted in time according to changes in data processing tasks, and each task is optimized to obtain sufficient resources as needed to obtain resource configuration information.

7. The data processing method according to claim 6, characterized in that: The calculation formula for evaluating resource allocation weights is: Among them, R is the resource configuration weight, T represents the task priority, w1 is the weight coefficient of the task priority, P represents the complexity of the task processing requirements, w2 is the weight coefficient of the processing requirement complexity, D represents the data volume, w3 is the weight coefficient of the data volume, S represents the total system resources, and w4 is the weight coefficient of the total system resources.

8. The data processing method according to claim 1, characterized in that: Based on the resource configuration information, according to the data characteristics and processing objectives, a matching processing algorithm is selected to process the data, and the steps of obtaining optimized processing information are specifically as follows: Based on the resource configuration information, analyze data characteristics, change the data processing flow, adjust the data segmentation size and the number of parallel processing, match the processing target, optimize the processing efficiency, and obtain the processing parameter configuration; Based on the processing parameter configuration, a matching processing algorithm is selected to process the data, and the processing parameters are adjusted to reduce processing delays, optimize the response speed and accuracy of data processing, and obtain optimized configuration information; Based on the optimized configuration information, the load distribution of the processing nodes is adjusted, and the data transmission between the differentiated nodes is changed to optimize the continuity and efficiency of data processing, thereby obtaining optimized processing information.

9. The data processing method according to claim 1, characterized in that: Based on the optimization processing information, the data results are sorted, the processing results are visualized, and the output format of the data processing results is adjusted to obtain the processing output information in the following steps: Based on the optimization processing information, the data processing results are sorted, the processed data is sorted and indexed, the key processing results are classified and marked, and a data result index is obtained; Based on the data result index, visualizing the results, selecting a matching chart type to display the data processing results, adjusting the design elements of the chart, including color depth and display ratio, optimizing the visual presentation of information, and obtaining result visualization information; Based on the result visualization information, the output format of the data processing result is adjusted to match the subsequent use target, and the processing output information is obtained.

10. A data processing device, characterized in that: According to any one of claims 1 to 9, the data processing method, the processing device comprises: The data access module verifies the identification parameters and access timestamp of the data source based on the multi-source data to be processed, connects the data source through a security protocol, extracts the access data through the data interface, verifies that the data type matches the format, and stores it in a temporary buffer to obtain access verification data; The data verification module verifies the data based on the access verification data, checks the data integrity and consistency, excludes the data that does not meet the rules, and re-requests the data, evaluates the data integrity index, verifies the data quality and security, and obtains the aggregated complete data; The data noise protection module classifies data based on the aggregated complete data, identifies sensitive information and non-sensitive information, applies random noise to sensitive information, uses asymmetric encryption technology to process personal identification information, re-encodes data, optimizes the security of data during transmission, and obtains noise-added data; The data processing implementation module dynamically allocates processor and memory resources based on the noise-added data and task priorities, configures the data processing path, adjusts the data flow and the load of the processing node, selects a matching processing algorithm for data processing based on data characteristics and processing objectives, and adjusts processing parameters to obtain optimized processing information; The processing result output module organizes the data results based on the optimized processing information, performs processing result visualization, displays the data processing results through charts and graphs, adjusts the output format of the data processing results, optimizes the readability and comprehensibility of the data processing results, and obtains processing output information.

Citation Information

Patent Citations

  • Data integration system and method

    CN102043837A

  • Account book database management method based on electronic invoice voucher

    CN117788190A

  • Real-time data migration method and computer readable medium

    CN118193502A

  • Whole-process optimal scheduling method and system for meteorological satellite data

    CN118819781A

  • Full-dimensional battlefield situation information presentation device and terminal equipment

    CN118885253A

Cited By

  • Verification method, system and equipment for sensitive data in database and storage medium

    CN120470632A

  • Multi-source data integration method based on data space

    CN120578657A

  • A multi-source data integration method based on data space

    CN120578657B