Distributed intelligent data calculation method and system based on source sticking mode

By adopting a distributed intelligent data computing method based on the source-attached model in the healthcare field, and deploying source-attached computing nodes to process data in parallel, the problem of computing latency and throughput reduction caused by centralized data collection is solved. This achieves efficient and accurate data processing and rapid response, adapting to the dynamic expansion of medical equipment and business needs.

CN121349680AInactive Publication Date: 2026-01-16XINJIANG GREATSOFT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511492560.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-01-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies are not computationally efficient in the dynamic expansion of large-scale, multi-source data, especially in the healthcare field. Centralized collection and batch processing frameworks lead to computational latency and reduced throughput, failing to meet the needs of second-level emergency response and rapid response to public health emergencies.

Method used

A distributed intelligent data computing method based on the source-attached model is adopted. By deploying source-attached computing nodes in business scenarios, the system receives data from local sensors and external interfaces and performs parallel computing. It generates standardized results based on a distributed scheduling framework, performs classification storage and validity determination to avoid invalid data accumulation and computing delays, and generates comprehensive data indicators by comparing the correlation of node data.

Benefits of technology

It enables efficient processing of large-scale, multi-source data, meets the needs of emergency response in seconds and rapid response to public health emergencies, improves computing efficiency and output accuracy, ensures data quality and consistency, and adapts to the diverse data formats of newly added medical equipment and the needs of frequent business iterations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121349680A_ABST
    Figure CN121349680A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed intelligent data calculation method and system based on a source sticking mode, and relates to the technical field of data calculation. According to the method, in a source attaching mode, source attaching computing nodes are deployed according to a service scene, local sensor monitoring data and external interface interaction data are received, parallel computing and primary processing are synchronously executed on the two types of data, and cross-regional bandwidth pressure is reduced from the source; a data index optimization algorithm is issued through a distributed scheduling framework, standardized calculation results are generated and stored in a classified mode, validity is judged, and calculation delay caused by invalid data is avoided; all standardized calculation results passing validity judgment are collected, data formats are unified and then uploaded to a central platform, comprehensive data indexes are generated, efficient processing and accurate output of large-scale multi-source data are achieved, the real-time performance, throughput and dynamic expansibility of data processing are improved, and the data processing efficiency is improved. And therefore, the problem of low calculation efficiency of the large-scale multi-source data in the dynamic expansion process in the prior art is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data computing technology, and in particular to a distributed intelligent data computing method and system based on a source-attached model. Background Technology

[0002] With the accelerated digitalization of the healthcare sector, multi-source data such as electronic health records, electronic medical records, diagnostic and treatment data, and public health monitoring data are experiencing explosive growth. This data encompasses diverse sources including primary healthcare service platforms, medical insurance settlement systems, and wearable devices, exhibiting complex formats and stringent real-time requirements (e.g., emergency room data). Existing technologies primarily employ centralized data acquisition tools to batch-collect information from various data sources. These tools then perform data cleaning, format conversion, and standardization processes such as extraction, transformation, and loading before storing the standardized data in a unified data warehouse. Subsequently, offline batch processing frameworks are used to summarize and statistically analyze historical data, or simple real-time computing components are used to process some high-frequency data. Finally, based on a pre-defined fixed algorithm model, basic analysis is performed on the processed data to output standardized conclusions such as diagnostic and treatment data reports and public health statistical results, supporting the basic data application needs of the healthcare sector.

[0003] For example, Chinese Invention Patent CN117407407B discloses a method, apparatus, device, and computer medium for updating a dataset from multiple heterogeneous data sources. This includes: collecting data from multiple heterogeneous data sources into an associated source database based on the data information required to construct the dataset, thus obtaining an initial dataset from multiple heterogeneous data sources; generating a first task mapping table corresponding to the data update information in response to receiving data update information from the aforementioned multiple heterogeneous data sources; collecting updated data into the source database based on the first task mapping table; and updating the initial dataset from multiple heterogeneous data sources based on the updated data in the source database, thus obtaining the dataset from multiple heterogeneous data sources.

[0004] For example, the Chinese invention patent application CN120723742A discloses an automated local knowledge base construction system based on multi-source data acquisition and distributed computing, which includes: a multi-source data collector and a distributed data processing and parsing module; the multi-source data collector realizes the access, data acquisition, data transmission and incremental update of multi-source data, and generates incremental change records; the distributed data processing and parsing module includes multiple distributed nodes, which are used to perform distributed processing and parsing of data files based on incremental change records on the data collected by the multi-source data collector, and establish a local knowledge base.

[0005] In scenarios where healthcare data needs to respond quickly to public health emergencies and support dynamic adjustments to personalized treatment plans, the continuous addition of new data sources (such as the integration of new smart medical devices) and frequent iterations of business requirements (such as the addition of new infectious disease monitoring indicators) necessitate dynamic expansion of the data processing architecture and analysis models to adapt to these changes. However, existing technologies rely on centralized collection and batch processing frameworks. When the data scale grows exponentially (such as tens of millions of new medical records added daily) or real-time requirements surge (such as emergency decisions requiring second-level response), single-node resource bottlenecks and rigid task scheduling lead to a sharp increase in computational latency. This results in a decrease in data processing throughput and a lag in the output of analysis results during dynamic expansion, highlighting the problem of low computational efficiency for large-scale, multi-source data during dynamic expansion. Summary of the Invention

[0006] To address the problem of low computational efficiency in existing technologies for large-scale, multi-source data during dynamic expansion, this invention provides a distributed intelligent data computing method and system based on a source-attached model. The technical solution is as follows: On the one hand, a distributed intelligent data computing method based on the source-attached model is provided. This method includes: Step 1, in the source-attached model, i.e., under the distributed data acquisition architecture, source-attached computing nodes deployed at preset data acquisition points in the business scenario receive monitoring data collected in real time from local sensors and interactive data transmitted from external interfaces, and perform parallel computing and preliminary processing to reduce bandwidth pressure caused by cross-regional data transmission; Step 2, according to the established distributed scheduling framework, a preset data indicator optimization algorithm is issued to each source-attached computing node to generate standardized computing results that meet the business needs of each source-attached computing node, while performing classified storage and validity judgment to avoid computing delays caused by invalid data accumulation; Step 3, the standardized computing results after validity judgment are summarized and uploaded to the established central platform in a unified data format. By comparing the correlation of data from different nodes to check for calculation deviations, a comprehensive data indicator adapted to the business needs of each source-attached computing node is generated to achieve efficient processing and accurate output of large-scale multi-source data.

[0007] On the other hand, a distributed intelligent data computing system based on the source-attached model is provided. This system applies a distributed intelligent data computing method based on the source-attached model. The system includes: a parallel computing and processing module for each source-attached computing node, which is used to receive monitoring data collected in real time from local sensors and interactive data transmitted from external interfaces through source-attached computing nodes deployed at preset data collection points in the business scenario under the source-attached model, i.e., the distributed data acquisition architecture, and perform parallel computing and preliminary processing to reduce the bandwidth pressure caused by cross-regional data transmission volume; a standardized computing and validity judgment module, which is used to issue preset data indicator optimization algorithms to each source-attached computing node according to the constructed distributed scheduling framework to generate standardized computing results that meet the business needs of each source-attached computing node, and at the same time perform classification storage and validity judgment to avoid computing delays caused by invalid data accumulation; and a source-attached computing node indicator generation module, which is used to summarize the standardized computing results after validity judgment, upload them to the established central platform in a unified data format, and generate comprehensive data indicators that adapt to the business needs of each source-attached computing node by comparing the correlation of data from different nodes to identify computing deviations, so as to achieve efficient processing and accurate output of large-scale multi-source data.

[0008] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. This invention, through a distributed data acquisition architecture with a source-attached mode, deploys source-attached computing nodes according to preset collection points in healthcare business scenarios. These nodes directly receive real-time data from local medical sensors and interactive data from external interfaces, processing them in parallel. This reduces cross-regional data transmission volume at the source, alleviates bandwidth pressure, and avoids transmission congestion caused by adding new data sources in a centralized architecture. Relying on a distributed scheduling framework, data indicator optimization algorithms are distributed to each node, generating standardized results, which are then categorized, stored, and their validity determined. This avoids computational delays caused by data accumulation and solves the problem of rigid task scheduling in batch processing frameworks. Valid results are summarized and uploaded to the central platform in a unified format. By comparing the correlation between node data to identify discrepancies, comprehensive data indicators are generated. This step enables efficient processing of large-scale, multi-source data, meeting the needs of second-level emergency response and rapid response to public health emergencies. It solves the problems of decreased data throughput and delayed analysis results during dynamic expansion, improving computational efficiency and output accuracy.

[0009] 2. By verifying integrity and numerical reasonableness, a data quality score is derived using the geometric mean. This score is then combined with the average quality deviation over consecutive periods to determine data compliance. This approach accurately filters invalid data, corrects deviations, and adapts to the diverse data formats of newly added medical equipment. It avoids resource consumption by low-quality data in centralized processing. If the quality is substandard, the sampling frequency is dynamically adjusted to optimize acquisition. If it still fails to meet the standards, sensor calibration is prompted, ensuring the quality of source data and reducing subsequent computational resource waste. Qualified data is split into tasks based on type and time window, and distributed to various source computing nodes for parallel processing. Cross-validation is then used to assess the consistency of node computations. This distributed computing model improves data processing throughput, meeting the second-level response requirements of emergency departments. Inconsistent results are subject to task reassignment and secondary calculation to ensure result reliability. This supports the dynamic expansion of high-frequency iterative services such as infectious disease monitoring, resolving latency issues caused by the rigid scheduling of centralized batch processing frameworks, and ensuring data processing efficiency and accuracy during public health emergencies.

[0010] 3. By first extracting the timestamps and values ​​of the calculation results from each source computing node when acquiring parallel computing evaluation values, unifying the format and sorting them, and then interpolating to fill in missing data according to a preset granularity, time series alignment is achieved, adapting to the different sampling frequencies of newly added medical equipment. Next, values ​​with the same timestamp are matched, the differences between nodes and their deviations from the average are calculated, generating a difference sequence. A rate of change sequence is obtained through differential operations, and its standard deviation is used as the evaluation value. This process, through time alignment and difference analysis, accurately quantifies the consistency of node calculations, supporting the dynamic expansion of services such as infectious disease monitoring. Distributed parallel processing avoids centralized bottlenecks, provides second-level response to emergency data needs, solves the throughput reduction problem under tens of millions of medical records, ensures high efficiency in data processing during dynamic expansion, and provides a reliable basis for verifying calculation results in response to public health emergencies.

[0011] 4. By first matching the parallel calculation results of each feeder computing node with a preset business indicator template when generating standardized calculation results, the system outputs results containing basic indicator values ​​and business judgment conclusions. This adapts to the frequently iterating business needs such as infectious disease monitoring, allowing for rapid integration of new indicators without refactoring the architecture. Next, the system first identifies and calibrates any data acquisition drift in the feeder computing nodes. Using historically available results as a benchmark, it calculates offset compensation and correction through regression analysis to reduce measurement errors and accommodate the accuracy differences of newly added intelligent medical devices. If the results are still unsatisfactory after calibration, they are marked for manual review and recorded as abnormal, avoiding interference from invalid data. The entire process relies on a distributed scheduling framework, eliminating the need for centralized batch processing. It can efficiently process tens of millions of medical records, meet the second-level response requirements of emergency departments, solve the problems of decreased throughput and delayed results during dynamic expansion, and improve the efficiency of large-scale multi-source data processing.

[0012] 5. By classifying and storing data and determining its validity, qualified results are first mapped to three storage partitions—real-time, historical, and abnormal—based on business tags and data types. This facilitates rapid data retrieval and adapts to the diverse data needs during public health emergencies. The positive / negative timestamp deviation is then calculated, and the validity assessment result is obtained through harmonic averaging. This ensures precise control over the time dimension quality of data and prevents the accumulation of invalid data. If the result meets the standards, storage is completed; otherwise, repair is triggered: if latency is the primary factor, the upload frequency is adjusted via API; if advancement is the primary factor, the collection logic is optimized; if the time base is unstable, the node clock is calibrated; if the results still do not meet the standards after repair, the data is archived for review and an alarm is triggered, ensuring data validity. The entire process relies on a distributed architecture, eliminating the need for centralized batch processing. It meets the requirements for second-level emergency response, solves the problems of decreased throughput and delayed results during dynamic scaling, and improves the efficiency of large-scale, multi-source data processing.

[0013] 6. When generating comprehensive data indicators, the standardized results of each feed source calculation node are first treated as a multi-dimensional vector. Cosine similarity combined with a standard deviation weighting factor is used to calculate the optimized correlation similarity between nodes, accurately quantifying data synergy and adapting to the multi-dimensional data comparison needs of newly added medical equipment. If the similarity meets the standard, the similarity increment is calculated; if it does not, abnormal nodes are isolated and rolled back to avoid data drift interfering with the results. If aggregation still fails to meet the standard, feed source collaboration optimization is initiated: core deviation nodes are located, and synchronization cycles and upload gaps are allocated based on historical synchronization error peaks to reduce transmission conflicts; if collaboration meets the standard for two consecutive cycles, the configuration is solidified; otherwise, the clock is calibrated. If it still fails to meet the standard after repair, an alert is issued. The degree of node collaboration is measured by the product of timestamp alignment rate and effective interaction rate, ensuring data consistency across multiple nodes. The entire process relies on distributed collaboration to efficiently support responses to public health emergencies. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart illustrating a distributed intelligent data computing method based on a source-attaching pattern provided in an embodiment of the present invention; Figure 2 A flowchart for generating comprehensive data indicators provided in this embodiment of the invention; Figure 3 This is a flowchart of source-attachment collaborative optimization provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a distributed intelligent data computing system based on the source-attachment mode provided in an embodiment of the present invention. Detailed Implementation

[0016] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0017] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0018] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0019] This invention provides a distributed intelligent data computing method based on a source-attached model, such as... Figure 1 The flowchart shown is for a distributed intelligent data computing method based on the source-attached model. The processing flow of this method may include the following steps: Step one: In the source-based mode, i.e., under the distributed data acquisition architecture, source-based computing nodes deployed at preset data acquisition points in the business scenario receive real-time monitoring data from local sensors and interactive data transmitted from external interfaces, and perform parallel computing and preliminary processing to reduce bandwidth pressure caused by cross-regional data transmission. Step two: Based on the established distributed scheduling framework, preset data indicator optimization algorithms are distributed to each source-based computing node to generate standardized calculation results that meet the business needs of each node. Simultaneously, the results are categorized, stored, and validated to avoid computational delays caused by the accumulation of invalid data. The preset data indicator optimization algorithm is a set of algorithms adapted to the business needs of each post source node. It contains three core categories: first, data cleaning algorithm, which removes null values ​​and corrects format errors; second, indicator conversion algorithm, which converts raw data into standard business indicators; and third, anomaly detection algorithm, which identifies out-of-range data according to thresholds. In the third step, the standardized calculation results after validity judgment are summarized and uploaded to the established central platform in a unified data format. By comparing the correlation of data from different nodes, calculation deviations are investigated, and comprehensive data indicators adapted to the business needs of each post source calculation node are generated to achieve efficient processing and accurate output of large-scale multi-source data.

[0020] This method focuses on alleviating bandwidth pressure, avoiding computational latency, and improving the efficiency of large-scale multi-source data processing. Its value can be fully demonstrated by combining it with a medical and health scenario example: During the peak of influenza season, a tertiary hospital needs to urgently connect 50 new intelligent body temperature monitoring devices to expand the monitoring scope. At the same time, it needs to add two infectious disease monitoring indicators: the frequency of fever patients seeking medical treatment and the infection rate in the ward. In addition, 12 million new medical records are added every day. In the emergency scenario, doctors need to obtain patient data within 3 seconds to support decision-making. The traditional centralized architecture can no longer meet the needs due to transmission congestion and slow processing. This method employs the following steps: Step 1 involves deploying source-attached computing nodes at key data collection points such as outpatient triage desks and inpatient wards. These nodes directly receive real-time temperature data from monitors and patient interaction data from the electronic medical record system, simultaneously performing parallel computation and preliminary cleaning (e.g., removing abnormal fever data below 37℃ and correcting records with missing patient IDs). Only valid data is transmitted, resolving bandwidth congestion issues associated with centralized data collection and ensuring stable access for 50 new devices. Step 2 utilizes a distributed scheduling framework to precisely distribute fever patient identification algorithms to each node, generating standardized results tailored to the needs of infectious disease and emergency departments. These results are stored according to their intended use, while invalid data (e.g., duplicate temperature records) is removed to avoid computational delays. Step 3 aggregates data to the central platform according to the HL7 FHIR medical standard, comparing the correlation between outpatient and ward data to identify discrepancies. Within 3 seconds, influenza monitoring indicators are generated, including the fever rate in each ward and 24-hour visit trends. This satisfies real-time emergency needs and supports subsequent personalized treatment (e.g., adjusting medication based on patient temperature changes), completely resolving the issues of reduced throughput and delayed results in centralized architectures.

[0021] Furthermore, the parallel computing and preliminary processing specifically include raw data preprocessing for removing invalid data and correcting data deviations, and multi-node parallel computing for synchronously calculating basic business indicators after the raw data preprocessing is qualified.

[0022] The specific process of raw data preprocessing is as follows: Based on the integrity verification results and numerical reasonableness verification results of the key fields corresponding to the specified batch of raw data, the data quality is quantitatively evaluated, and a data quality score is output. The data quality score is obtained by calculating the geometric mean of the integrity verification pass rate and the numerical reasonableness verification pass rate. The geometric mean is less sensitive to extreme values ​​(such as the pass rate of a certain indicator being close to 0 or 100%) than the arithmetic mean, to prevent abnormal fluctuations of a single indicator from distorting the overall score. For example, if the integrity verification pass rate is 90%, the numerical reasonableness verification pass rate is 10%, the arithmetic mean is 50%, and the geometric mean is... The geometric mean, approximately 30%, indicates that a weakness in any one indicator can lower the overall score. This more accurately reflects the true quality of data across both dimensions. Furthermore, this calculation method highlights the synergistic effect of the two indicators; a higher score can only be obtained when both completeness and reasonableness meet the standards, reducing business risks caused by a single dimension meeting the standard while the other is deficient. The completeness verification pass rate represents the proportion of data records in a specified batch of raw data where key fields (such as timestamps and device identifiers) are complete and correctly formatted. The numerical reasonableness verification pass rate represents the proportion of data records in a specified batch of raw data where the data values ​​are within the preset device range. These data values ​​can be specific quantitative values ​​collected by the source computing node that reflect the status of the monitored object, such as body temperature in a medical setting.

[0023] The data quality score is compared with a preset data quality score to calculate the quality score deviation. The average quality score deviation over a preset number of data collection periods (usually two consecutive periods) is calculated. If the average quality score deviation is within the preset deviation range, the batch of raw data is marked as qualified raw data and enters the multi-node parallel computing stage. The preset data quality score is represented by the sum and average of historical data quality scores from the historical distributed intelligent data computing process in the database. Otherwise, a data collection anomaly alarm is triggered, and data collection optimization is performed in the next collection period to increase the data sampling frequency by using the average quality score deviation. When the average quality score deviation exceeds the preset deviation range by 10%, the sampling frequency is increased by 20%-30% (e.g., from 1 second / time to 0.7-0.8 seconds / time). The increased frequency increases the data sample size and reduces the impact of single collection anomalies on the overall quality, thereby improving the integrity and rationality of the original data. After data collection optimization, if the re-acquired average quality score deviation is within the preset deviation range, the multi-node parallel computing stage is entered; otherwise, the preset personnel are prompted to perform sensor calibration, such as checking and reconfiguring the data transmission interface. The preset deviation range is based on the historical data quality score deviation values ​​of the past 3 months in the database, and its 95% confidence interval is taken as the preset deviation range.

[0024] The specific process of multi-node parallel computing is as follows: Qualified raw data is divided into a preset number of computing tasks according to the collection type and time window, and allocated to each feed source computing node for parallel computing, outputting the parallel computing results; the parallel computing results are cross-validated to output a parallel computing evaluation value used to quantify the logical consistency of computing results among different feed source computing nodes; if the parallel computing evaluation value is within the preset parallel computing evaluation range, the parallel computing results are encapsulated into a standard data format and input into the constructed distributed scheduling framework for subsequent operations. The preset parallel computing evaluation range is: statistically analyzing the parallel computing evaluation values ​​that meet the logical consistency standard within the past 6 months, taking the minimum and maximum values ​​to form an interval as the preset evaluation range, ensuring the rationality of the judgment standard; conversely, for feed source computing nodes that are not within the preset parallel computing evaluation range, after task reassignment and secondary parallel computing, if the re-acquired parallel computing evaluation value is still not within the preset parallel computing evaluation range, a computing anomaly alarm is issued; otherwise, the corresponding parallel computing result is input into the constructed distributed scheduling framework.

[0025] Specifically, the process for obtaining the parallel computing evaluation value is as follows: extract the timestamps and corresponding values ​​from the results of each source computing node, convert timestamps of different formats (such as records containing milliseconds and records without milliseconds) to the same format, and then sort them in chronological order. This step eliminates comparison barriers caused by differences in time formats, ensuring that all node data are comparable in the time dimension.

[0026] Using a preset time granularity (e.g., 1 second) as a benchmark, a continuous standard timeline is generated. For missing time points on the standard timeline at each patch source calculation node, linear interpolation is used to fill in the missing data. This involves calculating the intermediate missing value based on the known data before and after the missing point, proportional to the time interval. This addresses the issue of inconsistent sampling frequencies among different patch source calculation nodes (e.g., some nodes sample twice per second, while others sample once per second), aligning all node data on the same time scale and ensuring a consistent time benchmark for subsequent calculations.

[0027] Based on the aligned timestamps, the values ​​of each node are matched one by one. First, the difference between the values ​​of any two source computing nodes at the same time point is calculated to form a difference sequence between source computing nodes. Then, the average value of all source computing nodes at the same time point is used as the benchmark to calculate the difference between the value of each source computing node and the average value, forming a difference sequence between the source computing nodes and the benchmark. This not only presents the direct deviation between source computing nodes intuitively, but also anchors the relative level of deviation through the overall average value, thus comprehensively reflecting the data synergy.

[0028] The difference sequence between the node and the baseline is differentiated to obtain the rate of change sequence of the difference over time. The standard deviation of this rate of change sequence is then calculated as an evaluation value for parallel computing. The smaller the standard deviation, the more synchronized and stable the trend of deviation changes among the source computing nodes (e.g., the numerical fluctuation amplitude and direction of multiple devices are consistent), indicating higher logical consistency of the parallel computing results. This process, by amplifying subtle deviation changes and quantifying the degree of dispersion, can accurately identify anomalies in the calculations of source computing nodes in large-scale data, providing a reliable basis for collaborative verification in decision-making in scenarios such as healthcare.

[0029] In this embodiment, during parallel computing and preliminary processing, the raw data preprocessing uses a geometric mean score for integrity and numerical rationality verification to accurately quantify data quality, avoid distorting scores due to single-indicator anomalies, reduce business risks, and dynamically optimize sampling frequency and trigger sensor calibration based on deviation values, effectively improving the integrity and rationality of raw data. Multi-node parallel computing divides tasks by type and time window and allocates them to each node. Combined with cross-validation and standard deviation quantification of parallel computing evaluation values, it ensures logical consistency of calculation results. In case of anomalies, tasks are reassigned for secondary calculations, ensuring accurate synchronous calculation of basic business indicators. Compared with existing technologies, this process solves the problems of difficult data quality control, low computational efficiency, and rigid anomaly handling in centralized processing. Raw data preprocessing reduces latency caused by invalid data accumulation, and multi-node parallel computing overcomes single-node resource bottlenecks, achieving efficient processing of large-scale, multi-source data. This provides a reliable data foundation for real-time decision-making in scenarios such as healthcare (e.g., emergency data support), improving data collaboration and business response speed.

[0030] Furthermore, standardized calculation results that meet the business needs of each feed source computing node are generated. Specifically, the parallel calculation results of each feed source computing node in the constructed distributed scheduling framework are matched with a preset business indicator template to output standardized calculation results. The standardized calculation results include quantitative values ​​of basic indicators reflecting the data status within the current collection period, and business rule judgment results reflecting the business status judgment conclusion based on preset rules. These include qualified and unqualified results. Qualified results refer to those that meet the business scenario requirements (e.g., a patient's body temperature of 36-37.2℃ in a medical scenario), while unqualified results refer to abnormal data trends (e.g., a sudden rise or fall in a patient's body temperature within a short period) and mismatched related data (e.g., timestamps and business events are misaligned). The output standardized calculation results are judged based on preset qualified judgment criteria. If the quantified value of the basic indicator is within the preset range and the business rule judgment result is qualified, the output standardized calculation result will be marked as a usable result. If the quantified value of the basic indicator exceeds the preset range or the business rule judgment result is unqualified, it indicates that there is acquisition drift on the source side of the corresponding feed source calculation node, and feed source offset calibration will be performed to reduce the interference of measurement error of the feed source calculation node. The preset range is: the quantified value of the basic indicator in the historical usable results of the corresponding feed source calculation node in the past 6 months is statistically analyzed, and the 95% confidence interval is taken as the preset range. If the quantified value of the basic indicator exceeds the preset range and the business rule judgment result is unqualified, a standardized calculation warning will be issued.

[0031] Specifically, the source offset calibration involves: using the historical available results of the corresponding source calculation node as a benchmark, calculating the corresponding source offset through regression analysis, and subtracting the source offset from the current basic index quantification value as the reverse compensation amount to correct the current standardized calculation results; if the result is still unqualified after offset calibration correction, the batch of standardized calculation results is marked as pending manual review and an anomaly record is triggered.

[0032] Specifically, the distributed scheduling framework is constructed as follows: First, the business requirements and computing capabilities of each source computing node are analyzed, and task priorities are assigned. Then, a node communication network is built, and a task allocation module (supporting dynamic allocation based on node load), a result collection module (receiving calculation results in real time), and an anomaly monitoring module (tracking task progress) are deployed. Finally, historical data interfaces are integrated, and an indicator template library is integrated to form a distributed scheduling framework that can dynamically adapt to the addition or removal of nodes. The process of matching results with preset business indicator templates involves: extracting the parallel computing results of each node from the distributed scheduling framework and mapping the fields according to template fields (such as data type, unit, and statistical dimension); automatically converting data with inconsistent formats (such as inconsistent units), and filling missing non-critical fields with default values; after matching, standardized calculation results are output according to the display format specified by the template (such as table or line chart data).

[0033] In this embodiment, standardized results are generated by matching with preset business indicator templates, ensuring that the data format of each node is uniform and the meaning is clear, meeting the needs of different business scenarios. The qualification judgment standard combines numerical range and rule judgment, providing dual protection for data validity and reducing the risk of misjudgment from a single dimension. The source offset calibration uses historical available results as a benchmark and accurately corrects acquisition drift through regression analysis, reducing measurement errors and improving data accuracy. Abnormal situations are handled in a graded manner (early warning, manual review), which not only responds to serious problems in a timely manner but also retains abnormal data for traceability. Compared with traditional methods, this significantly improves the degree of data standardization and reliability.

[0034] Further, the data is categorized and stored, and its validity is determined. The specific process is as follows: After the standardized calculation results are generated, all qualified results are statistically analyzed, and based on their corresponding business tags and data types, they are mapped to the corresponding storage partitions. The storage partitions include a real-time data area, a historical data area, and an abnormal data area. The real-time data area is used to support real-time business queries and decisions, such as real-time vital sign data of emergency patients in medical scenarios; the historical data area is used for trend analysis and backtracking, such as monthly diagnosis and treatment data statistics; the abnormal data area is used to avoid interfering with the storage and use of normal data, such as monitoring data with excessive timestamp deviation. The validity dimension calculation is performed on the data in the storage partitions to obtain the corresponding validity evaluation results. The validity dimension calculation includes timestamp positive deviation calculation and timestamp negative deviation calculation. The timestamp positive deviation calculation indicates: the actual time of data collection from the source calculation node is extracted. The timestamp is compared with the reference deadline of the preset collection period. The difference between the actual timestamp and the reference deadline is calculated, and the proportion of this difference to the preset collection period length is calculated, reflecting the degree of delay in uploading the source data. The timestamp negative deviation calculation is: the actual timestamp of the data collected by the source calculation node is compared with the reference start time of the preset collection period. The difference between the actual timestamp and the reference start time is calculated, and the proportion of this difference to the preset collection period length is calculated, reflecting the degree of advance of the source data collection. The validity evaluation result is the result of harmonic averaging the results of the timestamp positive deviation calculation and the timestamp negative deviation calculation. The harmonic average has low sensitivity to extreme deviation values, can balance the impact of timestamp positive and negative deviations, avoid excessive distortion of the overall validity evaluation result by a single deviation, and more objectively reflect the comprehensive qualification of the data time dimension.

[0035] If the validity assessment result is less than the preset validity assessment result, it is determined to be valid and classified storage is completed. Otherwise, it is determined to be invalid and the timestamp repair process is triggered to correct the time deviation of the data collected by the source computing node and ensure the accuracy of the data time dimension. The preset validity assessment result is represented by the sum and average of the historical validity assessment results in the historical distributed intelligent data computing result classification storage process in the database.

[0036] The triggering of the timestamp repair process is as follows: If the positive timestamp deviation is greater than the negative timestamp deviation, the time synchronization interface of the source computing node is called based on the first timestamp deviation to adjust the data upload scheduling frequency to compress latency; if the positive timestamp deviation is less than the negative timestamp deviation, the preset personnel are prompted to optimize the data collection triggering logic of the source computing node based on the second timestamp deviation to avoid premature data collection; if the positive timestamp deviation is equal to the negative timestamp deviation, it indicates that the corresponding source computing node has an unstable time base, which cannot guarantee the accuracy and consistency of time recording. Therefore, the preset personnel need to be prompted to calibrate the node's local clock first to restore the reliability of the time base; the first timestamp deviation and the second timestamp deviation quantify the data collection degree of the corresponding source computing node, respectively. The first timestamp deviation represents the difference between the positive and negative timestamp deviations, and the second timestamp deviation represents the difference between the negative and positive timestamp deviations; after triggering the timestamp repair process, the validity assessment result is re-acquired. If the re-acquired validity assessment result is not less than the preset validity assessment result, the classification and storage are completed; otherwise, it is marked as pending manual review and archived to the historical abnormal data archive area, and an abnormal alarm is triggered.

[0037] In this embodiment, the preset collection cycle is usually 1 minute to 1 hour (e.g., 1 minute / time for medical vital sign monitoring), and the reference cutoff time is usually the end of the collection cycle (e.g., the cutoff time for a 1-minute cycle is 01 or 02 minutes on the hour). Both are set according to business needs and data usage frequency. When there is a sudden business (e.g., a public health event) that requires high-frequency monitoring, or when the equipment load is too high and the collection frequency needs to be reduced, the cycle and cutoff time can be fine-tuned.

[0038] This example maps qualified results to real-time, historical, and abnormal data areas based on business tags and data types, achieving hierarchical data storage. This ensures efficient real-time business queries, facilitates historical data backtracking and analysis, and prevents abnormal data from interfering with normal storage. It considers positive and negative timestamp deviations because positive deviations reflect delayed data uploads (affecting real-time decision-making), while negative deviations indicate premature data collection (potentially leading to data misalignment). Both can compromise the accuracy of the time dimension. By using a harmonic average to calculate the validity assessment results, the impact of both types of deviations can be objectively balanced, avoiding the distortion of the assessment by a single deviation. The timestamp repair process precisely handles deviations according to their magnitude (adjusting upload frequency, optimizing trigger logic, and calibrating the clock). If the data is still unqualified after repair, it is archived and an alarm is triggered. This not only corrects time deviations to ensure data accuracy but also reduces the impact of abnormal data on business through hierarchical processing, thus improving the overall reliability of data storage and use.

[0039] like Figure 2The flowchart for generating comprehensive data metrics shown below has the following design logic: First, it determines whether the obtained optimized association similarity exceeds a preset value. If it does, it obtains the similarity increment and further determines its relationship with the preset value. If it does not exceed the preset value, it triggers a node anomaly warning. Based on the similarity increment judgment result, it decides whether to generate a weighted average metric or dynamically adjust the weight allocation. After the weight adjustment, based on whether the improvement in similarity after compensation meets the standard, it decides whether to combine the compensation result to obtain comprehensive data metrics or trigger an aggregation anomaly warning for source-attachment collaborative optimization to ensure the accuracy and stability of data processing.

[0040] Further understanding is needed regarding the specific steps for generating comprehensive data indicators: performing multi-node data correlation analysis on the standardized calculation results after validity assessment, and calculating the data matching degree between each feed source computing node according to the cosine similarity formula, denoted as the optimized correlation similarity; if the optimized correlation similarity is greater than the preset optimized correlation similarity, the difference between the optimized correlation similarity and the preset optimized correlation similarity is used as the similarity increment to reflect the degree to which the data collaboration of the feed source computing nodes exceeds the baseline level; if the optimized correlation similarity is not greater than the preset optimized correlation similarity, a node anomaly warning is triggered, determining that there is data drift on the source side of the corresponding feed source computing node, automatically isolating the node's data and rolling it back to the previous valid state. The preset optimized correlation similarity is calculated by statistically analyzing the optimized correlation similarity of each feed source computing node when data collaboration is normal within the past 3 months, taking its 80th percentile as the preset value to ensure coverage of most normal collaboration scenarios.

[0041] If the similarity increment is not greater than the preset similarity increment, a comprehensive data index for the current scenario is directly generated based on the weighted average of the standardized calculation results of each feed source calculation node. The preset similarity increment is calculated based on the similarity increment of the degree of collaboration of node data in historical data, and the average value is increased by 10% as the preset value to reserve reasonable space for collaborative improvement. If the similarity increment is greater than the preset similarity increment, the aggregation correction coefficient compensation amount is obtained based on the deviation ratio between the similarity increment and the preset similarity increment. By dynamically adjusting the weight distribution among the feed source calculation node data, the degree of aggregation between multi-node data is improved. The deviation ratio represents the ratio of the difference between the similarity increment and the preset similarity increment to the preset similarity increment. After compensation, if the improvement in similarity obtained is not less than the preset improvement, a comprehensive data index is obtained by combining the compensation result; otherwise, an aggregation anomaly warning is triggered to optimize feed source collaboration. The preset improvement is calculated by analyzing the improvement in similarity after historical aggregation correction and taking the median as the preset value to ensure that the improvement in similarity after correction achieves a practical effect.

[0042] The aggregation correction coefficient represents a dynamic coefficient used to adjust the weights of data at each node. Its value is positively correlated with the stability and historical collaborative performance of the node data (the higher the stability and the better the collaboration, the larger the coefficient). The deviation ratio is linearly mapped from 0 to 1 to an aggregation correction coefficient compensation amount of 0.1-0.5 (the larger the ratio, the higher the compensation amount). Based on this compensation amount, the original weights of each node are corrected: for nodes with high stability, the weight = original weight × (1 + compensation amount × 0.8); for nodes with low stability, the weight = original weight × (1 - compensation amount × 0.5). After correction, the total weight remains 1. By tilting the weights towards high-quality nodes, the interference from abnormal nodes is reduced, improving the consistency and reliability of multi-node data aggregation.

[0043] like Figure 3 The flowchart shown illustrates the source-linking collaborative optimization process. Its design logic is as follows: First, the core nodes are located, and the time synchronization cycle and data upload interval are reallocated accordingly, laying the foundation for subsequent operations. Next, it is determined whether the node collaboration level exceeds a preset value within two consecutive collection cycles. If it does, the current state is maintained; otherwise, a prompt is made to calibrate the local clock reference. Then, the similarity improvement rate is re-acquired, and it is determined whether it reaches the preset improvement rate. If it meets the target, a comprehensive data indicator is obtained based on the node collaboration level monitoring results; if it does not meet the target, a collaboration anomaly warning is triggered, thereby achieving optimized control of node collaboration and ensuring stable system operation.

[0044] It's important to understand that the specific process of source-based collaborative optimization is as follows: First, acquire the collection logs and synchronization records of the corresponding source-based computing nodes to locate the core nodes where data collaboration deviations occur. The core node represents the source-based computing node where the data collaboration deviation occurs. Second, based on the acquired historical synchronization error peak value of the core node, reallocate the time synchronization cycle and data upload interval of the core node to reduce data transmission conflicts between nodes. The historical synchronization error peak value represents the maximum deviation between the actual synchronization time and the theoretical synchronization time of the core node in past collection cycles. Third, monitor the node collaboration level. If the corresponding node collaboration level is greater than the preset node collaboration level within two consecutive collection cycles, maintain the current node collaboration state; otherwise, prompt the preset personnel to recalibrate the local clock reference of the core node. Fourth, if the similarity improvement after monitoring the node collaboration level is not less than the preset improvement, obtain comprehensive data indicators based on the node collaboration level monitoring results; otherwise, issue a collaboration anomaly warning. The node collaboration level represents the product of the proportion of aligned timestamps of each core node to the total data volume and the proportion of effective cross-node data interaction to the total interactive data volume, intuitively reflecting the consistency level of multi-node data.

[0045] The optimized method for obtaining association similarity is as follows: First, the standardized calculation results of each source node are transformed into multi-dimensional vectors. That is, the number of dimensions of each node vector is consistent with the number of basic indicators, and the value of each dimension corresponds to a quantified value of a basic indicator output by that node (such as body temperature, heart rate, and blood pressure in a medical scenario). Next, the cosine similarity between any two node vectors is calculated: by dividing the vector dot product by the product of the magnitudes of the two vectors, the association similarity reflecting the consistency of the data direction between the two nodes is obtained (the closer the value is to 1, the more synchronized the data trend). Each pair of nodes participating in the collaborative analysis corresponds to one association similarity result. Based on this, a standard deviation weighting factor is introduced, which is inversely weighted according to the standard deviation of the corresponding basic indicator value sequence (the smaller the standard deviation, the higher the indicator stability, and the greater the weight), to highlight the contribution of stable indicators to the association similarity. Finally, the arithmetic mean of the weighted association similarities of all node pairs is taken to obtain the final optimized association similarity. This result reflects both the overall synergy of the node data and strengthens the influence of stable indicators through weight allocation, making the similarity calculation more aligned with actual business needs.

[0046] In this embodiment, the correlation similarity is optimized by calculating the weighting factor of multi-dimensional vector similarity and standard deviation. This differs from traditional single-dimensional matching, reflecting both data trend consistency and highlighting the weight of stable indicators, making the evaluation of node collaboration more accurate. Abnormal nodes are automatically isolated and rolled back to avoid cascading effects. When the similarity increment exceeds the limit, the weight is dynamically adjusted through deviation ratio mapping, solving the problem that traditional weighted averages are insensitive to collaboration fluctuations. Source-attached collaboration optimization, combined with log location of core nodes, adjusts the synchronization cycle according to historical error peaks, which is more targeted than existing methods that simply calibrate clocks, and continuous monitoring ensures collaboration stability. Overall, it realizes dynamic optimization of the entire process from data correlation analysis to anomaly handling, improving the accuracy and reliability of comprehensive data indicators, especially suitable for complex scenarios of multi-node collaboration.

[0047] This invention provides a distributed intelligent data computing system based on a source-attaching model, such as... Figure 4The diagram shows the structure of a distributed intelligent data computing system based on a source-based model. This system includes: a parallel computing and processing module for each source-based computing node, used in a source-based model (distributed data acquisition architecture) to receive real-time monitoring data from local sensors and interactive data transmitted from external interfaces via source-based computing nodes deployed at preset data acquisition points in the business scenario. This module performs parallel computing and preliminary processing to reduce bandwidth pressure caused by cross-regional data transmission. A standardized computing and validity determination module is used to distribute preset data indicator optimization algorithms to each source-based computing node according to the established distributed scheduling framework. This generates standardized computing results that meet the business needs of each source-based computing node, while also classifying and storing the results and determining their validity to avoid computational delays caused by invalid data accumulation. Finally, a source-based computing node indicator generation module is used to summarize the standardized computing results after validity determination, upload them to the established central platform in a unified data format, compare the correlation between data from different nodes to identify computational deviations, and generate comprehensive data indicators adapted to the business needs of each source-based computing node. This enables efficient processing and accurate output of large-scale, multi-source data.

[0048] In this embodiment, the parallel computing and processing modules of each source-attached computing node first process data locally at the source-attached node, reducing cross-regional transmission volume to alleviate bandwidth pressure and laying a low-latency data foundation for subsequent modules. The standardization calculation and validity determination module, based on the results of the preceding processing, generates standardized data through algorithmic generation and completes the determination and storage, filtering invalid data from the source to avoid dragging down the efficiency of subsequent indicator generation. The indicator generation module of each source-attached computing node summarizes the standardized data, checks for deviations, generates comprehensive indicators, and verifies the processing quality of the preceding modules. Compared with traditional independent modules, the three are closely connected, both dividing the work to solve bandwidth, latency, and accuracy issues and supporting each other for optimization, ultimately achieving efficient and accurate processing of large-scale multi-source data to meet business decision-making needs.

[0049] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. A computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. Semiconductor media can be solid-state drives.

[0050] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.

[0051] In various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0052] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0053] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A distributed intelligent data computing method based on the Pasteur mode, characterized in that, The method comprises: Step one, in the mode of sticking to the source, through the sticking source computing node deployed in the preset data collection point in the business scene, the monitoring data collected by the local sensor in real time and the interaction data transmitted by the external interface are received, and parallel computing and preliminary processing are carried out to reduce the bandwidth pressure caused by the cross-regional data transmission amount; Step two, according to the distributed scheduling framework constructed, the preset data index optimization algorithm is issued to each sticking source computing node to generate standardized calculation results meeting the business needs of each sticking source computing node, and classification storage and effectiveness determination are carried out to avoid calculation delay caused by invalid data accumulation; Step three, the standardized calculation results after effectiveness determination are summarized to upload to the established central platform in a unified data format, the calculation deviation is checked by comparing the correlation of data of different nodes, the comprehensive data index adapting to the business needs of each sticking source computing node is generated, and efficient processing and accurate output of multi-source data are realized.

2. The method of claim 1, wherein the method further comprises: The parallel computing and preliminary processing specifically include raw data preprocessing for eliminating invalid data and correcting data deviation, and multi-node parallel computing for synchronously accounting basic business indexes after the raw data preprocessing is qualified; The specific process of the raw data preprocessing is as follows: Based on the completeness verification result and the numerical reasonableness verification result of the specified batch of raw data corresponding to the key field, the data quality is quantitatively evaluated, and a data quality score value is output; The data quality score value is obtained by calculating the geometric mean of the completeness verification pass rate and the numerical reasonableness verification pass rate; The completeness verification pass rate represents the proportion of the number of data records with complete key fields and correct format in the total number of data records in the specified batch of raw data; The numerical reasonableness verification pass rate represents the proportion of the number of data records with data values within the preset equipment range in the total number of data records in the specified batch of raw data; The data quality score value is compared with the preset data quality score value, and a quality score deviation value is calculated; The average quality score deviation value of a continuous preset number of collection periods is calculated, if the average quality score deviation value is within the preset deviation range, the batch of raw data is marked as qualified raw data, and enters the multi-node parallel computing link; Otherwise, trigger data collection exception alarm and execute data collection optimization in the next collection period to improve data sampling frequency and improve raw data quality by improving average quality score deviation value; After data collection optimization, if the newly obtained average quality score deviation value is within the preset deviation range, the batch of raw data is marked as qualified raw data, and enters the multi-node parallel computing link, otherwise, the preset personnel is prompted to calibrate the sensor.

3. The method of claim 2, wherein the method further comprises: The specific process of the multi-node parallel computing is as follows: The qualified raw data is divided into a preset number of calculation tasks according to the collection type and time window, and is distributed to each sticking source computing node for parallel computing, and the parallel computing result is output; The parallel computing result is cross-validated to output a parallel computing evaluation value for quantifying the logical consistency of the calculation results of different sticking source computing nodes; If the parallel computing evaluation value is within the preset parallel computing evaluation range, the parallel computing result is packaged into a standard data format and input into the constructed distributed scheduling framework; If the parallel computing evaluation value is not within the preset parallel computing evaluation range, the task is redistributed and secondary parallel computing is performed on the source-side computing node, and if the reacquired parallel computing evaluation value is still not within the preset parallel computing evaluation range, an abnormality alarm is given, otherwise, the parallel computing result is packaged into a standard data format and input into the constructed distributed scheduling framework.

4. The method of claim 3, wherein the method further comprises: The specific process of obtaining the parallel computing evaluation value is as follows: The time stamp and corresponding value in the parallel computing result of each source-side computing node are extracted, the time stamp format is unified and sorted in ascending order of time; The parallel computing results of each source-side computing node are completed by linear interpolation with a preset time granularity as a reference, and the time sequence alignment of all parallel computing results is completed; The node values at the same time stamp are matched, the value difference between nodes at each time point is calculated, and a complete calculation result difference sequence is generated; The average value of the node values is taken as a reference value, the difference between the parallel computing result of each source-side computing node at the same time point and the corresponding reference value is obtained, and is taken as the calculation result difference sequence; The calculation result difference sequence is differentiated to obtain the difference rate sequence, and the parallel computing evaluation value is obtained based on the standard deviation of the difference rate sequence.

5. The method of claim 1, wherein the method further comprises: The specific process of generating a standardized calculation result that meets the business requirements of each source-side computing node is as follows: The parallel computing result of each source-side computing node in the constructed distributed scheduling framework is matched with a preset business index template to output a standardized calculation result, which includes a basic index quantization value for reflecting the data state in the current collection period, and a business rule judgment result for reflecting the business state judgment conclusion based on the preset rules; The output standardized calculation result is judged based on a preset qualified judgment standard: If the basic index quantization value is within the preset value range and the business rule judgment result is a qualified result, the output standardized calculation result is marked as a usable result; If the basic index quantization value exceeds the preset value range or the business rule judgment result is an unqualified result, it indicates that there is a collection drift on the source side of the corresponding source-side computing node, and a source-side offset calibration is performed to reduce the interference of the measurement error of the source-side computing node; If the basic index quantization value exceeds the preset value range and the business rule judgment result is an unqualified result, a standardized calculation warning is given. The specific process of performing source-side offset calibration is as follows: The historical usable result of the corresponding source-side computing node is taken as a reference, and the corresponding source-side offset is calculated by regression analysis to reverse compensate and correct the current standardized calculation result; If the offset calibration is still unqualified after correction, the batch of standardized calculation results is marked as pending manual review, and an abnormal record is triggered.

6. The method of claim 5, wherein the method further comprises: The specific process of performing classification storage and validity determination is as follows: After the standardized calculation result is generated, all qualified results are counted and mapped to corresponding storage partitions according to corresponding business labels and data types, the storage partitions including real-time data area, historical data area and abnormal data area; Validity dimension calculation is performed on the data in the storage partition to obtain corresponding validity evaluation result; The validity dimension calculation includes timestamp positive deviation calculation and timestamp negative deviation calculation; The timestamp positive deviation calculation means that the actual timestamp of the data collected by the source computing node is compared with the reference cutoff time of the preset collection period, the difference between the actual timestamp and the reference cutoff time is counted, the proportion of the difference in the length of the preset collection period is calculated, and the degree of delay in uploading the source data is reflected; The timestamp negative deviation calculation means that the actual timestamp of the data collected by the source computing node is compared with the reference start time of the preset collection period, the difference between the actual timestamp and the reference start time is counted, the proportion of the difference in the length of the preset collection period is calculated, and the degree of advance in collecting the source data is reflected; The validity evaluation result represents the results of the timestamp positive deviation calculation and the timestamp negative deviation calculation, and the result of the harmonic average processing; If the validity evaluation result is less than the preset validity evaluation result, it is determined to be valid and the classified storage is completed, otherwise, it is determined to be invalid and a timestamp repair process is triggered to correct the time deviation of the data collected by the source computing node.

7. The method of claim 6, wherein the method further comprises: The timestamp repair process is triggered, specifically: If the timestamp positive deviation is greater than the timestamp negative deviation, the time synchronization interface of the source computing node is called based on the first timestamp deviation to adjust the data upload scheduling frequency to compress the delay; If the timestamp positive deviation is less than the timestamp negative deviation, the preset personnel is prompted to optimize the source computing node collection trigger logic based on the second timestamp deviation to avoid starting data collection in advance; If the timestamp positive deviation is equal to the timestamp negative deviation, it means that the corresponding source computing node has a time reference instability problem, and the preset personnel is prompted to prioritize the calibration of the node local clock; The first timestamp deviation and the second timestamp deviation respectively quantify the data collection degree of the corresponding source computing node; After triggering the timestamp repair process, the validity evaluation result is reacquired, if the reacquired validity evaluation result is not less than the preset validity evaluation result, the classified storage is completed, otherwise, it is marked as pending manual review and archived to the historical abnormal data archive area, and an abnormal alarm is triggered.

8. The distributed intelligent data computing method based on the source-attaching pattern as described in claim 1, characterized in that, The specific generation steps of the comprehensive data index include: Multi-node data correlation analysis is performed on the standardized calculation results after validity determination, and the data matching degree between each source computing node is calculated according to the cosine similarity formula, which is denoted as optimized correlation similarity; If the optimized correlation similarity is greater than the preset optimized correlation similarity, the difference between the optimized correlation similarity and the preset optimized correlation similarity is used as a similarity increment for reflecting the degree to which the data collaboration of the source computing node exceeds the reference level. If the optimized correlation similarity is not greater than the preset optimized correlation similarity, a node anomaly early warning is triggered, it is determined that there is data drift on the source side of the corresponding source computing node, and the data of the node is automatically isolated and rolled back to the last valid state; If the similarity increment is not greater than the preset similarity increment, the comprehensive data index under the current scene is directly generated based on the weighted average of the standardized calculation results of each source computing node; If the similarity increment is greater than the preset similarity increment, an aggregation correction coefficient compensation amount is obtained based on the deviation proportion mapping of the similarity increment and the preset similarity increment, the weight distribution between the data of the source computing nodes is dynamically adjusted, and the aggregation degree between the data of multiple nodes is improved; After compensation, if the similarity improvement amplitude obtained by reacquiring is not less than the preset improvement amplitude, the comprehensive data index is obtained in combination with the compensation result, otherwise, an aggregation anomaly early warning is triggered for source collaborative optimization.

9. The method of claim 8, wherein the method further comprises: The specific process of the source collaborative optimization is as follows: The acquisition log and synchronization record of the corresponding source computing node are obtained to locate the core node of the data collaboration deviation, and the core node represents the source computing node in which the data collaboration deviation occurs; Based on the obtained historical synchronization error peak value of the core node, the time synchronization period and the data upload gap of the core node are re-distributed to reduce the data transmission conflict between nodes, and the historical synchronization error peak value represents the maximum deviation value between the actual synchronization time and the theoretical synchronization time of the core node in the past acquisition period; The node collaboration degree is monitored, if the corresponding node collaboration degree in two continuous acquisition periods is greater than the preset node collaboration degree, the current node collaboration state is maintained, otherwise, a preset personnel is prompted to recalibrate the local clock reference of the core node; If the similarity improvement amplitude obtained by reacquiring is not less than the preset improvement amplitude after monitoring the node collaboration degree, the comprehensive data index is obtained in combination with the node collaboration degree monitoring result, otherwise, a collaboration anomaly early warning is performed; The node collaboration degree represents the product of the proportion of the number of aligned data time stamps of each core node to the total data amount and the proportion of the cross-node effective data interaction amount to the total interaction data amount, and directly reflects the consistency level of the multi-node data; The acquisition method of the optimized correlation similarity is as follows: The standardized calculation results of each source computing node are regarded as a multi-dimensional vector respectively, and each dimension corresponds to the value of a basic index; The cosine similarity between any two node vectors is calculated to obtain the correlation similarity of the node pair, and the node pair represents the combination of two source computing nodes participating in data collaboration analysis; On the basis of the cosine similarity, a standard deviation weight factor is introduced, and the optimized correlation similarities of all node pairs are averaged to obtain the final optimized correlation similarity.

10. A distributed intelligent data computing system based on the stickiness pattern, applying the distributed intelligent data computing method based on the stickiness pattern as claimed in any one of claims 1-9, characterized in that, It comprises: A parallel computing and processing module of each source computing node is used to receive the monitoring data collected by the local sensor in real time and the interaction data transmitted by the external interface through the source computing node deployed at the preset data acquisition point in the business scene under the source mode, and perform parallel computing and preliminary processing to reduce the bandwidth pressure caused by the cross-region data transmission amount; The standardized calculation and effectiveness determination module is configured to distribute preset data index optimization algorithms to each source computing node according to the constructed distributed scheduling framework, to generate standardized calculation results meeting the business requirements of each source computing node, and to simultaneously perform classified storage and effectiveness determination, so as to avoid calculation delay caused by invalid data accumulation. The index generation module of each source computing node is configured to collect the standardized calculation results after effectiveness determination, to upload the unified data format to the established central platform, to compare the correlation of data of different nodes to check calculation deviation, to generate comprehensive data indexes adapted to the business requirements of each source computing node, and to realize efficient processing and accurate output of multi-source data.

Citation Information

Patent Citations

  • Method, device, equipment and computer medium for updating data sets of multiple heterogeneous data sources

    CN117407407B

  • Local knowledge base automatic construction system based on multi-source acquisition and distributed computing

    CN120723742A