Cloud computing based data aggregation optimization method and system
Patent Information
- Application Number
- CN202610850693.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-09-11
AI Technical Summary
本发明的目的在于提供基于云计算的数据聚合优化方法和系统,以解决在云边协同架构下,多源异构数据因协议语义失配与跨源时序偏差相互耦合,导致云端聚合层在高并发场景下吞吐效能劣化,数据聚合一致性差的问题
1.通过在边缘节点对多源异构数据源的原始消息进行协议语义失配特征量化评估,识别高阻抗风险数据源并在本地生成优化协议转换路径,将原本集中压制于云端聚合层的协议转换计算开销前移至边缘节点分散承接,有效消除了高并发场景下多源异构数据因协议语义失配导致的云端聚合层吞吐效能劣化问题,使云端聚合层在业务高峰期仍能保持稳定的消息处理吞吐能力。
Smart Images

Figure CN122734441A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud-edge collaborative data processing technology, specifically to a data aggregation and optimization method and system based on cloud computing. Background Technology
[0002] With the in-depth application of cloud computing technology, business travel systems are migrating to the cloud. Multiple data sources, such as flight booking engines, online business travel platforms, and travel website order systems, continuously generate large-scale business data. The cloud aggregation layer needs to aggregate and process multi-source data in real time to support the stable operation of high-frequency business scenarios such as flight inquiries, order management, and itinerary changes.
[0003] However, online business travel platform data sources and airline order system data sources have significantly different application layer protocol systems, with substantial differences in field naming conventions, data type definitions, and business semantic expression methods. When the cloud aggregation layer directly receives multi-source data, it needs to independently perform protocol parsing and semantic conversion for each data source. During peak business periods such as concentrated flight changes and holiday travel, the concurrent influx of multi-source data causes a sharp accumulation of protocol conversion calculation tasks. A large amount of computing resources are occupied by redundant protocol conversion processes, resulting in a significant decrease in the actual processing throughput of the cloud aggregation layer, becoming a hidden bottleneck restricting the overall system performance. Furthermore, due to differences in network link conditions, their own data production rhythms, and system load states, the arrival time of messages at the cloud aggregation layer from heterogeneous data sources inherently has a significant time sequence deviation. Moreover, the time sequence deviation is not a fixed value but fluctuates continuously with the dynamic changes in business load. When the cloud aggregation layer aggregates messages from different data sources at the same time, it cannot guarantee that each message corresponds to the same logical time point at the business semantic level. This leads to the aggregation result incorrectly merging data from different time segments, resulting in data quality defects with inconsistent cross-source time sequences. Summary of the Invention
[0004] (1) Technical problems to be solved The purpose of this invention is to provide a data aggregation optimization method and system based on cloud computing, in order to solve the problem that in the cloud-edge collaborative architecture, multi-source heterogeneous data are coupled with each other due to protocol semantic mismatch and cross-source timing deviation, resulting in the degradation of throughput performance of the cloud aggregation layer in high-concurrency scenarios and poor data aggregation consistency.
[0005] (2) Technical solution To achieve the above objectives, in one aspect, the present invention provides a data aggregation optimization method based on cloud computing, the method comprising: S1. In the cloud-edge collaborative architecture, each edge node establishes a connection with the corresponding heterogeneous data source and continuously receives the original messages pushed by each heterogeneous data source.
[0006] S2. Each edge node extracts semantic fields from the original messages of each heterogeneous data source, and performs field mapping and semantic alignment processing on the semantic fields according to a unified private message protocol to obtain a standardized message.
[0007] S3. Statistically calculate the message arrival interval of each heterogeneous data source within the historical period, calculate the timing deviation compensation value of each heterogeneous data source relative to the reference time axis, and pre-correct the timestamp of the standardized message according to the timing deviation compensation value to obtain a timing pre-aligned standardized message stream.
[0008] S4. Upload the time-series pre-aligned standardized message stream to the cloud aggregation layer; the cloud aggregation layer divides the standardized message streams from different edge nodes into corresponding logical time windows according to the pre-corrected timestamps of each standardized message, and performs aggregation calculations on the multi-source standardized messages within the same logical time window to obtain the aggregation result.
[0009] Furthermore, the method for each edge node to extract semantic fields from the original messages of each heterogeneous data source, and to obtain standardized messages by mapping and aligning the semantic fields according to a unified private message protocol, includes: Each edge node pre-constructs a field mapping table corresponding to each heterogeneous data source; when mapping the fields of the original messages of each heterogeneous data source, the proportion of mismatched messages and the proportion of mismatched fields of each heterogeneous data source are used as feature components to construct the protocol semantic mismatch feature vector corresponding to each heterogeneous data source.
[0010] The protocol conversion pressure of each heterogeneous data source is quantitatively evaluated based on the protocol semantic mismatch feature vector to obtain the protocol conversion pressure evaluation value of each heterogeneous data source; heterogeneous data sources whose protocol conversion pressure evaluation value exceeds the preset pressure alarm threshold are marked as high impedance risk data sources; the high impedance risk data sources are locally generated and cached on the edge node to obtain the optimized protocol conversion path corresponding to each high impedance risk data source.
[0011] The raw messages subsequently received from high-impedance risk data sources are processed by optimizing the protocol conversion path for field mapping and semantic alignment. For non-high-impedance risk data sources, field mapping and semantic alignment are processed using the original field mapping relationship table. The results of both types of processing are encapsulated according to the message structure specified by the unified private message protocol to generate standardized messages that conform to the format of the unified private message protocol.
[0012] Furthermore, the method for quantitatively evaluating the protocol conversion pressure of each heterogeneous data source based on the protocol semantic mismatch feature vector to obtain the protocol conversion pressure assessment value of each heterogeneous data source includes: The protocol semantic mismatch intensity value of each heterogeneous data source in the current statistical period is calculated based on the protocol semantic mismatch feature vector.
[0013] Obtain the message arrival rate of each heterogeneous data source within the current statistical period, and calculate the baseline load value of protocol conversion for each heterogeneous data source under the current message arrival rate based on the message arrival rate and the baseline computational overhead estimated for performing a complete protocol conversion on a single original message.
[0014] The protocol conversion pressure assessment value of each heterogeneous data source in the current statistical period is calculated based on the protocol semantic mismatch strength value and the protocol conversion baseline load value.
[0015] Furthermore, the method for calculating the protocol conversion pressure assessment value of each heterogeneous data source within the current statistical period based on the protocol semantic mismatch strength value and the protocol conversion baseline load value includes: The initial protocol conversion pressure value is calculated based on the protocol semantic mismatch strength value and the corresponding protocol conversion baseline load value of each heterogeneous data source in the current statistical period; the initial protocol conversion pressure value of each heterogeneous data source in several consecutive historical statistical periods is extracted to form a historical sequence of initial protocol conversion pressure values.
[0016] The initial protocol conversion pressure value historical sequence is subjected to fast Fourier transform to obtain each frequency component and its corresponding amplitude. The period length corresponding to the frequency component whose amplitude exceeds the preset significance amplitude benchmark is extracted and marked as the main fluctuation period.
[0017] Based on the main fluctuation cycle, the initial protocol conversion pressure value in the current statistical period is periodically aligned and corrected to obtain the protocol conversion pressure correction value; based on the protocol conversion pressure correction value, the protocol conversion pressure correction value in the current statistical period and the previous statistical period is weighted and smoothed using the weighted moving average method to obtain the protocol conversion pressure assessment value of each heterogeneous data source in the current statistical period.
[0018] Furthermore, the method for generating and caching simplified protocol conversion paths locally at the edge node to obtain optimized protocol conversion paths corresponding to each high-impedance risk data source includes: Iterate through all field mapping operations in the field mapping relationship table corresponding to each high-impedance risk data source, and merge multiple field mapping operations with the same field type and consistent semantic verification rules into a single batch mapping operation to obtain the merged batch mapping operation set.
[0019] Based on the business necessity level of each standard field in the unified private message protocol, the semantic verification steps in the batch mapping operation set are graded and labeled. Semantic verification steps with a business necessity level lower than the preset necessity level benchmark are marked as skippable verification items, while semantic verification steps with a business necessity level not lower than the preset necessity level benchmark are marked as mandatory retention verification items. All verification execution steps corresponding to skippable verification items are removed from the batch mapping operation set. The mandatory retention verification items and batch mapping operations are reordered according to the actual appearance order of the fields in the original message to obtain a simplified operation execution sequence. The simplified operation execution sequence is then compiled into an optimized protocol conversion path.
[0020] Furthermore, the method by which the cloud aggregation layer divides the standardized message streams from different edge nodes into corresponding logical time windows based on the pre-corrected timestamps of each standardized message includes: The cloud aggregation layer statistically analyzes the arrival interval sequence of standardized message pre-correction timestamps for each heterogeneous data source within a certain number of historical receiving cycles. The average value of the arrival interval sequence is used to obtain the message arrival interval benchmark value for each heterogeneous data source. The maximum value among the message arrival interval benchmark values is used as the candidate benchmark for dividing the logical time window step size.
[0021] Obtain the cross-source timestamp dispersion of all standardized messages within the current reception period, and calculate the logical time window division step size of the current reception period based on the cross-source timestamp dispersion and the candidate benchmark.
[0022] Starting with the minimum value of the pre-corrected timestamps of all standardized messages within the current reception period, the overall time span of the multi-source standardized message stream is continuously divided according to the step size of the logical time window, generating several non-overlapping logical time windows.
[0023] The pre-correction timestamps of each standardized message are compared with the start and end boundaries of each logical time window. Standardized messages that fall within the same logical time window range are assigned to this logical time window; standardized messages whose pre-correction timestamps fall at the boundaries of adjacent logical time windows are assigned to the subsequent logical time window.
[0024] Furthermore, the method for obtaining the cross-source timestamp dispersion of all standardized messages within the current reception period, and calculating the logical time window division step size of the current reception period based on the cross-source timestamp dispersion and the candidate benchmark, includes: Based on the identifiers of each heterogeneous data source, the median of all standardized message pre-correction timestamps of each heterogeneous data source in the current reception period is extracted. The difference between the maximum and minimum median of the pre-correction timestamps of each heterogeneous data source is taken as the cross-source timestamp dispersion of the current reception period.
[0025] The dispersion ratio of the current reception period is obtained by dividing the cross-source timestamp dispersion by the candidate benchmark; the dispersion ratio of each reception period in a series of historical consecutive reception periods is extracted and arranged in chronological order to form a historical dispersion ratio sequence; the average value of the historical dispersion ratio sequence is obtained by taking the first difference; the look-ahead estimate is calculated based on the dispersion ratio of the current reception period and the average value of the period; and the dispersion ratio of the current reception period and the look-ahead estimate are weighted and summed to obtain the comprehensive dispersion ratio.
[0026] The initial logical time window partitioning step size is estimated based on the candidate benchmark and the comprehensive dispersion ratio, and the logical time window partitioning step size for the current reception period is obtained by applying constraints on both sides.
[0027] Furthermore, the bilateral constraints include a lower bound constraint value and an upper bound constraint value; Calculate the timestamp range between the maximum and minimum values of the pre-correction timestamps of all multi-source standardized messages within the same logical time window in each of several consecutive reception periods. Arrange the timestamp ranges in chronological order according to the reception periods to obtain a historical sequence. Use the mean of the historical timestamp range sequence as the lower bound constraint value for the logical time window division step size.
[0028] Extract the logical time window division step size actually used in each reception cycle within a certain number of historical reception cycles, arrange them according to the chronological order of the reception cycles to form a historical division step size sequence, take the average of the historical division step size sequence and multiply it by a preset expansion coefficient to obtain the upper bound constraint value of the logical time window division step size; the preset expansion coefficient is calibrated offline based on the complete coverage of cross-source standardized messages by the logical time window within the historical reception cycle.
[0029] On the other hand, based on the same inventive concept, this invention also provides a cloud computing-based data aggregation and optimization system, the system comprising: The edge node data receiving module is used in the cloud-edge collaborative architecture to establish connections between each edge node and its corresponding heterogeneous data source, and continuously receive raw messages pushed by each heterogeneous data source.
[0030] The message standardization module is used by each edge node to extract semantic fields from the original messages of each heterogeneous data source, and then perform field mapping and semantic alignment processing on the semantic fields according to a unified private message protocol to obtain standardized messages.
[0031] The timing pre-alignment module is used to count the message arrival intervals of each heterogeneous data source within the historical period, calculate the timing deviation compensation value of each heterogeneous data source relative to the reference time axis, and pre-correct the timestamps of standardized messages based on the timing deviation compensation value to obtain a timing pre-aligned standardized message stream.
[0032] The cloud aggregation computing module is used to upload time-series pre-aligned standardized message streams to the cloud aggregation layer. The cloud aggregation layer divides the standardized message streams from different edge nodes into corresponding logical time windows based on the pre-corrected timestamps of each standardized message. Within the same logical time window, it performs aggregation computing on the multi-source standardized messages to obtain the aggregation result.
[0033] (3) Beneficial effects Compared with the prior art, the beneficial effects of the present invention are: 1. By quantitatively evaluating the protocol semantic mismatch characteristics of raw messages from multi-source heterogeneous data sources at edge nodes, high-impedance risk data sources are identified and optimized protocol conversion paths are generated locally. The protocol conversion computation overhead, which was originally concentrated in the cloud aggregation layer, is moved forward to the edge nodes for distributed processing. This effectively eliminates the problem of cloud aggregation layer throughput degradation caused by protocol semantic mismatch of multi-source heterogeneous data in high-concurrency scenarios, enabling the cloud aggregation layer to maintain stable message processing throughput during peak business periods.
[0034] 2. By pre-calculating the time-series deviation compensation values of each heterogeneous data source at the edge node and pre-correcting the timestamps of standardized messages, time-series alignment is completed before the messages are uploaded to the cloud aggregation layer. This eliminates the risk of cross-source time-series inconsistency caused by the coupling of time-series deviations when multi-source data arrives at the cloud aggregation layer, ensuring that standardized messages from different data sources can be converged in the correct logical time dimension.
[0035] 3. With the combined effect of throughput performance and timing consistency, the cloud aggregation layer performs adaptive logical time window division on multi-source standardized message streams based on pre-corrected timestamps, and completes aggregation calculation within the same logical time window. This solves the system performance degradation dilemma caused by the coupling of protocol semantic mismatch and cross-source timing deviation, and achieves comprehensive optimization of the performance and consistency of multi-source heterogeneous data aggregation results in high-concurrency scenarios of general business travel. Attached Figure Description
[0036] Figure 1 This is a flowchart of the cloud computing-based data aggregation optimization method of the present invention.
[0037] Figure 2 This is a schematic diagram of the module composition of the cloud computing-based data aggregation and optimization system of the present invention. Detailed Implementation
[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0039] Example 1: As Figure 1 As shown, this embodiment provides a data aggregation optimization method based on cloud computing, the method including: S1. In the cloud-edge collaborative architecture, each edge node establishes a connection with the corresponding heterogeneous data source and continuously receives the original messages pushed by each heterogeneous data source.
[0040] For example, in a real-world business travel scenario, the flight booking engine, online business travel platform, and travel website order system belong to different system vendors and each uses independently designed application layer protocols, resulting in significant differences in field names, data types, and the way business semantics are expressed. For instance, for the same flight departure time, the flight booking engine stores Unix timestamp format data in the "dep_time" field, while the travel website order system stores an ISO 8601 format string in the "flight_departure" field; the two are inconsistent in both field names and data types. Edge nodes, located on the data source side, are responsible for accessing and preprocessing raw messages locally, thereby distributing the protocol conversion burden from the cloud aggregation layer to the local processing of each edge node.
[0041] S2. Each edge node extracts semantic fields from the original messages of each heterogeneous data source, and performs field mapping and semantic alignment processing on the semantic fields according to a unified private message protocol to obtain a standardized message.
[0042] The method for each edge node to extract semantic fields from the original messages of each heterogeneous data source and then perform field mapping and semantic alignment of the semantic fields according to a unified private message protocol to obtain a standardized message includes: Each edge node pre-constructs a field mapping table corresponding to each heterogeneous data source; when mapping the fields of the original messages of each heterogeneous data source, the proportion of mismatched messages and the proportion of mismatched fields of each heterogeneous data source are used as feature components to construct the protocol semantic mismatch feature vector corresponding to each heterogeneous data source.
[0043] The protocol conversion pressure of each heterogeneous data source is quantitatively evaluated based on the protocol semantic mismatch feature vector to obtain the protocol conversion pressure evaluation value of each heterogeneous data source.
[0044] The method for quantitatively evaluating the protocol conversion pressure of each heterogeneous data source based on the protocol semantic mismatch feature vector to obtain the protocol conversion pressure evaluation value of each heterogeneous data source includes: The protocol semantic mismatch intensity value of each heterogeneous data source in the current statistical period is calculated based on the protocol semantic mismatch feature vector.
[0045] Obtain the message arrival rate of each heterogeneous data source within the current statistical period, and calculate the baseline load value of protocol conversion for each heterogeneous data source under the current message arrival rate based on the message arrival rate and the baseline computational overhead estimated for performing a complete protocol conversion on a single original message.
[0046] The protocol conversion pressure assessment value of each heterogeneous data source in the current statistical period is calculated based on the protocol semantic mismatch strength value and the protocol conversion baseline load value.
[0047] The method for calculating the protocol conversion pressure assessment value of each heterogeneous data source within the current statistical period based on the protocol semantic mismatch strength value and the protocol conversion baseline load value includes: The initial protocol conversion pressure value is calculated based on the protocol semantic mismatch strength value and the corresponding protocol conversion baseline load value of each heterogeneous data source in the current statistical period; the initial protocol conversion pressure value of each heterogeneous data source in several consecutive historical statistical periods is extracted to form a historical sequence of initial protocol conversion pressure values.
[0048] The initial protocol conversion pressure value historical sequence is subjected to fast Fourier transform to obtain each frequency component and its corresponding amplitude. The period length corresponding to the frequency component whose amplitude exceeds the preset significance amplitude benchmark is extracted and marked as the main fluctuation period.
[0049] Based on the main fluctuation cycle, the initial protocol conversion pressure value in the current statistical period is periodically aligned and corrected to obtain the protocol conversion pressure correction value; based on the protocol conversion pressure correction value, the protocol conversion pressure correction value in the current statistical period and the previous statistical period is weighted and smoothed using the weighted moving average method to obtain the protocol conversion pressure assessment value of each heterogeneous data source in the current statistical period.
[0050] Heterogeneous data sources whose protocol conversion pressure assessment values exceed a preset pressure alarm threshold are marked as high-impedance risk data sources; the high-impedance risk data sources are locally generated and cached on the edge nodes to simplify the protocol conversion path and obtain the optimized protocol conversion path corresponding to each high-impedance risk data source.
[0051] The method for generating and caching simplified protocol conversion paths locally at edge nodes to obtain optimized protocol conversion paths for each high-impedance risk data source includes: Iterate through all field mapping operations in the field mapping relationship table corresponding to each high-impedance risk data source, and merge multiple field mapping operations with the same field type and consistent semantic verification rules into a single batch mapping operation to obtain the merged batch mapping operation set.
[0052] Based on the business necessity level of each standard field in the unified private message protocol, the semantic verification steps in the batch mapping operation set are graded and labeled. Semantic verification steps with a business necessity level lower than the preset necessity level benchmark are marked as skippable verification items, while semantic verification steps with a business necessity level not lower than the preset necessity level benchmark are marked as mandatory retention verification items. All verification execution steps corresponding to skippable verification items are removed from the batch mapping operation set. The mandatory retention verification items and batch mapping operations are reordered according to the actual appearance order of the fields in the original message to obtain a simplified operation execution sequence. The simplified operation execution sequence is then compiled into an optimized protocol conversion path.
[0053] The raw messages subsequently received from high-impedance risk data sources are processed by optimizing the protocol conversion path for field mapping and semantic alignment. For non-high-impedance risk data sources, field mapping and semantic alignment are processed using the original field mapping relationship table. The results of both types of processing are encapsulated according to the message structure specified by the unified private message protocol to generate standardized messages that conform to the format of the unified private message protocol.
[0054] For example, each edge node pre-constructs a corresponding field mapping table for each heterogeneous data source. This table records the name mapping relationship, data type conversion rules, and semantic verification rules between each field in the original protocol of the heterogeneous data source and its corresponding standard field in the unified private message protocol. During the field mapping process for the original messages from each heterogeneous data source, the edge node continuously calculates the proportion of messages with field mismatches in the current statistical period to the total number of messages received in that period (i.e., the mismatch message proportion), and the proportion of mismatched fields to the total number of fields in the original message (i.e., the mismatch field proportion). These two feature components are combined to form the protocol semantic mismatch feature vector of the heterogeneous data source in the current statistical period. Specifically, let the proportion of mismatched messages from a certain heterogeneous data source in the current statistical period be... The proportion of mismatched fields is Then the protocol semantic mismatch feature vector of this heterogeneous data source Represented as: ;in, This is the ratio of the number of messages with field mismatches in the current statistical period to the total number of messages received in that statistical period, with a value range of [value missing]. This reflects the frequency of field mismatch issues in the data source messages; This is the ratio of the number of fields with mismatches in a single original message to the total number of fields in the original message, with a value range of [value range missing]. This reflects the coverage density of protocol-mismatched fields in a single original message. Both components are dimensionless scale values, collectively describing the severity of the mismatch between the data source protocol and the unified private messaging protocol.
[0055] Based on this, the edge nodes quantitatively assess the protocol conversion pressure on various heterogeneous data sources. The protocol semantic mismatch strength value for each heterogeneous data source within the current statistical period is calculated based on the protocol semantic mismatch feature vector. The calculation method is as follows: ;in, and Preset weighting coefficients are used to reflect the relative contributions of the proportion of mismatched messages and the proportion of mismatched fields to the overall protocol conversion burden. The sum of the two is 1, and they can be calibrated offline according to specific business scenarios. Because... and All values are dimensionless proportional values; the weighted summation yields the semantic mismatch strength value of the protocol. Also a dimensionless value, its range is... A higher protocol semantic mismatch strength value indicates a greater semantic distance between the protocol of the original message from the data source and the unified private message protocol, and more computational resources are required when performing protocol conversion.
[0056] Edge nodes obtain the message arrival rate of each heterogeneous data source within the current statistical period. (Unit: messages / second), combined with an estimated baseline computational cost required to perform a complete protocol conversion on a single raw message. (Unit: milliseconds / message, pre-calibrated by offline benchmark testing), calculate the protocol conversion baseline load value for each heterogeneous data source at the current message arrival rate. ;in, The unit is stripes per second. The unit is milliseconds per message. The unit is milliseconds per second, which represents the protocol conversion computation time required per second under the current message arrival rate, reflecting the theoretical intensity of the data source's current message flow's occupation of computing resources.
[0057] Multiplying the protocol semantic mismatch strength value by the protocol conversion baseline load value yields the initial protocol conversion pressure value for each heterogeneous data source within the current statistical period. This represents the overall intensity of the data source's occupancy of protocol conversion computing resources under the combined effects of the current mismatch level and the current message arrival rate.
[0058] Considering the significant periodicity of message traffic in business travel transactions—such as peak travel times on weekday mornings and concentrated order inflows before and after holidays—judging the data source's pressure status solely based on the initial protocol conversion pressure value within a single statistical period is somewhat unpredictable. Therefore, a historical periodic analysis mechanism is introduced. Edge nodes extract the initial protocol conversion pressure values of each heterogeneous data source over several consecutive historical statistical periods, arranging them in chronological order to form a historical sequence of initial protocol conversion pressure values. From each frequency component, frequency components with amplitudes exceeding a preset significant amplitude benchmark are selected, and their corresponding period lengths are designated as the main fluctuation period of the initial protocol conversion pressure value of this heterogeneous data source. The physical meaning of the main fluctuation cycle is: the protocol conversion pressure of the data source is... The cycle exhibits a significant fluctuation pattern. For example, if the main fluctuation cycle of a data source corresponds to 24 statistical cycles, it indicates that the protocol conversion pressure of that data source exhibits a clear periodic high-low alternation within a time frame of 24 statistical cycles as a complete cycle. The current statistical cycle number is then correlated with the main fluctuation cycle. The modulus is used to determine the relative position of the current statistical period within the main fluctuation period. Using the average of the initial protocol conversion pressure values within the same historical statistical period as a reference, the current initial protocol conversion pressure value is additively corrected to obtain the protocol conversion pressure correction value. The unit of the protocol conversion pressure correction value is the same as the initial protocol conversion pressure value, which is milliseconds per second. The purpose of the correction operation is to eliminate the systematic bias introduced by the periodic fluctuations in business traffic, so that the current pressure assessment more accurately reflects the true load status of the data source during similar business time periods.
[0059] Based on the protocol conversion pressure correction value, a weighted moving average method is used to smooth the protocol conversion pressure correction values for the current statistical period and several previous statistical periods, thus obtaining the protocol conversion pressure assessment value for each heterogeneous data source in the current statistical period. ;in, For the current statistical period, going back to the first Protocol conversion pressure correction value for each statistical period, in milliseconds per second; For the corresponding weighting coefficients, and The weighting coefficients vary with The weighting increases monotonically, meaning that more recent statistical periods are assigned higher weights. Protocol transition stress assessment value. The unit is milliseconds per second, and when compared with the preset pressure alarm threshold, the preset pressure alarm threshold is also calibrated in milliseconds per second. The introduction of weighted moving average can effectively suppress the temporary anomalies in pressure values caused by occasional traffic surges or network jitter within a single statistical period, keeping the protocol conversion pressure assessment value stable over time and avoiding misleading conclusions about high impedance risk identification.
[0060] High-impedance-risk data sources refer to data sources where, under current business load conditions, performing a complete protocol conversion line by line according to the original field mapping table would incur computational overhead that could potentially reduce the throughput of the cloud aggregation layer. For data sources marked as high-impedance-risk, edge nodes generate and cache a simplified, optimized protocol conversion path locally to replace the complete protocol conversion execution process in the original field mapping table.
[0061] The generation of optimized protocol conversion paths involves traversing all field mapping operations in the field mapping relationship table corresponding to high-impedance risk data sources. Multiple field mapping operations with the same field type and consistent semantic verification rules are merged into a single batch mapping operation, resulting in a merged set of batch mapping operations. For example, if the field mapping relationship table contains three fields that are all string types and only require non-empty verification, the corresponding three independent mapping operations are merged into a single batch mapping operation, completing the type conversion and non-empty verification of the three fields at once, thereby reducing the additional overhead caused by the number of operation schedulings. Based on the business necessity level of each standard field in the unified private message protocol, each semantic verification step in the batch mapping operation set is graded and labeled. The business necessity level is determined by the actual role of the field in subsequent aggregation calculations and business decisions: core business fields such as flight query results, order status, and itinerary changes are assigned a high business necessity level and marked as mandatory retained verification items; steps with low business necessity level, such as format specification verification and extended attribute integrity verification, which are not closely related to core business logic, are assigned a low business necessity level and marked as skippable verification items. All skippable validation steps are removed from the batch mapping operation set. The forced-retain validation items and batch mapping operations are then reordered according to the actual order of field appearance in the original message, forming a simplified operation execution sequence. This simplified sequence is then compiled into an optimized protocol conversion path and cached locally on the edge nodes. Through this process, the optimized protocol conversion path significantly reduces the number of computational steps required for each original message to complete the protocol conversion, while preserving the integrity of critical business fields.
[0062] For subsequent incoming raw messages from high-impedance risk data sources, the edge node directly invokes the cached optimized protocol conversion path to perform field mapping and semantic alignment. For non-high-impedance risk data sources, the original field mapping table is used for complete field mapping and semantic alignment. Both processing results are encapsulated according to the message structure defined by the unified private message protocol, generating standardized messages conforming to the unified private message protocol format. The names, data types, and semantic meanings of each field in the standardized message are completely consistent with the unified private message protocol. After the above processing, there are no longer any format differences between raw messages from different data sources at the protocol level.
[0063] S3. Statistically calculate the message arrival intervals of each heterogeneous data source within the historical period, calculate the timing deviation compensation value of each heterogeneous data source relative to the reference time axis, and pre-correct the timestamps of standardized messages based on the timing deviation compensation value to obtain a time-pre-aligned standardized message stream. Because different data sources differ in network link conditions, data production rhythms, and their own system load states, the deviation between the arrival time of messages from each data source at the edge node and the actual occurrence time of the business event varies, and this deviation fluctuates continuously with the dynamic changes in business load. The pre-corrected timestamps are aligned to a unified logical time base at the business semantic level, eliminating the cross-source time axis misalignment problem introduced by the inconsistent timing deviations of different data sources.
[0064] S4. Upload the time-pre-aligned standardized message stream to the cloud aggregation layer; the cloud aggregation layer divides the standardized message streams from different edge nodes into corresponding logical time windows based on the pre-corrected timestamps of each standardized message. The method by which the cloud aggregation layer divides the standardized message streams from different edge nodes into corresponding logical time windows based on the pre-corrected timestamps of each standardized message includes: The cloud aggregation layer statistically analyzes the arrival interval sequence of standardized message pre-correction timestamps for each heterogeneous data source within several consecutive historical reception periods. The average value of the arrival interval sequence is taken to obtain the message arrival interval benchmark value of each heterogeneous data source. The maximum value among the message arrival interval benchmark values is used as the candidate benchmark for dividing the logical time window step size. Obtain the cross-source timestamp dispersion of all standardized messages within the current reception period, and calculate the logical time window division step size of the current reception period based on the cross-source timestamp dispersion and the candidate benchmark. The method for obtaining the cross-source timestamp dispersion of all standardized messages within the current reception period, and calculating the logical time window division step size of the current reception period based on the cross-source timestamp dispersion and the candidate benchmark, includes: Based on the identifiers of each heterogeneous data source, the median of all standardized message pre-correction timestamps of each heterogeneous data source in the current reception period is extracted. The difference between the maximum and minimum values of the median of the pre-correction timestamps of each heterogeneous data source is taken as the cross-source timestamp dispersion of the current reception period. The dispersion factor of the current reception period is obtained by dividing the cross-source timestamp dispersion by the candidate benchmark; the dispersion factor of each reception period in a series of historical consecutive reception periods is extracted and arranged in chronological order to form a historical dispersion factor sequence; the average value of the historical dispersion factor sequence is obtained by taking the first difference; the look-ahead estimate is calculated based on the dispersion factor of the current reception period and the average value of the period; the dispersion factor of the current reception period and the look-ahead estimate are weighted and summed to obtain the comprehensive dispersion factor. The initial logical time window partitioning step size is estimated based on the candidate benchmark and the comprehensive dispersion ratio, and the logical time window partitioning step size for the current reception period is obtained by applying constraints on both sides.
[0065] The bilateral constraints include lower bound constraint values and upper bound constraint values; Calculate the timestamp range between the maximum and minimum values of the pre-correction timestamps of all multi-source standardized messages within the same logical time window in each of several consecutive reception periods in history. Arrange the timestamp ranges in chronological order according to the reception periods to obtain a historical sequence. Use the mean of the historical timestamp range sequence as the lower bound constraint value of the logical time window division step size. Extract the logical time window division step size actually used in each reception cycle within a certain number of historical reception cycles, arrange them according to the chronological order of the reception cycles to form a historical division step size sequence, take the average of the historical division step size sequence and multiply it by a preset expansion coefficient to obtain the upper bound constraint value of the logical time window division step size; the preset expansion coefficient is calibrated offline based on the complete coverage of cross-source standardized messages by the logical time window within the historical reception cycle.
[0066] Starting from the minimum value of the pre-corrected timestamps of all standardized messages within the current reception period, the overall time span of the multi-source standardized message stream is continuously divided according to the step size of the logical time window, generating several non-overlapping logical time windows. The pre-correction timestamps of each standardized message are compared with the start and end boundaries of each logical time window. Standardized messages that fall within the same logical time window range are assigned to this logical time window; standardized messages whose pre-correction timestamps fall at the boundaries of adjacent logical time windows are assigned to the subsequent logical time window.
[0067] For example, the maximum value among all data source message arrival interval benchmarks is used as a candidate benchmark for the logical time window division step size. The unit is milliseconds. The candidate benchmark reflects the slowest message generation rhythm among all data sources. Using this benchmark as a benchmark ensures that the logical time window span is not less than the message interval of the sparsest data source, avoiding the inability of some data sources to be aggregated within the same window due to an excessively small window step.
[0068] To ensure that the logical time window segmentation step size can adapt to the actual dispersion of cross-source message timestamps within the current reception period, the cloud aggregation layer further calculates the cross-source timestamp dispersion of all standardized messages within the current reception period. All standardized messages received within the current reception period are grouped according to the identifiers of each heterogeneous data source. The median of the pre-corrected timestamps of all standardized messages from each heterogeneous data source within the current reception period is extracted. The difference between the maximum and minimum median of the pre-corrected timestamps from all data sources is taken as the cross-source timestamp dispersion for the current reception period. The unit is milliseconds. Using the median, rather than the mean or extreme values, as the representative timestamp value for each data source can effectively reduce the interference of individual, occasional delayed messages on the cross-source timestamp dispersion estimation results.
[0069] Cross-source timestamp dispersion Divided by candidate benchmark The dispersion factor of the current receiving period is obtained. Dispersion ratio This characterizes the amplification ratio of the cross-source timestamp dispersion within the current reception period relative to the message arrival interval baseline. When the dispersion ratio... A larger value indicates that messages from different data sources are scattered across the timeline, requiring a corresponding increase in the logical time window step size to ensure that messages from different data sources but corresponding to the same business logic moment can be grouped into the same logical time window.
[0070] To avoid significant jumps in the logical time window step size due to short-term fluctuations in the dispersion ratio of the current reception period, the cloud aggregation layer introduces a look-ahead prediction mechanism. This mechanism extracts the dispersion ratios corresponding to each reception period within a consecutive historical period and arranges them chronologically to form a historical dispersion ratio sequence. The periodic average rate of change of the dispersion multiple is obtained by taking the mean after first-order difference of the historical series of dispersion multiples. : ;in, The average rate of change of the period is the difference in dispersion ratio between adjacent receiving periods. All differences are dimensionless values. Both are dimensionless values, representing the average variation of the dispersion factor between adjacent reception periods.
[0071] Based on the dispersion factor of the current reception period Compared with the periodic average rate of change Calculate forward forecasts Forward-looking estimates This is a dimensionless value, reflecting the expected level of the dispersion factor in the next receiving cycle under the assumption of continued historical trends. The comprehensive dispersion factor is obtained by weighting and summing the measured dispersion factor of the current receiving cycle with the forward-looking estimate according to preset weights. ;in, The weighting coefficient for the current measured dispersion ratio, with a value range of [value range missing]. The message coverage rate can be calibrated offline based on the actual performance of the message coverage rate within the logical time window of the historical reception period. (Comprehensive dispersion ratio) As a dimensionless value, it takes into account the current measured dispersion while incorporating forward-looking judgments on recent trends, making the dynamic adjustment of the logical time window step size more stable.
[0072] Candidate benchmarks With the overall dispersion ratio Multiplying these values yields an estimated step size for the initial logical time window. To prevent the initial logical time window partitioning step size estimate from being too small, resulting in incomplete window coverage, or too large, resulting in excessively coarse aggregation granularity, constraints are applied to the initial logical time window partitioning step size estimate on both sides. Lower bound constraint value. The physical meaning is: the logical time window step size should be at least no less than the average range of message pre-correction timestamps within the same logical time window historically; otherwise, messages that should belong to the same logical time window will be split into different logical time windows. Upper bound constraint value. The calculation method involves taking the average of the historical step-size sequence and multiplying it by a preset expansion coefficient. Obtain the upper bound constraint value and preset the expansion coefficient. The logical time window is determined based on offline calibration results of cross-source standardized message complete coverage within historical reception cycles, ensuring that the logical time window can completely cover all cross-source messages arriving at the same business logic moment in the vast majority of reception cycles. The logical time window step size for the current reception cycle is determined. The calculation method for the constraints on both sides is as follows: .
[0073] The aggregation result is obtained by aggregating standardized messages from multiple sources within the same logical time window. The timestamps of each standardized message have been pre-corrected for timing deviations at the edge nodes before entering the cloud aggregation layer. The standardized messages from multiple sources gathered within the same logical time window correspond to the same logical time interval at the business semantic level, thereby ensuring the timing consistency of the aggregation result and effectively avoiding data quality defects caused by erroneous merging of data from different time segments.
[0074] Example 2: Based on the same inventive concept, such as Figure 2 As shown, this embodiment also provides a cloud computing-based data aggregation and optimization system, the system comprising: The edge node data receiving module is used in the cloud-edge collaborative architecture to establish connections between each edge node and its corresponding heterogeneous data source, and continuously receive raw messages pushed by each heterogeneous data source.
[0075] The message standardization module is used by each edge node to extract semantic fields from the original messages of each heterogeneous data source, and then perform field mapping and semantic alignment processing on the semantic fields according to a unified private message protocol to obtain standardized messages.
[0076] The timing pre-alignment module is used to count the message arrival intervals of each heterogeneous data source within the historical period, calculate the timing deviation compensation value of each heterogeneous data source relative to the reference time axis, and pre-correct the timestamps of standardized messages based on the timing deviation compensation value to obtain a timing pre-aligned standardized message stream.
[0077] The cloud aggregation computing module is used to upload time-series pre-aligned standardized message streams to the cloud aggregation layer. The cloud aggregation layer divides the standardized message streams from different edge nodes into corresponding logical time windows based on the pre-corrected timestamps of each standardized message. Within the same logical time window, it performs aggregation computing on the multi-source standardized messages to obtain the aggregation result.
[0078] It should be noted that the specific methods by which each module performs operations in the system described in the above embodiments have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0079] Finally, it should be noted that although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A data aggregation optimization method based on cloud computing, characterized in that, The method includes: In the cloud-edge collaborative architecture, each edge node establishes a connection with its corresponding heterogeneous data source and continuously receives raw messages pushed by each heterogeneous data source. Each edge node extracts semantic fields from the original messages of each heterogeneous data source, and then performs field mapping and semantic alignment on the semantic fields according to a unified private message protocol to obtain a standardized message. The message arrival intervals of each heterogeneous data source within the historical period are statistically analyzed, the timing deviation compensation value of each heterogeneous data source relative to the reference time axis is calculated, and the timestamps of standardized messages are pre-corrected based on the timing deviation compensation value to obtain a time-pre-aligned standardized message stream. The time-prealigned standardized message streams are uploaded to the cloud aggregation layer. The cloud aggregation layer divides the standardized message streams from different edge nodes into corresponding logical time windows based on the pre-corrected timestamps of each standardized message. Within the same logical time window, the multi-source standardized messages are aggregated to obtain the aggregation result.
2. The data aggregation and optimization method based on cloud computing according to claim 1, characterized in that, The method for each edge node to extract semantic fields from the original messages of each heterogeneous data source and then perform field mapping and semantic alignment of the semantic fields according to a unified private message protocol to obtain a standardized message includes: Each edge node pre-builds a field mapping table corresponding to each heterogeneous data source; when mapping the original messages of each heterogeneous data source to fields, the proportion of mismatched messages and the proportion of mismatched fields of each heterogeneous data source are used as feature components to construct the protocol semantic mismatch feature vector corresponding to each heterogeneous data source. The protocol conversion pressure of each heterogeneous data source is quantitatively evaluated based on the protocol semantic mismatch feature vector to obtain the protocol conversion pressure evaluation value of each heterogeneous data source; heterogeneous data sources whose protocol conversion pressure evaluation value exceeds the preset pressure alarm threshold are marked as high impedance risk data sources; the high impedance risk data sources are locally generated and cached on the edge node to obtain the optimized protocol conversion path corresponding to each high impedance risk data source. The raw messages subsequently received from high-impedance risk data sources are processed by optimizing the protocol conversion path for field mapping and semantic alignment. For non-high-impedance risk data sources, field mapping and semantic alignment are processed using the original field mapping relationship table. The results of both types of processing are encapsulated according to the message structure specified by the unified private message protocol to generate standardized messages that conform to the format of the unified private message protocol.
3. The data aggregation and optimization method based on cloud computing according to claim 2, characterized in that, The method for quantitatively evaluating the protocol conversion pressure of each heterogeneous data source based on the protocol semantic mismatch feature vector to obtain the protocol conversion pressure evaluation value of each heterogeneous data source includes: The protocol semantic mismatch intensity value of each heterogeneous data source in the current statistical period is calculated based on the protocol semantic mismatch feature vector. Get the message arrival rate of each heterogeneous data source in the current statistical period, and calculate the baseline load value of protocol conversion for each heterogeneous data source under the current message arrival rate based on the message arrival rate and the baseline computational overhead estimated for performing a complete protocol conversion on a single original message. The protocol conversion pressure assessment value of each heterogeneous data source in the current statistical period is calculated based on the protocol semantic mismatch strength value and the protocol conversion baseline load value.
4. The data aggregation and optimization method based on cloud computing according to claim 3, characterized in that, The method for calculating the protocol conversion pressure assessment value of each heterogeneous data source within the current statistical period based on the protocol semantic mismatch strength value and the protocol conversion baseline load value includes: The initial protocol conversion pressure value is calculated based on the protocol semantic mismatch strength value and the corresponding protocol conversion baseline load value of each heterogeneous data source in the current statistical period; the initial protocol conversion pressure value of each heterogeneous data source in several consecutive historical statistical periods is extracted to form a historical sequence of initial protocol conversion pressure values. The initial protocol conversion pressure value historical sequence is subjected to fast Fourier transform to obtain each frequency component and its corresponding amplitude. The period length corresponding to the frequency component whose amplitude exceeds the preset significant amplitude benchmark is extracted and marked as the main fluctuation period. Based on the main fluctuation cycle, the initial protocol conversion pressure value in the current statistical period is periodically aligned and corrected to obtain the protocol conversion pressure correction value; based on the protocol conversion pressure correction value, the protocol conversion pressure correction value in the current statistical period and the previous statistical period is weighted and smoothed using the weighted moving average method to obtain the protocol conversion pressure assessment value of each heterogeneous data source in the current statistical period.
5. The data aggregation and optimization method based on cloud computing according to claim 2, characterized in that, The method for generating and caching simplified protocol conversion paths locally at edge nodes to obtain optimized protocol conversion paths for each high-impedance risk data source includes: Iterate through all field mapping operations in the field mapping relationship table corresponding to each high impedance risk data source, and merge multiple field mapping operations with the same field type and consistent semantic verification rules into a batch mapping operation to obtain the merged batch mapping operation set. Based on the business necessity level of each standard field in the unified private message protocol, the semantic verification steps in the batch mapping operation set are graded and labeled. Semantic verification steps with a business necessity level lower than the preset necessity level benchmark are marked as skippable verification items, while semantic verification steps with a business necessity level not lower than the preset necessity level benchmark are marked as mandatory retention verification items. All verification execution steps corresponding to skippable verification items are removed from the batch mapping operation set. The mandatory retention verification items and batch mapping operations are reordered according to the actual appearance order of the fields in the original message to obtain a simplified operation execution sequence. The simplified operation execution sequence is then compiled into an optimized protocol conversion path.
6. The data aggregation and optimization method based on cloud computing according to claim 1, characterized in that, The method by which the cloud aggregation layer divides the standardized message streams from different edge nodes into corresponding logical time windows based on the pre-corrected timestamps of each standardized message includes: The cloud aggregation layer statistically analyzes the arrival interval sequence of standardized message pre-correction timestamps for each heterogeneous data source within several consecutive historical reception periods. The average value of the arrival interval sequence is taken to obtain the message arrival interval benchmark value of each heterogeneous data source. The maximum value among the message arrival interval benchmark values is used as the candidate benchmark for dividing the logical time window step size. Obtain the cross-source timestamp dispersion of all standardized messages within the current reception period, and calculate the logical time window division step size of the current reception period based on the cross-source timestamp dispersion and the candidate benchmark. Starting from the minimum value of the pre-corrected timestamps of all standardized messages within the current reception period, the overall time span of the multi-source standardized message stream is continuously divided according to the step size of the logical time window, generating several non-overlapping logical time windows. The pre-correction timestamps of each standardized message are compared with the start and end boundaries of each logical time window. Standardized messages that fall within the same logical time window range are assigned to this logical time window; standardized messages whose pre-correction timestamps fall at the boundaries of adjacent logical time windows are assigned to the subsequent logical time window.
7. The data aggregation and optimization method based on cloud computing according to claim 6, characterized in that, The method for obtaining the cross-source timestamp dispersion of all standardized messages within the current reception period, and calculating the logical time window division step size of the current reception period based on the cross-source timestamp dispersion and the candidate benchmark, includes: Based on the identifiers of each heterogeneous data source, the median of all standardized message pre-correction timestamps of each heterogeneous data source in the current reception period is extracted. The difference between the maximum and minimum values of the median of the pre-correction timestamps of each heterogeneous data source is taken as the cross-source timestamp dispersion of the current reception period. The dispersion factor of the current reception period is obtained by dividing the cross-source timestamp dispersion by the candidate benchmark; the dispersion factor of each reception period in a series of historical consecutive reception periods is extracted and arranged in chronological order to form a historical dispersion factor sequence; the average value of the historical dispersion factor sequence is obtained by taking the first difference; the look-ahead estimate is calculated based on the dispersion factor of the current reception period and the average value of the period; the dispersion factor of the current reception period and the look-ahead estimate are weighted and summed to obtain the comprehensive dispersion factor. The initial logical time window partitioning step size is estimated based on the candidate benchmark and the comprehensive dispersion ratio, and the logical time window partitioning step size for the current reception period is obtained by applying constraints on both sides.
8. The data aggregation and optimization method based on cloud computing according to claim 7, characterized in that, The bilateral constraints include lower bound constraint values and upper bound constraint values; Calculate the timestamp range between the maximum and minimum values of the pre-correction timestamps of all multi-source standardized messages within the same logical time window in each of several consecutive reception periods in history. Arrange the timestamp ranges in chronological order according to the reception periods to obtain a historical sequence. Use the mean of the historical timestamp range sequence as the lower bound constraint value of the logical time window division step size. Extract the logical time window division step size actually used in each reception cycle within a certain number of historical reception cycles, arrange them according to the chronological order of the reception cycles to form a historical division step size sequence, take the average of the historical division step size sequence and multiply it by a preset expansion coefficient to obtain the upper bound constraint value of the logical time window division step size; the preset expansion coefficient is calibrated offline based on the complete coverage of cross-source standardized messages by the logical time window within the historical reception cycle.
9. A cloud computing-based data aggregation and optimization system, characterized in that, The system includes: The edge node data receiving module is used in the cloud-edge collaborative architecture to establish connections between each edge node and its corresponding heterogeneous data source, and continuously receive raw messages pushed by each heterogeneous data source. The message standardization module is used by each edge node to extract semantic fields from the original messages of each heterogeneous data source, and then perform field mapping and semantic alignment of the semantic fields according to a unified private message protocol to obtain a standardized message. The timing pre-alignment module is used to count the message arrival interval of each heterogeneous data source within the historical period, calculate the timing deviation compensation value of each heterogeneous data source relative to the reference time axis, and pre-correct the timestamp of the standardized message according to the timing deviation compensation value to obtain the timing pre-aligned standardized message stream. The cloud aggregation computing module is used to upload time-series pre-aligned standardized message streams to the cloud aggregation layer. The cloud aggregation layer divides the standardized message streams from different edge nodes into corresponding logical time windows based on the pre-corrected timestamps of each standardized message. Within the same logical time window, it performs aggregation computing on the multi-source standardized messages to obtain the aggregation result.