Data processing and synchronizing method and system

By dynamically adjusting the data collection frequency, compression method, and storage strategy in the intelligent transportation system, and combining multi-dimensional data comparison and distributed transaction mechanisms, the problems of data synchronization and consistency maintenance are solved, achieving efficient data processing and storage, and improving the reliability and maintainability of the system.

CN121807908APending Publication Date: 2026-04-07ZUNYI BRANCH OF CHINA MOBILE GRP GUIZHOU COMPANY +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-25
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In intelligent transportation systems, traditional data processing and storage methods face efficiency and scalability issues, especially in data synchronization and consistency maintenance. Existing systems cannot process large amounts of changing data in real time, leading to a decline in data timeliness and accuracy.

Method used

By determining the importance and urgency scores of data, the system dynamically adjusts the data collection frequency, compression method, and storage strategy. It also adopts multi-dimensional data comparison and distributed transaction mechanisms to optimize the real-time propagation and consistency maintenance of data. Combined with containerized deployment and automated management, the system's reliability and maintainability are improved.

Benefits of technology

It effectively reduces network and storage pressure, improves data processing efficiency, ensures data timeliness and accuracy, enhances system reliability and maintainability, and enables efficient operation during peak resource demand periods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121807908A_ABST
    Figure CN121807908A_ABST
Patent Text Reader

Abstract

According to the data processing and synchronizing method and system, the importance score of the data is determined, the importance score is used for determining the frequency of data collection and the compression mode of data transmission, and the collection amount and the transmission amount of the data can be reduced. And meanwhile, the data is stored in a partitioned manner, so that the storage efficiency can be further improved, and convenience can be provided for subsequent data use. When the data is changed, the information in the plurality of databases is synchronously changed through the data change information, so that the consistency of the data is favorably kept. By changing the priority score, the sequence of the multiple pieces of data change information is determined, and when resources are insufficient, resources can be preferentially allocated to the data change information with the high priority. And meanwhile, different components are deployed in independent containers, so that automatic and batch management can be realized, the management efficiency of a plurality of components is improved, and the stability of the system is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of Internet of Things (IoT) technology, and in particular to a data processing and synchronization method and system. Background Technology

[0002] In current intelligent transportation systems, data is typically collected through terminal sensors, including speedometers, temperature sensors, and GPS locators. This data is usually transmitted via network to a central system for further processing and analysis. However, with the dramatic increase in data volume, traditional data processing and storage methods face efficiency and scalability issues. Particularly in terms of data synchronization and consistency maintenance, existing systems may be unable to handle large amounts of changing data in real time, leading to a decline in data timeliness and accuracy. Summary of the Invention

[0003] This disclosure provides a data processing and synchronization method and system. By determining the acquisition frequency, compression method, and storage method, the amount of data transmission is reduced, and multiple databases are synchronized through data change information to solve the problems of data timeliness and synchronization.

[0004] In one aspect, this embodiment provides a data processing and synchronization method, the method comprising: determining an importance score for the data, the importance score being used at least to determine the data collection frequency and compression method; storing the data based on the data access frequency; determining a data processing decision based on the data urgency score, the urgency score being used at least to reflect the degree of data abnormality; and synchronizing information in the first database and the second database based on the data change information when data change information in the first database is captured.

[0005] In embodiments of this disclosure, the method further includes: performing full compression on the data when the importance score of the data is greater than a first importance score threshold, wherein the importance score of the data is determined at least based on the type weight of the data and the evaluation function of the data; or performing simplified compression on the data when the importance score of the data is less than or equal to the first importance score threshold.

[0006] In embodiments of this disclosure, the method further includes: determining the data acquisition frequency, wherein the acquisition frequency is determined at least based on the basic acquisition frequency, the importance score, and the second importance score threshold.

[0007] In embodiments of this disclosure, storing data based on data access frequency includes: determining the data access frequency, the access frequency including at least historical access frequency or predicted access frequency; caching the data when the data access frequency is greater than an access frequency threshold, the access frequency threshold being determined at least based on the maximum cache capacity and the average access frequency; or partitioning the data storage based on the data access frequency and the data type when the data access frequency is less than or equal to the access frequency threshold.

[0008] In embodiments of this disclosure, determining a data processing decision based on the urgency score of the data includes: comparing the urgency score of the data with an urgency score threshold, wherein the data includes multidimensional attributes, and the urgency score of the data is determined at least based on the attribute weights of the multidimensional attributes and the urgency evaluation function of the multidimensional attributes; and processing the data based on the processing decision if the urgency score is greater than the urgency score threshold.

[0009] In embodiments of this disclosure, the method further includes: performing multidimensional data comparison on the data, wherein the comparison includes at least comparison based on multidimensional data hash, and the multidimensional data hash is determined at least based on the multidimensional attributes and attribute weights of the data.

[0010] In the embodiments of this disclosure, when data change information of the first database is captured, the information of the first database and the second database is synchronized based on the data change information, including: obtaining data change information of the first database, the data change information being used to record at least the write information, change information, and operation information of the data in the first database; and changing the data information in multiple second databases based on the data change log, so as to keep the data in the first database synchronized with the data in the second database.

[0011] In embodiments of this disclosure, data information in multiple second databases is changed based on data change logs, including: determining change priority scores for multiple data change information; and determining a synchronous change order based on multiple change priority scores to execute multiple data change information, wherein the change priority scores are determined at least based on the change urgency score and change frequency of the data change information.

[0012] In embodiments of this disclosure, the method further includes allocating resources based on a resource allocation model, wherein the resource allocation model is determined at least based on memory utilization and CPU utilization.

[0013] On the other hand, embodiments of this disclosure provide a data processing and synchronization system, which includes multiple components, including at least a scoring determination component, a data storage component, and a calculation decision component. The scoring determination component is used to determine the importance score of the data, and the importance score is used at least to determine the data collection frequency and compression method. The data storage component is used to store the data based on the data access frequency. The calculation decision component is used to determine the data processing decision based on the data urgency score, so as to process the data. The processing decision at least includes multi-dimensional data comparison, and the urgency score is used at least to reflect the degree of data anomaly.

[0014] In embodiments of this disclosure, multiple components are deployed on multiple independent containers, which operate through automated management and performance monitoring.

[0015] On the other hand, embodiments of this disclosure provide a network device, including: a memory for storing computer-readable instructions; and a processor for executing the computer-readable instructions, causing the network device to perform the data processing and synchronization method.

[0016] In another aspect, embodiments of this disclosure provide a non-transitory computer-readable storage medium for storing computer-readable instructions that, when executed by a processor, cause the processor to perform the aforementioned data processing and synchronization methods.

[0017] In another aspect, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the aforementioned data processing and synchronization method. Attached Figure Description

[0018] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0019] Figure 1 The schematic diagram illustrates an environmental application according to an embodiment of the present disclosure.

[0020] Figure 2 A flowchart illustrating a data processing and synchronization method according to an embodiment of the present disclosure is shown.

[0021] Figure 3 The flowchart illustrating the data storage process based on data access frequency according to an embodiment of the present disclosure is shown in the illustration.

[0022] Figure 4A flowchart illustrating a data processing decision according to an embodiment of the present disclosure is shown.

[0023] Figure 5 The flowchart illustrating the synchronization of information between the first and second databases based on data change information according to an embodiment of the present disclosure is shown in the illustration.

[0024] Figure 6 A block diagram illustrating a data processing and synchronization system according to an embodiment of the present disclosure is shown.

[0025] Figure 7 A block diagram of a network device according to an embodiment of the present disclosure is shown schematically.

[0026] Figure 8 A block diagram illustrating a non-transitory computer-readable storage medium according to an embodiment of the present disclosure is shown.

[0027] Figure 9 A block diagram illustrating a computer program product according to an embodiment of the present disclosure is shown schematically. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.

[0029] In current intelligent transportation systems, data is typically collected through terminal sensors, including speedometers, temperature sensors, and GPS locators. This data is usually transmitted to a central system via a network for further processing and analysis. However, with the dramatic increase in data volume, traditional data processing and storage methods face efficiency and scalability issues. Particularly in terms of data synchronization and consistency maintenance, existing systems may be unable to handle large amounts of changed data in real time, leading to decreased data timeliness and accuracy. Furthermore, with increasing system complexity, using traditional deployment and management methods can result in configuration errors and deployment delays, impacting the overall reliability and maintainability of the system and potentially causing resource bottlenecks during peak resource demands.

[0030] Based on this, this disclosure effectively reduces network and storage pressure by dynamically adjusting data collection frequency, compression methods, and storage methods. Urgency scoring effectively improves data processing efficiency. Simultaneously, when data changes, the Change Data Capture (CDC) mechanism and distributed transactions optimize real-time data propagation and consistency maintenance. Furthermore, containerized deployment and automated management of the system improve deployment efficiency and configuration consistency, thereby enhancing system reliability and maintainability. Finally, real-time monitoring and resource optimization ensure efficient system operation during peak resource demand periods, optimizing overall resource utilization.

[0031] The following section will combine... Figures 1-9 The data processing and synchronization methods of the embodiments of this disclosure will be described in detail.

[0032] Figure 1 The schematic diagram illustrates an environmental application according to an embodiment of the present disclosure.

[0033] like Figure 1 As shown, computing device 101 and multiple terminal devices 102 can perform real-time data transmission.

[0034] Terminal device 102 can acquire network data on the terminal side in real time. For example, in an intelligent transportation system, terminal device 102 may include sensors such as temperature sensors, pressure sensors, and GPS locators, or network devices such as smartphones, tablets, and intelligent transportation vehicles. Through terminal device 102, the movement status of terminals such as cars and pedestrians can be captured in real time.

[0035] The computing device 101 can receive real-time data collected by the terminal device 102, and store, analyze, and process the real-time data. The computing device 101 can be a single server or composed of multiple servers, such as a distributed server.

[0036] The following section will describe the data processing and synchronization process of computing device 101.

[0037] Figure 2 A flowchart illustrating a data processing and synchronization method according to an embodiment of the present disclosure is shown.

[0038] like Figure 2 As shown, a data processing and synchronization method according to an embodiment of this disclosure includes S201, S202, S203, and S204: S201. Determine the importance score of the data. The importance score is used at least to determine the data collection frequency and compression method.

[0039] In intelligent transportation system data management, data from various types of sensors can be collected in real time, including temperature sensors, humidity sensors, speedometers, accelerometers, and GPS locators, covering environmental and motion parameters. With the development of technology, the volume of traffic data is constantly increasing. Therefore, before acquiring sensor data, data compression and the determination of different data collection frequencies can be performed to reduce the volume of transmitted data. For example, the importance score of the data can be used to determine the data collection frequency and compression method, such as full compression or simplified compression.

[0040] In embodiments of this disclosure, the data transmitted in a single transmission includes data collected by multiple sensors, such as speed, temperature, and pressure. Furthermore, the importance score of the data can be determined by different sensor data, sensor data weights, and sensor data evaluation functions.

[0041] In embodiments of this disclosure, such as Figure 2 The method also includes: performing full compression on the data when the importance score of the data is greater than the first importance score threshold, wherein the importance score of the data is determined at least based on the data type weights and the data evaluation function; and performing simplified compression on the data when the importance score of the data is less than or equal to the first importance score threshold.

[0042] The importance score of the data can be determined using the following formula: +...+

[0043] Where I represents the importance score of the data, w i f is the weight of the i-th sensor data type. i (x i ) is the evaluation function for the i-th sensor data type.

[0044] In some embodiments, the first importance scoring threshold I min This can be a global lower bound parameter to ensure that low-importance data is not completely discarded during transmission. The first importance scoring threshold I... min The determination can be made in the following ways: based on historical statistics (such as normal fluctuation ranges from long-term monitoring); based on business requirements (such as security data requiring a minimum transmission frequency); or based on dynamic optimization (automatic adjustment in conjunction with network load and caching strategies).

[0045] If the importance score of the data is greater than the first importance threshold, it indicates that the data being transmitted is relatively important. In this case, all fields of the data can be serialized and compressed to preserve complete data information, such as using full compression (Protobuf compression). This method consumes some CPU computing power for compression / decompression, but the amount of data transmitted is significantly reduced compared to the original data. If the importance score of the data is less than or equal to the first importance threshold, it indicates that the data being transmitted is relatively unimportant. Only the core key fields of the data (such as extreme values ​​and status indicators) can be extracted, discarding unnecessary fields. Full compression is not necessary; for example, a simplified compression method (transmitting key data points) can be used. This method consumes almost no compression computing power, and the amount of data transmitted is only 10%-30% of that of full compression, with extremely low network and storage resource consumption.

[0046] By basing data compression on its importance, m determines the compression method. For critical data, priority is given to ensuring "data integrity," sacrificing a small amount of computing resources to ensure that high-value data (core data influencing decision-making and analysis) is not lost. For relatively unimportant data, priority is given to ensuring "resource utilization," discarding low-value redundant information and completing basic data transmission with minimal resource consumption to avoid waste. This achieves the goal of neither losing critical data in pursuit of resource conservation nor wasting resources by retaining all data, ensuring that data of varying importance is transmitted in the optimal way.

[0047] In embodiments of this disclosure, such as Figure 2 The method also includes: determining the data collection frequency, which is determined at least based on the base collection frequency, importance score, and second importance score threshold.

[0048] The data acquisition frequency can refer to the frequency at which the computing device acquires sensor data. The data acquisition frequency can be determined through the following methods:

[0049] Where R is the adjusted data acquisition frequency, and R0 is the basic data acquisition frequency. To adjust the coefficient, I thresh I represents the second importance score threshold, and I represents the importance score of the data.

[0050] Compared with the first importance score threshold I min Similarly, the second importance rating threshold I thresh Alternatively, it can be determined in the following ways: based on historical statistics (such as normal fluctuation ranges from long-term monitoring); based on business requirements (such as security data requiring a minimum transmission frequency); or based on dynamic optimization (automatic adjustment in conjunction with network load and caching strategies).

[0051] For example, for various sensors in an intelligent transportation system, the following evaluation functions and weights for sensor data types can be set: Sensor 1 is set as a speedometer (monitoring changes in vehicle speed), with a weight of w1=0.5; Sensor 1 is set as a temperature sensor (monitoring ambient temperature), with a weight of w2=0.2; and Sensor 3 is set as a GPS locator (position change), with a weight of w3=0.3.

[0052] Furthermore, the evaluation function for the speedometer is as follows:

[0053] Where, x 1,prev This represents the vehicle speed at the previous moment.

[0054] The evaluation function for the temperature sensor is shown below:

[0055] Where, x z This is a temperature threshold, such as 30°C.

[0056] The evaluation function for a GPS locator is shown below:

[0057] Where, x 3,prev This represents the position at the previous moment.

[0058] Therefore, based on these parameters, the importance score of the data can be calculated using the following formula:

[0059] For example, for various sensors in an intelligent transportation system, the basic acquisition frequency R0 can be set to once per second, and the adjustment coefficient can be... The second importance score threshold is 0.5. thresh The value is 0.7. Therefore, the adjusted data acquisition frequency R is: R = 1 × (1 + 0.5(I - 0.7)) For example, let I be the first importance score threshold. min If the importance score is 0.3, then for the data importance score, if I is greater than 0.3, then full compression (using protobuf compression) can be performed, and if I is less than or equal to 0.3, then simplified data (only sending key data points, such as extreme temperatures or location changes) can be performed.

[0060] Understandably, the weights and evaluation functions for different types of sensors can be set and adjusted in real time according to different situations. For example, for a moving car, different sensors can collect data such as its speed, temperature at different parts, and position in real time. The importance score I of this data can then be determined by the weights and evaluation functions of the sensors measuring speed, temperature, and position. Similarly, for a stationary car, different sensors can collect data such as temperature and pressure at different parts in real time. The importance score I of this data can then be determined by the weights and evaluation functions of the sensors measuring temperature and pressure.

[0061] In this way, when sensor data shows significant changes (such as rapid speed changes, abnormal temperature, or significant location changes), the system increases the acquisition frequency and ensures that the data is fully compressed and transmitted. For routine or unimportant data, the system reduces the acquisition frequency and may only send simplified data, thereby optimizing the use of network and storage resources.

[0062] In the embodiments of this disclosure, after S201, data can be collected and transmitted using a determined collection frequency and compression method. Before transmission, the data can also be encrypted, such as using Transport Layer Security (TLS). During transmission, Message Queuing Telemetry Transport (MQTT) can be used. This is a lightweight message transmission protocol suitable for low-bandwidth and unstable network environments, and is particularly suitable for ensuring data security during transmission for IoT devices.

[0063] S202. Store data based on data access frequency.

[0064] For data transmitted to computing devices, the computing devices can store it. By determining the data access frequency, a corresponding storage strategy can be determined to meet the data access requirements. For example, frequently accessed data (such as the vehicle's current location and speed) can be cached in Redis for fast access. Data in Cassandra can be partitioned according to its type and access frequency. For example, timestamp-based partitioning allows for convenient storage and retrieval of data within a specific time period.

[0065] Figure 3 The flowchart illustrating the data storage process based on data access frequency according to an embodiment of the present disclosure is shown in the illustration.

[0066] like Figure 3As shown, S202 above includes S301, S302a and S302b: S301. Determine the access frequency of the data. The access frequency shall include at least the historical access frequency or the predicted access frequency.

[0067] Access frequency can be historical or predicted. It can also be determined by combining both, for example, by adding the product of historical and predicted access frequencies to their respective weights. Access frequency can also be determined by counting the number of accesses within a given period. This period could be a day, an hour, or a similar time window, and the number of accesses is the number of times the data was requested during that period. Predicted access frequency can also be based on historical access frequency. Access frequency f d It can be determined using the following formula:

[0068] In some embodiments, the predicted access frequency can also be determined based on a prediction model, such as an autoregressive integrated moving average (ARIMA) model and historical data {y}. t-1 y t-2 , ..., y t-p The predicted access frequency can be determined using the following formula:

[0069] Among them, f(d), This represents the predicted future access frequency, specifically the data value at time t+1. c is a constant term used to balance the average level of the data. If the data has a fixed trend or a non-zero average, this parameter will adjust the predicted value. p is the order of the autoregressive (AR) component, representing how many past moments {y} are used. t-1 y t-2 , ..., y t-p To predict future values; The coefficients of the autoregressive component represent the i-th lag term. For predicted values The magnitude of the impact; y t-i This represents the actual value at time t-1, i.e., the data value at the i-th time in the past; q is the order of the moving average (MA) component, indicating how many past prediction error terms are used. t 1, t 2,..., t q To correct the prediction; θ is the coefficient of the moving average component, representing the j-th prediction error term. For predicted values The correction range; The prediction error at time tj is used to represent the actual value. Compared with the predicted value The difference between them is defined as:

[0070] S302a. When the access frequency of data is greater than the access frequency threshold, cache the data. The access frequency threshold is determined based at least on the maximum cache capacity and the average access frequency.

[0071] The access frequency threshold can be dynamically adjusted based on the system's cache capacity and access distribution. The access frequency threshold can be determined using the following formula:

[0072] During periods of high access volume, the access frequency threshold can be increased to reduce the amount of cached data; during periods of low access volume, the access frequency threshold can be decreased to fully utilize the cache. This approach is particularly suitable for systems with limited cache capacity and fluctuating access load.

[0073] S302b: When the data access frequency is less than or equal to the access frequency threshold, the data is stored in partitions based on the data access frequency and the data type.

[0074] In embodiments of this disclosure, the data storage strategy can be determined using the following formula:

[0075] Where C(d) determines whether data d should be cached, and f(d) is a function that determines the access frequency of the data or predicts the future access frequency. This is the access frequency threshold.

[0076] In other words, if the access frequency of data or the predicted future access frequency is greater than the access frequency threshold, the data can be cached; if the access frequency of data or the predicted future access frequency is less than or equal to the access frequency threshold, the data can be partitioned for storage.

[0077] For example, the implementation of a caching decision function determines which data should be cached in Redis for fast access. For instance, the current location and speed of a vehicle are frequently accessed data, typically used in real-time traffic monitoring and navigation services. Access frequency evaluation f(d): If a piece of data (such as vehicle location) is accessed more than 10 times per minute, it is considered frequently accessed data. Access frequency threshold. Set to 5 times / minute. Decision: If f(d) > If so, then C(d) = 1, meaning the data is cached.

[0078] When partitioning data, you can use the time information of the data for partitioning, for example, you can partition using the following formula:

[0079] Where p(d) represents the result of data partitioning; n is the number of partitions; d.timestamp is the timestamp attribute of the data; and hash is the hash function.

[0080] First, perform a hash operation on the timestamp d.timestamp of the data d to obtain a hash value; then divide this hash value by n and take the remainder. The final result p(d) is the target partition number of the data in the partitioned storage.

[0081] For example, data partitioning functions are implemented to optimize the data storage structure in a Cassandra database, improving query efficiency through data partitioning. Partitioning facilitates the management of large amounts of data generated daily, especially for time-sensitive queries. Partitioning strategy: Partitioning is based on the timestamp attribute of the data, using the formula: P(d) = hash(d.timestamp) mod n, where n is the number of partitions determined by the data volume and query load, for example, 24 (one partition per hour).

[0082] Caching data accelerates access to high-frequency data and reduces database load. Caching frequently accessed hot data (such as real-time vehicle location, traffic congestion status, and high-priority change data) in Redis leverages the significantly faster memory read / write speeds compared to disk, reducing query response time in traffic control centers from milliseconds to microseconds. This also reduces direct access to the Cassandra master database, preventing it from becoming a performance bottleneck due to high-concurrency queries and ensuring the stability of data writes and change synchronization. Furthermore, partitioning data improves the efficiency of time-series data queries. Traffic data (such as vehicle location and speed) is partitioned by timestamp. When querying traffic data for a specific time period or region, only the target partition is scanned, avoiding full data scanning and eliminating the need to traverse all historical data, significantly reducing I / O overhead and improving response speed. This also facilitates subsequent processing, such as multi-dimensional data comparison, enabling rapid retrieval of similar data from corresponding regions, eliminating the tedious process of first acquiring and then filtering data.

[0083] S203. Based on the urgency score of the data, determine the data processing decision to process the data. The urgency score is used to reflect at least the degree of abnormality of the data.

[0084] By determining the urgency score of data, it is possible to decide whether to process the data. This allows for the urgent handling of abnormal data while ignoring normal data, thereby reducing the processing volume and improving processing efficiency.

[0085] Figure 4 A flowchart illustrating a data processing decision according to an embodiment of the present disclosure is shown.

[0086] like Figure 4 As shown, S203 above includes S401 and S402: S401. Compare the urgency score of the data with the urgency score threshold. The data includes multidimensional attributes. The urgency score of the data is determined based at least on the attribute weights of the multidimensional attributes and the urgency evaluation function of the multidimensional attributes.

[0087] The urgency score of data can be calculated based on a comprehensive score of attributes and weights across multiple dimensions. Let x represent a real-time data point, and x = {x1, x2, ..., x...} m Let} be a data vector containing m attributes. Then, the urgency score of data x can be defined as:

[0088] Where g(x) is the urgency score of data x; m is the number of attributes of the data; w i f is the attribute weight of the i-th attribute, representing the importance of this attribute to urgency, satisfying that the weights of all attributes are equal to 1; i (xi ) is the urgency evaluation function for the i-th attribute, used to map the original attribute value to the urgency score (such as normalization, threshold judgment).

[0089] For different attributes of intelligent transportation systems, f i (x i The urgency evaluation function can be different, for example: vehicle speed x i Its urgency evaluation function f i (x i ) can be:

[0090] That is, when the speed x1 exceeds the set speed limit value v limit In cases where the speeding rate is exceeded, the urgency score is calculated based on the speeding value; otherwise, the score is 0.

[0091] Acceleration x2, its urgency evaluation function f i (x i ) can be:

[0092] The larger the absolute value of acceleration, the more unstable the vehicle's state. Therefore, the absolute value of acceleration is directly used as the urgency score.

[0093] Temperature x3, its urgency evaluation function f i (x i ) can be:

[0094] That is, if the temperature exceeds the set threshold T max If the urgency score is 1, the urgency score is returned; otherwise, it is 0.

[0095] The rate of change of location x4, and its urgency evaluation function f i (x i ) can be:

[0096] That is, the rate of change of position is based on the maximum rate of change d. max Normalization indicates the relative level of urgency.

[0097] The urgency score threshold θ is the threshold that triggers emergency handling, determining when special processing is applied to real-time data. The urgency score threshold θ can be determined as follows: Based on historical distribution settings, by statistically analyzing the distribution of urgency scores in historical data, an urgency score threshold is set. θ Set it to a percentile (e.g., 90%), which can be determined using the following formula:

[0098] This means setting it to the 90th percentile of the historical data urgency score to ensure that only the 10% of data that is most urgent is given special treatment.

[0099] Alternatively, based on business needs, fixed values ​​can be preset according to different business scenarios, such as θ=50, meaning that emergency handling is triggered when g(x)>50. Or, the urgency scoring threshold can be dynamically adjusted in real time based on relevant system metrics. θ That is, it can be determined by the following formula: ×Load factor in, The basic urgency score threshold is used; the load factor can be used to reflect relative indicators of the current system load (such as queue length and CPU utilization); k is an adjustment coefficient used to control the magnitude of change in the urgency score threshold.

[0100] S402. When the urgency score is greater than the urgency score threshold, process the data based on the processing decision.

[0101] The logic for processing real-time data can be implemented using the following formula:

[0102] Where S(x) is the processing decision for real-time data x, g(x) is the urgency score of data x, and θ is the urgency score threshold θ.

[0103] For example, real-time data processing logic is implemented using Kafka to process real-time data streams and respond to emergencies. For instance, real-time monitoring of abnormal vehicle behavior, such as sudden deceleration which may indicate a traffic accident. Emergency assessment g(x): determined based on a sudden decrease in vehicle speed. The emergency score threshold θ is set to a speed reduction exceeding 50%. Decision: If g(x) ≥ θ, then emergency response procedures are executed.

[0104] In some embodiments, S(x) can be a comprehensive processing decision, which may include different operations such as multidimensional hash comparison, cache writing, and alarm push. Processing data with different urgency levels through different operations is beneficial to improving data processing efficiency.

[0105] According to embodiments of this disclosure, determining whether to process data through urgency scoring facilitates rapid data prioritization, allowing resources to be allocated to high-urgency data processing first, avoiding excessive resource consumption by low-urgency data, thereby improving overall data processing efficiency. For example, in intelligent transportation scenarios, high-urgency data such as accident alarms can be prioritized for processing to ensure rapid emergency response, while routine road condition monitoring data can be processed sequentially.

[0106] In embodiments of this disclosure, such as Figure 4 The method shown also includes: performing multidimensional data comparison on the data, the comparison including at least comparison based on multidimensional data hash, the multidimensional data hash being determined at least based on the multidimensional attributes and attribute weights of the data.

[0107] Multidimensional data hash comparison is an efficient method that maps multidimensional traffic data (such as vehicle speed, location, and temperature) to hash values ​​using a hash function, and then quickly compares and filters target data using these hash values. Multidimensional data hash comparison does not require comparing the original data dimension by dimension; it can quickly filter out data that meets certain conditions (such as data related to abnormal vehicle movement or congested road sections) simply by matching hash values. Multidimensional data hash comparison can be determined using the following formula:

[0108] Where H(x) is the hash value of the multidimensional data x. x i w is the i-th dimension attribute of the data. i It is the attribute weight of the corresponding dimension.

[0109] For example, multidimensional hashing is implemented using Spark for multidimensional data comparison of large datasets. Traffic data can be compared to identify patterns or anomalies, such as monitoring traffic congestion or accidents using speed and location data. Multidimensional hashing: For example, considering a vehicle's speed x1 and location x2, with weights of 0.6 and 0.4 respectively, then H(x) = 0.6 × hash(x1) + 0.4 × hash(x2).

[0110] Multidimensional data hash comparison can significantly improve processing efficiency: hash value comparison is a "constant-level" operation, which avoids the high I / O overhead of comparing each field of multidimensional raw data. It is especially suitable for real-time processing of massive traffic data (such as thousands of vehicle data per second), which greatly improves the efficiency and accuracy of large-scale data processing.

[0111] In the embodiments disclosed herein, Spark's distributed computing capabilities can be leveraged to perform efficient multidimensional comparisons of collected data, particularly suitable for large-scale datasets. Spark uses a distributed cluster to split massive traffic data (such as city-wide vehicle trajectories and multi-time-period sensor data) across multiple nodes for parallel computation, avoiding single-node computing power bottlenecks. Even when processing tens or hundreds of millions of data points, it can quickly complete multidimensional comparisons (such as identifying congested road sections and abnormally moving vehicles). Real-time streaming processing of data is achieved through Kafka, suitable for monitoring vehicle status and responding to emergencies. Kafka can receive real-time vehicle data uploaded by sensors with high throughput (such as thousands of speed and location data points per second) and forward them to the processing stage with low latency, ensuring that the traffic control center has real-time access to vehicle dynamics without data backlog or delay.

[0112] By combining Spark distributed multidimensional comparison with Kafka real-time stream processing, we can achieve dual guarantees of "efficient processing of large-scale data" and "real-time response to core business", perfectly adapting to the massive data and emergency response needs of intelligent transportation systems.

[0113] S204. When data change information of the first database is captured, the information of the first database and the second database is synchronized based on the data change information.

[0114] According to embodiments of this disclosure, by determining a data importance score, which is used to determine the frequency of data collection and the compression method for transmitted data, it is beneficial to reduce the amount of data collected and transmitted, thereby improving data processing efficiency. Simultaneously, partitioning the data for storage can further improve storage efficiency and facilitate subsequent data use. When data changes, information in multiple databases is synchronously updated through data change information, which helps maintain data consistency.

[0115] Figure 5 The flowchart illustrating the synchronization of information between the first and second databases based on data change information according to an embodiment of the present disclosure is shown in the illustration.

[0116] like Figure 5 As shown, S204 above includes S501 and S502: S501. Obtain data change information from the first database. The data change information is used to record at least the write information, change information, and operation information of data in the first database.

[0117] During data synchronization, a change data capture (CDC) mechanism can be used to achieve real-time data propagation, and distributed transactions can be used to ensure data consistency. For example, the "change log" of the primary database can be continuously monitored to obtain information on data changes in the primary database. The primary database can be a database storing core business data (such as a Cassandra database used to store sensor data). Data change information can refer to the "write operation log" of the primary database, which fully records detailed information about every data change in the database (including change type, primary key of the changed data, field values ​​before and after modification, change timestamp, etc.). CDC tools or custom monitoring services can be used to connect to the database log interface to ensure that once a data change generates a log, it can be detected immediately, avoiding omissions or delays in change reporting.

[0118] S502. Based on the data change log, change the data information in multiple second databases to keep the data in the first database synchronized with the data in the second database.

[0119] Furthermore, the computing device can publish data change logs to a pre-defined synchronization channel, triggering full replica synchronization. That is, after each data change is processed, it is immediately published to a message queue, ensuring that the change is instantly perceived by the subscribed second database, which can be a replica database. Through "publish / subscribe," data change information is propagated in real time to all database replicas (e.g., primary database → standby database, local database → off-site disaster recovery database), ensuring that the data in all replicas is consistent with the source database. During synchronization, distributed transactions, such as 2PC (two-phase commit protocol), are used to ensure the atomicity, consistency, isolation, and durability of changes, avoiding situations where some replicas synchronize successfully while others fail.

[0120] In the embodiments of this disclosure, S502 includes: determining a change priority score for multiple data change information; and determining a synchronous change order based on the multiple change priority scores to execute the multiple data change information, wherein the change priority score is determined at least based on the change urgency score and change frequency of the data change information.

[0121] When there are many data change messages and / or the system load is high, the change priority score of multiple data change messages can be determined, and the change order of multiple data change messages can be determined based on the change priority score.

[0122] Change priority scoring can be determined using the following formula:

[0123] in, Rate the priority of data changes; Score the urgency of the change; To change the frequency; Regulatory factors.

[0124] For data changes related to traffic accidents or sudden speed drops, a high change urgency score can be set. For frequently updated location data, set a high change frequency. For example, emergency incident reports It could be 1.0, and The value is low for regular location updates.

[0125] The urgency score for a change can be determined based on the severity of the event causing the change (such as a traffic accident or abnormal speed change), and can be defined in the following range: Low urgency: ∈[0, 0.3], suitable for routine position updates and minor state changes with no impact; moderate urgency: ∈[0.3, 0.7], applicable to significant state changes (such as speed reduction, vehicle deviation from path); high urgency: ∈[0.7, 1], applicable to major events (such as traffic accidents, sudden decrease in speed).

[0126] The change frequency can be determined based on the update frequency of data changes (such as the GPS location update frequency), and the following range can be defined: Low frequency: ∈[0, 0.3], suitable for data that is updated infrequently (such as abnormal temperature alarms); medium frequency: ∈[0.3, 0.7], suitable for data with medium update frequency (such as vehicle speed updates); high frequency: ∈[0.7, 1], suitable for frequently updated data (such as high-frequency GPS location updates).

[0127] In some embodiments, the adjustment factor It can also be dynamically adjusted. For example, under normal circumstances (system load is normal, no major events), the priority score can be changed. More attention will be paid to frequency and adjustment factors. The priority score can be set between [0.3, 0.5] to balance the impact of urgency and frequency. In emergency situations (such as traffic accidents or high system load), the priority score can be changed. It tends to favor urgency, moderating factors It can be set between [0.7, 0.9] to ensure a rapid response to emergency data.

[0128] In embodiments of this disclosure, such as Figure 2 The method also includes allocating resources based on a resource allocation model, which is determined at least based on memory utilization and CPU utilization.

[0129] When the system load is high, NGINX is used for load balancing and to monitor system resources in real time, dynamically adjusting resource allocation. If a server is determined to be overloaded, NGINX can redirect subsequent new requests (such as CDC synchronization requests and Kafka consumption requests) to servers with lower loads, achieving linkage between resource allocation and request scheduling. The resource allocation decision model is as follows:

[0130] Where A(t) is the resource allocation decision at time t, C(t) is the CPU utilization at time t, M(t) is the memory utilization at time t, and α is the adjustment factor.

[0131] For example, if a Kafka compute node has C(t) = 0.85 and M(t) = 0.75, then A(t) = 0.5 × 0.85 + 0.5 × 0.75 = 0.8. A higher value indicates a more urgent need for resources on that node. If the calculated A(t) value exceeds a preset threshold (e.g., 0.7, set according to system stability requirements), it is determined that Kafka resources are insufficient, triggering resource allocation actions. For example, the system can automatically start additional server instances or redirect some requests to servers with lower load.

[0132] According to embodiments of this disclosure, resource scheduling based on memory and CPU utilization can ensure that high-priority tasks run first, preventing critical tasks from being delayed due to insufficient resources. Simultaneously, resources are dynamically adjusted according to the load, avoiding waste of CPU, memory, and other resources, and preventing idleness, thus ensuring the overall efficient operation of the system.

[0133] Figure 6 A block diagram illustrating a data processing and synchronization system according to an embodiment of the present disclosure is shown.

[0134] like Figure 6 As shown, a data processing and synchronization system 600 includes multiple components, including at least a scoring determination component 601, a data storage component 602, and a calculation and decision component 603.

[0135] The scoring component 601 is used to determine the importance score of the data, and the importance score is used at least to determine the data collection frequency and compression method; the data storage component 602 is used to store the data based on the data access frequency; and the calculation decision component 603 is used to determine the data processing decision based on the data urgency score, so as to process the data, and the processing decision includes at least multi-dimensional data comparison of the data, and the urgency score is used at least to reflect the degree of data anomaly.

[0136] In embodiments of this disclosure, multiple components are deployed on multiple independent containers, which operate through automated management and performance monitoring.

[0137] In the aforementioned System 600, each component, such as the database, application server, and front-end, is deployed in an independent container. Simultaneously, Ansible automated deployment scripts and configuration management are used to ensure consistent and up-to-date configurations across all environments. This significantly improves deployment efficiency and system maintainability, enabling efficient and scalable service deployment while maintaining consistency and ease of configuration management.

[0138] For example, deploy critical components such as databases (e.g., PostgreSQL), application servers (e.g., Java Spring Boot applications), front-end services (e.g., React applications), and data processing services (e.g., Apache Kafka, Spark) in independent containers. Use Docker Compose or Kubernetes to orchestrate these container deployments, ensuring each component runs in its dedicated environment, unaffected by the state of other components. Containerized deployment allows multiple components to be deployed in isolation, avoiding mutual interference. This isolation prevents a failure in one component from affecting other components, ensuring uninterrupted data acquisition, no storage loss, and smooth change synchronization, effectively enhancing system reliability. It also ensures environmental consistency, reducing deployment risks and lowering the runtime environment requirements for data acquisition, storage, and change components.

[0139] Containerized deployment also supports rapid scaling up / down, for example, by using a container deployment decision model to achieve scaling up / down:

[0140] Where D(i) is the deployment decision, W(i) is the workload prediction for service i, and R(i) is the current resource utilization. It is a normalization function used to determine whether service instances need to be expanded or reduced.

[0141] For example, if W(i) / R(i) is greater than a predetermined threshold (future load is much higher than the current resource capacity), the decision is to expand the instance and increase the number of containers to cope with the upcoming high load. Conversely, if W(i) / R(i) is less than a predetermined threshold (current resources are largely idle, and the load is much lower than expected), the decision is to shrink the instance and reduce the number of containers to save resources. Thus, by dynamically comparing the predicted load with current resources, the system automatically decides on container scaling, ensuring that services always match business loads with reasonable resource allocation, guaranteeing performance while avoiding resource waste.

[0142] In the embodiments disclosed herein, Ansible automates the management of configurations and deployments across multiple environments (development, testing, and production), ensuring configuration consistency. Ansible playbooks are written to automate deployment processes and application configurations, such as automatically configuring network settings, database connections, and dependency management. All configurations (network, database connections, dependencies) and deployment steps (installation, startup, verification) are written into standardized playbooks (YAML scripts) and executed uniformly across environments, ensuring 100% configuration consistency. This helps eliminate failures caused by environmental inconsistencies while simplifying operations and maintenance and traceability.

[0143] In the embodiments disclosed herein, performance monitoring can be implemented based on the ELK Stack and Prometheus. The ELK Stack can collect system logs and application logs, index them, and enable rapid retrieval and analysis of log data. ELK is a combination of three tools: Elasticsearch, Logstash, and Kibana, which work together to process logs: Logstash can collect scattered log data (such as error logs and runtime logs) in batches from systems (such as servers and Docker containers) and applications (such as MQTT services and CDC synchronization tools), and then transmit them in a unified manner. Elasticsearch can classify and index the collected logs according to rules. Kibana can provide a visualization interface, and when a system error occurs, Kibana can be used to quickly locate the fault location.

[0144] Prometheus can monitor real-time performance metrics such as CPU utilization, memory utilization, and network traffic, and visualize the data using Grafana. Specifically, Prometheus can obtain real-time performance metrics for systems and applications, including CPU utilization, memory utilization, network traffic, database read / write latency, and service response time. Grafana can visually display the metrics collected by Prometheus in the form of charts, dashboards, and other formats.

[0145] The ELK Stack and Prometheus provide powerful log management and performance monitoring capabilities, allowing you to quickly pinpoint the root cause through logs, monitor performance metrics in real time, and proactively identify risks such as resource shortages and service lag, which helps maintain stable system operation.

[0146] In the embodiments of this disclosure, resource optimization can also be performed using the following resource optimization model:

[0147] Where O(t) represents the resource optimization operation, M(t,j) is the value of the j-th monitoring metric at time t, and U(j) is the importance weight of the metric. It is a decision function.

[0148] First, all monitored metrics (such as CPU utilization, memory utilization, network traffic, etc.) are weighted by multiplying their values ​​by their respective weights. Then, a decision function analyzes this weighted sum, ultimately outputting specific resource optimization actions. This ensures that resource optimization is based on the weighted results of the importance of multi-dimensional metrics, providing a comprehensive overview while highlighting the impact of core metrics, resulting in more rational and efficient resource allocation. The decision function is responsible for translating the "weighted sum of monitoring metrics" into specific resource optimization actions; for example, when the weighted sum exceeds a certain threshold, the decision function outputs a capacity expansion operation.

[0149] Figure 7 A block diagram of a network device according to an embodiment of the present disclosure is schematically illustrated; like Figure 7 As shown, the network device 700 of this embodiment includes a memory 701 and a processor 702.

[0150] Memory 701 is used to store computer-readable instructions. Processor 702 is used to execute the computer-readable instructions, causing the network device to perform the data processing and synchronization method.

[0151] Figure 8 A block diagram illustrating a non-transitory computer-readable storage medium according to an embodiment of the present disclosure is shown schematically. like Figure 8 As shown, a non-transitory computer-readable storage medium 800 according to an embodiment of the present disclosure is used to store computer-readable instructions 801, which, when executed by a processor, cause the processor to perform the data processing and synchronization method described above.

[0152] Figure 9 A block diagram illustrating a computer program product according to an embodiment of the present disclosure is shown schematically. like Figure 9 As shown, a computer program product 900 according to an embodiment of this disclosure includes a computer program 901, which, when executed by a processor, implements the data processing and synchronization method described above.

[0153] The above description, with reference to the accompanying drawings, illustrates a data processing and synchronization method and system according to embodiments of the present disclosure. By determining the importance score of data, which is used to determine the frequency of data collection and the compression method of transmitted data, it is beneficial to reduce the amount of data collected and transmitted, thereby improving the efficiency of data processing. Simultaneously, partitioning and storing data can further improve storage efficiency and facilitate subsequent data use. When data changes, information in multiple databases is updated synchronously through data change information, which helps maintain data consistency. By using change priority scoring to determine the order of multiple data change information, resources can be allocated preferentially to high-priority data change information when resources are scarce. Furthermore, deploying different components in independent containers facilitates automated and batch management, improves the management efficiency of multiple components, and enhances system stability.

[0154] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0155] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0156] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.

[0157] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.

[0158] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described above can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.

[0159] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0160] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A data processing and synchronization method, characterized in that, The method includes: The importance score of the data is determined, and the importance score is used at least to determine the data collection frequency and compression method; The data is stored based on the access frequency of the data. Based on the urgency score of the data, a processing decision is determined to process the data, wherein the urgency score at least reflects the degree of abnormality of the data; and When data change information is captured in the first database, the information in the first database and the second database is synchronized based on the data change information.

2. The data processing and synchronization method according to claim 1, characterized in that, Also includes: If the importance score of the data is greater than the first importance score threshold, the data is fully compressed. The importance score of the data is determined at least based on the type weight of the data and the evaluation function of the data. or If the importance score of the data is less than or equal to the first importance score threshold, simplified compression is performed on the data.

3. The data processing and synchronization method according to claim 2, characterized in that, Also includes: The data collection frequency is determined, and the collection frequency is determined based at least on the basic collection frequency, the importance score, and the second importance score threshold.

4. The data processing and synchronization method according to claim 1, characterized in that, The storage of the data based on the access frequency of the data includes: Determine the access frequency of the data, wherein the access frequency includes at least the historical access frequency or the predicted access frequency; If the access frequency of the data exceeds an access frequency threshold, the data is cached, whereby the access frequency threshold is determined at least based on the maximum cache capacity and the average access frequency; or If the access frequency of the data is less than or equal to the access frequency threshold, the data is stored in partitions based on the access frequency and the type of the data.

5. The data processing and synchronization method according to claim 1, characterized in that, The process of determining the data processing decision based on the urgency score of the data includes: The urgency score of the data is compared with an urgency score threshold. The data includes multidimensional attributes, and the urgency score is determined based at least on the attribute weights of the multidimensional attributes and the urgency evaluation function of the multidimensional attributes. If the urgency score is greater than the urgency score threshold, the data is processed based on the processing decision.

6. The data processing and synchronization method according to claim 5, characterized in that, Also includes: The data is subjected to multidimensional data comparison, which includes at least a comparison based on a multidimensional data hash, and the multidimensional data hash is determined based at least on the multidimensional attributes of the data and the attribute weights.

7. The data processing and synchronization method according to claim 1, characterized in that, When data change information of the first database is captured, synchronizing the information of the first database and the second database based on the data change information includes: Obtain data change information from the first database, wherein the data change information is used to record at least the write information, change information, and operation information of the data in the first database; and Based on the data change log, data information in multiple second databases is modified to keep the data in the first database synchronized with the data in the second database.

8. The data processing and synchronization method according to claim 7, characterized in that, The modification of data information in multiple second databases based on the data change log includes: Determine the change priority score for multiple data change information; and Based on multiple change priority scores, a synchronous change order is determined to execute multiple data change messages, wherein the change priority scores are determined at least based on the change urgency score and change frequency of the data change messages.

9. The data processing and synchronization method according to claim 7, characterized in that, Also includes: Resources are allocated based on a resource allocation model, which is determined at least based on memory utilization and CPU utilization.

10. A data processing and synchronization system, characterized in that, The system includes multiple components, including at least a scoring determination component, a data storage component, and a calculation and decision-making component; wherein... The scoring component is used to determine the importance score of the data, which is used at least to determine the data collection frequency and compression method; the data storage component is used to store the data based on the access frequency of the data; and the calculation decision component is used to determine the processing decision of the data based on the urgency score of the data, so as to process the data, the processing decision includes at least multi-dimensional data comparison of the data, and the urgency score is used at least to reflect the degree of abnormality of the data.

11. The data processing and synchronization system according to claim 10, characterized in that, The multiple components are deployed based on multiple independent containers, which are operated through automated management and performance monitoring.