Data processing method and device, storage medium and electronic equipment

Through dynamic sharding and traffic shaping mechanisms, queue depth and sharding priority are adjusted according to network performance load, which solves the problem that fixed queue depth cannot adapt to traffic bursts, improves the stability and reliability of data processing, and reduces packet loss rate and resource occupation.

CN120358196APending Publication Date: 2025-07-22CHINA MOBILE COMM GRP CHONGQING CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510606309.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, data cache with fixed queue depth cannot adapt to traffic burst scenarios, resulting in reduced data transmission stability and reliability, high packet loss rate, and lack of active prediction and plastic surgery capabilities for traffic abnormalities.

Method used

The coordinated mechanism of dynamic sharding and traffic shaping is adopted to dynamically adjust the queue depth of the data cache queue according to the network performance load of the target host, and coordinate the data sharding through DSCP tagging and queue scheduling priority to realize the dynamic allocation and transmission of data shards.

Benefits of technology

It effectively reduces the packet loss rate during data processing, improves the stability and reliability of data processing, reduces the use of protocol stack resources by shard reorganization, and improves the efficiency and resource utilization of data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358196A_ABST
    Figure CN120358196A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, a storage medium and electronic equipment. The method comprises the following steps: in response to a data acquisition request sent by a target application, determining to-be-processed data according to data acquisition information in the data acquisition request; acquiring data attribute information corresponding to the to-be-processed data and a network performance load corresponding to the target host; under the condition that the to-be-processed data is dynamic data, the to-be-processed data is divided into a plurality of data fragments according to the data attribute information, and the dynamic data is used for representing communication service data stored by the target host; determining the queue depth of a data cache queue according to the network performance load of the target host; and caching the plurality of data fragments into a data cache queue with queue depth. According to the scheme provided by the embodiment of the invention, the stability and reliability of data processing can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of cloud computing and network traffic management, and particularly relates to a data processing method, apparatus, storage medium, and electronic device. Background Art

[0002] With the rapid development of cloud computing technology, more and more enterprises and organizations deploy their services in the cloud. In the cloud computing resource pool, the collection and transmission of monitoring data are key links for the stable operation of the cloud platform.

[0003] In the related art, a static queue with a fixed queue depth is usually used to cache packets, which cannot adapt to scenarios of traffic bursts, thereby increasing the packet loss rate of data transmission and reducing the stability of data transmission. Summary of the Invention

[0004] Embodiments of this application provide a data processing method, apparatus, storage medium, and electronic device, which can effectively improve the stability and reliability of data processing.

[0005] In a first aspect, embodiments of this application provide a data processing method, which includes: in response to a data collection request sent by a target application, determining data to be processed according to the data collection information in the data collection request; obtaining data attribute information corresponding to the data to be processed and the network performance load corresponding to the target host; in the case where the data to be processed is dynamic data, dividing the data to be processed into multiple data shards according to the data attribute information, where the dynamic data is used to represent communication service data stored in the target host; determining the queue depth of the data cache queue according to the network performance load of the target host; and caching the multiple data shards into the data cache queue with the queue depth.

[0006] In a second aspect, embodiments of this application provide a data processing apparatus, which includes: a data determination module, configured to determine data to be processed according to the data collection information in the data collection request in response to a data collection request sent by a target application; a parameter acquisition module, configured to obtain data attribute information corresponding to the data to be processed and the network performance load corresponding to the target host; a sharding module, configured to divide the data to be processed into multiple data shards according to the data attribute information in the case where the data to be processed is dynamic data, where the dynamic data is used to represent communication service data stored in the target host; a queue adjustment module, configured to determine the queue depth of the data cache queue according to the network performance load of the target host; and a data caching module, configured to cache the multiple data shards into the data cache queue with the queue depth.

[0007] In a third aspect, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the data processing method described in the first aspect is implemented.

[0008] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the data processing method described in the first aspect is implemented.

[0009] Fifthly, an embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute the data processing method described in the first aspect.

[0010] As can be seen from the above, in the embodiment of the present application, the queue depth of the data cache queue is not fixed, but dynamically changes. Moreover, the queue depth is dynamically matched with the network performance load of the target host in real time. Compared with the data cache queue with a fixed queue depth in the related art, it can adapt to the scenario of traffic bursts, effectively reduce the packet loss rate in the data processing process, and improve the stability and reliability of data processing. In addition, in the embodiment of the present application, the data to be processed is fragmented by using the data attribute information corresponding to the data to be processed, realizing dynamic fragmentation of the data, making the size of the data fragments more reasonable, reducing the protocol stack resources occupied by fragmentation and recombination, and further reducing the impact of instantaneous traffic caused by fragmentation of large packets, and improving the stability and reliability of data processing.

[0011] It can be seen that the method provided by the embodiment of the present application can effectively improve the stability and reliability of data processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0013] Figure 1 is a schematic structural diagram of a data processing system in the related art;

[0014] Figure 2 is a schematic structural diagram of a data processing system provided by an embodiment of the present application;

[0015] Figure 3 is a schematic flowchart of a data processing method provided by an embodiment of the present application;

[0016] Figure 4 is an overall process architecture diagram of a data processing method provided by an embodiment of the present application;

[0017] Figure 5 is a schematic structural diagram of a data processing device provided by another embodiment of the present application;

[0018] Figure 6 It is a schematic structural diagram of an electronic device provided by another embodiment of the present application. Detailed implementation manners

[0019] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.

[0020] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or sequence between these entities or operations. Moreover, the term "comprises", "comprising" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0021] For ease of understanding, before explaining the solution provided by the present application, the background of the solution provided by the present application will be explained first.

[0022] With the expansion of the scale of cloud computing, the cloud platform data collection faces network stability problems caused by high concurrency and large message transmission.

[0023] In the cloud computing resource pool, the monitoring data collection and transmission is the core link for the stable operation of the platform. Traditional technical solutions are usually as Figure 1 the structure of the data processing system shown, such as Figure 1As shown, a traditional data processing system includes a cloud platform component, for example, Clivia, which includes an interface component Clivia API and a data processing component Clivia MGT, wherein the interface component can send an application's data collection request to the data processing component, so that the data processing component can collect host data in a resource pool, and transmit the collected host data to the interface component in the form of a message queue RabbitMQ, and then send it to the application via HTTP (Hypertext Transfer Protocol).

[0024] Based on the traditional system structure, a full data synchronization mechanism is usually used to implement data processing. Specifically, the cloud platform component (e.g., clivia) returns all collected data (e.g., host CPU, network card performance, etc.) at one time through the full data interface (e.g., items_collect). When the resource pool scales up, the messages generated by the full data exceed the MTU (Maximum Transmission Unit) limit, and the fragmentation and reassembly occupy a large amount of protocol stack resources, causing instantaneous traffic impact. For example, 145,000 packets within 100ms, resulting in high pressure on fragmentation message reassembly, triggering queue overflow and packet loss. For example, RabbitMQ transmits large packets on port 25672, resulting in tap-DYMANA packet loss.

[0025] In addition, in related technologies, fixed queue depths (such as) are usually used to cache messages. For example, DVS (Digital Video Server) uses a queue depth of 1024KB by default. When burst traffic exceeds the queue capacity (i.e., queue depth), subsequent messages are directly discarded. For example, the DVS queue suffers from continuous packet loss due to virtual machine processing delays, with a packet loss rate as high as 1.7%. It can be seen that fixed queue depths cannot adapt to burst traffic scenarios, resulting in a linear increase in packet loss rate as the resource pool scales.

[0026] In addition, in related technologies, traffic priorities are differentiated through DSCP (Differentiated Services Code Point) marking, but there is a lack of dynamic scheduling capabilities for fragmented messages, which makes fragmented message scheduling disconnected from resource status. For example, the fragmentation strategy and queue parameters are not dynamically optimized in combination with real-time indicators such as CPU utilization and network throughput, resulting in low resource utilization; for another example, when performance collection data and signaling data share a queue, the transmission order is not dynamically adjusted according to the importance of the fragmented message, resulting in the discarding of key messages.

[0027] Moreover, in the related art, a passive alarm triggering mechanism is usually adopted, that is, the alarm is triggered only after packet loss occurs, lacking the ability of active prediction and shaping of traffic anomalies, relying on manual post-analysis and processing, and having low efficiency.

[0028] To solve the problems of the prior art, embodiments of the present application provide a data processing method, device, storage medium and electronic device.

[0029] The solution provided by the embodiments of the present application is based on Figure 2 the system structure of the data processing system shown to implement data processing. In Figure 2 it, the cloud platform component 10 is deployed in the target host 20 to implement operations such as data collection of the target application 30 on the data in the target host 20. The cloud platform component includes a dynamic sharding collection unit 101, a traffic shaping control unit 102, and a sharded packet scheduler 103.

[0030] The dynamic sharding collection component 101 includes a static data cache unit 1010 and a dynamic data sharding rule engine 1011. The static data cache unit 1010 is used to implement the collection of static data, and the dynamic data sharding rule engine 1011 is used to implement the sharding of dynamic data.

[0031] The traffic shaping control unit 102 includes a real-time metric collector 1020 and a dynamic queue adjustment unit 1021. The real-time metric collector 1020 is used to collect the dynamic data of the target host in real time; the dynamic queue adjustment unit 1021 is used to adjust the queue depth of the data cache queue for caching dynamic data.

[0032] The sharded packet scheduler 103 includes a DSCP marking engine 1030 and a queue scheduler 1031. The DSCP marking engine 1030 is used to mark data shards to determine the service level of data shards; the queue scheduler 1031 is used to schedule data shards and cache the data shards into the data cache queue.

[0033] As can be seen from the above, in the embodiments of the present application, the system architecture shown in Figure 2 is adopted. The load pressure of data collection is reduced through the collaborative mechanism of dynamic sharding and traffic shaping, load balancing is achieved by using dynamic sharding collection, and finally closed-loop feedback optimization is achieved through the collaborative mechanism of sharding and shaping.

[0034] The method provided by the embodiments of the present application will be introduced below with the cloud platform component in the data processing system as the execution subject.

[0035] Figure 3 The flowchart of the data processing method provided by an embodiment of the present application is shown. As Figure 3 shown, the method includes the following steps S301 to S305:

[0036] Step S301: In response to a data collection request sent by a target application, determine the data to be processed according to the data collection information in the data collection request.

[0037] In step S301, the target application can be an application that has an access requirement for the data in the resource pool, and it can collect data by calling the platform interface. When the target application calls the platform interface, it needs to carry sharding parameters (X-Slice-Params: {"slice_id": 1, "total_slices": 10, "data_type": "cpu_usage"}) in the HTTP Header. Among them, slice_id is the identifier of the data shard, total_slices is the total number of data shards, data_type is the type of the data shard, and cpu_usage is the CPU usage rate. In the embodiment of the present application, the cloud platform component automatically verifies the legality of the shard number. If slice_id > total_slices, an HTTP 400 error is returned.

[0038] In step S301, the data collection information may include, but is not limited to, information such as the data type, data volume, and time span corresponding to the data to be collected by the target application. Among them, the data type includes dynamic data and static data. The dynamic data is used to represent the communication service data stored in the target host. For example, fault diagnosis data, statistical data of periodic network throughput, CPU load, network bandwidth, etc.; the static data is data that remains unchanged for a long time or is updated at a low frequency, and may include the hardware data of the target host, such as the hardware model of the target host, network topology, etc.

[0039] Step S302: Obtain the data attribute information corresponding to the data to be processed and the network performance load corresponding to the target host.

[0040] In step S302, the data attribute information of the data to be processed may include, but is not limited to, the total amount of the data to be processed and the time span determined by the timestamp of the data to be processed. For example, if the range of the timestamp corresponding to the data in the data to be processed is from 0:00 to 0:10, the time span corresponding to the data to be processed is 10 minutes.

[0041] In step S302, the network performance load at least includes the memory utilization rate and network card throughput of the target host.

[0042] Step S303: In the case where the data to be processed is dynamic data, divide the data to be processed into multiple data shards according to the data attribute information.

[0043] In step S303, the data sharding is obtained by dividing according to the data attribute information of the data to be processed. That is, for different data to be processed, the number of data shards obtained and the amount of data corresponding to each data shard are different. By dividing the data shards through the data attribute information, the dynamic division of the data shards is realized, making the size of the data shards more reasonable, reducing the protocol stack resources occupied by shard recombination, thereby reducing the impact of the instantaneous traffic caused by large message sharding, and improving the stability and reliability of data processing.

[0044] Step S304, determine the queue depth of the data cache queue according to the network performance load of the target host.

[0045] In step S304, the queue depth of the data cache queue is not fixed, but dynamically changes. Moreover, the queue depth is dynamically matched with the network performance load of the target host in real time. Compared with the data cache queue with a fixed queue depth in the related art, it can adapt to the scenario of traffic burst, effectively reducing the packet loss rate in the data processing process and improving the stability and reliability of data processing.

[0046] Step S305, cache multiple data shards into the data cache queue with a queue depth.

[0047] After dynamically dividing the data to be processed through step S303 and dynamically adjusting the queue depth of the data cache queue through step S304, the cloud platform component can cache the data shards into the data cache queue and send them to the target application, thereby realizing the acquisition of the data to be processed by the target application.

[0048] Based on the solution defined in the above steps S301 to S305, it can be known that in the embodiment of the present application, the queue depth of the data cache queue is not fixed, but dynamically changes. Moreover, the queue depth is dynamically matched with the network performance load of the target host in real time. Compared with the data cache queue with a fixed queue depth in the related art, it can adapt to the scenario of traffic burst, effectively reducing the packet loss rate in the data processing process and improving the stability and reliability of data processing. In addition, in the embodiment of the present application, the data to be processed is sharded by using the data attribute information corresponding to the data to be processed, realizing the dynamic sharding of the data, making the size of the data shards more reasonable, reducing the protocol stack resources occupied by shard recombination, thereby reducing the impact of the instantaneous traffic caused by large message sharding, and improving the stability and reliability of data processing.

[0049] It can be seen that the method provided by the embodiment of the present application can effectively improve the stability and reliability of data processing.

[0050] The implementation process of the method provided by the embodiments of the present application is introduced below.

[0051] For dynamic data, in the embodiments of the present application, a sharding processing method is adopted to implement the processing of the data to be processed.

[0052] In some embodiments, the cloud platform component shards the data to be processed according to the data attribute information of the data to be processed. Specifically, the cloud platform component calculates the ratio of the total amount of the data to be processed to the first data volume threshold of the pre-sharded data shards to obtain a first value; and calculates the ratio between the time span and the preset time window to obtain a second value; then, determines the number of shards according to the first value and the second value; and divides the data to be processed into the number of data shards corresponding to the number of shards.

[0053] In the embodiments of the present application, the dynamic data sharding rule engine can implement the sharding of the data to be processed, where the number of shards of the data to be processed can be represented by formula (1):

[0054]

[0055] In formula (1), N is the number of shards of the data to be processed; N0 is the total amount of the data to be processed; N1 is the first data volume threshold, for example, 1MB; T0 is the time span corresponding to the data to be processed; T1 is the preset time window, for example, 5min.

[0056] In an example, when the total amount of CPU data within 1 hour is 8MB, according to formula (1), the dynamic data sharding rule engine is automatically split into 12 data shards, that is, N = 12, and each data shard contains 5 minutes of data.

[0057] Through formula (1), based on the time window (with T1 as the granularity) and the data type (such as CPU, network, memory), the dynamic data is split into multiple data shards, and the upper limit of the single shard capacity is N1, so as to avoid the problem of MTU shard recombination. When T1 is 5min and N1 is 1MB, it can ensure that the transmission time of a single data shard is less than 10ms, reduce the data transmission delay, and improve the data transmission efficiency.

[0058] In the embodiments of the present application, multiple data shards have at least two different service levels, and the service level is used to characterize the influence degree of the data shard on the service. Among them, the DSCP marking engine can mark the data shards, and the identifier corresponding to each data shard is used to characterize the service level corresponding to the data shard.

[0059] In one example, the DSCP marking engine embeds DSCP priority tags in the RabbitMQ message header. The priority tags need to be strongly associated with the service SLA (Service Level Agreement). Shard data corresponding to key metrics such as CPU over-threshold alarms and memory exhaustion events is marked as EF (DSCP = 46); shard data with data timeliness requirements, for example, shard data with high real-time requirements (data sensitive to second-level latency), can also be marked as EF; shard data directly affecting service availability (such as fault diagnosis data) can also be marked as EF.

[0060] In addition, the DSCP marking engine marks ordinary shard data of regular monitoring data, such as statistical data of periodic network throughput, as AF41 (DSCP = 34).

[0061] In the embodiment of the present application, the queue scheduler caches multiple shard data into the data cache queue in sequence according to the preset cache order; during the process of caching multiple shard data, the current queue utilization rate of the data cache queue is calculated every preset time period; when the current queue utilization rate is greater than the preset utilization rate threshold, the following steps are iteratively executed until the remaining memory of the data cache queue is less than or equal to the preset memory threshold: determine the shard data with the highest current service level from the uncached shard data; cache the shard data with the highest current service level into the data cache queue.

[0062] In one example, after the queue scheduler caches shard data of a preset time period or a preset data volume into the data cache queue each time, the current queue utilization rate of the data cache queue is calculated once. If the current queue usage is greater than the usage threshold (for example, 70%), the queue scheduler preferentially caches the shard data with the highest service level into the data cache queue. For example, when the current queue utilization rate is greater than 70%, only EF-level shard data is cached, and AF41-level shard data is temporarily stored or discarded. If all EF-level shard data is cached into the data cache queue and there is still some remaining space in the data cache queue, shard data of the next level can be cached into the data cache queue, and the above process is repeated until the remaining space in the data cache queue is insufficient.

[0063] During the process of caching the shard data with the highest current service level into the data cache queue, to reduce the risk of queue overflow, the queue scheduler performs shard demotion on the shard data with the highest current service level, divides the shard data with the highest current service level among the uncached shard data into multiple sub-shard data, and caches the multiple sub-shard data into the data cache queue.

[0064] In the above embodiments, the second data volume threshold corresponding to the sub-data shard is less than the first data volume threshold. For example, the data shard at the EF level is 1 MB, that is, the first data volume threshold is 1 MB, and the second data volume threshold can be 256 KB.

[0065] In one example, when the queue is congested, for example, the utilization rate of the queue > 70%, the queue scheduler splits the data shard with the highest current service level from 1 MB into 4 sub-data shards of 256 KB each to reduce the risk of transmission failure of a single data shard. The downgraded sub-data shards carry the metadata of the original shard, which is convenient for the receiving end to reconstruct. Moreover, by downgrading the shards and adjusting the shard granularity, the transmission success rate in high-load scenarios can be improved.

[0066] It should be noted that, to improve the transmission success rate of data shards, in the embodiments of the present application, an ACK (Acknowledge character) confirmation retransmission mechanism can also be adopted. After each data shard is transmitted, wait for the ACK signal feedback from the receiving end (for example, the target application). If the confirmation times out (the default is 500 ms), the cloud platform component triggers a retransmission, and the number of retry attempts ≤ 3 times. When the traffic shaping control unit detects that the throughput of the network card is greater than 8 Gbps, it automatically reduces the shard size from 1 MB to 512 KB and discards the expired shards through the TTL (Time To Live) mechanism of RabbitMQ.

[0067] In the embodiments of the present application, the dynamic queue adjustment unit can determine the queue depth of the data cache queue according to the host parameters of the target host. Specifically, the dynamic queue adjustment unit determines the load adjustment coefficient based on the network card throughput; determines the memory adjustment coefficient based on the memory utilization rate; and adjusts the preset initial queue depth based on the memory adjustment coefficient and the load adjustment coefficient to obtain the queue depth of the data cache queue.

[0068] In one embodiment, the real-time metric collector collects the CPU utilization rate C of the target host every second through Linux kernel interfaces, such as / proc / net / dev, / proc / stat current and the network card throughput T rx . The dynamic queue adjustment unit calculates the queue depth using a non-linear weight adjustment formula, and the adjustment of the queue depth can be achieved through formula (2):

[0069]

[0070] In formula (2), Q depth is the queue depth; Q is the initial queue depth, for example, 1024 MB; is the load adjustment coefficient; T rx is the network card throughput; T rx0is the reference bandwidth, for example, 10 Gbps; α is the bandwidth adjustment coefficient, for example, 0.5; is the memory adjustment coefficient; C current is the CPU utilization rate; C current0 is the preset queue usage threshold, for example, 90%.

[0071] The queue depth can be limited within a reasonable range of 512 - 4096 through formula (2).

[0072] In some embodiments, after caching multiple data shards into a data cache queue with a queue depth, the shard packet scheduler obtains the network bandwidth corresponding to the target host and the initial traffic shaping factor corresponding to the target host; modifies the network bandwidth through the initial traffic shaping factor to obtain the data transmission rate of multiple data shards; and transmits the data shards in the cache queue to the target application according to the data transmission rate of multiple data shards. For example, if the reference bandwidth of the network is B and the initial traffic shaping factor is γ, the data transmission rate can be expressed by formula (3):

[0073] S = γ·B (3)

[0074] In formula (3), S is the data transmission rate; γ is the initial traffic shaping adjustment factor; B is the reference bandwidth of the network. For example, when the reference bandwidth B is 1 Gbps, if γ = 0.8, then S is 800 Mbps.

[0075] After transmitting the data shards in the cache queue to the target application according to the data transmission rate of multiple data shards, the shard packet scheduler also obtains the data packet loss rate within a continuous plurality of data transmission cycles; and reduces the data transmission rate when the data packet loss rates in a plurality of consecutive cycles are greater than the preset packet loss rate threshold.

[0076] Specifically, the shard packet scheduler determines the traffic shaping factor correction value based on the deviation degree between the data packet loss rate and the preset packet loss rate threshold; performs a weighted calculation on the traffic shaping factor correction value and the initial traffic shaping factor to obtain the target traffic shaping factor; and then adjusts the data transmission rate based on the target traffic shaping factor to obtain the target data transmission rate. Among them, the target traffic shaping factor is less than the initial traffic shaping factor.

[0077] In the embodiments of the present application, the shard packet scheduler monitors the real-time load by mirroring the traffic through ERSPAN (Encapsulated Remote Switch Port Analyzer), and statistically counts the rate of shard arrival and the queue usage rate every 100 ms. When the packet loss rate P is detected in 3 consecutive cycles lossWhen it is > 0.1%, the traffic shaping factor is automatically down-regulated, and the adjustment of the traffic shaping factor can be achieved through formula (4):

[0078]

[0079] In formula (4), γ new is the target traffic shaping factor; γ old is the initial traffic shaping factor; β0 is the weighting coefficient of the initial traffic shaping factor; P loss is the data packet loss rate; P0 is the packet loss rate threshold; is the traffic shaping correction factor; β1 is the weighting coefficient of the traffic shaping correction factor; β0 + β1 = 1.

[0080] It should be noted that the traffic shaping factor is the core parameter for dynamically adjusting the traffic output rate, and it optimizes the network transmission stability through closed-loop feedback.

[0081] In an application scenario, when the packet loss rate is found to exceed the threshold P loss > 0.1% in three consecutive data transmission cycles (each cycle is 100 ms), the adjustment of the traffic shaping factor is triggered. When β0 = 0.9, β1 = 0.1, and P0 = 0.5%, formula (4) can be represented by formula (5):

[0082]

[0083] In formula (5), the historical weight of the initial traffic shaping factor is 0.9, indicating that 90% of the historical factor value is retained to avoid traffic oscillation caused by mutation. The dynamic correction is 0.1, indicating that the remaining 10% weight is adjusted according to the deviation degree between the current packet loss rate and the threshold (0.5%).

[0084] When P loss = 0.1% (i.e., the lower limit of the threshold), the correction term is 0.1·(1 - 0.2) = 0.08, and the traffic shaping factor γ decreases slightly; when P loss = 0.5% (i.e., the upper limit of the correction), the correction term is 0.1·(1 - 1) = 0, and the traffic shaping factor γ decreases the most, γ new = 0.9γ old .

[0085] It should be noted that in addition to affecting the data transmission rate, the adjusted traffic shaping factor γ new also affects the priority of queue scheduling and the compensation for fragmentation degradation.

[0086] For the priority of queue scheduling, the γ of high-priority fragments (e.g., EF level) remains unchanged, and only the γ of ordinary fragments (AF level) is adjusted; when the packet loss rate increases, the γ of ordinary fragments is decreased to give priority to ensuring the bandwidth of EF-level fragments.

[0087] For shard degradation compensation, after the traffic shaping factor is down-regulated, shard splitting is triggered. For example, 1MB of data shards are split into 4 data shards of 256KB each, and the sending interval of sub-data shards is restricted by γ to avoid burst traffic exacerbating congestion.

[0088] In one example, in the scenario of the coordination of ERSPAN and traffic shaping, ERSPAN provides real-time traffic mirroring, counts the shard arrival rate and queue utilization rate, and determines whether the γ adjustment condition is met based on the monitored data. For example, if the packet loss rate > 0.1%, then γ is adjusted. The adjusted γ is used for traffic shaping to change the shard sending strategy, and finally the effect is verified through ERSPAN monitoring to form a closed-loop optimization.

[0089] In another example, when γ old = 1.0, the reference bandwidth is 1Gbps, and the allowed rate is 1Gbps. If the packet loss rate for 3 consecutive cycles is 0.3%, then that is, the allowed data transmission rate is reduced to 940Mbps, and the sending volume of ordinary shards is preferentially reduced. Whether to continue to adjust γ is determined by monitoring whether the packet loss rate drops in the next cycle through ERSPAN.

[0090] The above completes the introduction to the processing of dynamic data.

[0091] For static data, that is, in the case where the data type of the data to be processed is static data, the cloud platform component can send the static data stored in the data storage unit of the target host to the target application.

[0092] In the embodiment of the present application, non-real-time data such as the host hardware model and network topology can be stored in a static data cache unit (for example, Redis cache library), and the materialized view technology (Materialized View) is used to implement low-frequency full-volume queries (such as once every 10 minutes) to reduce the pressure of repeated acquisition. If the requested data type is hardware information, the cache result is directly read from Redis (response time ≤ 2ms).

[0093] In one embodiment, Figure 4 shows the overall process architecture diagram of the method provided by the embodiment of the present application. As Figure 4 shown, the target application can read the static data and dynamic data in the target host through the cloud platform component. For static data, the target application can directly read the data in the Redis cache library; for dynamic data, the cloud platform component implements the transmission of dynamic data by sharding the dynamic data and adopting the dynamic data shard degradation and ACK confirmation retransmission mechanism.

[0094] As can be seen from the above, the solution provided by the embodiments of this application adopts a dynamic sharding acquisition mechanism, and realizes dynamic sharding and data transmission of data based on data types and time windows; adopts a traffic shaping and queue adaptive algorithm, and combines network throughput and CPU utilization rate to realize dynamic adjustment of the queue depth of the data cache queue; adopts sharded packet priority collaborative scheduling, and realizes the transmission of sharded data based on DSCP marking and queue status.

[0095] Compared with the existing full-volume data acquisition and static queue solutions, adopting the solution provided by the embodiments of this application reduces the amount of data transmitted in a single transmission of dynamic data by 60%, and reduces the time-consuming of RabbitMQ sharding recombination to 5ms (original 15ms), thus improving the transmission efficiency of data sharding. In the solution provided by the embodiments of this application, by dynamically adjusting the queue depth of the data cache queue, the packet loss rate in the burst traffic scenario is reduced from 1.7% to 0.08%, thus reducing the risk of data cache queue overflow. In the solution provided by the embodiments of this application, by suppressing burst traffic through the traffic shaping factor γ, the CPU utilization rate fluctuation is reduced by 50%, supporting the stable operation of a resource pool at the ten-thousand-node level, and realizing the optimization of resource utilization.

[0096] The embodiments of this application also provide a data processing device, as Figure 5 shown. The device 500 includes: a data determination module 501, a parameter acquisition module 502, a sharding module 503, a queue adjustment module 504, and a data cache module 505.

[0097] The data determination module 501 is configured to determine the data to be processed according to the data acquisition information in the data acquisition request in response to the data acquisition request sent by the target application.

[0098] The parameter acquisition module 502 is configured to acquire the data attribute information corresponding to the data to be processed and the network performance load of the target host.

[0099] The sharding module 503 is configured to divide the data to be processed into multiple data shards according to the data attribute information when the data to be processed is dynamic data, and the dynamic data is used to represent the communication service data stored in the target host.

[0100] The queue adjustment module 504 is configured to determine the queue depth of the data cache queue according to the network performance load of the target host.

[0101] The data cache module 505 is configured to cache multiple data shards into a data cache queue with a queue depth.

[0102] In some embodiments, the data attribute information at least includes the total amount of data to be processed and the time span determined by the timestamp of the data to be processed. The sharding module is specifically configured to calculate the ratio of the total amount of data to the first data volume threshold of the pre-sharded data shards to obtain a first value; calculate the ratio of the time span to the preset time window to obtain a second value; determine the number of shards according to the first value and the second value; and divide the data to be processed into the number of data shards equal to the number of shards.

[0103] In some embodiments, multiple data shards have at least two different service levels, and the service level is used to characterize the degree of influence of the data shard on the service. The data caching module includes: an initial caching module, a utilization rate calculation module, and an iteration module. Among them, the initial caching module is configured to cache multiple data shards into the data caching queue in sequence according to the preset caching order; the utilization rate calculation module is configured to calculate the current queue utilization rate of the data caching queue every preset time period during the process of caching multiple data shards; the iteration module is configured to, when the current queue utilization rate is greater than the preset utilization rate threshold, iteratively execute the following steps until the remaining memory of the data caching queue is less than or equal to the preset memory threshold: determine the data shard with the highest current service level from the un-cached data shards; cache the data shard with the highest current service level into the data caching queue.

[0104] In some embodiments, the iteration module is further configured to divide the shard data with the highest current service level among the un-cached data shards into multiple sub-data shards, where the second data volume threshold corresponding to the sub-data shards is less than the first data volume threshold; and cache the multiple sub-data shards into the data caching queue.

[0105] In some embodiments, the network performance load at least includes memory utilization and network card throughput. The queue adjustment module is specifically configured to determine a load adjustment coefficient based on the network card throughput; determine a memory adjustment coefficient based on the memory utilization; and adjust the preset initial queue depth based on the memory adjustment coefficient and the load adjustment coefficient to obtain the queue depth of the data caching queue.

[0106] In some embodiments, the data processing device further includes: a data transmission module, configured to, after caching multiple data shards into the data caching queue with a queue depth, obtain the network bandwidth corresponding to the target host and the initial traffic shaping factor corresponding to the target host; correct the network bandwidth through the initial traffic shaping factor to obtain the data transmission rate of multiple data shards; and transmit the data shards in the caching queue to the target application according to the data transmission rate of multiple data shards.

[0107] In some embodiments, the data processing device further includes: a packet loss rate acquisition module and a rate adjustment module. Among them, the packet loss rate acquisition module is configured to acquire the data packet loss rate within a continuous plurality of data transmission cycles after transmitting the data shards in the buffer queue to the target application according to the data transmission rate of the plurality of data shards; the rate adjustment module is configured to reduce the data transmission rate when the plurality of data packet loss rates are continuously greater than a preset packet loss rate threshold.

[0108] In some embodiments, the rate adjustment module is specifically configured to determine a traffic shaping factor correction value based on the deviation degree between the data packet loss rate and the preset packet loss rate threshold; perform weighted calculation on the traffic shaping factor correction value and the initial traffic shaping factor to obtain a target traffic shaping factor, where the target traffic shaping factor is less than the initial traffic shaping factor; adjust the data transmission rate based on the target traffic shaping factor to obtain a target data transmission rate.

[0109] In some embodiments, the data processing device further includes: a static data processing module, configured to, when the data type of the data to be processed is static data, send the static data stored in the data storage unit of the target host to the target application, where the static data includes at least the hardware data of the target host.

[0110] The data processing device provided in the embodiments of the present application can implement each process implemented by the foregoing method embodiments. To avoid repetition, it will not be elaborated here.

[0111] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0112] Figure 6 FIG. shows a schematic hardware structure diagram of an electronic device provided in an embodiment of the present application.

[0113] The electronic device may include a processor 601 and a memory 602 storing computer program instructions.

[0114] Specifically, the above-mentioned processor 601 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits for implementing the embodiments of the present application.

[0115] The memory 602 may include a mass storage for data or instructions. By way of example and not limitation, the memory 602 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disc, a magneto-optical disc, a magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more of these. In a suitable case, the memory 602 may include a removable or non-removable (or fixed) medium. In a suitable case, the memory 602 may be internal or external to the integrated gateway disaster recovery device. In a specific embodiment, the memory 602 is a non-volatile solid state memory.

[0116] The memory may include a read only memory (ROM), a random access memory (RAM), a magnetic disk storage media device, an optical storage media device, a flash memory device, an electrical, optical, or other physical / tangible memory storage device. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0117] The processor 601 reads and executes the computer program instructions stored in the memory 602 to implement any of the data processing methods in the above embodiments.

[0118] In one example, the electronic device may further include a communication interface 603 and a bus 610. Among them, as Figure 6 shown, the processor 601, the memory 602, and the communication interface 603 are connected through the bus 610 to complete the communication with each other.

[0119] The communication interface 603 is mainly used to implement the communication between each module, device, unit, and / or device in the embodiments of the present application.

[0120] The bus 610 includes hardware, software, or both, and couples components of the electronic device to each other. By way of example and not limitation, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses or a combination of two or more of these. Where appropriate, the bus 610 may include one or more buses. Although embodiments of the present application describe and illustrate specific buses, the present application contemplates any suitable bus or interconnect.

[0121] In addition, in combination with the data processing method in the above embodiments, embodiments of the present application may be implemented by providing a computer-readable storage medium. Computer program instructions are stored on the computer-readable storage medium; when the computer program instructions are executed by a processor, any one of the data processing methods in the above embodiments is implemented.

[0122] In addition, in combination with the data processing method in the above embodiments, embodiments of the present application may be implemented by providing a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is caused to execute any one of the data processing methods as in the above embodiments.

[0123] It should be clear that the present application is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and illustrated as examples. However, the method process of the present application is not limited to the specific steps described and illustrated, and those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present application.

[0124] The functional modules shown in the above-described structural block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present application are programs or code segments for performing the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted via a data signal carried in a carrier wave over a transmission medium or a communication link. "Machine-readable medium" can include any medium that can store or transmit information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical discs, hard disks, fiber optic media, radio frequency (RF) links, and so on. The code segment can be downloaded via a computer network such as the Internet, an intranet, and so on.

[0125] It should also be noted that in the exemplary embodiments mentioned in the present application, some methods or systems are described based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be executed in the order mentioned in the embodiments, or different from the order in the embodiments, or several steps can be executed simultaneously.

[0126] Aspects of the present disclosure have been described above with reference to the flowcharts and / or block diagrams of the data processing methods, devices, storage media, and electronic devices according to the embodiments of the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and the combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware for performing the specified functions or actions, or by a combination of dedicated hardware and computer instructions.

[0127] As described above, this is only the specific implementation manner of the present application. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated herein. It should be understood that the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present application.

Claims

1. A data processing method, characterized in that, Including: In response to a data collection request sent by a target application, determining data to be processed according to the data collection information in the data collection request; Obtaining data attribute information corresponding to the data to be processed and the network performance load corresponding to the target host; In the case where the data to be processed is dynamic data, dividing the data to be processed into multiple data shards according to the data attribute information, where the dynamic data is used to represent communication service data stored in the target host; Determining the queue depth of a data cache queue according to the network performance load of the target host; Caching the multiple data shards into the data cache queue with the queue depth.

2. The method according to claim 1, characterized in that The data attribute information at least includes the total amount of the data to be processed and the time span determined by the timestamp of the data to be processed. Dividing the data to be processed into multiple data shards according to the data attribute information includes: Calculating a ratio of the total amount of the data to a first data volume threshold of a pre-sharded data shard to obtain a first value; Calculating a ratio of the time span to a preset time window to obtain a second value; Determining the number of shards according to the first value and the second value; Dividing the data to be processed into the number of shards of the determined number of shards.

3. The method according to claim 2, wherein The multiple data shards have at least two different service levels, and the service level is used to represent the degree of influence of the data shard on the service. Caching the multiple data shards into the data cache queue with the queue depth includes: Caching the multiple data shards into the data cache queue in sequence according to a preset caching order; During the process of caching the multiple data shards, calculating the current queue utilization rate of the data cache queue every preset time period; In the case where the current queue utilization rate is greater than a preset utilization rate threshold, iteratively execute the following steps until the remaining memory of the data cache queue is less than or equal to a preset memory threshold: Determining the data shard with the highest current service level from the uncached data shards; Caching the data shard with the highest current service level into the data cache queue.

4. The method according to claim 3, wherein Caching the data shard with the highest current service level into the data cache queue includes: Dividing the shard data with the highest current service level among the uncached data shards into multiple sub-data shards, where the second data volume threshold corresponding to the sub-data shard is less than the first data volume threshold; Caching multiple sub-data shards into the data cache queue.

5. The method according to any one of claims 1-4, characterized in that, The network performance load at least includes memory utilization and network card throughput. Determining the queue depth of the data cache queue according to the host parameters of the target host includes: Determining a load adjustment coefficient based on the network card throughput; Determining a memory adjustment coefficient based on the memory utilization; Adjusting a preset initial queue depth based on the memory adjustment coefficient and the load adjustment coefficient to obtain the queue depth of the data cache queue.

6. The method according to any one of claims 1 to 4, characterized in that, After caching the multiple data shards into the data cache queue with the queue depth, the method further includes: Obtain the network bandwidth corresponding to the target host and the initial traffic shaping factor corresponding to the target host; Correct the network bandwidth by the initial traffic shaping factor to obtain the data transmission rate of the multiple data shards; Transmit the data shards in the cache queue to the target application according to the data transmission rate of the multiple data shards.

7. The method according to claim 6, wherein After transmitting the data shards in the cache queue to the target application according to the data transmission rate of the multiple data shards, the method further includes: Obtain the data packet loss rate within a continuous plurality of data transmission cycles; In the case where a plurality of the data packet loss rates are continuously greater than a preset packet loss rate threshold, reduce the data transmission rate.

8. The method according to claim 7, wherein The reducing the data transmission rate includes: Determine a traffic shaping factor correction value based on the deviation degree between the data packet loss rate and the preset packet loss rate threshold; Perform weighted calculation on the traffic shaping factor correction value and the initial traffic shaping factor to obtain a target traffic shaping factor, where the target traffic shaping factor is less than the initial traffic shaping factor; Adjust the data transmission rate based on the target traffic shaping factor to obtain a target data transmission rate.

9. The method according to any one of claims 1 to 4, characterized in that, The method further includes: In the case where the data type of the data to be processed is static data, send the static data stored in the data storage unit of the target host to the target application, where the static data at least includes the hardware data of the target host.

10. A data processing device, characterized in that, The apparatus includes: A data determination module, configured to determine data to be processed according to data collection information in the data collection request in response to a data collection request sent by a target application; A parameter acquisition module, configured to acquire data attribute information corresponding to the data to be processed and the network performance load corresponding to the target host; A sharding module, configured to divide the data to be processed into multiple data shards according to the data attribute information in the case where the data to be processed is dynamic data, where the dynamic data is used to represent communication service data stored in the target host; A queue adjustment module, configured to determine the queue depth of the data cache queue according to the network performance load of the target host; A data caching module, configured to cache the multiple data shards into a data cache queue having the queue depth.

11. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the data processing method according to any one of claims 1-9 is implemented.

12. A computer-readable storage medium, characterized in that, Computer program instructions are stored on a computer-readable storage medium, and when the computer program instructions are executed by a processor, the data processing method according to any one of claims 1-9 is implemented.

Citation Information

Cited By

  • Dynamic multi-source security system and security method

    CN120729871A

  • Transport layer data packet management method based on fragmentation and recombination

    CN121309497A