Automated balancing deduplication method, system, device and medium based on message backlog

CN115689002BActive Publication Date: 2026-10-09GUANGZHOU XUANWU WIRELESS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211322303.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-27
Publication Date
2026-10-09
Estimated Expiration
2042-10-27

AI Technical Summary

Technical Problem

[0002]目前,随着企业用户量、业务量的不断增长,以往的单台服务器架构已经不能满足业务的增长速度,单台服务器资源的提升成本也越来越贵,而且单台服务器会存在单点故障问题,当此台服务器硬件或者网络发生问题时会导致用户服务彻底不可用,在服务体验越来越严格的今天,服务不可用会导致用户业务的流失

Benefits of technology

[0035] This invention provides an automatic balancing deduplication method, system, device, and medium based on message backlog. Based on the message backlog situation in real-time data, the invention uses a DSP algorithm to predict the trend of message backlog and message receiving speed within a future window period. The distributed deduplication threshold is dynamically adjusted based on the predicted data, and a deduplication strategy is deployed to ensure that message backlog can be effectively processed and submitted for deduplication verification even with elastic processing. This ensures high-performance message submission on the integrated messaging platform while minimizing duplicate message submissions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115689002B_ABST
    Figure CN115689002B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of information processing, and discloses an automatic balance deduplication method, system, device and medium based on message backlog, wherein the method is applied to the interaction process of a fusion message processing platform and an operator channel, and comprises the following steps: when a starting process for processing a to-be-submitted message is deployed in a non-single-node mode, real-time data is acquired, the real-time data is predicted and processed according to a DSP algorithm, and predicted data is acquired; a distributed deduplication threshold is dynamically adjusted according to the predicted data; the current message backlog is compared with the distributed deduplication threshold, if the current message backlog is greater than the distributed deduplication threshold, the to-be-submitted message is submitted to the operator channel in real time, and if the current message backlog is smaller than the distributed deduplication threshold, the to-be-submitted message is subjected to the distributed deduplication processing. The application not only guarantees the high-performance submission of the fusion message platform message, but also ensures that the message is not repeatedly submitted as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information processing technology, and in particular to an automatic balancing and deduplication method, system, device and medium based on message backlog. Background Technology

[0002] Currently, with the continuous growth of enterprise users and business volume, the traditional single-server architecture can no longer meet the needs of business growth. Upgrading single-server resources is becoming increasingly expensive, and single servers suffer from single points of failure. When a server experiences hardware or network problems, user service becomes completely unavailable. In today's increasingly demanding service experience environment, service unavailability can lead to the loss of user business. Therefore, to improve business service processing speed, reduce single-machine costs, minimize the probability of failure, and enhance user experience, multiple servers are needed to collaboratively process the same business. A converged message processing platform is built on a distributed microservice architecture. The system consists of numerous microservice modules, and there are inevitably many network calls between these modules. The platform's core responsibility is to receive messages from the business side, assemble and dispatch them according to internal platform types, and quickly submit them to different types of operator channels. Therefore, the converged message processing platform needs to have extremely fast submission performance, and the channel receiving end requires that the results of a single request or multiple requests for the same message be consistent, avoiding different results due to multiple submissions—in other words, it needs to possess a certain degree of idempotency. Therefore, the platform needs a system that can both ensure high-speed message submission and guarantee that messages are not likely to be submitted repeatedly, meeting the availability requirement. Summary of the Invention

[0003] This invention provides an automatic balancing and deduplication method, system, device, and medium based on message backlog, which enables high-performance message submission from a converged messaging platform while ensuring that messages are not submitted repeatedly as much as possible.

[0004] To achieve the above objectives, the first aspect of the present invention provides an automatic balancing and deduplication method based on message backlog, the method being applied to the interaction process between a converged message processing platform and operator channels, the method comprising:

[0005] When the current process handling messages to be submitted is deployed in a non-single-node manner, real-time data is acquired, and the real-time data is used for prediction processing according to the DSP algorithm to obtain prediction data; wherein, the real-time data includes the current message submission speed, the current message backlog, and the current message receiving speed; the prediction data includes the predicted message submission speed, the predicted message backlog, and the predicted message receiving speed within a future window period, wherein the future window period includes at least the current prediction period and the next prediction period;

[0006] The distributed deduplication threshold is dynamically adjusted based on the predicted data; wherein, the distributed deduplication threshold is the optimal message backlog reached by the current channel when the message to be submitted is subjected to distributed deduplication processing; the distributed deduplication processing is based on Redis and is used to deduplicate the message to be submitted according to its type; the type of message to be submitted includes duplicate messages and non-duplicate messages;

[0007] The current message backlog is compared with the distributed deduplication threshold. If the current message backlog is greater than the distributed deduplication threshold, the message to be submitted is immediately submitted to the operator channel. If the current message backlog is less than the distributed deduplication threshold, the message to be submitted is processed using the distributed deduplication method.

[0008] Further, the message to be submitted is deduplicated according to its type, including:

[0009] The type of the message to be submitted is determined. If it is a duplicate message, the message to be submitted is discarded; if it is a non-duplicate message, the message to be submitted is immediately submitted to the operator channel.

[0010] Furthermore, the real-time data is processed using a DSP algorithm to obtain predicted data, including:

[0011] The real-time data is preprocessed, including filling in missing data and removing abnormal data;

[0012] The main cycle for acquiring preprocessed real-time data;

[0013] The predicted data is obtained by transforming the real-time data within the main cycle using the DSP algorithm.

[0014] Furthermore, the main cycle for acquiring preprocessed real-time data includes:

[0015] By performing a fast Fourier transform on the preprocessed real-time data, a periodogram providing energy at each frequency is obtained. When the frequency energy exceeds a preset threshold, the period corresponding to that frequency is taken as a candidate period.

[0016] Calculate the autocorrelation function of the preprocessed real-time data, and select the candidate period with the largest peak value as the main period on the ACF plot.

[0017] Furthermore, the predicted data is obtained by transforming the real-time data within the main cycle using the DSP algorithm, including:

[0018] The real-time data within the main cycle is converted into frequency domain data by Fast Fourier Transform and then denoised.

[0019] The denoised frequency domain data is converted into time domain data using inverse fast Fourier transform, and used as the prediction data.

[0020] Furthermore, the distributed deduplication threshold is determined using the distributed deduplication threshold calculation formula:

[0021]

[0022] In the formula, T is the distributed deduplication threshold, D is the pacing factor, RS is the predicted message receiving rate for the next prediction period, RS2 is the predicted message receiving rate for the current prediction period, B is the predicted message backlog for the next prediction period, B2 is the predicted message backlog for the current prediction period, D(i) is the initial value of the pacing factor, B(i) is the fixed factor, CEIL is the integer part of the quotient of the predicted message backlog and the fixed factor, and MOD is the remainder of the quotient of the predicted message backlog and the fixed factor.

[0023] Furthermore, dynamically adjusting the distributed deduplication threshold based on the predicted data also includes:

[0024] Compare the trends of the forecast data in the current forecast period and the next forecast period;

[0025] When the predicted message receiving speed and the predicted message backlog of the current channel increase, the distributed deduplication threshold is reduced and the pace is increased to increase the submission speed of the current channel.

[0026] When the predicted message receiving speed of the current channel decreases and the predicted message backlog decreases, the distributed deduplication threshold is increased and the pace is increased so that the submission speed of the current channel decreases steadily until the current message backlog is 0.

[0027] When the predicted message receiving speed and the predicted message backlog of the current channel increase, the distributed deduplication threshold is reduced and the pace is slowed down so as to increase the submission speed of the current channel.

[0028] A second aspect of the present invention provides an automatic message backlog balancing and deduplication system, the system being applied to the interaction process between a converged message processing platform and operator channels, characterized in that the system comprises:

[0029] The data acquisition module is used to acquire real-time data when the current process for processing messages to be submitted is deployed in a non-single-node manner, and to perform predictive processing on the real-time data according to the DSP algorithm to obtain predicted data; wherein, the real-time data includes the current message submission speed, the current message backlog, and the current message receiving speed; the predicted data includes the predicted message submission speed, the predicted message backlog, and the predicted message receiving speed within a future window period, wherein the future window period includes at least the current prediction period and the next prediction period;

[0030] A threshold adjustment module is used to dynamically adjust the distributed deduplication threshold based on the predicted data; wherein, the distributed deduplication threshold is the optimal message backlog reached by the current channel when the message to be submitted is subjected to distributed deduplication processing; the distributed deduplication processing is based on Redis and is used to perform deduplication processing on the message to be submitted according to the type of the message to be submitted; the type of the message to be submitted includes duplicate messages and non-duplicate messages;

[0031] The message deduplication module is used to compare the current message backlog with the distributed deduplication threshold. If the current message backlog is greater than the distributed deduplication threshold, the message to be submitted is immediately submitted to the operator channel; if the current message backlog is less than the distributed deduplication threshold, the message to be submitted is processed by the distributed deduplication.

[0032] A third aspect of the present invention provides an electronic device including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements an automatic balancing and deduplication method based on message backlog as described in any of the first aspects above.

[0033] A fourth aspect of the present invention provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform an automatic balancing and deduplication method based on message backlog as described in any one of the first aspects above.

[0034] Compared with the prior art, the beneficial effects of the embodiments of the present invention are as follows:

[0035] This invention provides an automatic balancing deduplication method, system, device, and medium based on message backlog. Based on the message backlog situation in real-time data, the invention uses a DSP algorithm to predict the trend of message backlog and message receiving speed within a future window period. The distributed deduplication threshold is dynamically adjusted based on the predicted data, and a deduplication strategy is deployed to ensure that message backlog can be effectively processed and submitted for deduplication verification even with elastic processing. This ensures high-performance message submission on the integrated messaging platform while minimizing duplicate message submissions. Attached Figure Description

[0036] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is a diagram illustrating the interaction process between the integrated message processing platform and the operator's channel side of this invention.

[0038] Figure 2 This is a flowchart of an automatic balancing and deduplication method based on message backlog provided in a certain embodiment of the present invention;

[0039] Figure 3 This is a flowchart of the message deduplication process of the integrated message processing platform of this invention;

[0040] Figure 4 This is a data prediction flowchart provided in a certain embodiment of the present invention;

[0041] Figure 5 This is a flowchart of distributed deduplication processing provided in a certain embodiment of the present invention;

[0042] Figure 6 This is a device diagram of an automatic balancing and deduplication system based on message backlog provided in a certain embodiment of the present invention;

[0043] Figure 7 This is a structural diagram of an electronic device provided in a certain embodiment of the present invention. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings and examples. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] It should be understood that the step numbers used in the text are for ease of description only and are not intended to limit the order in which the steps are performed.

[0046] It should be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0047] The terms “comprising” and “including” indicate the presence of the described feature, whole, step, operation, element and / or component, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.

[0048] The term “and / or” refers to any combination of one or more of the associated listed items, as well as all possible combinations, and includes these combinations.

[0049] Explanation of some proper nouns:

[0050] Idempotency: Idempotency is a mathematical concept that means performing the same operation multiple times yields the same result as performing it once. For example, if we decrement the existing inventory by 1 when consuming a message, consuming two identical messages would decrement the inventory by 2, which is not idempotent. However, if the processing logic after consuming a message sets the inventory to 0, or decrements it by 1 if the current inventory is 10, then consuming multiple messages will yield the same result, which is idempotent.

[0051] CAP theorem: The CAP theorem states that a distributed system can only satisfy at most two of the three properties: consistency, availability, and partition tolerance.

[0052] ① Consistency (C): For each read operation by a client, either the latest data is read, or the read fails. In other words, consistency is a promise to clients accessing the system from the perspective of a distributed system: either I return an error, or I return absolutely consistent and up-to-date data. It is easy to see that it emphasizes data correctness.

[0053] ②Availability A: Any client request will receive a response without errors. In other words, availability, from the perspective of a distributed system, is another promise to clients accessing the system: I will definitely return data to you, and I won't return errors, but I don't guarantee the data is up-to-date; the emphasis is on preventing errors.

[0054] ③ Partition tolerance P: Since distributed systems communicate over a network, and the network is unreliable, the system will continue to provide service and will not crash even if any number of messages are lost or delayed. In other words, partition tolerance is another promise from the perspective of the distributed system to clients accessing the system: I will keep running regardless of any data synchronization problems that occur internally; the emphasis is on not crashing.

[0055] For a distributed system, P is a prerequisite and must be guaranteed because any network interaction will inevitably involve latency and data loss, which we must accept and ensure the system does not crash. Therefore, only C and A remain as options. We must either guarantee data consistency (ensuring absolute data correctness) or guarantee availability (ensuring the system does not fail). Thus, a trade-off must be made between C and A.

[0056] When C (Consistency) is selected, which ensures CP, the system will return an error or timeout if certain information cannot be guaranteed to be up-to-date due to network partitions.

[0057] When A (Availability) is selected, which means ensuring AP, the system will always process client queries and attempt to return the latest available version of information, even if it cannot be guaranteed to be up-to-date due to network partitions.

[0058] The converged message processing platform is built on a microservice architecture. The platform requires an availability mechanism that ensures both high-speed message submission and a high probability of non-duplicate submission. Based on the distributed CAP theorem, extremely high submission speed equates to high availability (A in CAP), while ensuring a high probability of non-duplicate submission equates to message consistency (in the event of a network partition, blocking or retransmission until the network connection is restored, ensuring message delivery, C in CAP). As explained in the terminology, CP and AP can only guarantee one pair of satisfactions at a time. However, for platform message submission, the choice between CP and AP is dynamically variable. For example, when the message backlog is small or nonexistent, CP is prioritized to ensure absolute idempotency of message submission; when the message backlog is large, AP should be chosen to prioritize high-performance message submission, ensuring messages are delivered to carrier channels as quickly as possible. Therefore, an internally customized message deduplication solution needs to be developed.

[0059] The interaction process between the integrated messaging platform and operator channels is as follows: Figure 1 As shown, specifically, the business side pushes the messages to be submitted to the MQ for temporary storage; the channel integration module subscribes to the corresponding TOPIC in the MQ, consumes the messages to be submitted, and performs deduplication processing before submission (each message has a unique identifier); after deduplication, it is promptly submitted to the corresponding operator channel.

[0060] In one embodiment, such as Figure 2 As shown, the first aspect of the present invention provides an automatic balancing and deduplication method based on message backlog, which is applied to the interaction process between a converged message processing platform and operator channels, including:

[0061] S1. When the startup process currently processing messages to be submitted is deployed in a non-single-node mode, real-time data is obtained, and the real-time data is predicted according to the DSP algorithm to obtain prediction data; wherein, the real-time data includes the current message submission speed, the current message backlog, and the current message receiving speed; the prediction data includes the predicted message submission speed, the predicted message backlog, and the predicted message receiving speed within the future window period, and the future window period includes at least the current prediction period and the next prediction period;

[0062] Specifically, the message deduplication process of the integrated message processing platform is as follows: Figure 3 As shown, the business side pushes messages to the converged message processing platform. The messages enter the MQ (Message Queue) and wait in the message queue. Subsequently, the channel integration module in the converged message processing platform subscribes to the corresponding TOPIC in the MQ and polls to check if there are any messages to be committed, thus consuming the messages to be committed. When consuming messages to be committed, it is necessary to dynamically select CP (Contribution) and AP (Action) based on the distributed CAP theorem, that is, to dynamically adjust the message deduplication strategy.

[0063] Before deduplicating messages, the deployment mode of the currently running process is determined. If it's a single-node deployment, single-process in-memory deduplication is used, meaning deduplication is performed based on the type of the message to be submitted, which includes duplicate and non-duplicate messages. In a specific embodiment, the type of the message to be submitted is determined; if it's a duplicate message, it's discarded; if it's a non-duplicate message, it's immediately submitted to the operator's channel. Single-process in-memory deduplication, which deduplicates messages by category, is convenient and fast, and while it has high performance, it doesn't support distributed deduplication and cannot be used in distributed architectures.

[0064] When the current process handling messages to be submitted is deployed in a non-single-node manner, real-time data is acquired, and predictive processing is performed on the real-time data using a DSP algorithm to obtain predicted data. Real-time data can be collected from the converged message processing platform via the Prometheus operations and maintenance monitoring platform; this real-time data includes the current message submission rate, the current message backlog, and the current message reception rate.

[0065] When collecting system metrics such as current message submission speed, current message backlog, and current message reception speed for each channel of the Prometheus operations and maintenance monitoring platform, the platform typically uses a fixed sampling frequency, such as every 30 or 60 seconds (currently 60 seconds). This results in a time series of discrete points, stored in the Prometheus operations and maintenance monitoring platform as Timestamp:Value key-value pairs. Based on business monitoring metric data collected from a real production environment, some data snippets are shown below:

[0066] Timestamp, Value

[0067] 1656790000, 562

[0068] 1656790060, 594

[0069] 1656790120, 685

[0070] Since most regular online services exhibit regular tidal fluctuations in their time series, the collected business monitoring indicator data, after being processed based on time series and discrete points, has the following main characteristics:

[0071] ①Limited length: Whether it is data from a day, a month, or a year, it is of finite length.

[0072] ② Data discreteness: Indicator collection is usually done at fixed intervals. For example, the default configuration of Prometheus scrape_interval is one minute, which means data is collected once per minute, resulting in 1440 discrete data points per day.

[0073] ③ Cyclical fluctuations: Business volume exhibits cyclicality, with distinct peak and trough characteristics at different times within a cycle.

[0074] The load on a converged messaging platform exhibits tidal characteristics. Looking at metrics like traffic and message backlog over time reveals distinct peaks and troughs. Further observation reveals that these volatile services naturally exhibit periodicity over time, especially those directly or indirectly serving people. This periodicity is determined by people's work patterns. For example, office workers often send notification messages at the start of the workday (9-10 AM is typically the peak message traffic); after get off work, they stop sending messages and block some, resulting in significantly lower request volumes at night. For such services, the solution is to predict future data based on past data. With predicted data (e.g., message submission volume for the next hour), we can implement proactive strategies for message deduplication.

[0075] Specifically, for complex time series variations, this application employs a DSP algorithm to perform predictive processing on real-time data to obtain predictive data. This predictive data includes the predicted message submission rate, predicted message backlog, and predicted message reception rate within a future window period, which includes at least the current prediction period and the next prediction period. During the operation of the fused message processing platform, the Prometheus operations and monitoring platform collects time-domain data, and the fused message processing platform extracts features from these observations and uses these features to make predictions.

[0076] In a specific embodiment, the data prediction process is as follows: Figure 4 As shown, it includes:

[0077] S11. Preprocess the real-time data. Preprocessing includes filling in missing data and removing outlier data. Preprocessing the real-time data can ensure the diversity and accuracy of the data.

[0078] S12, Obtain the main cycle of preprocessed real-time data;

[0079] In a specific embodiment, a periodogram providing energy at each frequency is first obtained by performing a fast Fourier transform on the preprocessed real-time data. When the frequency energy exceeds a preset threshold, the period corresponding to the frequency is taken as a candidate period. Then, the autocorrelation function of the preprocessed real-time data is calculated, and the candidate period with the largest peak value is taken as the main period on the ACF graph to determine the main period of the preprocessed real-time data.

[0080] Specifically, the acquisition of the principal period includes two stages: First, a periodogram is obtained by performing a Fast Fourier Transform on the preprocessed real-time data (let's say the length is N). This periodogram provides the energy at each frequency k / N. When the frequency energy exceeds a preset threshold, the corresponding period T = N / k (e.g., N = 8d, k = 8, then T = 1d) is selected as a candidate period. Second, the Autocorrelation Function (ACF) of the preprocessed real-time data is calculated. The ACF is the cross-correlation between a signal and itself at different time points; simply put, it's a function of the similarity between two observations and the time difference between them. If M is the period of a sequence, its ACF value at point M must be a local high. Based on this characteristic, the candidate periods obtained in the first stage are confirmed on the ACF plot, and finally, the point located at the "highest peak" is selected as the principal period of the sequence (i.e., the fundamental period). That is, the candidate period with the largest peak value on the ACF plot is selected as the principal period. The concept of the principal period refers to the projection of a signal's frequency domain into the time domain in a DSP algorithm. A periodic signal cycles continuously along the time axis with a fundamental frequency period. During one revolution of the fundamental frequency signal around the complex plane, there are n sampling points, with each sampling interval being exactly the same. Summing these points according to the sampling order yields the characteristic information for each frequency. Determining the principal period helps extract frequency characteristics from real-time data, which can then be used for predicting data within future time windows.

[0081] S13. The real-time data within the main cycle is transformed and processed using DSP algorithms to obtain the prediction data;

[0082] In a specific embodiment, real-time time-series data collected by Prometheus within a predetermined time window is processed using a DSP algorithm to perform time-domain and frequency-domain transformation to obtain prediction data. Specifically, real-time data within the main period is converted to frequency-domain data using a Fast Fourier Transform (FFT) and then denoised. The denoised frequency-domain data is then converted back to time-domain data using an Inverse Fast Fourier Transform (IFT) as the prediction data. This prediction data serves as an indicator for analyzing the message backlog within future window periods and adjusting the distributed deduplication threshold accordingly.

[0083] S2. Dynamically adjust the distributed deduplication threshold based on the predicted data; whereby the distributed deduplication threshold is the optimal message backlog reached by the current channel when performing distributed deduplication processing on the messages to be submitted; the distributed deduplication processing is based on Redis and is used to perform deduplication processing on the messages to be submitted according to their type; the types of messages to be submitted include duplicate messages and non-duplicate messages;

[0084] Specifically, the platform needs to dynamically balance performance and data uniqueness, allowing for a small probability of multiple submissions. Therefore, it needs to dynamically adjust the distributed deduplication threshold in advance based on predicted data. Since the current deployment method for processing messages to be submitted is not single-node, a distributed message deduplication verification process is required. However, if every message needs to be deduplicated and checked in the distributed cache, frequent network access will result in significant bandwidth overhead and extremely low performance, which cannot meet actual needs. Therefore, this application dynamically adjusts the distributed deduplication threshold based on predicted data, and then dynamically adjusts the deduplication strategy based on the comparison between the current message backlog and the distributed deduplication processing threshold. The distributed deduplication threshold is the optimal message backlog reached by the current channel when distributed deduplication processing of messages to be submitted is required.

[0085] Specifically, the distributed deduplication process is as follows: Figure 5 As shown, the distributed caching component Redis provides a Set data structure to store unique data sets. It allows for quick determination of whether an element exists in the set. The command used in this Redis Set structure is `SADD`. `SADD key msgID`: Adds one or more `msgID` elements to the set `key`. `msgID` elements already existing in the set are ignored. If `key` does not exist, a set containing only `msgID` elements is created. If the `member` element is not in the set, it returns 1, meaning the message is a non-duplicate message; if the `member` element already exists in the set, it returns 0, meaning the message is a duplicate message. Each message has a unique identifier. When submitted to the distributed cache, if it doesn't exist, it exists in the Set; if it already exists, it returns a response indicating that it already exists.

[0086] In one embodiment, the distributed deduplication threshold is determined using a distributed deduplication threshold calculation formula:

[0087]

[0088] In the formula, T is the distributed deduplication threshold, D is the pacing factor, RS is the predicted message receiving rate for the next prediction period, RS2 is the predicted message receiving rate for the current prediction period, B is the predicted message backlog for the next prediction period, B2 is the predicted message backlog for the current prediction period, D(i) is the initial value of the pacing factor, B(i) is the fixed factor (10 to the power of i, with the number of decimal places determined by i), CEIL is the integer part of the quotient of the predicted message backlog and the fixed factor, and MOD is the remainder of the quotient of the predicted message backlog and the fixed factor. Determining the distributed deduplication threshold using this formula can avoid performance degradation due to untimely threshold adjustments, thus preventing failures caused by large message backlogs and delayed submissions.

[0089] In one embodiment, the distributed deduplication threshold can be dynamically adjusted based on the predicted data by comparing the trend of the predicted data. Specifically:

[0090] By comparing the trend of the predicted data in the current prediction period and the next prediction period: when the predicted message receiving speed and the predicted message backlog of the current channel increase, the distributed deduplication threshold is decreased and the pace is increased to increase the submission speed of the current channel; when the predicted message receiving speed and the predicted message backlog of the current channel decrease, the distributed deduplication threshold is increased and the pace is increased to ensure that the submission speed of the current channel decreases steadily until the current message backlog reaches 0; when the predicted message receiving speed and the predicted message backlog of the current channel increase, the distributed deduplication threshold is decreased and the pace is decreased to increase the submission speed of the current channel. Adjusting the distributed deduplication threshold through trend comparison analysis of the predicted data, thereby deduplicating messages to be submitted, can improve message processing efficiency and prevent excessive message backlog from damaging the server.

[0091] Optionally, the distributed deduplication threshold can be adjusted based on correlation analysis. Specifically: when the value of (RS-RS2)*(B-B2) is positive, the pace factor is negatively correlated with the distributed deduplication threshold; the submission speed of the current channel is negatively correlated with the distributed deduplication threshold; and the predicted message backlog is negatively correlated with the distributed deduplication threshold.

[0092] S3. Compare the current message backlog with the distributed deduplication threshold. If the current message backlog is greater than the distributed deduplication threshold, the message to be submitted will be submitted to the operator channel immediately. If the current message backlog is less than the distributed deduplication threshold, the message to be submitted will be processed for distributed deduplication.

[0093] Specifically, the current message backlog is compared with a distributed deduplication threshold, and the messages to be submitted are processed based on the comparison result. If the current message backlog is greater than the distributed deduplication threshold, the messages to be submitted are directly submitted to the operator's channel immediately; if the current message backlog is less than the distributed deduplication threshold, distributed deduplication processing is performed on the messages to be submitted. This application automatically deduplicates messages based on message submission speed and message backlog, making it possible for the converged message processing platform to meet the availability requirement of ensuring both high-speed message submission and a high probability of messages not being submitted repeatedly.

[0094] The advantages and disadvantages of this application compared with existing technologies are shown in the table below:

[0095]

[0096] To ensure high-performance message submission on the converged messaging platform while minimizing duplicate submissions, this application provides an automatic balancing deduplication method based on message backlog. This method is applied to the interaction between the converged messaging platform and operator channels. It includes: when the current message processing process is deployed in a non-single-node manner, acquiring real-time data and performing predictive processing on the real-time data using a DSP algorithm to obtain predicted data; dynamically adjusting the distributed deduplication threshold based on the predicted data; comparing the current message backlog with the distributed deduplication threshold; if the current message backlog is greater than the distributed deduplication threshold, immediately submitting the message to the operator channel; if the current message backlog is less than the distributed deduplication threshold, performing distributed deduplication processing on the message to be submitted. This invention, based on message submission speed and message backlog, uses a DSP algorithm to predict the message backlog trend within a future window period. It can detect and adjust the deduplication strategy in advance before the preset threshold is reached, ensuring that message backlog can be effectively processed using elastic handling for submission and deduplication verification. This ensures high-performance message submission on the converged messaging platform while minimizing duplicate submissions.

[0097] It should be noted that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order requirement for the execution of these steps, and they can be executed in other orders.

[0098] In another embodiment, such as Figure 6As shown, the second aspect of the present invention provides an automatic message backlog balancing and deduplication system, applied to the interaction process between a converged message processing platform and operator channels, comprising: a data acquisition module 10, used to acquire real-time data when the current process for processing messages to be submitted is deployed in a non-single-node manner, and to perform predictive processing on the real-time data according to a DSP algorithm to acquire predicted data; wherein, the real-time data includes the current message submission speed, the current message backlog, and the current message receiving speed; the predicted data includes the predicted message submission speed, the predicted message backlog, and the predicted message receiving speed within a future window period, the future window period including at least the current prediction period and the next prediction period; and a threshold adjustment module. 20 is used to dynamically adjust the distributed deduplication threshold based on predicted data; where the distributed deduplication threshold is the optimal message backlog reached by the current channel when performing distributed deduplication processing on the messages to be submitted; the distributed deduplication processing is based on Redis and is used to perform deduplication processing on the messages to be submitted according to their type; the types of messages to be submitted include duplicate messages and non-duplicate messages; the message deduplication module 30 is used to compare the current message backlog with the distributed deduplication threshold. If the current message backlog is greater than the distributed deduplication threshold, the messages to be submitted are immediately submitted to the operator channel; if the current message backlog is less than the distributed deduplication threshold, the messages to be submitted are processed using distributed deduplication.

[0099] In this embodiment of the application, to ensure high-performance message submission by the converged messaging platform while minimizing duplicate message submissions, an automatic message backlog balancing and deduplication system based on message backlog is designed. This system is applied to the interaction between the converged messaging platform and operator channels, and includes a data acquisition module, a threshold adjustment module, and a message deduplication module. When the message backlog is small or nonexistent, the converged messaging platform prioritizes CP (Content Processing) to ensure absolute idempotency of message submission; when the message backlog is large, AP (Action Processing) should be selected to prioritize high-performance message submission and deliver messages to the operator channel side as quickly as possible. This improves message submission speed and performance while reducing failures caused by excessive message backlog.

[0100] It should be noted that the modules in the aforementioned automatic message backlog balancing deduplication system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module. For specific limitations regarding an automatic message backlog balancing deduplication system, please refer to the limitations of an automatic message backlog balancing deduplication method described above; both have the same function and role, and will not be repeated here.

[0101] A third aspect of the present invention provides an electronic device comprising:

[0102] Processor, memory, and bus;

[0103] The bus is used to connect the processor and the memory;

[0104] The memory is used to store operation instructions;

[0105] The processor is configured to execute an operation corresponding to an automatic balancing deduplication method based on message backlog as shown in the first aspect of this application by invoking the operation instructions.

[0106] In one alternative embodiment, an electronic device is provided, such as Figure 7 As shown, Figure 7 The illustrated electronic device 5000 includes a processor 5001 and a memory 5003. The processor 5001 and the memory 5003 are connected, for example, via a bus 5002. Optionally, the electronic device 5000 may also include a transceiver 5004. It should be noted that in practical applications, the transceiver 5004 is not limited to one type, and the structure of this electronic device 5000 does not constitute a limitation on the embodiments of this application.

[0107] Processor 5001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 5001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0108] Bus 5002 may include a path for transmitting information between the aforementioned components. Bus 5002 may be a PCI bus or an EISA bus, etc. Bus 5002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 7 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0109] The memory 5003 may be a ROM or other type of static storage device capable of storing static information and instructions, RAM or other type of dynamic storage device capable of storing information and instructions, or it may be an EEPROM, CD-ROM or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.

[0110] The memory 5003 is used to store application code that executes the scheme of this application, and its execution is controlled by the processor 5001. The processor 5001 is used to execute the application code stored in the memory 5003 to implement the content shown in any of the foregoing method embodiments.

[0111] Among them, electronic devices include, but are not limited to: mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers.

[0112] The fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements an automatic balancing and deduplication method based on message backlog as shown in the first aspect of the present application.

[0113] Another embodiment of this application provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.

[0114] Furthermore, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0115] In summary, this invention discloses an automatic message backlog balancing and deduplication method, system, device, and medium. The method is applied to the interaction process between a converged message processing platform and operator channels, including: when the current process handling messages to be submitted is deployed in a non-single-node manner, acquiring real-time data and performing predictive processing on the real-time data using a DSP algorithm to obtain predicted data; dynamically adjusting the distributed deduplication threshold based on the predicted data; comparing the current message backlog with the distributed deduplication threshold; if the current message backlog is greater than the distributed deduplication threshold, immediately submitting the message to be submitted to the operator channel; if the current message backlog is less than the distributed deduplication threshold, performing distributed deduplication processing on the message to be submitted. This invention ensures high-performance message submission from the converged message platform while minimizing duplicate message submissions.

[0116] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

[0117] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.

Claims

1. An automatic balancing and deduplication method based on message backlog, the method being applied to the interaction process between a converged message processing platform and operator channels, characterized in that... The method includes: When the current process handling messages to be submitted is deployed in a non-single-node manner, real-time data is acquired, and the real-time data is used for prediction processing according to the DSP algorithm to obtain prediction data; wherein, the real-time data includes the current message submission speed, the current message backlog, and the current message receiving speed; the prediction data includes the predicted message submission speed, the predicted message backlog, and the predicted message receiving speed within a future window period, wherein the future window period includes at least the current prediction period and the next prediction period; The distributed deduplication threshold is dynamically adjusted based on the predicted data; wherein, the distributed deduplication threshold is the optimal message backlog reached by the current channel when the message to be submitted is subjected to distributed deduplication processing; the distributed deduplication processing is based on Redis and is used to deduplicate the message to be submitted according to its type; the type of message to be submitted includes duplicate messages and non-duplicate messages; The current message backlog is compared with the distributed deduplication threshold. If the current message backlog is greater than the distributed deduplication threshold, the message to be submitted is immediately submitted to the operator channel. If the current message backlog is less than the distributed deduplication threshold, the message to be submitted is processed by the distributed deduplication method. The step of performing prediction processing on the real-time data according to the DSP algorithm to obtain prediction data includes: The real-time data is preprocessed, including filling in missing data and removing abnormal data; The main cycle for acquiring preprocessed real-time data; The predicted data is obtained by transforming the real-time data within the main cycle using the DSP algorithm. The distributed deduplication threshold is determined using the distributed deduplication threshold calculation formula: In the formula, T is the distributed deduplication threshold, D is the pacing factor, RS is the prediction message receiving rate for the next prediction period, RS2 is the prediction message receiving rate for the current prediction period, B is the prediction message backlog for the next prediction period, and B2 is the prediction message backlog for the current prediction period. Initialize the pace factor to a value. CEIL is the integer part of the quotient of the predicted message backlog and the fixed factor, and MOD is the remainder part of the quotient of the predicted message backlog and the fixed factor.

2. The automatic balancing and deduplication method based on message backlog as described in claim 1, characterized in that, The step of deduplicating the message to be submitted according to its type includes: The type of the message to be submitted is determined. If it is a duplicate message, the message to be submitted is discarded; if it is a non-duplicate message, the message to be submitted is immediately submitted to the operator channel.

3. The automatic balancing and deduplication method based on message backlog as described in claim 1, characterized in that, The main cycle for acquiring preprocessed real-time data includes: By performing a fast Fourier transform on the preprocessed real-time data, a periodogram providing energy at each frequency is obtained. When the frequency energy exceeds a preset threshold, the period corresponding to that frequency is taken as a candidate period. Calculate the autocorrelation function of the preprocessed real-time data, and select the candidate period with the largest peak value as the main period on the ACF plot.

4. The automatic balancing and deduplication method based on message backlog as described in claim 1, characterized in that, The step of transforming the real-time data within the main cycle using the DSP algorithm to obtain the predicted data includes: The real-time data within the main cycle is converted into frequency domain data by Fast Fourier Transform and then denoised. The denoised frequency domain data is converted into time domain data using inverse fast Fourier transform, and used as the prediction data.

5. The automatic balancing and deduplication method based on message backlog as described in claim 1, characterized in that, The step of dynamically adjusting the distributed deduplication threshold based on the predicted data further includes: Compare the trends of the forecast data in the current forecast period and the next forecast period; When the predicted message receiving speed and the predicted message backlog of the current channel increase, the distributed deduplication threshold is reduced and the pace is increased to increase the submission speed of the current channel. When the predicted message receiving speed of the current channel decreases and the predicted message backlog decreases, the distributed deduplication threshold is increased and the pace is increased so that the submission speed of the current channel decreases steadily until the current message backlog is 0. When the predicted message receiving speed and the predicted message backlog of the current channel increase, the distributed deduplication threshold is reduced and the pace is slowed down so as to increase the submission speed of the current channel.

6. An automatic message backlog balancing and deduplication system, the system being applied to the interaction process between a converged message processing platform and operator channels, characterized in that... The system includes: The data acquisition module is used to acquire real-time data when the current process for processing messages to be submitted is deployed in a non-single-node manner, and to perform predictive processing on the real-time data according to the DSP algorithm to obtain predicted data; wherein, the real-time data includes the current message submission speed, the current message backlog, and the current message receiving speed; the predicted data includes the predicted message submission speed, the predicted message backlog, and the predicted message receiving speed within a future window period, wherein the future window period includes at least the current prediction period and the next prediction period; A threshold adjustment module is used to dynamically adjust the distributed deduplication threshold based on the predicted data; wherein, the distributed deduplication threshold is the optimal message backlog reached by the current channel when the message to be submitted is subjected to distributed deduplication processing; the distributed deduplication processing is based on Redis and is used to perform deduplication processing on the message to be submitted according to the type of the message to be submitted; the type of the message to be submitted includes duplicate messages and non-duplicate messages; The message deduplication module is used to compare the current message backlog with the distributed deduplication threshold. If the current message backlog is greater than the distributed deduplication threshold, the message to be submitted is immediately submitted to the operator channel; if the current message backlog is less than the distributed deduplication threshold, the message to be submitted is processed by the distributed deduplication. The step of performing prediction processing on the real-time data according to the DSP algorithm to obtain prediction data includes: The real-time data is preprocessed, including filling in missing data and removing abnormal data; The main cycle for acquiring preprocessed real-time data; The predicted data is obtained by transforming the real-time data within the main cycle using the DSP algorithm. The distributed deduplication threshold is determined using the distributed deduplication threshold calculation formula: In the formula, T is the distributed deduplication threshold, D is the pacing factor, RS is the prediction message receiving rate for the next prediction period, RS2 is the prediction message receiving rate for the current prediction period, B is the prediction message backlog for the next prediction period, and B2 is the prediction message backlog for the current prediction period. Initialize the pace factor to a value. CEIL is the integer part of the quotient of the predicted message backlog and the fixed factor, and MOD is the remainder part of the quotient of the predicted message backlog and the fixed factor.

7. An electronic device, characterized in that, The system includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the automatic balancing and deduplication method based on message backlog as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the automatic balancing and deduplication method based on message backlog as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Distributed message equalization processing method and device, electronic equipment and storage medium

    CN108769162A

  • Message queue-based data processing method and device, computer equipment and medium

    CN112612607A