Quantitative transaction data real-time distribution method and device based on cloud transmission

By clustering and differential data compression of financial transaction data streams, the problem of excessive compression time caused by traditional encoding is solved, and efficient data transmission is achieved.

CN120416348AActive Publication Date: 2025-08-01CHONGQING QUARK QUANTUM NETWORK TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510905293.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The prior art in financial transactions has caused the compression time to be too long due to the direct encoding of all data, which increases the data transmission delay and affects the real-time nature of data transmission.

Method used

By performing cluster analysis on real-time transaction data streams, the data cluster center and individual difference data are determined, and only the cluster center and individual difference data are compressed to generate and distribute the compressed data stream.

Benefits of technology

This greatly reduces the amount of compressed data, improves compression efficiency, reduces transmission delay, and ensures real-time data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416348A_ABST
    Figure CN120416348A_ABST
Patent Text Reader

Abstract

The invention discloses a quantitative transaction data real-time distribution method and device based on cloud transmission, and the method comprises the steps: carrying out the clustering analysis of a real-time transaction data flow during compression, and dividing the real-time transaction data flow into a plurality of data class clusters; then, individual difference data between all the class cluster centers and all the remaining transaction data in the corresponding data class clusters are determined, and then the multiple data class clusters are simplified into an expression form of the class cluster centers and the individual difference data; then, the class cluster center and the individual difference data are compressed, so that a compressed data stream corresponding to the real-time transaction data stream can be obtained; and finally, the compressed data stream is sent to each transaction device, and real-time distribution of the transaction data can be completed. Therefore, only the common part and the difference part in the transaction data are compressed, so that the compressed data volume is greatly reduced, and therefore, the compression efficiency can be improved, and the transmission delay caused by overlong compression time is further reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data transmission, and particularly relates to a real-time distribution method and device for quantitative trading data based on cloud transmission. Background Art

[0002] With the rapid development of Internet technology, big data processing technology is required in more and more fields. Especially in the financial industry, due to the large number of categories such as stocks, funds, bonds, and wealth management in financial transactions, in order to enable customers to grasp market information in real time, it is necessary to send market transaction data of many categories to the client in real time. And because data transmission has high timeliness requirements, and with the rapid growth of the number of products in each category, the amount of market transaction data to be sent each time is getting larger and larger. Traditional technologies directly send a large amount of data to the client. This data transmission method faces problems such as bandwidth bottlenecks and delays. And effective data compression algorithms can significantly reduce the data volume, alleviate these problems, and improve transmission efficiency and resource utilization.

[0003] Currently, traditional lossless and lossy compression technologies directly encode all data. Although they can achieve data compression and reduce transmission delay, due to the increasing amount of market transaction data to be sent each time, the compression time is getting longer and longer. In this way, it will still increase the transmission delay of transaction data, thus affecting the real-time nature of data transmission. Therefore, how to provide a real-time distribution method for quantitative trading data with high compression efficiency to ensure the real-time nature of data transmission has become an urgent problem to be solved. Summary of the Invention

[0004] The purpose of the present invention is to provide a real-time distribution method and device for quantitative trading data based on cloud transmission, so as to solve the problem that the existing technology directly encodes all data, resulting in too long compression time and further increasing the data transmission delay.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, a real-time distribution method for quantitative trading data based on cloud transmission is provided, including: Obtain a real-time trading data stream; Perform clustering processing on each trading data in the real-time trading data stream to obtain at least one data cluster; Determine the cluster center of each data cluster, and determine the individual difference data between each cluster center and the remaining trading data in the corresponding data cluster, as well as the difference positions corresponding to each individual difference data; Perform data compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center; Using each compressed cluster center, the compressed individual difference data corresponding to each cluster center, and the difference positions corresponding to each individual difference data, a compressed data stream corresponding to the real-time transaction data stream is formed; The compressed data stream is distributed to each trading device to complete the real-time distribution of the real-time transaction data stream.

[0006] Based on the above-disclosed content, after obtaining the real-time transaction data stream, the present invention first performs clustering processing on the transaction data in the real-time transaction data stream to obtain at least one data cluster; then, determines the cluster centers of each data cluster, and determines the individual difference data between each cluster center and the remaining transaction data in the corresponding data cluster, as well as the difference positions corresponding to each individual difference data; then, performs compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain each compressed cluster center and compressed individual difference data; then, using the difference positions of the aforementioned individual difference data, the compressed cluster centers, and the compressed individual difference data, a compressed data stream can be formed; finally, the compressed data stream is sent to each trading device to complete the real-time distribution of the real-time transaction data stream.

[0007] Through the above design, when performing compression, the present invention performs clustering analysis on the real-time transaction data stream, thereby dividing the real-time transaction data stream into multiple data clusters; then, determines the individual difference data between each cluster center and the remaining transaction data in the corresponding data cluster, and further simplifies the multiple data clusters into a representation form of a cluster center plus individual difference data; then, performs compression on the cluster center and the individual difference data to obtain a compressed data stream corresponding to the real-time transaction data stream; finally, sends the compressed data stream to each trading device to complete the real-time distribution of the transaction data; thus, the present invention only compresses the common part and the difference part in the transaction data, thereby greatly reducing the amount of compressed data. Based on this, the compression efficiency can be improved, and further the transmission delay caused by too long compression time can be reduced. Therefore, the present invention is very suitable for large-scale application and promotion in the field of data transmission.

[0008] In a possible design, performing clustering processing on each transaction data in the real-time transaction data stream to obtain at least one data cluster includes: Calculating the data distance between each transaction data in the real-time transaction data stream; Determining a plurality of initial clustering centers according to the data distance between each transaction data; Using a clustering optimization algorithm to perform optimization processing on the plurality of initial clustering centers to obtain a plurality of optimal initial clustering centers of the real-time transaction data stream; Based on multiple optimal initial clustering centers, perform clustering processing on the real-time transaction data stream to obtain multiple initial data clusters; Perform clustering correction processing on the multiple initial data clusters to obtain the at least one data cluster.

[0009] In a possible design, determine multiple initial clustering centers according to the data distances between each transaction data, including: Use the data distances between each transaction data to construct a distance matrix, and sort each data distance in the distance matrix in ascending order to obtain a distance sequence; Determine a distance threshold based on the distance sequence; For any transaction data in the real-time transaction data stream, calculate the local density of the any transaction data according to the distance threshold and the data distances between the any transaction data and each specified data in the specified data set, and after polling all the transaction data in the real-time transaction data stream, obtain the local densities of each transaction data, where each specified data in the specified data set is each transaction data in the real-time transaction data stream except the any transaction data; For the any transaction data, screen out at least one target data from the real-time transaction data stream, where the local density of any target data is greater than the local density of the any transaction data; Screen out the minimum data distance from the data distances between each target data and the any transaction data as the calibration distance corresponding to the any transaction data, and after polling all the transaction data in the real-time transaction data stream, obtain the calibration distances corresponding to each transaction data; Use the local densities and calibration distances of each transaction data to determine multiple initial clustering centers from each transaction data.

[0010] In a possible design, use a clustering optimization algorithm to perform optimization processing on multiple initial clustering centers to obtain multiple optimal initial clustering centers of the real-time transaction data stream, including: Based on multiple initial clustering centers, construct a bat population, where each bat individual in the bat population corresponds to an initial position, an initial velocity, an initial pulse loudness, and an initial pulse emission frequency, and the initial position of any bat individual corresponds to an initial clustering center; Initialize the optimization times t, and perform clustering processing on the real-time transaction data stream based on the positions of each bat individual at the t-th optimization to obtain the clustering clusters at the t-th optimization, where each bat individual corresponds to a clustering cluster, the initial value of t is 1, and when t is 1, the position of any bat individual at the t-th optimization is the initial position of the any bat individual; Based on the clustering clusters at the t-th optimization, calculate the fitness of each bat individual at the t-th optimization. Among them, the greater the fitness of any bat individual, the higher the clustering accuracy of the clustering cluster corresponding to the bat individual. Judge whether the fitness of each bat individual at the t-th optimization is greater than the historical optimal fitness of each bat individual. If not, determine the maximum fitness, search weight factor, and pulse frequency at the t-th optimization. According to the maximum fitness, search weight factor, and pulse frequency, update the velocity of each bat individual at the t-th optimization to obtain the updated velocity corresponding to each bat individual, and use the updated velocity corresponding to each bat individual to update the position of the bat individual at the t-th optimization to obtain the updated position corresponding to each bat individual. Among them, when t is 1, the velocity of any bat individual at the t-th optimization is the initial velocity corresponding to the bat individual. Use the pulse emission frequency of each bat individual at the t-th optimization to perturb the updated position corresponding to each bat individual to obtain the perturbed position corresponding to each bat individual. Among them, when t is 1, the pulse emission frequency of any bat individual at the t-th optimization is the initial pulse emission frequency corresponding to the bat individual. Based on the pulse loudness of each bat individual at the t-th optimization, judge whether to retain the perturbed position corresponding to each bat individual, and when t is 1, the pulse loudness of any bat individual at the t-th optimization is the initial pulse loudness corresponding to the bat individual. If so, increment t by 1, and use the perturbed position corresponding to each bat individual as the position of each bat individual at the t-th optimization, and re-cluster the real-time transaction data stream based on the position of each bat individual at the t-th optimization until the fitness of each bat individual at the t-th optimization is greater than the historical optimal fitness of each bat individual. Then, determine multiple optimal initial clustering centers of the real-time transaction data stream according to the position of each bat individual at the t-th optimization.

[0011] In a possible design, determining the search weight factor at the t-th optimization includes: Determine the search weight factor at the t-th optimization according to the following formula (1); (1) In the above formula (1), represents the search weight factor at the t-th optimization, represents the maximum number of optimizations, successively represent the maximum search weight and the minimum search weight, and represents the weight exponent; Correspondingly, using the pulse emission frequencies of each bat individual during the t-th optimization, the updated positions corresponding to each bat individual are perturbed to obtain the perturbed positions corresponding to each bat individual, which includes: For any bat individual, generate a first random number and determine whether the first random number is greater than the pulse emission frequency of the any bat individual during the t-th optimization; If so, obtain the maximum perturbation factor, the minimum perturbation factor, and the total number of individuals in the bat population; According to the maximum perturbation factor, the minimum perturbation factor, and the total number of individuals, calculate the perturbation coefficient during the t-th optimization; Using the perturbation coefficient during the t-th optimization, perturb the updated position corresponding to the any bat individual to obtain the perturbed position corresponding to the any bat individual.

[0012] In a possible design, calculating the perturbation coefficient during the t-th optimization according to the maximum perturbation factor, the minimum perturbation factor, and the total number of individuals includes: Calculate the perturbation coefficient during the t-th optimization according to the following formula (2); (2) In the above formula (2), represents the perturbation coefficient during the t-th optimization, successively represent the maximum perturbation factor and the minimum perturbation factor, represents the total number of individuals.

[0013] In a possible design, performing clustering correction processing on multiple initial data clusters to obtain the at least one data cluster includes: According to each initial data cluster, screen out misclassified data from the real-time transaction data stream; Use all the screened misclassified data to form a set to be classified; Delete all misclassified data from the real-time transaction data stream, and use the remaining transaction data to form a correctly classified set; For any misclassified data in the set to be classified, determine the transaction data closest to the any misclassified data from the correctly classified set as the nearest neighbor data of the any misclassified data; According to the nearest neighbor data of the any misclassified data, reclassify the any misclassified data, and after polling all the misclassified data in the set to be classified, complete the clustering correction processing of multiple initial data clusters to obtain the at least one data cluster.

[0014] In a possible design, data compression processing is performed on each cluster center and the individual difference data corresponding to each cluster center to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center, including: Performing sampling compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain the sampled compression data corresponding to each cluster center and the sampled compression individual difference data corresponding to each cluster center; Performing lossless compression processing on the sampled compression data corresponding to each cluster center and the sampled compression individual difference data corresponding to each cluster center, so as to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center after the lossless compression processing.

[0015] In a second aspect, a real-time distribution device for quantitative trading data based on cloud transmission is provided, including: An acquisition unit, configured to acquire a real-time trading data stream; A clustering unit, configured to perform clustering processing on each trading data in the real-time trading data stream to obtain at least one data cluster; A compression unit, configured to determine the cluster center of each data cluster, and determine the individual difference data between each cluster center and the remaining each trading data in the corresponding data cluster, as well as the difference positions corresponding to each individual difference data; A compression unit, configured to perform data compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center; The compression unit is further configured to use each compressed cluster center, the compressed individual difference data corresponding to each cluster center, and the difference positions corresponding to each individual difference data to form a compressed data stream corresponding to the real-time trading data stream; A distribution unit, configured to distribute the compressed data stream to each trading device to complete the real-time distribution of the real-time trading data stream.

[0016] In a third aspect, another real-time distribution device for quantitative trading data based on cloud transmission is provided. Taking the device as an electronic device as an example, it includes a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the real-time distribution method for quantitative trading data based on cloud transmission as described in the first aspect or any possible design in the first aspect.

[0017] Fourthly, a storage medium is provided, on which instructions are stored. When the instructions run on a computer, they execute the real-time distribution method of quantitative trading data based on cloud transmission as described in the first aspect or any possible design in the first aspect.

[0018] Fifthly, a computer program product containing instructions is provided. When the instructions run on a computer, the computer is made to execute the real-time distribution method of quantitative trading data based on cloud transmission as described in the first aspect or any possible design in the first aspect.

[0019] Beneficial effects: When the present invention performs compression, it conducts clustering analysis on the real-time trading data stream, thereby dividing the real-time trading data stream into multiple data clusters; then, it determines the individual difference data between each cluster center and the remaining trading data in the corresponding data cluster, and further simplifies the multiple data clusters into a representation form of cluster center plus individual difference data; then, by compressing the cluster center and individual difference data, the compressed data stream corresponding to the real-time trading data stream can be obtained; finally, by sending the compressed data stream to each trading device, the real-time distribution of trading data can be completed; thus, the present invention only compresses the common part and the difference part in the trading data, thereby greatly reducing the amount of compressed data. Based on this, the compression efficiency can be improved, and further the transmission delay caused by too long compression time can be reduced. Therefore, the present invention is very suitable for large-scale application and promotion in the field of data transmission. Description of the drawings

[0020] Figure 1 It is a schematic flow chart of the steps of the real-time distribution method of quantitative trading data based on cloud transmission provided by an embodiment of the present invention; Figure 2 It is a schematic structural diagram of the real-time distribution device of quantitative trading data based on cloud transmission provided by an embodiment of the present invention; Figure 3 It is a schematic structural diagram of the electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the drawings and the descriptions of the embodiments or the prior art. Obviously, the following descriptions of the structures of the drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. It should be noted here that the descriptions of these embodiments are used to help understand the present invention, but do not constitute a limitation to the present invention.

[0022] It should be understood that although terms such as first and second may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, the first unit may be referred to as the second unit, and similarly, the second unit may be referred to as the first unit, without departing from the scope of the exemplary embodiments of the present invention.

[0023] It should be understood that for the term "and / or" that may appear herein, it is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and both A and B exist simultaneously; for the term " / and" that may appear herein, it is a description of another association object relationship, indicating that two relationships may exist. For example, A / and B may represent: A exists alone, and both A and B exist; in addition, for the character " / " that may appear herein, generally it represents that the front and rear associated objects are in an "or" relationship.

[0024] Embodiment: See Figure 1 As shown, for the real-time distribution method of quantitative trading data based on cloud transmission provided in this embodiment, when compressing, it performs clustering analysis on the real-time trading data stream, and determines the individual difference data between the data clusters obtained by clustering and each remaining trading data in the cluster. In this way, the real-time trading data stream can be simplified into a representation form of the cluster center plus individual difference data; then, compressing the aforementioned cluster center and individual difference data can obtain the compressed data stream corresponding to the real-time trading data stream. Finally, sending the compressed data stream to each trading device can complete the real-time distribution of trading data. Thus, this method only compresses the common part and the difference part in the trading data, thereby greatly reducing the amount of compressed data. Based on this, the compression efficiency can be improved, and further the transmission delay caused by too long compression time can be reduced. Therefore, this method is very suitable for large-scale application and promotion in the field of data transmission. Among them, for example, this method can but is not limited to running on the cloud server side. It can be understood that the aforementioned execution subject does not constitute a limitation to the embodiments of the present application. Correspondingly, the running steps of this method can but are not limited to the following steps S1 to S6.

[0025] S1. Obtain the real-time trading data stream; in this embodiment, for example, the real-time trading data stream can but is not limited to including market trading data of many categories such as stocks, funds, bonds, and wealth management, and is obtained by the cloud server crawling from various financial trading websites or platforms in real time at a preset time interval. Based on this, the real-time nature of the trading data can be ensured. Of course, the aforementioned example is only illustrative, and the types of trading data are not limited to the aforementioned example.

[0026] After obtaining the real-time transaction data stream, in order to transmit the collected real-time transaction data stream to each node (i.e., trading device) in the trading system with low latency, so as to ensure that market data and trading signals can be distributed to each trading device (i.e., trading terminal) in real time and stably, this embodiment provides an improved real-time compression method to reduce the compression time of the real-time transaction data stream, thereby reducing the problem that the traditional compression method compresses all data, resulting in too long compression time and further increasing the transmission delay.

[0027] Among them, for example, the compression process of the real-time transaction data stream can be but is not limited to the following steps S2 to S5.

[0028] S2. Cluster each transaction data in the real-time transaction data stream to obtain at least one data cluster; in this embodiment, by performing clustering analysis on the real-time transaction data stream, the similar transaction data in the real-time transaction data stream is clustered into a cluster, and then, the individual difference data between the cluster center of each cluster and the remaining transaction data in the corresponding cluster is determined; finally, only the cluster center and the individual difference data are compressed, so as to reduce the amount of data to be compressed and further achieve the purpose of reducing the compression time.

[0029] Optionally, the clustering process of the real-time transaction data stream can be but is not limited to the following steps S21 to S25.

[0030] S21. Calculate the data distance between each transaction data in the real-time transaction data stream; in this embodiment, each transaction data in the real-time transaction data stream can be encoded as a vector (such as using one-hot encoding), and then, using the vector distance formula, calculate the data distance between each transaction data; at the same time, for any transaction data in the real-time transaction data stream, calculate the data distance between the any transaction data and the remaining each transaction data in the real-time transaction data stream. In this way, after polling all transaction data, the data distance between each transaction data in the real-time transaction data stream can be obtained.

[0031] S22. Determine multiple initial cluster centers according to the data distance between each transaction data; in this embodiment, the local density of data points is used to determine multiple initial cluster centers, and the process can be but is not limited to the following steps S22a to S22f.

[0032] S22a. Using the data distances between individual transaction data, construct a distance matrix, and sort the data distances in the distance matrix in ascending order to obtain a distance sequence; in this embodiment, any row in the distance matrix is a transaction data, and the data distances between it and the remaining individual transaction data in the real-time transaction data stream. For example, if the first row in the distance matrix corresponds to the first transaction data, then this row is the data distances between the first transaction data and the remaining individual transaction data in the real-time transaction data stream; thus, after constructing the distance matrix in the foregoing manner, the elements in the distance matrix can be sorted in ascending order to obtain a distance sequence; then, based on the distance sequence, a distance threshold can be determined for subsequent calculation of the local density of each transaction data according to the distance threshold; the process of determining the distance threshold is shown in step S22b below.

[0033] S22b. Based on the distance sequence, determine the distance threshold; in this embodiment, filter out the data distances of the top 2% after sorting in the distance sequence, and then calculate the average value as the distance threshold.

[0034] After determining the distance threshold, based on this, the local density of each transaction data can be calculated, and the process is shown in step S22c below.

[0035] S22c. For any transaction data in the real-time transaction data stream, calculate the local density of the any transaction data according to the distance threshold and the data distances between the any transaction data and each specified data in the specified data set, and after polling all the transaction data in the real-time transaction data stream, obtain the local density of each transaction data, where each specified data in the specified data set is each transaction data in the real-time transaction data stream except the any transaction data; in specific applications, for example, but not limited to, use the following formula (3) to calculate the local density of the foregoing any transaction data.

[0036] (3) In the above formula (3), represents the local density of the any transaction data, represents the data distance between the any transaction data and the th specified data in the specified data set, represents the distance threshold, represents the total number of specified data, where is the local density function, and when is less than 0, the value of the local density function is 1, and when is greater than or equal to 0, the value of the local density function is 0.

[0037] Thus, based on the foregoing formula (3), the local density of each transaction data in the real-time transaction data stream can be calculated. Based on this, the initial clustering centers can be preliminarily screened based on the local density of the data points. The process is as shown in the following steps S22d to S22f.

[0038] S22d. For any one of the transaction data, at least one target data is screened out from the real-time transaction data stream, where the local density of any one of the target data is greater than the local density of any one of the transaction data; in a specific implementation, for any one of the foregoing transaction data, it is equivalent to screening out the transaction data whose local density is greater than the local density of any one of the transaction data, and then using it as the target data. Finally, the initial clustering centers can be preliminarily selected with the help of the determined target data. The process is as shown in the following steps S22e to S22f.

[0039] S22e. The smallest data distance is screened out from the data distances between each target data and any one of the transaction data as the calibration distance corresponding to any one of the transaction data. After polling all the transaction data in the real-time transaction data stream, the calibration distances corresponding to each transaction data are obtained; in this embodiment, this step is to screen out the high-density data points closest to any one of the transaction data; then, the distance between the high-density data point and any one of the foregoing transaction data, that is, the high-density distance, is used as the calibration distance. Finally, based on this calibration distance and the local density, multiple initial clustering centers can be determined. The process is as shown in the following step S22f.

[0040] S22f. Using the local density and calibration distance of each transaction data, multiple initial clustering centers are determined from each transaction data; in this embodiment, the standard deviation of the calibration distances corresponding to all the transaction data in the real-time transaction data stream and the density mean of the local densities corresponding to all the transaction data are first calculated; then, based on the standard deviation, a calibration threshold is determined (for example, the calibration threshold is 2 times the standard deviation); then, from each transaction data in the real-time transaction data stream, the transaction data whose calibration distance is greater than or equal to the calibration threshold is screened out as the preliminary selected initial clustering centers; finally, from each preliminary selected initial clustering center, the preliminary selected initial clustering center whose local density is greater than or equal to the density mean is screened out to use the screened preliminary selected initial clustering center as the initial clustering center.

[0041] Thus, this embodiment objectively selects the initial clustering centers by using the density of data points (i.e., local density) in the real-time transaction data stream. That is, there should be a large distance between the selected clustering centers. Therefore, the high-density distances (i.e., the aforementioned calibrated distances) of other data points in the clustering should be less than or equal to the aforementioned calibrated threshold. Based on this, the transaction data with a calibrated distance greater than the calibrated threshold can be used as the preselected initial clustering centers. At the same time, in actual use, there may be noise data points in the dataset with a large high-density distance but low local density. Once such noise data points are selected, it is very easy to interfere with the selection of other normal clustering centers. Therefore, it is also necessary to use local density to denoise the preselected initial clustering centers, so as to obtain the initial clustering centers after denoising.

[0042] Thus, through the aforementioned steps S22a to S22f, the selection of the initial clustering centers can be completed. Then, the initial clustering centers can be optimized to obtain at least one optimal initial clustering center. The reason for optimizing the initial clustering centers is that the clustering algorithm is sensitive to the initial clustering centers. If the initial clustering centers are not selected reasonably, it is easy for the clustering to fall into a local optimum and the convergence speed will also be slow. Therefore, this embodiment provides a clustering optimization algorithm to avoid the problem that the clustering falls into a local optimum due to the unreasonable selection of the initial clustering centers.

[0043] Among them, the optimization process of the initial clustering centers can be but is not limited to the following steps shown in S23.

[0044] S23. Use the clustering optimization algorithm to optimize multiple initial clustering centers to obtain multiple optimal initial clustering centers of the real-time transaction data stream. In specific applications, this embodiment provides an improved bat algorithm to perform the optimization process of the initial clustering centers, and its process can be but is not limited to the following steps shown in S23a to S23i.

[0045] S23a. Based on multiple initial clustering centers, construct a bat population. Each bat individual in the bat population corresponds to an initial position, an initial velocity, an initial pulse loudness, and an initial pulse emission frequency. And the initial position of any bat individual corresponds to an initial clustering center. In this embodiment, the number of bat individuals is the same as the number of initial clustering centers and they correspond one by one. Therefore, the initial clustering center corresponding to each bat individual is used as its initial position. Of course, the aforementioned initial velocity, initial pulse loudness, and initial pulse emission frequency are preset values, and the initial velocity, initial pulse loudness, and initial pulse emission frequency corresponding to each bat individual are different and can be specifically set according to actual use.

[0046] After constructing the bat population, the bat sonar of each bat individual can be used to detect objects, and the optimization of the initial clustering center can be carried out. The process is as shown in the following steps S23b to S23i.

[0047] S23b. Initialize the number of optimization times t, and based on the positions of each bat individual at the t-th optimization, cluster the real-time transaction data stream to obtain the clustering clusters at the t-th optimization. Among them, each bat individual corresponds to a clustering cluster. The initial value of t is 1. When t is 1, the position of any bat individual at the t-th optimization is the initial position of the any bat individual; in specific applications, it is equivalent to determining the clustering center of the real-time transaction data stream at the t-th optimization based on the positions of each bat individual at the t-th optimization. Then, according to the clustering center at the t-th optimization, cluster the real-time transaction data stream to obtain the clustering clusters at the t-th optimization. That is, as previously described, the initial position of the bat individual corresponds to an initial clustering center. Therefore, at the beginning of the optimization (i.e., the first optimization), the multiple initial clustering centers determined by the previous step S22 are used to cluster the real-time transaction data stream to obtain the clustering clusters at the first optimization; then, update the positions of each bat individual to obtain the clustering center at the second optimization. Next, use the clustering center at the second optimization to cluster the real-time transaction data stream. In this way, based on this principle, until the optimization ends, the optimal initial clustering center can be determined according to the positions of the bat individuals.

[0048] Furthermore, in this embodiment, for example, but not limited to, the K-means clustering algorithm can be used to cluster the real-time transaction data stream. Of course, K-means clustering is a common method for data clustering, and its principle will not be elaborated here.

[0049] After obtaining the clustering clusters at the t-th optimization, based on this, calculate the fitness of each bat individual at the t-th optimization, so as to judge whether to stop the optimization iteration based on the fitness in the subsequent process; among them, the calculation process of the fitness is as shown in the following step S23c.

[0050] S23c. Calculate the fitness of each bat individual at the t-th optimization based on the clustering clusters at the t-th optimization. Among them, the greater the fitness of any bat individual, the higher the clustering accuracy of the clustering cluster corresponding to the bat individual. In specific implementation, the fitness of any bat individual at the t-th optimization is used to represent the dissimilarity between each transaction data and the clustering center in the clustering cluster corresponding to the bat individual at the t-th optimization. And the smaller the dissimilarity, the greater the fitness. At the same time, in this embodiment, one bat individual corresponds to one clustering cluster. Therefore, taking any bat individual as an example, the calculation process of fitness is described as follows: First, calculate the distance between the clustering center of the clustering cluster corresponding to the bat individual and the rest of the data in the clustering cluster. Then, sum the distances between the clustering center and each data, and take the reciprocal of the sum result, then the fitness of the bat individual at the t-th optimization can be obtained.

[0051] For example, assume that the clustering center of the bat individual at the t-th optimization is transaction data C, and the corresponding clustering cluster is cluster 1. Then calculate the distance between each transaction data in cluster 1 and transaction data C. Then, sum the distances and take the reciprocal to obtain the fitness of the bat individual at the t-th optimization. In this way, it can be concluded from the above fitness calculation process that the more similar the data in the cluster is to the clustering center, the greater the fitness, that is, the smaller the dissimilarity, the greater the fitness.

[0052] Based on this, after calculating the fitness of each bat individual at the t-th optimization based on step S23, the judgment of stopping the optimization iteration can be carried out, and the process is as shown in step S23d below.

[0053] S23d. Judge whether the fitness of each bat individual at the t-th optimization is greater than the historical best fitness of each bat individual. In this embodiment, the historical best fitness of any bat individual is the maximum fitness obtained by the bat individual during the process before the t-th optimization. Of course, at the beginning of the optimization, such as at the first optimization, the fitness of each bat individual at the 1st optimization can be directly set as the corresponding historical best fitness. At this time, the optimization needs to continue. In this way, after the second optimization is completed, each bat individual calculates a fitness again. At this time, the fitness of each bat individual at the second optimization can be compared with the corresponding historical best fitness. If it is greater, the historical best fitness is updated to the fitness at the second optimization. Otherwise, it is not updated. Based on this, continuous optimization can continuously update the historical best fitness. Of course, in the above step S23d, if the iteration stop condition cannot be met, the optimization needs to continue, and the process is as shown in steps S23e - S23i below.

[0054] S23e. If not, determine the maximum fitness, search weight factor, and pulse frequency at the t-th optimization. In this embodiment, when the traditional bat algorithm updates the velocity, no weight is added, resulting in problems such as slow search speed or premature convergence during the search. Therefore, in this embodiment, a search weight factor is added to balance the global search and local search of the algorithm. Among them, for example, but not limited to, according to the following formula (1), determine the search weight factor at the t-th optimization.

[0055] (1) In the above formula (1), represents the search weight factor at the t-th optimization, represents the maximum number of optimizations, successively represent the maximum search weight and the minimum search weight, and represents the weight exponent; in this embodiment, the maximum search weight, the minimum search weight, and the weight exponent can be, for example, but not limited to, set to 0.9, 0.2, and 2 in sequence; of course, the specific numerical settings can be specifically set according to actual use and are not limited to the foregoing examples here.

[0056] Thus, as can be seen from the foregoing formula (1), in the early stage of optimization, the search weight factor is relatively large, and as the number of optimizations increases, the search weight factor gradually becomes smaller. Among them, the larger the value of the search weight factor, the stronger the global search ability, and vice versa, the stronger the local search ability; therefore, after adding the search weight factor, it can have a strong global search ability in the early stage to improve the convergence speed, and have a high local search in the later stage to improve the search accuracy.

[0057] After calculating the search weight factor at the t-th optimization, the pulse frequency at the t-th optimization can be calculated to update the velocity of the bat individual based on the pulse frequency subsequently.

[0058] In this embodiment, each of the foregoing bat individuals corresponds to a pulse frequency at the t-th optimization. Taking any bat individual as an example, for example, use the following formula (4) to calculate the pulse frequency corresponding to it at the t-th optimization.

[0059] (4) In the above formula (4), represents the pulse frequency of the said any bat individual at the t-th optimization, respectively represent the maximum pulse frequency and the minimum pulse frequency, while represents a random number on [0,1].

[0060] Thus, as can be seen from the above formula (4), since They are different, and the pulse frequency of each bat individual is also different during the t-th optimization; after calculating the search weight factor and pulse frequency during the t-th optimization, the speed and position can be updated, and the process is as shown in the following step S23f.

[0061] S23f. Update the speed of each bat individual during the t-th optimization according to the maximum fitness, search weight factor, and pulse frequency, obtain the updated speed corresponding to each bat individual, and use the updated speed corresponding to each bat individual to update the position of the bat individual during the t-th optimization to obtain the updated position corresponding to each bat individual. Among them, when t is 1, the speed of any bat individual during the t-th optimization is the initial speed corresponding to the bat individual.

[0062] In specific implementation, for example, but not limited to, first determine the optimal bat position during the t-th optimization according to the maximum fitness during the t-th optimization (that is, the position of the bat individual corresponding to the maximum fitness during the t-th optimization is used as the optimal bat position during the t-th optimization); then, for any bat individual, according to the optimal bat position, search weight factor, and pulse frequency, and use the following formula (5) to calculate the updated speed corresponding to the bat individual.

[0063] (5) In the above formula (5), represents the updated speed corresponding to the bat individual, represents the speed of the bat individual during the t-th optimization, represents the position of the bat individual during the t-th optimization, represents the optimal bat position, represents the pulse frequency of the bat individual during the t-th optimization, represents the search weight factor during the t-th optimization.

[0064] In this way, based on the foregoing formula (5), the speed update of the bat individual can be completed; then, the position update can be performed. Taking any bat individual as an example, it will be elaborated. First, for any bat individual, generate a random vector (in this embodiment, the elements in the random vector take values between [-1, 1], and the length is the same as the length of the position vector corresponding to the bat individual); then, determine the updated position corresponding to the bat individual according to the updated speed corresponding to the bat individual and the random vector.

[0065] Optionally, for example, use the following formula (6) to calculate the updated position corresponding to the bat individual.

[0066] (6) The above formula (6) represents the updated position corresponding to any bat individual, represents a random vector, represents the norm of the vector; based on the above formula (6), it can be seen that this embodiment introduces random variables to increase the ability of individual bat positions to change, thereby improving the diversity of the population.

[0067] In this way, based on the above formulas (5) and (6), the speed and position of each individual bat can be updated; then, a local search of the individual bat can be performed, that is, position disturbance, and the process is shown in the following step S23g.

[0068] S23g. Use the pulse emission frequency of each bat individual during the t-th optimization search to perturb the updated position corresponding to each bat individual to obtain the perturbed position corresponding to each bat individual, wherein when t is 1, the pulse emission frequency of any bat individual during the t-th optimization search is the initial pulse emission frequency of any bat individual.

[0069] In specific applications, the search step size during local search in the traditional bat algorithm is a fixed value (usually between [-1,1]), which is obviously unable to adapt to the changes in the algorithm during operation. At the same time, a larger step size is beneficial to improving the global exploration ability of the algorithm, and a smaller step size is beneficial to the local development ability of the algorithm and improves the optimization accuracy. Therefore, this embodiment provides a step size with adaptive adjustment to improve the adaptability of the algorithm.

[0070] In this embodiment, any individual bat is taken as an example to illustrate the position disturbance step, and the process may be but is not limited to the following steps S23g1 to S23g4.

[0071] S23g1. For any individual bat, generate a first random number, and determine whether the first random number is greater than the pulse emission frequency of any individual bat during the t-th optimization search; in this embodiment, when t is 1, the pulse emission frequency of any individual bat during the t-th optimization search is the initial pulse emission frequency of any individual bat; wherein, when the first random number is greater than the pulse emission frequency of any individual bat during the t-th optimization search, a local search can be performed, and the process is shown in the following steps S23g2 to S23g4.

[0072] S23g2. If so, obtain the maximum perturbation factor, the minimum perturbation factor, and the total number of individuals in the bat population; in this embodiment, both the maximum and minimum frequency factors are set values and are not specifically limited here; after obtaining the maximum perturbation factor, the minimum perturbation factor, and the total number of individuals in the bat population, the perturbation coefficient at the t-th optimization can be calculated, and the process is as shown in the following step S23g3.

[0073] S23g3. Calculate the perturbation coefficient at the t-th optimization according to the maximum perturbation factor, the minimum perturbation factor, and the total number of individuals; in this embodiment, for example, but not limited to, the perturbation coefficient at the t-th optimization can be calculated according to the following formula (2).

[0074] (2) In the above formula (2), represents the perturbation coefficient at the t-th optimization, represent the maximum perturbation factor and the minimum perturbation factor in sequence, represents the total number of individuals.

[0075] Thus, based on formula (2), it can be seen that in this embodiment, an exponentially decreasing factor is used to replace the fixed step size (i.e., the aforementioned perturbation coefficient) in the traditional technology, thereby improving the optimization accuracy.

[0076] After calculating the perturbation coefficient at the t-th optimization, based on this, the perturbation of the updated position corresponding to any bat individual can be performed, and the process is as shown in the following step S23g4.

[0077] S23g4. Use the perturbation coefficient at the t-th optimization to perform perturbation processing on the updated position corresponding to any bat individual to obtain the perturbed position corresponding to any bat individual; in this embodiment, for example, but not limited to, the following formula (7) can be used to perform the perturbation of the updated position of any bat individual.

[0078] (7) In the above formula (7), represents the perturbed position corresponding to any bat individual, represents the average pulse loudness of all bat individuals in the bat population at the t-th optimization; in this embodiment, when t is 1, the pulse loudness of any bat individual at the t-th optimization is the initial pulse loudness of this bat individual; thus, at the first optimization, the average pulse loudness of all bat individuals is the mean of the initial pulse loudnesses of all bat individuals; of course, at each iteration, the pulse loudness and the pulse emission frequency will also be continuously updated, and their update processes will be elaborated in detail below.

[0079] Thus, based on the foregoing steps S23g1 to S23g4, the position perturbation of each bat individual can be completed, that is, the local search of the position of each bat individual is realized; then, it is necessary to judge whether the position perturbation is acceptable, and the process is as shown in the following step S23h.

[0080] S23h. Based on the pulse loudness of each bat individual during the t-th optimization, judge whether to retain the perturbed position corresponding to each bat individual, and when t is 1, the pulse loudness of any bat individual during the t-th optimization is the initial pulse loudness of the any bat individual; in specific applications, for any bat individual, a second random number is generated, and then it is judged whether the second random number is less than the pulse loudness of the any bat individual during the t-th optimization. If so, calculate the fitness before the position perturbation and the fitness after the position perturbation of the any bat individual to obtain the new fitness and the original fitness respectively; then, judge whether the new fitness is greater than the original fitness; among them, if it is greater than the original fitness, accept the position perturbation, that is, use the perturbed position corresponding to the any bat individual as its position in the next iteration. For example, if this is the first optimization, then its corresponding perturbed position is the position of the any bat individual in the second optimization.

[0081] Of course, if the foregoing conditions are not met, that is, the second random number is greater than or equal to the pulse loudness of the any bat individual during the t-th optimization, or the new fitness is less than or equal to the original fitness, then the position perturbation is not accepted, that is, directly use the updated position corresponding to the any bat individual as the position in the next iteration.

[0082] At the same time, the following discloses the update process of the pulse loudness and pulse emission frequency of any bat individual. Still taking any bat individual as an example, it is described as follows: Among them, for example, but not limited to, the following formulas (8) and (9) can be used to update the pulse loudness and pulse emission frequency.

[0083] (8) (9) In the above formula (8), represents the pulse loudness of the any bat individual during the (t + 1)-th optimization, represents the pulse loudness of the any bat individual during the t-th optimization, is the attenuation coefficient, which takes values in [-1, 1].

[0084] In the above formula (9), represents the pulse emission frequency of the any bat individual during the t-th optimization, represents the frequency adjustment coefficient, is greater than 0, represents the initial pulse emission frequency of any one of the bat individuals.

[0085] Thus, after completing the judgment on whether to retain the perturbation position based on the foregoing step S23h, continuous iteration can be performed, and the process is as shown in the following step S23i.

[0086] S23i. If so, increment t by 1, and use the perturbation positions corresponding to each bat individual as the positions of each bat individual during the t-th optimization. Then, re-cluster the real-time transaction data stream based on the positions of each bat individual during the t-th optimization until the fitness of each bat individual during the t-th optimization is greater than the historical best fitness of each bat individual. At this time, determine multiple optimal initial clustering centers of the real-time transaction data stream according to the positions of each bat individual during the t-th optimization. In this embodiment, when t is 2, it is equivalent to using the perturbation positions corresponding to each bat individual during the first optimization as the positions of each bat individual during the second optimization. Then, re-execute the foregoing steps S23b to S23i until the foregoing iteration stop condition is met. At this time, the positions of each bat individual when the iteration stop condition is met can be used as the optimal initial clustering centers of the real-time transaction data stream. In addition, in this embodiment, if there is no corresponding transaction data in the real-time transaction data stream for the obtained optimal initial clustering center, at this time, the transaction data with the smallest distance from the optimal initial clustering center can be used as the optimal initial clustering center.

[0087] Of course, in the foregoing step S23g1, for any one bat individual, if the generated first random number is less than or equal to the pulse emission frequency of the any one bat individual during the t-th optimization, at this time, no local search is performed on the any one bat individual, and step S23i can be directly executed, that is, use its updated position as the position for the next iteration.

[0088] Thus, through the foregoing steps S23a to S23i, the optimization process of the initial clustering center can be completed. Then, based on the optimal initial clustering center, the real-time transaction data stream can be clustered, and the process is as shown in the following step S24.

[0089] S24. Cluster the real-time transaction data stream based on multiple optimal initial clustering centers to obtain multiple initial data clusters. In this embodiment, based on multiple optimal initial clustering centers, the K-means clustering algorithm is used to cluster the real-time transaction data stream, thereby obtaining multiple initial data clusters.

[0090] Since there may be a problem that some data objects in the clustering result may be misclassified when the K-means algorithm clusters data, this embodiment also sets a clustering correction step, and its process is as shown in step S25 below.

[0091] S25. Perform clustering correction processing on multiple initial data clusters to obtain the at least one data cluster; in this embodiment, for example, but not limited to, the following steps S25a to S25e can be used to complete the clustering correction.

[0092] S25a. According to each initial data cluster, screen out misclassified data from the real-time transaction data stream; in specific applications, for any transaction data in the real-time transaction data stream, the distance between the any transaction data and the cluster centers in each initial data cluster can be calculated first, and the calculated distances can be sorted in ascending order to obtain a sorted sequence; then, the first two distances in the sorted sequence are screened out, and it is determined whether the absolute value of the difference between the first two distances is less than a preset threshold; wherein, if so, the any transaction data is used as misclassified data, and when all transaction data in the real-time transaction data stream have been polled, all misclassified data is screened out from the real-time transaction data stream.

[0093] In this embodiment, if the absolute value of the difference between the distances of a data object from the two nearest cluster centers is less than the preset threshold, it is classified as misclassified data. In this way, after all misclassified data is screened out from the real-time transaction data stream, a set of data to be classified can be formed, and its process is as shown in step S25b below.

[0094] S25b. Use all the screened misclassified data to form a set of data to be classified; after obtaining the set of data to be classified, the misclassified data can be deleted from the real-time transaction data stream, so as to obtain correctly classified data, and the correctly classified data can be used to form a correctly classified set, and its process is as shown in step S25c below.

[0095] S25c. Delete all misclassified data from the real-time transaction data stream, and use the remaining transaction data to form a correctly classified set; after forming the correctly classified set based on this step, the misclassified data in the set of data to be classified can be reclassified based on this, and its process is as shown in step S25d and step S25e below.

[0096] S25d. For any misclassified data in the set to be classified, determine, from the correctly classified set, the transaction data that is closest in distance to the said any misclassified data as the nearest data to the said any misclassified data; in this embodiment, it is equivalent to screening, from the correctly classified set, the correctly classified transaction data that is most similar to the said any misclassified data as the nearest data to the said any misclassified data, and then, the said any misclassified data can be reclassified according to this nearest data, and its reclassification process is as shown in the following step S25e.

[0097] S25e. Reclassify the said any misclassified data according to the nearest data of the said any misclassified data, and after polling all the misclassified data in the set to be classified, complete the clustering correction process of multiple initial data clusters to obtain the said at least one data cluster; in this embodiment, directly classify the said any misclassified data into the initial data cluster corresponding to its nearest data; for example, assume that the said any misclassified data originally belongs to the initial data cluster 3, and the initial data cluster where its corresponding nearest data is located is cluster 2, then the said any misclassified data is divided into cluster 2; thus, the clustering correction of the said any misclassified data can be completed.

[0098] Thus, through the foregoing steps S25a - S25e, in this embodiment, by finding the nearest neighbors of each data object in the set to be classified and relying on the nearest neighbor idea, the data objects in the set to be classified are classified into the categories where their nearest neighbors are located. In this way, the data objects with misclassified categories in the initial clustering result can be corrected back to the correct classes, thereby improving the accuracy of clustering.

[0099] Thus, through the foregoing steps S21 - S25, the clustering process of the real - time transaction data stream can be completed, so as to group the similar data in the real - time transaction data stream into one category; then, the differences between the cluster centers and the remaining data in each data cluster can be determined, so as to represent the data cluster in the form of cluster center + individual differences; finally, only the cluster centers and the individual difference data are compressed to obtain the compressed data stream of the real - time transaction data stream.

[0100] Among them, the determination process of the individual difference data is as shown in the following step S3.

[0101] S3. Determine the cluster centers of each data cluster, and determine the individual difference data between each cluster center and the remaining transaction data in the corresponding data cluster, as well as the difference positions corresponding to each individual difference data; in this embodiment, when using the K-means clustering algorithm, the cluster centers of each data cluster will be determined, and then, the cluster centers can be compared with each transaction data in the data cluster to obtain the individual difference data between the cluster center and the remaining transaction data in the corresponding data cluster.

[0102] Optionally, the following uses an example to elaborate: Suppose the cluster center of data cluster 3 is: CGTGTACTGTGATACGTG (starting from position 0 and ending at position 17), and the transaction data included in the cluster are Q1 and Q2. Among them, Q1 is: CGACTGTACTATGATACG, and Q2 is: CGCGTACTGATTGATACGTG; then, by comparing the cluster center with Q1, it can be seen that the difference between the two is that: insert AC after the first position of the cluster center, replace G at the eighth position with A, and delete the 15th and 16th positions, then the transaction data Q1 can be obtained. Similarly, replace the second position of the cluster center with C, and insert AT after the eighth position to obtain the transaction data Q2. Based on this, the individual difference data between the cluster center and Q1 are AC, A, and GT, and their corresponding difference positions are the first position, the eighth position, and the 15th to 16th positions, and the corresponding difference operations are insertion operation, replacement operation, and deletion operation; of course, the principle of obtaining the individual difference data and difference positions between the cluster center and the transaction data Q2 is the same as the previous example, which will not be elaborated here; of course, the data content in the previous example is only for illustration and is not regarded as real transaction data.

[0103] In this way, after determining the individual difference data and difference positions between each cluster center and the remaining transaction data in the corresponding data cluster, the compression process of the cluster center and the individual difference data can be performed, and the process is as shown in the following step S4.

[0104] S4. Perform data compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center; in specific applications, first perform sampling compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain the sampling compressed data corresponding to each cluster center and the sampling compressed individual difference data corresponding to each cluster center; then, perform lossless compression processing on the sampling compressed data corresponding to each cluster center and the sampling compressed individual difference data corresponding to each cluster center, so as to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center after the lossless compression processing.

[0105] Optionally, when performing sampling compression processing, any class cluster center and the individual difference data corresponding to the class cluster center are encoded as column vectors. Then, a 128×256 observation matrix is used to perform matrix multiplication with the column vectors, so as to obtain a set of 128-dimensional column vectors as the compression result after compressed sampling (that is, the sampled compressed data and the sampled compressed individual difference data are substantially within one column vector); then, the LZW algorithm (a lossless data compression algorithm based on dynamic dictionary construction) is used to perform lossless compression on the aforementioned sampled compressed data and the sampled compressed individual difference data, so as to obtain the compressed any class cluster center and the compressed individual difference data corresponding to the class cluster center; of course, the compression processes of the remaining class cluster centers and their corresponding individual difference data are also like this, which will not be elaborated here.

[0106] Furthermore, the aforementioned LZW algorithm is a common technique for lossless compression, and its principle will not be elaborated.

[0107] In this way, after completing data compression based on the aforementioned step S4, the compressed data stream can be combined with the difference positions of the aforementioned individual difference data, and the process is as shown in the following step S5.

[0108] S5. Using each compressed class cluster center, the compressed individual difference data corresponding to each class cluster center, and the difference positions corresponding to each individual difference data, form the compressed data stream corresponding to the real-time transaction data stream; in this embodiment, it is still necessary to associate and record the difference operations corresponding to the difference positions together to obtain the compressed data stream, so as to perform data decompression based on the difference operations and positions in the subsequent process; and after obtaining the compressed data stream corresponding to the real-time transaction data stream, data distribution can be performed, and the process is as shown in the following step S6.

[0109] S6. Distribute the compressed data stream to each trading device to complete the real-time distribution of the real-time transaction data stream.

[0110] Thus, through the real-time distribution method of quantitative trading data based on cloud transmission described in detail in the foregoing steps S1 to S6, when the present invention performs compression, it performs clustering analysis on the real-time trading data stream, and determines the individual difference data between the data clusters obtained by clustering and the remaining trading data in each cluster. In this way, the real-time trading data stream can be simplified into a representation form of the cluster center plus individual difference data; then, by compressing the foregoing cluster center and individual difference data, the compressed data stream corresponding to the real-time trading data stream can be obtained. Finally, by sending the compressed data stream to each trading device, the real-time distribution of trading data can be completed; thus, the present invention only compresses the common part and the difference part in the trading data, thereby greatly reducing the amount of compressed data. Based on this, the compression efficiency can be improved, and further the transmission delay caused by too long compression time can be reduced. Therefore, the present invention is very suitable for large-scale application and promotion in the field of data transmission.

[0111] As Figure 2 shown, in the second aspect of this embodiment, there is provided a hardware device for implementing the real-time distribution method of quantitative trading data based on cloud transmission described in the first aspect of the embodiment, including: An acquisition unit, configured to acquire a real-time trading data stream.

[0112] A clustering unit, configured to perform clustering processing on each trading data in the real-time trading data stream to obtain at least one data cluster.

[0113] A compression unit, configured to determine the cluster center of each data cluster, and determine the individual difference data between each cluster center and the remaining trading data in the corresponding data cluster, as well as the difference positions corresponding to each individual difference data.

[0114] A compression unit, configured to perform data compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center.

[0115] The compression unit is further configured to use each compressed cluster center, the compressed individual difference data corresponding to each cluster center, and the difference positions corresponding to each individual difference data to form the compressed data stream corresponding to the real-time trading data stream.

[0116] A distribution unit, configured to distribute the compressed data stream to each trading device to complete the real-time distribution of the real-time trading data stream.

[0117] For the working process, working details and technical effects of the device provided in this embodiment, reference can be made to the first aspect of the embodiment, which will not be elaborated herein.

[0118] As Figure 3As shown, in the third aspect of this embodiment, another real-time distribution device for quantitative trading data based on cloud transmission is provided. Taking the device as an electronic device as an example, it includes: a memory, a processor, and a transceiver that are communicatively connected in sequence. Among them, the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer programs and execute the real-time distribution method for quantitative trading data based on cloud transmission as described in the first aspect of the embodiment.

[0119] Specifically, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first input first output (FIFO), and / or first in last out (FILO), etc.; specifically, the processor may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). At the same time, the processor may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state.

[0120] In some embodiments, the processor may integrate a GPU (Graphics Processing Unit). The GPU is responsible for rendering and drawing the content to be displayed on the display screen. For example, the processor may be, but is not limited to, a microprocessor of the STM32F105 series, a reduced instruction set computer (RISC) microprocessor, a processor with an X86 architecture, or a processor integrated with an embedded neural-network processing unit (NPU). The transceiver may be, but is not limited to, a Wi-Fi wireless transceiver, a Bluetooth wireless transceiver, a General Packet Radio Service (GPRS) wireless transceiver, a ZigBee (a low-power local area network protocol based on the IEEE 802.15.4 standard) wireless transceiver, a 3G transceiver, a 4G transceiver, and / or a 5G transceiver, etc. In addition, the device may also include, but is not limited to, a power module, a display screen, and other necessary components.

[0121] For the working process, working details, and technical effects of the electronic device provided in this embodiment, reference may be made to the first aspect of the embodiment, which will not be elaborated here.

[0122] The fourth aspect of this embodiment provides a storage medium storing instructions for the real-time distribution method of cloud-transmission-based quantitative trading data described in the first aspect of the embodiment, that is, instructions are stored on the storage medium. When the instructions run on a computer, the real-time distribution method of cloud-transmission-based quantitative trading data described in the first aspect of the embodiment is executed.

[0123] Among them, the storage medium refers to a carrier for storing data, which may be, but is not limited to, a floppy disk, an optical disc, a hard disk, a flash memory, a USB flash drive, and / or a Memory Stick, etc. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0124] For the working process, working details, and technical effects of the storage medium provided in this embodiment, reference may be made to the first aspect of the embodiment, which will not be elaborated here.

[0125] The fifth aspect of this embodiment provides a computer program product containing instructions. When the instructions run on a computer, the computer is made to execute the real-time distribution method of cloud-transmission-based quantitative trading data described in the first aspect of the embodiment. Among them, the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0126] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A real-time distribution method for quantitative trading data based on cloud transmission, characterized in that, Including: Obtain the real-time transaction data stream; Perform clustering processing on each transaction data in the real-time transaction data stream to obtain at least one data cluster; Determine the cluster centers of each data cluster, and determine the individual difference data between each cluster center and the remaining transaction data in the corresponding data cluster, as well as the difference positions corresponding to each individual difference data; Perform data compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center; Use each compressed cluster center, the compressed individual difference data corresponding to each cluster center, and the difference positions corresponding to each individual difference data to form the compressed data stream corresponding to the real-time transaction data stream; Distribute the compressed data stream to each transaction device to complete the real-time distribution of the real-time transaction data stream.

2. The method according to claim 1, wherein Performing clustering processing on each transaction data in the real-time transaction data stream to obtain at least one data cluster includes: Calculate the data distance between each transaction data in the real-time transaction data stream; Determine a plurality of initial clustering centers according to the data distance between each transaction data; Use the clustering optimization algorithm to perform optimization processing on a plurality of initial clustering centers to obtain a plurality of optimal initial clustering centers of the real-time transaction data stream; Based on a plurality of optimal initial clustering centers, perform clustering processing on the real-time transaction data stream to obtain a plurality of initial data clusters; Perform clustering correction processing on a plurality of initial data clusters to obtain the at least one data cluster.

3. The method according to claim 2, wherein Determining a plurality of initial clustering centers according to the data distance between each transaction data includes: Use the data distance between each transaction data to construct a distance matrix, and sort each data distance in the distance matrix in ascending order to obtain a distance sequence; Determine a distance threshold based on the distance sequence; For any transaction data in the real-time transaction data stream, calculate the local density of the any transaction data according to the distance threshold and the data distance between the any transaction data and each specified data in the specified data set, and after polling all the transaction data in the real-time transaction data stream, obtain the local density of each transaction data, where each specified data in the specified data set is each transaction data except the any transaction data in the real-time transaction data stream; For the any transaction data, screen out at least one target data from the real-time transaction data stream, where the local density of any target data is greater than the local density of the any transaction data; Screen out the smallest data distance from the data distances between each target data and the any transaction data as the calibration distance corresponding to the any transaction data, and after polling all the transaction data in the real-time transaction data stream, obtain the calibration distances corresponding to each transaction data; Use the local density and calibration distance of each transaction data to determine a plurality of initial clustering centers from each transaction data.

4. The method according to claim 2, characterized in that, Using a clustering optimization algorithm, optimize multiple initial clustering centers to obtain multiple optimal initial clustering centers for the real-time transaction data stream, including: Based on multiple initial clustering centers, construct a bat population. Each bat individual in the bat population corresponds to an initial position, an initial velocity, an initial pulse loudness, and an initial pulse emission frequency, and the initial position of any bat individual corresponds to an initial clustering center; Initialize the number of optimization times t, and perform clustering processing on the real-time transaction data stream based on the positions of each bat individual at the t-th optimization to obtain the clustering clusters at the t-th optimization. Each bat individual corresponds to a clustering cluster. The initial value of t is 1, and when t is 1, the position of any bat individual at the t-th optimization is the initial position of the any bat individual; Based on the clustering clusters at the t-th optimization, calculate the fitness of each bat individual at the t-th optimization. The greater the fitness of any bat individual, the higher the clustering accuracy of the clustering cluster corresponding to the any bat individual; Judge whether the fitness of each bat individual at the t-th optimization is greater than the historical optimal fitness of each bat individual; If not, determine the maximum fitness, search weight factor, and pulse frequency at the t-th optimization; According to the maximum fitness, search weight factor, and pulse frequency, update the velocities of each bat individual at the t-th optimization to obtain the updated velocities corresponding to each bat individual, and use the updated velocities corresponding to each bat individual to update the positions of the bat individuals at the t-th optimization to obtain the updated positions corresponding to each bat individual. When t is 1, the velocity of any bat individual at the t-th optimization is the initial velocity corresponding to the any bat individual; Use the pulse emission frequencies of each bat individual at the t-th optimization to perturb the updated positions corresponding to each bat individual to obtain the perturbed positions corresponding to each bat individual. When t is 1, the pulse emission frequency of any bat individual at the t-th optimization is the initial pulse emission frequency of the any bat individual; Based on the pulse loudness of each bat individual at the t-th optimization, judge whether to retain the perturbed position corresponding to each bat individual. When t is 1, the pulse loudness of any bat individual at the t-th optimization is the initial pulse loudness of the any bat individual; If so, increment t by 1, and use the perturbed positions corresponding to each bat individual as the positions of each bat individual at the t-th optimization, and re-perform clustering processing on the real-time transaction data stream based on the positions of each bat individual at the t-th optimization until the fitness of each bat individual at the t-th optimization is greater than the historical optimal fitness of each bat individual. Then, based on the positions of each bat individual at the t-th optimization, determine multiple optimal initial clustering centers for the real-time transaction data stream.

5. The method according to claim 4, wherein Determine the search weight factor at the t-th optimization, including: Determine the search weight factor at the t-th optimization according to the following formula (1); (1) In the above formula (1), represents the search weight factor during the t-th optimization, represents the maximum number of optimizations, successively represent the maximum search weight and the minimum search weight, and represents the weight exponent; Correspondingly, using the pulse emission frequencies of each bat individual during the t-th optimization, the updated positions corresponding to each bat individual are perturbed to obtain the perturbed positions corresponding to each bat individual, which includes: For any bat individual, generate a first random number and determine whether the first random number is greater than the pulse emission frequency of the any bat individual during the t-th optimization; If so, obtain the maximum perturbation factor, the minimum perturbation factor, and the total number of individuals in the bat population; According to the maximum perturbation factor, the minimum perturbation factor, and the total number of individuals, calculate the perturbation coefficient during the t-th optimization; Use the perturbation coefficient during the t-th optimization to perturb the updated position corresponding to the any bat individual to obtain the perturbed position corresponding to the any bat individual.

6. The method according to claim 5, wherein Calculating the perturbation coefficient during the t-th optimization according to the maximum perturbation factor, the minimum perturbation factor, and the total number of individuals includes: Calculate the perturbation coefficient during the t-th optimization according to the following formula (2); (2) In the above formula (2), represents the perturbation coefficient during the t-th optimization, successively represent the maximum perturbation factor and the minimum perturbation factor, represents the total number of the individuals.

7. The method according to claim 2, wherein Performing clustering correction processing on multiple initial data clusters to obtain the at least one data cluster, including: According to each initial data cluster, screen out misclassified data from the real-time transaction data stream; Use all the screened misclassified data to form a set to be classified; Delete all misclassified data from the real-time transaction data stream, and use the remaining transaction data to form a correctly classified set; For any misclassified data in the set to be classified, determine the transaction data closest to the any misclassified data from the correctly classified set as the nearest neighbor data of the any misclassified data; According to the nearest neighbor data of the any misclassified data, reclassify the any misclassified data, and after polling all the misclassified data in the set to be classified, complete the clustering correction processing of multiple initial data clusters to obtain the at least one data cluster.

8. The method according to claim 1, wherein Performing data compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center, including: Performing sampling compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain the sampled compression data corresponding to each cluster center and the sampled compression individual difference data corresponding to each cluster center; Perform lossless compression processing on the sampled compression data corresponding to each cluster center and the sampled compression individual difference data corresponding to each cluster center, so as to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center after the lossless compression processing.

9. A real-time distribution device for quantitative trading data based on cloud transmission, characterized in that, Includes: An acquisition unit for acquiring a real-time transaction data stream; A clustering unit for clustering each transaction data in the real-time transaction data stream to obtain at least one data cluster; A compression unit for determining the cluster center of each data cluster, and determining the individual difference data between each cluster center and the remaining transaction data in the corresponding data cluster and the difference positions corresponding to each individual difference data; A compression unit for performing data compression processing on each cluster center and the individual difference data corresponding to each cluster center to obtain each compressed cluster center and the compressed individual difference data corresponding to each cluster center; The compression unit is further configured to use each compressed cluster center, the compressed individual difference data corresponding to each cluster center, and the difference positions corresponding to each individual difference data to form a compressed data stream corresponding to the real-time transaction data stream; A distribution unit for distributing the compressed data stream to each trading device to complete the real-time distribution of the real-time transaction data stream.

10. A real-time distribution device for quantitative trading data based on cloud transmission, characterized in that, Comprising: A memory, a processor, and a transceiver that are communicatively connected in sequence, wherein the memory is used for storing computer programs, the transceiver is used for sending and receiving messages, and the processor is used for reading the computer programs and executing the real-time distribution method of quantitative trading data based on cloud transmission according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method for compressing and storing GPS data based on route clustering

    CN101894135A

  • Full-pulse data lossless compression method based on K-means clustering

    CN106452452A

  • Parameter selection method for support vector machine based on hybrid bat algorithm

    CN108121999A

  • Improved binary bat algorithm-based distributed system task scheduling method

    CN108694077A

  • Transfer function model parameter recognition method and device based on improved particle swarm algorithm

    CN111428849A