Data pre-distribution method and device based on self-attention mechanism
By using the Transformer model based on the self-attention mechanism in large-scale systems for data pre-distribution, the problems of data distribution delay and network pressure are solved, and efficient and low-latency data distribution effect is achieved.
Patent Information
- Application Number
- CN202411592676.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-08
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has obvious limitations in dealing with the latency and network pressure of data distribution in large-scale systems, especially when facing the demand for high-load data distribution, which becomes more prominent.
The Transformer model based on the self-attention mechanism is used for pre-distribution of data. Through in-depth analysis and learning of historical data, it predicts the data subscription requirements at future moments, and distributes the data to the required nodes in advance.
It significantly reduces data transmission delay, improves the efficiency and response speed of data distribution, effectively reduces network pressure, and enhances the stability and reliability of the system.
Smart Images

Figure CN120223752A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a data pre-distribution method and device based on a self-attention mechanism. Background Art
[0002] In many key fields such as military, aerospace, weather forecasting, and equipment R & D, extremely large systems (also known as ultra-large-scale systems) play a crucial role. Such systems are composed of a huge amount of hardware, simulation models, codes, users, and data volume, and are software-intensive systems. With the continuous progress of various disciplinary technologies, the technical complexity is increasing day by day, and the research on the simulation technology of extremely large systems has become particularly urgent. The real-time transmission and interaction of simulation data are the core elements to ensure the simulation speed and efficiency of extremely large systems.
[0003] Currently, the Data Distribution Service (DDS) has been widely used in data interaction of various simulation tasks and robot systems due to its excellent real-time performance, and has been accepted by the OMG organization as the standard for data distribution. In the simulation of extremely large systems, DDS, as a method of data transmission and interaction, can adapt to complex simulation requirements. There is extensive research on extremely large systems and DDS, including the improvement of DDS to enhance data transmission real-time performance, the research on data publication / subscription services that support end-to-end quality of service in wide area networks, and the reliability and timeliness analysis of fault-tolerant distributed publication / subscription systems.
[0004] Although DDS performs well in data distribution services, the existing technologies mainly focus on improving the distribution efficiency of data publishers to ensure the timeliness and reliability of data distribution, while the research on data subscribers is relatively less. In addition, with the development of computer technology, machine learning and deep learning methods have shown better performance than traditional methods in many fields, but the research on the interaction between these methods and simulation systems and data distribution still has gaps. The existing technologies have not fully considered improving data distribution efficiency and reducing latency from the perspective of subscribers through pre-distribution technology. Therefore, the existing technologies have obvious limitations in dealing with the latency and network pressure of data distribution in large-scale systems, especially when facing high-load data distribution requirements, these problems become more prominent. Summary of the Invention
[0005] Embodiments of the present application provide a data pre-distribution method and device based on a self-attention mechanism. To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a preamble to the subsequent detailed description.
[0006] In a first aspect, an embodiment of the present application provides a data pre-distribution method based on a self-attention mechanism, and the method includes:
[0007] Dividing all nodes of a distributed simulation network to obtain multiple simulation member groups, where the distributed simulation network is used to simulate the behavior of an extremely large system;
[0008] Determining the subscription topic historical data of a preset number of consecutive historical moments before the current moment as a historical time series;
[0009] Inputting the historical time series into a pre-trained subscription topic prediction model, and outputting the data topics to be subscribed at future moments corresponding to the historical time series. The pre-trained subscription topic prediction model is trained based on a Transformer neural network with a self-attention mechanism;
[0010] When there is a release of target data related to the data topics to be subscribed, data pre-distribution is performed for multiple simulation member groups.
[0011] Optionally, dividing all nodes of the distributed simulation network includes:
[0012] Obtaining the geographical locations of the nodes in the distributed simulation network and the data subscription relationships between the nodes;
[0013] Dividing all nodes in the distributed simulation network according to the geographical locations and data subscription relationships.
[0014] Optionally, generating the pre-trained subscription topic prediction model includes the following steps:
[0015] Using a Transformer neural network with a self-attention mechanism to create a self-attention mechanism model for time series prediction;
[0016] Obtaining the historical data of data distribution in the distributed simulation network;
[0017] Inputting the historical data into the self-attention mechanism model in an autoregressive manner, training the self-attention mechanism model to capture the temporal pattern of data distribution, and obtaining the pre-trained subscription topic prediction model.
[0018] Optionally, each simulation member group carries a central data distribution node, and the central data distribution node carries a pre-distributed data basic information table and a subscription information table of this member group;
[0019] Performing data pre-distribution for multiple simulation member groups includes:
[0020] Querying the target central data distribution node corresponding to the target data based on each pre-distributed data basic information table;
[0021] Pre-distribute the target data to the target central data distribution node;
[0022] Based on the local member group subscription information table carried by the target central data distribution node, distribute the target data to the target member nodes in the simulation member group corresponding to the target central data distribution node.
[0023] Optionally, based on the local member group subscription information table carried by the target central data distribution node, distributing the target data to the target member nodes in the simulation member group corresponding to the target central data distribution node includes:
[0024] Match the target data with the local member group subscription information table carried by the target central data distribution node to obtain a matching result;
[0025] When the matching result indicates the existence of target member nodes subscribing to the target data, distribute the target data to the target member nodes;
[0026] When the target data fails to match target member nodes subscribing to the target data within a preset period, clear the target data.
[0027] Optionally, determine the subscription topic historical data at a preset number of consecutive historical moments before the current moment as the historical time series, including:
[0028] Analyze the data publish / subscribe relationship of the distributed simulation network to generate a vocabulary list of subscription topics;
[0029] Obtain the subscription topic historical data at a preset number of consecutive historical moments before the current moment from the vocabulary list;
[0030] Use the subscription topic historical data at consecutive historical moments as the historical time series.
[0031] Optionally, the Transformer neural network includes a word embedding module, a position encoding module, a self-attention module, and a feed-forward network module;
[0032] Input the historical time series into a pre-trained subscription topic prediction model, and output the subscription data topics for future moments corresponding to the historical time series, including:
[0033] Tokenize the historical time series through the word embedding module to convert the word text into the form of word vectors, obtaining a word vector sequence;
[0034] Endow the word vector sequence with time and order information through the position encoding module to obtain a target word vector sequence;
[0035] Extract the feature information of the time series of the target word vector sequence through the self-attention module;
[0036] The vector dimension of the feature information is mapped to a high-dimensional space through a feed-forward network module, and after being processed by a non-linear activation function in the high-dimensional space, it is mapped to a low-dimensional space, and the data topic to be subscribed at the future moment corresponding to the historical time series is output.
[0037] Optionally, the calculation formula of the position encoding module is:
[0038]
[0039] Where PE(pos, 2i) is the value of the 2i-th dimension in the position encoding vector, and PE(pos, 2i + 1) is the value of the (2i + 1)-th dimension in the position encoding vector, pos represents the position index in the word vector sequence, i is the dimension index, d represents the dimension of the position encoding vector, 10000 is the scaling factor, sin is the trigonometric function used to calculate the value of the position encoding vector, the sine function is used for even dimensions, and the cosine function is used for odd dimensions.
[0040] Optionally, the calculation formula of the self-attention module is:
[0041]
[0042] Where Q is the representation of the target word vector sequence, used to represent the attention degree of each element in the sequence to other elements, K represents the features of each element in the sequence, V is used to output the corresponding feature value after determining the importance of each element, T is the transpose operation, QK T represents the dot product of the query matrix Q and the key matrix K, represents the square root of the dimension of the query matrix Q and the key matrix K vectors, used to prevent the result of the dot product from being too large.
[0043] In a second aspect, an embodiment of the present application provides a data pre-distribution device based on a self-attention mechanism, and the device includes:
[0044] A network partitioning module, configured to partition all nodes of the distributed simulation network to obtain a plurality of simulation member groups, and the distributed simulation network is used to simulate the behavior of an extremely large system;
[0045] A subscription topic historical data determination module, configured to determine the subscription topic historical data of a preset number of consecutive historical moments before the current moment as the historical time series;
[0046] A model prediction module, configured to input the historical time series into a pre-trained subscription topic prediction model, and output the data topic to be subscribed at the future moment corresponding to the historical time series. The pre-trained subscription topic prediction model is trained based on a Transformer neural network with a self-attention mechanism;
[0047] A model output module, configured to perform data pre-distribution for multiple simulation member groups when there is a target data publication related to the data topic to be subscribed.
[0048] The technical solution provided by the embodiments of the present application may include the following beneficial effects:
[0049] In the embodiments of the present application, a data pre-distribution technology based on the Transformer model is adopted. Through the self-attention mechanism, in-depth analysis and learning of historical data are carried out, enabling the model to accurately predict the data subscription requirements at future moments. This prediction ability enables the system to pre-distribute data to the required nodes in advance, thereby significantly reducing data transmission latency and improving the efficiency and response speed of data distribution. In addition, by optimizing network traffic and reducing unnecessary data transmission, this solution can also effectively relieve network pressure and enhance the stability and reliability of the system.
[0050] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0052] Figure 1 is a schematic flowchart of a data pre-distribution method based on the self-attention mechanism provided by the embodiments of the present application;
[0053] Figure 2 is a schematic diagram of the network structure of a Transformer neural network provided by the embodiments of the present application;
[0054] Figure 3 is a schematic diagram of the structure of a data pre-distribution device based on the self-attention mechanism provided by the embodiments of the present application;
[0055] Figure 4 is a schematic diagram of the structure of a device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The following description and the accompanying drawings fully illustrate the specific embodiments of the present invention, enabling those skilled in the art to practice them.
[0057] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0058] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0059] In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. In addition, in the description of the present invention, unless otherwise specified, "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0060] This application provides a data pre-distribution method and device based on the self-attention mechanism to solve the problems existing in the above-mentioned related technologies. In the embodiments of this application, a data pre-distribution technology based on the Transformer model is adopted. Through the self-attention mechanism, in-depth analysis and learning of historical data are carried out, enabling the model to accurately predict the data subscription requirements at future moments. This prediction ability enables the system to pre-distribute data to the required nodes in advance, thereby significantly reducing data transmission latency and improving the efficiency and response speed of data distribution. In addition, by optimizing network traffic and reducing unnecessary data transmission, this solution can also effectively relieve network pressure and enhance the stability and reliability of the system. The following will be described in detail with exemplary embodiments.
[0061] The following will be combined with the attached Figure 1 - attached Figure 3 , and a data pre-distribution method based on the self-attention mechanism provided by the embodiments of this application will be introduced in detail. This method can be implemented depending on a computer program and can run on a data pre-distribution device based on the von Neumann architecture and based on the self-attention mechanism. This computer program can be integrated into an application or run as an independent tool class application.
[0062] Please refer to Figure 1 , which is a schematic flowchart of a data pre-distribution method based on the self-attention mechanism provided by the embodiments of this application. As Figure 1 shown, the method of the embodiments of this application may include the following steps:
[0063] S101, divide all nodes of the distributed simulation network to obtain a plurality of simulation member groups, and the distributed simulation network is used to simulate the behavior of a large-scale system;
[0064] Among them, the distributed simulation network refers to a network composed of multiple nodes. These nodes are distributed at different locations and work collaboratively through network connections to simulate the behavior of extremely large systems. In the field of simulation, such a network can handle large-scale simulation tasks and improve simulation efficiency through parallel computing. In the distributed simulation network, a node refers to a single computing unit, which can be a computer, a server, or other devices. They are responsible for executing a part of the simulation task and exchanging data with other nodes. A simulation member group refers to a group of nodes obtained by partitioning in the distributed simulation network. These nodes are grouped together because of their common characteristics for more effective management and data exchange.
[0065] In some embodiments of the present application, the specific process of partitioning all nodes of the distributed simulation network includes: obtaining the geographical locations of the nodes in the distributed simulation network and the data subscription relationships between the nodes; partitioning all nodes in the distributed simulation network according to the geographical locations and data subscription relationships.
[0066] Among them, each simulation member group carries a central data distribution node, and the central data distribution node carries a basic pre-distributed data information table and a subscription information table of this member group. The basic pre-distributed data information includes data topics, QoS requirements. The subscription information table of this member group records the subscription information of all nodes in this node group.
[0067] S102, determine the historical data of the subscription topics at a preset number of consecutive historical moments before the current moment as the historical time series;
[0068] Among them, the current moment refers to a specific time point being considered in time series analysis. A period of time before this moment is used to analyze historical data to predict future trends. The preset number refers to the number of time points of historical data points determined in advance for analysis before time series analysis, preferably the past 10 moments. Consecutive historical moments refer to a series of consecutive time points, and the data collected at these time points is used to construct the historical time series. In the simulation system, the subscription topic refers to a specific category of data or information. Nodes in the simulation system may be interested in these topics and subscribe to relevant updates or messages. Historical data refers to the data that has occurred and been recorded in the past. The historical time series is a sequence of data points arranged in chronological order.
[0069] In some embodiments of the present application, the specific process of determining the historical data of subscription topics at a preset number of consecutive historical moments before the current moment as the historical time series includes: analyzing the data publishing / subscribing relationships in the distributed simulation network to generate a vocabulary of subscription topics; obtaining, from the vocabulary, the historical data of subscription topics at a preset number of consecutive historical moments before the current moment; and using the historical data of subscription topics at the consecutive historical moments as the historical time series.
[0070] For example, all data publishing and subscribing activities in the distributed simulation network are analyzed. Through this analysis, all unique subscription topics, that is, the data types or events that the nodes in the network are concerned about and subscribe to, can be identified. These unique subscription topics are aggregated to form a vocabulary. Each entry in the vocabulary represents a unique subscription topic. Once the vocabulary is available, the next step is to extract the historical data related to each subscription topic from the vocabulary. These data cover a certain number of consecutive historical moments counting backwards from the current moment. The subscription topic data at the extracted consecutive historical moments are organized into a historical time series. A time series is a set of data points arranged in chronological order for analysis and prediction.
[0071] S103. Input the historical time series into a pre-trained subscription topic prediction model, and output the data topics to be subscribed to at the future moment corresponding to the historical time series. The pre-trained subscription topic prediction model is trained based on the Transformer neural network with self-attention mechanism.
[0072] In the embodiments of the present application, the specific process of generating the pre-trained subscription topic prediction model is as follows: use the Transformer neural network with self-attention mechanism to create a self-attention mechanism model for time series prediction; obtain the historical data of data distribution in the distributed simulation network; and input the historical data into the self-attention mechanism model in an autoregressive manner to train the self-attention mechanism model to capture the temporal law of data distribution, thereby obtaining the pre-trained subscription topic prediction model.
[0073] Among them, Transformer was first used for processing text sequences in natural language processing tasks. In the present application, the structure of this network is used to predict the topic subscription status at future moments. The original Transformer neural network includes an encoder and a decoder part as Figure 2 shown. In the present application, only the encoder part needs to be used. The encoder part is composed of a word embedding module, a position encoding module, a self-attention module, and a feed-forward network module.
[0074] Among them, each simulation member group carries a central data distribution node, and the central data distribution node carries a pre-distributed data basic information table and a subscription information table of this member group.
[0075] In some embodiments of the present application, the specific process of inputting a historical time series into a pre-trained subscription topic prediction model and outputting the to-be-subscribed data topic at a future moment corresponding to the historical time series includes: tokenizing the historical time series through a word embedding module to convert the word text into the form of word vectors, obtaining a word vector sequence; endowing the word vector sequence with time and order information through a position encoding module to obtain a target word vector sequence; extracting the feature information of the time series of the target word vector sequence through a self-attention module; mapping the vector dimension of the feature information to a high-dimensional space through a feed-forward network module, and after processing with a non-linear activation function in the high-dimensional space and mapping to a low-dimensional space, outputting the to-be-subscribed data topic at a future moment corresponding to the historical time series.
[0076] Specifically, the calculation formula of the position encoding module is:
[0077]
[0078] Among them, PE(pos, 2i) is the value of the 2i-th dimension in the position encoding vector, and PE(pos, 2i + 1) is the value of the (2i + 1)-th dimension in the position encoding vector. pos represents the position index in the word vector sequence, i is the dimension index, d represents the dimension of the position encoding vector, 10000 is the scaling factor, sin is a trigonometric function used to calculate the value of the position encoding vector, the sine function is used for even dimensions, and the cosine function is used for odd dimensions.
[0079] Specifically, the calculation formula of the self-attention module is:
[0080]
[0081] Among them, Q is the representation of the target word vector sequence, used to represent the degree of attention of each element in the sequence to other elements, K represents the features of each element in the sequence, V is used to output the corresponding feature value after determining the importance of each element, T is the transpose operation, QK T represents the dot product of the query matrix Q and the key matrix K, represents the square root of the dimension of the query matrix Q and the key matrix K vectors, used to prevent the result of the dot product from being too large.
[0082] Among them, the Transformer model is characterized by being able to process time series in parallel without processing step by step in chronological order. This way improves work efficiency, but makes the sequence lack the sign of chronological order. For example, the data status at the 0th moment and the 3rd moment is the same. This way obviously does not conform to the actual situation and lacks important time information. Therefore, positional encoding endows the sequence with time and order information, which is an essential part for the model to recognize time series. This application uses cosine positional encoding, and the specific calculation formula is the calculation formula of the positional encoding module.
[0083] Among them, the self-attention mechanism is the core module of the Transformer model. It calculates the similarity between each element in the sequence and other elements as weights, and then performs weighted summation on the original information. The scaled dot-product attention mechanism used in this study is used to extract the feature information of time series, and the calculation formula is, for example, the calculation formula of the self-attention module. Where d represents the length of the word vector. The purpose of adding this term is to scale the softmax calculation because the softmax calculation contains the exponential operation of e^x. If the value is large, the features will be annihilated after the exponential operation. Therefore, the vector depth is used for scaling to obtain more objective data indicators.
[0084] Among them, the feed-forward network module consists of two fully connected layer neural networks. A non-linear activation function is added between the two fully connected layers, thus adding non-linear calculation to the model, making the model have better non-linear ability and enhancing the generalization ability of the model. The calculation formula of the feed-forward network is: FFN(x) = max(0, x·W1 + b1)·W2 + b2. The feed-forward network module uses two fully connected layers to first map the vector dimension of the intermediate variable to a high-dimensional space, processes it with a non-linear activation function in the high-dimensional space, and then maps it to a low-dimensional space. The purpose of doing this is because the vector differences in the high-dimensional space are easier to distinguish, which can increase the generalization ability of the model.
[0085] S104, when there is a target data publication related to the data topic to be subscribed, perform data pre-distribution for multiple simulation member groups.
[0086] Among them, each simulation member group carries a central data distribution node, and the central data distribution node carries a pre-distribution data basic information table and a subscription information table of this member group.
[0087] In some embodiments of the present application, for multiple simulation member groups, the specific process of performing data pre-distribution is as follows: Based on each pre-distribution data basic information table, query the target central data distribution node corresponding to the target data; pre-distribute the target data to the target central data distribution node; based on the subscription information table of this member group carried by the target central data distribution node, distribute the target data to the target member nodes in the simulation member group corresponding to the target central data distribution node.
[0088] Specifically, the specific process of distributing the target data to the target member nodes in the simulation member group corresponding to the target central data distribution node based on the subscription information table of this member group carried by the target central data distribution node includes: matching the target data with the subscription information table of this member group carried by the target central data distribution node to obtain a matching result; when the matching result indicates that there are target member nodes subscribing to the target data, distribute the target data to the target member nodes; when the target data does not match the target member nodes subscribing to the target data within a preset period, clear the target data.
[0089] In the embodiments of the present application, the extremely large scale and complexity of the system determine that its simulation is distributed. The present application improves the distributed simulation network of the extremely large system. The entire network is divided into multiple simulation member groups according to geographical location and subscription relationship. Each group has a central data distribution node, where there is a pre-distribution data basic information table and a subscription information table of this member group. The basic information table contains the basic information of the pre-distributed data, such as data topics, QoS requirements, etc., which is convenient for data distribution management; while the subscription information table contains the subscription information of all nodes in this node group for query use. The central data distribution node is responsible for caching and managing the subscribed data of each simulation member in this group (including distribution, transmission, discarding and destruction, etc.). If information or data is pre-distributed to the central service node, the data will be directly distributed to the nodes after matching the subscription information table; if the pre-distributed data is not subscribed even when it expires, the central service node will discard it. Among them, whether the data published by the data publisher needs to be pre-distributed to the specified central service node, the present application constructs a Transformer model based on the self-attention mechanism. By identifying and analyzing historical data, it predicts the simulation members or nodes that may subscribe to this topic at future moments and then performs pre-distribution.
[0090] In the embodiments of the present application, a data pre-distribution technology based on the Transformer model is adopted. Through the self-attention mechanism, in-depth analysis and learning of historical data are carried out, enabling the model to accurately predict the data subscription requirements at future moments. This prediction ability enables the system to pre-distribute data to the required nodes in advance, thereby significantly reducing data transmission latency and improving the efficiency and response speed of data distribution. In addition, by optimizing network traffic and reducing unnecessary data transmission, this solution can also effectively relieve network pressure and enhance the stability and reliability of the system.
[0091] The following are embodiments of the device of the present invention, which can be used to execute the method embodiments of the present invention. For details not disclosed in the embodiments of the device of the present invention, please refer to the method embodiments of the present invention.
[0092] Please refer to Figure 3 It shows a schematic structural diagram of a data pre-distribution device based on the self-attention mechanism provided by an exemplary embodiment of the present invention. The data pre-distribution device based on the self-attention mechanism can be implemented as all or part of a device through software, hardware, or a combination of both. The device 1 includes a network partitioning module 10, a subscription topic historical data determination module 20, a model prediction module 30, and a model output module 40.
[0093] The network partitioning module 10 is used to partition all nodes of the distributed simulation network to obtain multiple simulation member groups. The distributed simulation network is used to simulate the behavior of an extremely large system;
[0094] The subscription topic historical data determination module 20 is used to determine the subscription topic historical data of a preset number of consecutive historical moments before the current moment as the historical time series;
[0095] The model prediction module 30 is used to input the historical time series into a pre-trained subscription topic prediction model and output the data topic to be subscribed at the future moment corresponding to the historical time series. The pre-trained subscription topic prediction model is trained based on the Transformer neural network with the self-attention mechanism;
[0096] The model output module 40 is used to perform data pre-distribution for multiple simulation member groups when there is a target data release related to the data topic to be subscribed.
[0097] It should be noted that when the data pre-distribution device based on the self-attention mechanism provided in the above embodiments executes the data pre-distribution method based on the self-attention mechanism, only the division of the above functional modules is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the data pre-distribution device based on the self-attention mechanism provided in the above embodiments and the embodiments of the data pre-distribution method based on the self-attention mechanism belong to the same concept. The implementation process is detailed in the method embodiments and will not be repeated here.
[0098] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0099] In the embodiments of the present application, a data pre-distribution technology based on the Transformer model is adopted. Through the self-attention mechanism, in-depth analysis and learning of historical data are carried out, enabling the model to accurately predict the data subscription requirements at future moments. This prediction ability enables the system to pre-distribute data to the required nodes in advance, thereby significantly reducing data transmission latency and improving the efficiency and response speed of data distribution. In addition, by optimizing network traffic and reducing unnecessary data transmission, this solution can also effectively relieve network pressure and enhance the stability and reliability of the system.
[0100] The present invention also provides a computer-readable medium, on which program instructions are stored. When the program instructions are executed by a processor, the data pre-distribution method based on the self-attention mechanism provided in each of the above method embodiments is implemented.
[0101] The present invention also provides a computer program product containing instructions. When it runs on a computer, it enables the computer to execute the data pre-distribution method based on the self-attention mechanism in each of the above method embodiments.
[0102] Please refer to Figure 4 , which is a schematic structural diagram of a device provided by an embodiment of the present application. As Figure 4 shown, the device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.
[0103] Among them, the communication bus 1002 is used to realize the connection and communication between these components.
[0104] Among them, the user interface 1003 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface.
[0105] Among them, the network interface 1004 may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface).
[0106] Among them, the processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire device 1000 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling the data stored in the memory 1005, it executes various functions of the device 1000 and processes data. Optionally, the processor 1001 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1001 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 1001 and may be implemented separately by a single chip.
[0107] Among them, the memory 1005 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory 1005 may also be at least one storage device located far from the aforementioned processor 1001. As Figure 4 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a data pre-distribution application program based on the self-attention mechanism.
[0108] In Figure 4 in the device 1000 shown, the user interface 1003 is mainly used to provide an interface for the user to input and obtain the data input by the user; while the processor 1001 can be used to call the data pre-distribution application program stored in the memory 1005 based on the self-attention mechanism, and specifically perform the following operations:
[0109] Divide all nodes of the distributed simulation network to obtain multiple simulation member groups, and the distributed simulation network is used to simulate the behavior of an extremely large system;
[0110] Determine the subscription topic historical data of a preset number of consecutive historical moments before the current moment as the historical time series;
[0111] Input the historical time series into a pre-trained subscription topic prediction model, and output the data topics to be subscribed at the future moment corresponding to the historical time series. The pre-trained subscription topic prediction model is trained based on the Transformer neural network with self-attention mechanism;
[0112] When there is a release of target data related to the data topic to be subscribed, perform data pre-distribution for multiple simulation member groups.
[0113] In one embodiment, when the processor 1001 executes dividing all nodes of the distributed simulation network, it specifically performs the following operations:
[0114] Obtain the geographical locations of the nodes in the distributed simulation network and the data subscription relationships between the nodes;
[0115] Divide all nodes in the distributed simulation network according to the geographical locations and data subscription relationships.
[0116] In one embodiment, when the processor 1001 executes generating the pre-trained subscription topic prediction model, it specifically performs the following operations:
[0117] Adopt a Transformer neural network with self-attention mechanism to create a self-attention mechanism model for realizing time series prediction;
[0118] Obtain the historical data of the distributed simulation network for data distribution;
[0119] Input the historical data into the self-attention mechanism model in an autoregressive manner, and train the self-attention mechanism model to capture the timing law of data distribution, so as to obtain the pre-trained subscription topic prediction model.
[0120] In one embodiment, when the processor 1001 executes performing data pre-distribution for multiple simulation member groups, it specifically performs the following operations:
[0121] Query the target central data distribution node corresponding to the target data based on each pre-distribution data basic information table;
[0122] Pre-distribute the target data to the target central data distribution node;
[0123] Based on the local member group subscription information table carried by the target central data distribution node, distribute the target data to the target member nodes in the simulation member group corresponding to the target central data distribution node.
[0124] In one embodiment, when the processor 1001 executes to distribute the target data to the target member nodes in the simulation member group corresponding to the target central data distribution node based on the local member group subscription information table carried by the target central data distribution node, the following operations are specifically performed:
[0125] Match the target data with the local member group subscription information table carried by the target central data distribution node to obtain a matching result;
[0126] When the matching result indicates that there are target member nodes subscribing to the target data, distribute the target data to the target member nodes;
[0127] When the target data does not match the target member nodes subscribing to the target data within a preset period, clear the target data.
[0128] In one embodiment, when the processor 1001 executes to determine the subscription topic historical data of a preset number of consecutive historical moments before the current moment as the historical time series, the following operations are specifically performed:
[0129] Analyze the data publishing / subscribing relationship of the distributed simulation network to generate a vocabulary list of subscription topics;
[0130] Obtain the subscription topic historical data of a preset number of consecutive historical moments before the current moment from the vocabulary list;
[0131] Use the subscription topic historical data of the consecutive historical moments as the historical time series.
[0132] In one embodiment, when the processor 1001 executes to input the historical time series into a pre-trained subscription topic prediction model and output the subscription data topics for future moments corresponding to the historical time series, the following operations are specifically performed:
[0133] Tokenize the historical time series through a word embedding module to convert the word text into the form of word vectors, obtaining a word vector sequence;
[0134] Assign time and order information to the word vector sequence through a position encoding module to obtain a target word vector sequence;
[0135] Extract the feature information of the time series of the target word vector sequence through the self-attention module;
[0136] Map the vector dimension of the feature information to a high-dimensional space through the feed-forward network module, and after processing with a non-linear activation function in the high-dimensional space, map it to a low-dimensional space, and output the data subscription topic for the future moment corresponding to the historical time series.
[0137] In the embodiment of the present application, a data pre-distribution technology based on the Transformer model is adopted, and the historical data is deeply analyzed and learned through the self-attention mechanism, so that the model can accurately predict the data subscription requirements at future moments. This prediction ability enables the system to pre-distribute data to the required nodes in advance, thereby significantly reducing data transmission latency and improving the efficiency and response speed of data distribution. In addition, by optimizing network traffic and reducing unnecessary data transmission, this solution can also effectively relieve network pressure and enhance the stability and reliability of the system.
[0138] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program for data pre-distribution based on the self-attention mechanism can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory or a random access memory, etc.
[0139] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A data pre-distribution method based on self-attention mechanism, characterized in that: The method comprises: Partitioning all nodes of a distributed simulation network to obtain a plurality of simulation member groups, wherein the distributed simulation network is used to simulate the behavior of a very large system; Determine the subscription topic historical data of a preset number of consecutive historical moments before the current moment as a historical time series; Input the historical time series into a pre-trained subscription topic prediction model, and output the data topic to be subscribed at a future moment corresponding to the historical time series, wherein the pre-trained subscription topic prediction model is obtained by training a Transformer neural network based on a self-attention mechanism; When there is target data release related to the data subject to be subscribed, data pre-distribution is performed for the multiple simulation member groups.
2. The method according to claim 1, characterized in that The dividing of all nodes of the distributed simulation network includes: Obtaining the geographical locations of nodes in the distributed simulation network and data subscription relationships between nodes; All nodes in the distributed simulation network are divided according to the geographical locations and the data subscription relationships.
3. The method according to claim 1, characterized in that Follow these steps to generate a pre-trained subscription topic prediction model, including: Adopt the Transformer neural network based on the self-attention mechanism to create a self-attention mechanism model for time series prediction; Acquire historical data of data distribution performed by the distributed simulation network; The historical data is input into the self-attention mechanism model by using an autoregressive method, and the self-attention mechanism model is trained to capture the temporal regularity of data distribution, thereby obtaining a pre-trained subscription topic prediction model.
4. The method according to claim 1, characterized in that: Each simulation member group carries a central data distribution node, and the central data distribution node carries a pre-distributed data basic information table and a member group subscription information table; The performing of data pre-distribution for the plurality of simulation member groups comprises: Based on each pre-distributed data basic information table, query the target central data distribution node corresponding to the target data; Pre-distributing the target data to the target central data distribution node; Based on the member group subscription information table carried by the target central data distribution node, the target data is distributed to the target member nodes in the simulation member group corresponding to the target central data distribution node.
5. The method according to claim 4, characterized in that The method of distributing the target data to the target member nodes in the simulation member group corresponding to the target central data distribution node based on the member group subscription information table carried by the target central data distribution node includes: Matching the target data with the member group subscription information table carried by the target central data distribution node to obtain a matching result; When the matching result indicates that there is a target member node that subscribes to the target data, distributing the target data to the target member node; When the target data is not matched to a target member node that subscribes to the target data within a preset period, the target data is cleared.
6. The method according to claim 1, characterized in that The determining of the subscription topic historical data of a preset number of consecutive historical moments before the current moment as a historical time series includes: Analyzing the data publish / subscribe relationship of the distributed simulation network to generate a vocabulary of subscription topics; From the vocabulary, obtain subscription topic history data of a preset number of consecutive historical moments before the current moment; The subscription topic historical data of the continuous historical moments is taken as a historical time series.
7. The method according to claim 1, characterized in that The Transformer neural network includes a word embedding module, a position encoding module, a self-attention module, and a feedforward network module; The step of inputting the historical time series into a pre-trained subscription topic prediction model and outputting a data topic to be subscribed at a future moment corresponding to the historical time series includes: Tokenizing the historical time series through the word embedding module to convert the word text into a word vector form to obtain a word vector sequence; The position encoding module is used to assign time and sequence information to the word vector sequence to obtain a target word vector sequence; Extracting feature information of the time series of the target word vector sequence through the self-attention module; The vector dimension of the feature information is mapped to a high-dimensional space through the feedforward network module, and is mapped to a low-dimensional space after being processed using a nonlinear activation function in the high-dimensional space, and the data topic to be subscribed at a future moment corresponding to the historical time series is output.
8. The method according to claim 7, characterized in that The calculation formula of the position encoding module is: Among them, PE(pos,2i) is the value of the 2i-th dimension in the position encoding vector, and PE(pos,2i+1) is the value of the 2i+1-th dimension in the position encoding vector, pos represents the position index in the word vector sequence, i is the dimension index, d represents the dimension of the position encoding vector, 10000 is the scaling factor, sin is a trigonometric function used to calculate the value of the position encoding vector, sine function is used for even dimensions, and cosine function is used for odd dimensions.
9. The method according to claim 7, characterized in that: The calculation formula of the self-attention module is: Among them, Q is the representation of the target word vector sequence, which is used to indicate the degree of attention of each element in the sequence to other elements. K represents the feature of each element in the sequence. V is used to output the corresponding feature value after determining the importance of each element. T is the transposition operation. QK T represents the dot product of the query matrix Q and the key matrix K, Represents the square root of the dimension of the query matrix Q and the key matrix K vector, used to prevent the result of the dot product from being too large.
10. A data pre-distribution device based on a self-attention mechanism, characterized in that: The device comprises: A network partitioning module, used to partition all nodes of a distributed simulation network to obtain multiple simulation member groups, wherein the distributed simulation network is used to simulate the behavior of a very large system; A subscription topic historical data determination module is used to determine the subscription topic historical data of a preset number of consecutive historical moments before the current moment as a historical time series; A model prediction module is used to input the historical time series into a pre-trained subscription topic prediction model, and output the data topic to be subscribed at a future moment corresponding to the historical time series, wherein the pre-trained subscription topic prediction model is obtained by training a Transformer neural network based on a self-attention mechanism; The model output module is used to perform data pre-distribution for the multiple simulation member groups when there is target data release related to the data subject to be subscribed.