State storage method of stream computing system, state storage system and electronic equipment
By pre-creating multiple state storage systems in the stream computing system and using hash tables and first linked lists to store different versions of data, the problem of low state storage performance in stream computing systems is solved, achieving more efficient data management and querying.
Patent Information
- Application Number
- CN202410524314.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-28
- Publication Date
- 2025-10-28
AI Technical Summary
Existing stream computing systems have low state storage performance, making it difficult to meet user needs.
By pre-creating multiple state storage systems in the key-value storage system, managing data in a multi-concurrency manner, and using hash tables and first linked lists to store different versions of data, the computation speed and performance are improved.
It improves the state storage performance of the stream computing system, reduces disk read and lookup time, saves memory, and enhances the overall system performance.
Smart Images

Figure CN120848952A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of stream computing, and more specifically, to a state storage method, state storage system, and electronic device for a stream computing system. Background Technology
[0002] With the development of stream computing systems, the amount of data generated has also increased. Much of this data will be used later, so the storage of massive amounts of data is an urgent problem to be solved. In general, state storage can be used to store data. State storage is an important component of stream computing systems and a key factor that determines its throughput, latency, CPU / memory consumption, etc. However, the throughput and performance of commonly used state storage are relatively low, making it difficult to meet user needs.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] This application provides a state storage method, state storage system, and electronic device for a stream computing system, in order to at least solve the technical problem of low performance of state storage in related technologies.
[0005] According to one aspect of the embodiments of this application, a state storage method for a stream computing system is provided, comprising: responding to a state processing instruction sent by the stream computing system, determining a target state storage system corresponding to the state processing instruction from a plurality of state storage systems pre-created for state tables stored in a key-value storage system, wherein the state processing instruction includes: a state query instruction, a state write instruction, and a persistence instruction; and performing data management on a hash table and / or a first linked list in the target state storage system based on the state processing instruction, wherein the hash table stores at least one second linked list corresponding to identifier data, the second linked list is configured to store different versions of data, and the first linked list is used to store data that has not been persistently stored in the key-value storage system.
[0006] According to one aspect of the embodiments of this application, a state storage system is provided, comprising: multiple interfaces connected to a stream computing system, different interfaces being used to receive different instructions sent by the stream computing system, wherein the different instructions include: a state write instruction, a state query instruction, and a persistence instruction; a first linked list for storing data that is not persisted to a key-value storage system; and a hash table containing at least one second linked list, different second linked lists corresponding to different identifier data, and different second linked lists being configured to store different versions of data, version information of different versions of data, and persistence information of whether different versions of data are persisted to a key-value storage system.
[0007] According to one aspect of the embodiments of this application, a storage system for a stream computing system is provided, comprising: a key-value storage system storing multiple state tables, wherein different state tables are used to store data of different data types generated by the stream computing system; and a set of multiple state storage systems corresponding one-to-one with the multiple state tables, wherein the set of state storage systems includes at least one state storage system.
[0008] According to another aspect of the embodiments of this application, a state storage device for a stream computing system is also provided, comprising: a first determining module, configured to determine a target state storage system corresponding to the state processing instruction from a plurality of state storage systems pre-created for a state table stored in a key-value storage system in response to a state processing instruction sent by the stream computing system, wherein the state processing instruction includes: a state query instruction, a state write instruction, and a persistence instruction; and a data management module, configured to perform data management on a hash table and / or a first linked list in the target state storage system based on the state processing instruction, wherein the hash table stores at least one second linked list corresponding to identifier data, the second linked list is configured to store different versions of data, and the first linked list is used to store data that has not been persistently stored in the key-value storage system.
[0009] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of this application when it runs.
[0010] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of this application.
[0011] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods of various embodiments of this application.
[0012] In this embodiment, in response to a state processing instruction sent by the stream computing system, a target state storage system corresponding to the state processing instruction is determined from multiple state storage systems pre-created for the state tables stored in the key-value storage system. The state processing instruction includes a state query instruction, a state write instruction, and a persistence instruction. Based on the state processing instruction, data management is performed on the hash table and / or the first linked list in the target state storage system. The hash table stores at least one second linked list corresponding to the identifier data. The second linked list is configured to store different versions of data, and the first linked list is used to store data that has not been persistently stored in the key-value storage system. It is noteworthy that multiple state storage systems can be pre-created for the state tables stored in the key-value storage system. This allows for data management using a multi-concurrency approach, improving computation speed. Furthermore, different versions of data are stored using hash tables in the target state storage system. Upon receiving a state processing instruction, the target state storage system corresponding to the instruction can be determined promptly. Based on the instruction, data management is performed on the hash table and / or the first linked list in the target state storage system. Storing at least one identifier using a hash table not only saves memory but also acts as a cache, reducing disk reads and lookups of the target state. This improves the performance of the target state storage system and solves the technical problem of low performance in state storage in related technologies.
[0013] It is worth noting that the general description above and the detailed description that follow are merely for illustrative purposes and do not constitute a limitation on this application. Attached Figure Description
[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0015] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a state storage method for a stream computing system according to an embodiment of this application;
[0016] Figure 2 This is a structural block diagram of a computing environment according to an embodiment of this application;
[0017] Figure 3 This is a structural block diagram of a service mesh according to an embodiment of this application;
[0018] Figure 4 This is a flowchart of the state storage method of the stream computing system according to Embodiment 1 of this application;
[0019] Figure 5This is a schematic diagram of a state storage system for a stream computing system according to an embodiment of this application;
[0020] Figure 6 This is a schematic diagram of a state storage system according to an embodiment of this application;
[0021] Figure 7 This is a schematic diagram of the storage system of a stream computing system according to Embodiment 3 of this application;
[0022] Figure 8 This is a schematic diagram of a state storage device for a stream computing system according to Embodiment 4 of this application;
[0023] Figure 9 This is a structural block diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:
[0027] Stream computing, also known as streaming computing, is a computing model used to process and analyze real-time data streams. Streaming data can come from various sources, such as sensors, log files, network events, or online transactions. Compared to traditional batch processing models, stream computing does not require waiting for the complete collection and storage of the dataset, but processes data elements on an instantaneous basis, usually as soon as they arrive.
[0028] State storage system: also known as State Store, refers to the component or mechanism used in a stream computing system to store and manage state information. State refers to information accumulated from past events. It provides memory for stream computing applications, allowing state tracking and updates when processing real-time data streams.
[0029] State table: The data in the state storage can be stored in the form of a table in a key-value store system / database, and this table is called a state table.
[0030] Checkpoints: In stream computing, checkpoints in state storage are a fault-tolerance mechanism used to periodically capture snapshots of the application state. Checkpoints allow the stream computing system to recover from a known consistent state in the event of a failure, rather than starting from scratch. This mechanism is crucial for ensuring the reliability of stream computing tasks and the accuracy of data.
[0031] Key-value storage, also known as key-value pair storage, is a key-value pair storage method, in which each key-value pair has a unique key and a corresponding value.
[0032] Example 1
[0033] According to an embodiment of this application, a state storage method for a stream computing system is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0034] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a state storage method for a stream computing system, according to an embodiment of this application. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0035] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0036] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method in the embodiments of this application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby implementing the method in the above embodiments. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0037] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0038] The display can be, for example, a touchscreen liquid crystal display (LCD), which allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0039] Figure 1The hardware structure block diagram shown can serve not only as an exemplary block diagram of the aforementioned computer terminal 10 (or mobile device), but also as an exemplary block diagram of the aforementioned server. In one optional embodiment, Figure 2 The use of the above is illustrated in a block diagram. Figure 1 The computer terminal 10 (or mobile device) shown is an embodiment of a computing node in computing environment 201. Figure 2 This is a structural block diagram of a computing environment according to an embodiment of this application, such as... Figure 2 As shown, computing environment 201 includes multiple computing nodes (such as servers) running on a distributed network (represented as 210-1, 210-2, ..., in the diagram). Each computing node contains local processing and memory resources, and end user 202 can remotely run applications or store data within computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 within computing environment 201, representing services "A", "D", "E", and "H", respectively.
[0040] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or requests of end user 202 can be provided to ingress gateway 230. Ingress gateway 230 may include a corresponding agent to handle the provisioning and / or requests for services (one or more services provided in computing environment 201).
[0041] Services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services may be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. VM-based virtualization can simulate a real computer by initializing a virtual machine, executing programs and applications without directly accessing any actual hardware resources. While the machine is virtualized by a virtual machine, container-based virtualization can launch containers to virtualize an entire operating system (OS), allowing multiple workloads to run on a single OS instance.
[0042] In one embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). For example, such as Figure 2As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively referred to as Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively referred to as containers). One or more containers within a Pod handle requests related to one or more corresponding functions of the service. Proxy 245 typically controls service-related network functions such as routing and load balancing. Other services can also be equipped with similar Pods.
[0043] During operation, executing a user request from end user 202 may require invoking one or more services in computing environment 201, and executing one or more functions of one service may require invoking one or more functions of another service. For example... Figure 2 As shown, service "A" 220-1 receives user requests from terminal user 202 from ingress gateway 230. Service "A" 220-1 can call service "D" 220-2, and service "D" 220-2 can request service "E" 220-3 to perform one or more functions.
[0044] The aforementioned computing environment can be a cloud computing environment, where resource allocation is managed by cloud services, allowing functionality development without needing to consider implementation, adjustment, or server scaling. This computing environment allows developers to execute event-responsive code without building or maintaining complex infrastructure. Services can be partitioned into a set of functions that can automatically and independently scale, rather than scaling a single hardware device to handle potential loads.
[0045] In another alternative embodiment, Figure 3 The use of the above is illustrated in a block diagram. Figure 1 The computer terminal 10 (or mobile device) shown is an embodiment of a service mesh. Figure 3 This is a structural block diagram of a service mesh according to an embodiment of this application, such as... Figure 3 As shown, the service mesh 300 is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to the decomposition of an application into multiple smaller services or instances, which are distributed across different clusters / machines.
[0046] like Figure 3 As shown, a microservice may include application service instance A and application service instance B, which together form the functional application layer of service mesh 300. In one implementation, application service instance A runs as a container / process 308 on machine / workload container group 314 (Pod), and application service instance B runs as a container / process 310 on machine / workload container group 316 (Pod).
[0047] In one implementation, application service instance A can be a product query service, and application service instance B can be a product order placement service.
[0048] like Figure 3 As shown, application service instance A and grid agent (sidecar) 303 coexist in machine workload container group 314, and application service instance B and grid agent 305 coexist in machine workload container 316. Grid agents 303 and 305 form the data plane layer of service mesh 300. Grid agents 303 and 305 run as containers / processes 304 and 306 respectively, and can receive requests 312 for product query services. Grid agent 303 and application service instance A can communicate bidirectionally, and grid agent 305 and application service instance B can also communicate bidirectionally. Furthermore, grid agents 303 and 305 can also communicate bidirectionally with each other.
[0049] In one implementation, traffic from application service instance A is routed to the appropriate destination via mesh proxy 303, and network traffic from application service instance B is routed to the appropriate destination via mesh proxy 305. It should be noted that the network traffic mentioned here includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representative State Transfer (REST), high-performance, general-purpose open-source frameworks (Google Remote Procedure Call, gRPC), and open-source in-memory data structure storage systems (Redis).
[0050] In one implementation, the functionality of the extended data plane layer can be achieved by writing custom filters for the agents (Envoy) in service mesh 300. The service mesh agent configuration can enable the service mesh to correctly proxy service traffic, achieving service interoperability and service governance. Mesh agents 303 and 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.
[0051] like Figure 3As shown, the service mesh 300 also includes a control plane layer. This control plane layer can consist of a set of services running in a dedicated namespace, hosted by a managed control plane component 301 within machine / workload container groups (machine / Pods) 302. Figure 3 As shown, the managed control plane component 301 communicates bidirectionally with grid agents 303 and 305. The managed control plane component 301 is configured to perform several control and management functions. For example, the managed control plane component 301 receives telemetry data transmitted by grid agents 303 and 305 and can further aggregate this telemetry data. In addition to these services, the managed control plane component 301 can also provide a user-facing Application Programming Interface (API) to facilitate easier manipulation of network behavior and to provide configuration data to grid agents 303 and 305.
[0052] Under the aforementioned operating environment, this application provides the following: Figure 4 The state storage method of the stream computing system is shown. Figure 4 This is a flowchart of the state storage method of the stream computing system according to Embodiment 1 of this application, as follows: Figure 4 As shown, after the stream computing system issues a state processing instruction, the following operations can be performed.
[0053] Step S402: In response to the state processing instruction sent by the stream computing system, determine the target state storage system corresponding to the state processing instruction from multiple state storage systems pre-created for the state tables stored in the key-value storage system. The state processing instruction includes: state query instruction, state write instruction, and persistence instruction.
[0054] The aforementioned stream computing system can be a computing system for processing real-time data streams. It can extract, process, and analyze data from the data stream in real time, and continuously output computing results as the data stream is constantly updated. Stream computing systems typically feature low latency, high throughput, and scalability. They can handle massive amounts of real-time data and support complex data processing and analysis tasks, such as real-time prediction and real-time alerts. Application areas of stream computing systems include finance, e-commerce, the Internet of Things, and smart manufacturing. They can help enterprises understand and respond to data streams in real time, thereby achieving more intelligent and efficient operations.
[0055] The aforementioned state write command can be issued by the stream computing system. The state write command can be used to trigger the data writing process. Optionally, the state write command can specify what the data to be written is, that is, the state write command can contain the data to be written.
[0056] The aforementioned state query command can be issued by the stream computing system to trigger the data query process, that is, to trigger a query to see if the data that the user wants to query exists in the state storage system. Optionally, the state query command can include the identifier data of the data to be queried.
[0057] The aforementioned state persistence instructions can be used to trigger a persistence process that stores at least one data in the state storage system to the key-value storage system, and to trigger a process that updates the preset version information of the state storage system, wherein the preset version information can be the global version number of the state storage system.
[0058] The aforementioned target state storage system can be a component used in a stream computing system to save and manage state information. That is, the target state storage system can realize functions such as data writing, data querying, and data persistence. The aforementioned data can be any type of data generated by the stream computing system, such as sensor data, log files, network events, or online transaction data.
[0059] The aforementioned target state storage system can be used to store different versions of data generated by the stream computing system at different working stages based on state processing instructions issued by the stream computing system, or to query already stored data based on state processing instructions issued by the stream computing system, or to store all non-persistent data to a persistent key-value storage system based on persistence instructions issued by the stream computing system. It should be noted that the stream computing system can contain multiple state storage systems. Optionally, the underlying state storage system corresponds to a state table in a persistent key-value storage system, with different state tables used to store different states of the stream computing system. Optionally, the persistent key-value storage system can adopt a row-based database based on a log-structured merge tree database (LSM database). The underlying storage of a single-machine key-value storage system should use a distributed file system to facilitate the migration of the state storage system between nodes. The state storage system has a hash table in memory. The keys of the hash table correspond to the keys input from the upper layer, and the values are a multi-version linked list. The multi-version linked list stores multiple versions of the value corresponding to the key. Optionally, as the upper layer is continuously updated, the same key will generate multiple versions of the value. The multi-version linked list also records the version information of different versions of the data and whether it is persisted. The hash table stores two types of data: one is the data that has just been written by the upper layer but has not yet been persisted, and the other is the data that has been persisted, that is, the cache, which can be evicted when memory is tight.
[0060] In one optional embodiment, for a state table, to improve computation speed, multiple concurrent read and write operations are employed, and the concurrency level may change at any time. Therefore, different state types can correspond to different state tables. Thus, multiple state storage systems can be pre-created for different state tables. After obtaining a state processing instruction, the target storage system corresponding to the state processing instruction can be determined based on the data corresponding to the instruction. Optionally, if the state processing instruction is a state query instruction, the target state storage system can be determined based on the identifier data of the data to be queried; if the state processing instruction is a state write instruction, the target state storage system can be determined based on the identifier data of the data to be written; and if the state processing instruction is a persistence instruction, all state storage systems can be identified as the target state storage system.
[0061] In one optional embodiment, the user can view and manage status processing instructions through the management interface or command-line tool of the stream computing system, thereby triggering the stream computing system to send status processing instructions. Optionally, this can be done in the following ways: First, log in to the management interface of the stream computing system or use the command-line tool to enter the system management interface. Second, the user can view the relevant options or commands for status processing instructions. Optionally, the management interface may have a specific page or tab for viewing and managing status processing instructions, while the command-line tool may have specific commands for querying and managing status processing instructions. Third, the user can view and manage status writing instructions according to the options or commands provided by the system. This may include viewing sent status processing instructions, creating new status writing instructions, modifying existing status processing instructions, etc. Finally, the user can perform the corresponding operations as needed to obtain the status processing instructions sent by the stream computing system.
[0062] In another alternative embodiment, the stream computing system can generate and send status processing instructions based on its operating status. For example, when the stream computing system starts processing data, it can send a status query instruction to obtain the data to be processed. After the stream computing system finishes processing the data, it can send a status write instruction to store the processed data in the key-value storage system.
[0063] Step S404: Based on the state processing instructions, perform data management on the hash table and / or the first linked list in the target state storage system. The hash table stores at least one second linked list corresponding to the identifier data. The second linked list is configured to store different versions of data. The first linked list is used to store data that has not been persistently stored in the key-value storage system.
[0064] The first linked list mentioned above can be used to store data that is not persistently stored in a key-value storage system.
[0065] The second linked list mentioned above can be a linked list used in the target state storage system to store data corresponding to state write instructions or state query instructions. Optionally, different versions of data with the same key can be stored in the same linked list. To distinguish between different versions of data, the second linked list can also store version information of different versions of data, including but not limited to version number or timestamp, for example, version information can be version number 1.1, 1.2, etc.
[0066] In an optional embodiment, when the state processing instruction is a state query instruction, data management can be performed on the hash table in the target state storage system based on the state query instruction to determine whether the hash table stores the data to be queried corresponding to the state query instruction. Optionally, the second linked list corresponding to the identifier data to be queried in the hash table can be queried based on the identifier data of the data to be queried contained in the state query instruction. When the state processing instruction is a state write instruction, data management can be performed on the hash table and the first linked list in the target state storage system based on the state write instruction. Optionally, the data to be written can be written to the second linked list corresponding to the identifier data in the hash table based on the identifier data of the data to be written contained in the state write instruction. In certain cases, the data to be written also needs to be written to the first linked list. When the state processing instruction is a persistence instruction, since this instruction is used to store all data not stored in the key-value storage system into the key-value storage system, data management can be performed on the first linked list. Optionally, all data in the first linked list can be stored into the key-value storage system.
[0067] Optionally, the identifier data carried in the state processing instruction can be obtained by the data identifier carried in the state processing instruction, and then its hash value can be calculated by a hash function. Next, the hash value is used as an index, or the corresponding data calculation is performed on the hash value and the result is used as an index. Thus, the corresponding position can be located from the hash table of the target state storage system according to the index. Further, at that position, the stored second linked list is obtained. Finally, data can be searched or inserted into the second linked list.
[0068] In another optional embodiment, after receiving the status query instruction sent by the stream computing system, a query process for querying the data to be queried can be triggered. Optionally, the system can query whether there is target identifier data in the target state storage system that is the same as the identifier data of the data to be queried. If the target identifier data exists in the hash table of the target state storage system, the data can be obtained from the second linked list corresponding to the target identifier data.
[0069] In another optional embodiment, after receiving the state write instruction sent by the stream computing system, a write process for writing the data to be written can be triggered. Optionally, a second linked list corresponding to the identifier data of the data to be written can be determined from the target state storage system, and the data to be written can be written to the second linked list. In addition, if the data to be written is not persisted to the key-value storage system, the data to be written can be further written to the first linked list.
[0070] In another optional embodiment, after receiving the persistence instruction sent by the stream computing system, the data persistence storage process can be triggered. Optionally, data that has not been persisted to the key-value storage system stored in the first linked list can be read, thereby storing all the read data into the key-value storage system.
[0071] It should be noted that before managing data in the hash table and / or the first linked list in the target state storage system based on state processing instructions, the validity and permissions of the state processing instructions must first be verified. If the state processing instructions are valid and have the corresponding permissions, the target state storage system can proceed to the next step. Then, the target state storage system manages data in the hash table and / or the first linked list in the target state storage system according to the state processing instructions and returns the processing results to the stream computing system.
[0072] Figure 5 This is a schematic diagram of a state storage system for a stream computing system according to an embodiment of this application, such as... Figure 5 As shown, this system may include a central memory control module and multiple state storage systems. Figure 5 This example uses four state storage systems (state storage system 1, state storage system 2, state storage system 3, and state storage system 4) as examples, along with the state table of at least one persistent key-value storage system. Figure 5 This example uses two state tables (State Table 1 and State Table 2) and a distributed file system. The central memory control module controls the memory usage of the state storage system, such as by collecting memory information and performing memory eviction, thereby avoiding the overhead of global lock conflicts and memory overuse caused by multiple state stores. State storage systems 1 and 2 are based on State Table 1 (key-value persistent storage), while state storage systems 3 and 4 are based on State Table 2 (key-value persistent storage). These two state tables can be used to gradually store temporarily stored data in the state storage system into the distributed file system. Optionally, the target state storage system can contain a hash table (not shown in the diagram) and a non-persistent linked list (i.e., the first linked list mentioned above). The hash table contains multiple different multi-version linked lists (i.e., the second linked list mentioned above), for example... Figure 5The state storage system 1 contains a hash table and a non-persistent linked list. The hash table contains three multi-version linked lists. The first multi-version linked list stores the values 1 and 2 corresponding to key 1. The second multi-version linked list stores the value 3 corresponding to key 2. The third multi-version linked list stores the value "deleted" corresponding to key 3, that is, the value corresponding to key 3 has been deleted. The non-persistent linked list stores the values 2 and 3.
[0073] In this embodiment, in response to a state processing instruction sent by the stream computing system, a target state storage system corresponding to the state processing instruction is determined from multiple state storage systems pre-created for the state tables stored in the key-value storage system. The state processing instruction includes a state query instruction, a state write instruction, and a persistence instruction. Based on the state processing instruction, data management is performed on the hash table and / or the first linked list in the target state storage system. The hash table stores at least one second linked list corresponding to the identifier data. The second linked list is configured to store different versions of data, and the first linked list is used to store data that has not been persistently stored in the key-value storage system. It is noteworthy that multiple state storage systems can be pre-created for the state tables stored in the key-value storage system. This allows for data management using a multi-concurrency approach, improving computation speed. Furthermore, different versions of data are stored using hash tables in the target state storage system. Upon receiving a state processing instruction, the target state storage system corresponding to the instruction can be determined promptly. Based on the instruction, data management is performed on the hash table and / or the first linked list in the target state storage system. Storing at least one identifier using a hash table not only saves memory but also acts as a cache, reducing disk reads and lookups of the target state. This improves the performance of the target state storage system and solves the technical problem of low performance in state storage in related technologies.
[0074] In the above embodiments of this application, when the state processing instruction includes a state query instruction, data management is performed on the hash table and / or the first linked list in the target state storage system based on the state processing instruction, including: determining the identifier data of the data to be queried carried in the state query instruction; determining whether there is target identifier data in the target state storage system that is the same as the identifier data of the data to be queried; if the target identifier data exists in the hash table of the target state storage system, determining the second linked list corresponding to the target identifier data from the target state storage system, wherein the second linked list also stores version information of different versions of data; obtaining the data corresponding to the largest version information in the second linked list corresponding to the target identifier data as the data to be queried, and sending the data to be queried to the stream computing system.
[0075] The aforementioned target identification data can be the key of the data stored in the target state storage system.
[0076] The data corresponding to the maximum version information mentioned above can be the latest version of the different versions of data stored in the second linked list. For example, the maximum version information can be the maximum version number or the latest timestamp.
[0077] In an optional embodiment, after the state processing instruction sent by the stream computing system includes a state query instruction, a query process for querying the data to be queried can be triggered. Optionally, since the state query instruction contains the identifier data of the data to be queried, the system can query whether there is target identifier data in the target state storage system that is the same as the identifier data of the data to be queried. If the target identifier data exists in the hash table of the target state storage system, the second linked list corresponding to the target identifier data can be determined from the target state storage system, and the data corresponding to the largest version information in the second linked list can be used as the data to be queried. That is, the first version data is determined as the data to be queried. Further, the data to be queried can be sent to the stream computing system. That is, for the key queried by the upper layer, the system first checks whether the key exists in the hash table. If it exists, the latest version value is returned.
[0078] In the above embodiments of this application, when the target identifier data does not exist in the hash table of the target state storage system, the method further includes: sending a query instruction to the key-value storage system, wherein the query instruction is used to query whether the key-value storage system stores the data to be queried; if the data to be queried is received from the key-value storage system, sending the data to be queried to the stream computing system; if the data to be queried is not received from the key-value storage system, sending a prompt message to the stream computing system, wherein the prompt message is used to indicate that the key-value storage system does not store the data to be queried.
[0079] In one optional embodiment, if the target identifier data does not exist in the hash table of the target state storage system, a query command can be sent to the underlying persistent key-value storage system. If the query data is received from the key-value storage system, the query data can be sent to the stream computing system. That is, if the target identifier data does not exist in the hash table of the target state storage system, a query is made to the underlying persistent key-value storage system. If the target identifier data exists in the underlying system, the value of the corresponding target identifier data is returned. If the query data is not received from the key-value storage system, a prompt message can be sent to the stream computing system. That is, if the underlying system returns that the key does not exist, it means that the key is not in the state storage system, and a prompt message indicating that the key does not exist will be returned.
[0080] In the above embodiments of this application, when receiving the data to be queried sent by the key-value storage system, the method further includes: storing the target identifier data and the data to be queried in the hash table of the target state storage system, and marking the persistence information of the data to be queried as already persisted.
[0081] In one optional embodiment, if the key-value storage system sends the data to be queried, it means that the underlying persistent key-value storage system has the data to be queried. However, the target state storage system does not have the data to be queried in its hash table. Therefore, the target identifier data and the data to be queried can be stored in the hash table of the target state storage system. That is, if the key is not in the hash table, it needs to be inserted into the hash table along with its corresponding value. For a non-existent key, the value is a deleted marker, and the state of the value is marked as persistent. That is, the persistence information of the data to be queried is marked as persistent.
[0082] In the above embodiments of this application, when the state processing instruction includes a state write instruction, data management is performed on the hash table and / or the first linked list in the target state storage system based on the state processing instruction, including: determining the data to be written carried in the state write instruction and the identifier data of the data to be written; determining the second linked list corresponding to the identifier data of the data to be written from the hash table of the target state storage system, wherein the second linked list also stores version information of different versions of data; obtaining the data corresponding to the maximum version information from the second linked list corresponding to the identifier data of the data to be written, and determining the target state of the data corresponding to the maximum version information, wherein the target state includes: not triggered persistence, not completed persistence, and already persisted; and writing the data to be written to the target state storage system based on the target state.
[0083] In one optional embodiment, after determining the target state of the data corresponding to the maximum version information, the data to be written can be controlled to be written to the target state storage system based on the target state of the data corresponding to the maximum version information. Optionally, during the storage process, assuming the target state is "not triggered for persistence," it means that the data stored in the target state storage system is non-persistent data, and modifications to this data will not affect the persistence process. Therefore, the data to be written can be directly stored in the target state storage system. Assuming the target state is "already persisted," it means that the data stored in the target state storage system has been persisted to the key-value storage system, and modifications to this data will not affect the persistence process. Therefore, the data to be written can be directly stored in the target state storage system. Assuming the target state is "not yet persisted," it means that data stored in the target state storage system is in the process of being stored to the key-value storage system. Therefore, to avoid errors in the persisted data in the key-value storage system, the data corresponding to the maximum version information cannot be directly overwritten, and new data needs to be added to the second linked list.
[0084] In one optional embodiment, the target state can be stored as a value of different versions of data in a second linked list, thereby retrieving the target state of the data corresponding to the maximum version information from the second linked list corresponding to the identifier data of the data to be written. In another optional embodiment, the second linked list can also store persistence information of different versions of data. This persistence information is used to characterize whether different versions of data are persistently stored in the key-value storage system. Therefore, based on the persistence information of the data corresponding to the maximum version information, it can be determined whether the target state of the data corresponding to the maximum version information has been persisted. Since persistence is an asynchronous operation and does not affect data querying and writing, it can be further combined with the preset version information of the target state storage system to determine whether the target state of the data corresponding to the maximum version information has not been persisted.
[0085] In the above embodiments of this application, the second linked list also stores version information of different versions of data, and persistence information of whether different versions of data are persistently stored in the key-value storage system; determining the target state of the data corresponding to the maximum version information includes: obtaining the target persistence information and target version information of the data corresponding to the maximum version information from the second linked list corresponding to the identifier data of the data to be written; if the target version information is the same as the preset version information of the target state storage system, the target state is determined to be non-permanent; if the target version information is different from the preset version information, and the target persistence information is that the data corresponding to the maximum version information is not persistent, the target state is determined to be non-permanent; if the target version information is different from the preset version information, and the target persistence information is that the data corresponding to the maximum version information is persistent, the target state is determined to be persistent.
[0086] The aforementioned target persistence information can be whether the data corresponding to the maximum version information is persistently stored in the key-value storage system.
[0087] The target version information mentioned above can be the version information of the data corresponding to the maximum version information, such as the version number or timestamp.
[0088] The aforementioned preset version information can be the global version number of the target state storage system. Optionally, the global version number can be stored in the memory of the state storage system.
[0089] In an optional embodiment, the second linked list also stores version information of different versions of data, such as version numbers of different versions of data, and persistence information of whether these different versions of data are persistently stored in the key-value storage system. When determining the target state of the data corresponding to the maximum version information, the target persistence information and target version information of the data corresponding to the maximum version information can be obtained first, and the target version information can be compared with the preset version information of the target state storage system. Optionally, if the target version information is the same as the preset version information of the target state storage system, the target state can be considered as not having triggered persistence. Optionally, if the target version information is different from the preset version information, and the target persistence information is that the first version of the data is not persisted, the target state can be considered as not having completed persistence. Optionally, if the target version information is different from the preset version information, and the target persistence information is that the first version of the data is persisted, the target state can be considered as having been persisted.
[0090] In the above embodiments of this application, writing the data to be written to the target state storage system based on the target state includes: when the target state is that persistence has not been triggered or has been persisted, overwriting the data corresponding to the maximum version information with the data to be written; when the target state is that persistence has not been completed, storing the data to be written to the second linked list corresponding to the identifier data of the data to be written, and storing the preset version information as the version information of the data to be written to the second linked list corresponding to the identifier data of the data to be written.
[0091] In one optional embodiment, if the target state is "not persisted," it can be assumed that the data corresponding to the maximum version information has not yet been stored in the key-value storage system, and no persistence instruction to store the data corresponding to the maximum version information in the key-value storage system has been received. Therefore, the data to be written can overwrite the data in the key-value storage system. If the target state is "already persisted," it can be determined that the data corresponding to the maximum version information has been stored in the key-value storage system. However, since each piece of data written has a higher version than the previous piece of data written, the data to be written with a higher version can directly overwrite the data corresponding to the maximum version information. Optionally, if the target state is "not persisted," it means that the data stored in the target state storage system is being persistently stored in the key-value storage system. Therefore, the data to be written can first be stored in the second linked list corresponding to the identifier data of the data to be written, as a new value in the second linked list, and the preset version information can be used as the version information of the data to be written and stored together in the second linked list corresponding to the identifier data of the data to be written.
[0092] In the above embodiments of this application, the method further includes: storing the data to be written to the first linked list of the target state storage system when the target state is either not yet persisted or has already been persisted.
[0093] In one optional embodiment, if the target state is either not yet persisted or has already been persisted, the data line to be written can be stored in the first linked list of the target state storage system, and then the data stored in the first linked list of different state storage systems can be persisted and stored in the key-value storage system gradually.
[0094] In the above embodiments of this application, when the state processing instruction includes a persistence instruction, data management is performed on the hash table and / or the first linked list in the target state storage system based on the state processing instruction, including: updating the preset version information of the target state storage system; obtaining at least one piece of data stored in the first linked list of the target state storage system; storing at least one piece of data in a key-value storage system; and clearing the first linked list.
[0095] The aforementioned state persistence instruction can be used to trigger the process of storing at least one data that is not persistently stored in the target state storage system into the key-value storage system, and to trigger the process of updating the preset version information of the target state storage system.
[0096] The aforementioned preset version information can be the global version number of the target storage state system.
[0097] In one optional embodiment, if a state persistence instruction is received from the stream computing system, the preset version information of the target state storage system can be updated based on the state persistence instruction, and at least one data stored in the first linked list of the target state storage system can be stored in the key-value storage system. Furthermore, in order to avoid data duplication persistence, the first linked list can be cleared after at least one data is successfully stored in the key-value storage system.
[0098] Optionally, after each data storage of at least one data stored in the first linked list of the target state storage system is transferred to the key-value storage system, the global version number of the target state storage system needs to be incremented. That is, after the upper layer triggers persistence and transfers at least one data storage of the first linked list to the key-value storage system, the state storage system increments the global version number by one.
[0099] In the above embodiments of this application, after storing at least one data in a key-value storage system, the method further includes: modifying the persistence information of at least one data stored in the target state storage system to indicate that at least one data has been persistently stored in the key-value storage system.
[0100] In one optional embodiment, after storing at least one piece of data in a key-value storage system, the persistence information of at least one piece of data stored in the target state storage system can be modified. Optionally, the persistence information of at least one piece of data stored in the target state storage system can be modified to indicate that at least one piece of data has been persistently stored in the key-value storage system.
[0101] In the above embodiments of this application, after storing at least one data in a key-value storage system, the method further includes: determining a second linked list corresponding to the identifier data of at least one data from a hash table in the target state storage system, wherein the second linked list also stores version information of different versions of data; if at least one data is the data corresponding to the maximum version information in the second linked list corresponding to the identifier data of at least one data, then using at least one data as a cache item in the target state storage system; if at least one data is not the data corresponding to the maximum version information in the second linked list corresponding to the identifier data of at least one data, then releasing at least one data.
[0102] In one optional embodiment, all data in the unpersisted linked list is retrieved and written to the underlying persistent key-value storage system. Simultaneously, the unpersisted linked list is cleared. The writing process is asynchronous and does not affect the query and write processes. After writing is complete, the written value is marked as persistent. If this value is already the latest version of the key (i.e., the data corresponding to the maximum version information mentioned above), then the key can be used as a cached item, and... Figure 5The central memory control module shown here performs memory replacement based on the memory status. The specific replacement process is described in detail later. If the value is not the latest version of the key at this time, the value can be released.
[0103] In the above embodiments of this application, determining the target state storage system corresponding to a state processing instruction from multiple state storage systems pre-created for the state table stored in the key-value storage system includes: when the state processing instruction includes a state query instruction, determining the target state storage system from the multiple state storage systems based on the identifier data of the data to be queried carried in the state query instruction; when the state processing instruction includes a state query instruction, determining the target state storage system from the multiple state storage systems based on the identifier data of the data to be written carried in the state write instruction; and when the state processing instruction includes a persistence instruction, determining the multiple state storage systems as the target state storage system.
[0104] In one optional embodiment, for a state table, in order to improve the calculation speed, multiple concurrent read and write operations are adopted, and the concurrency level may be changed at any time. Therefore, different state types can correspond to different state tables. When data needs to be written, the target state storage system can be determined from multiple state storage systems corresponding to the state table based on the identifier data of the data to be written. When data needs to be queried, the target state storage system can be determined from multiple state storage systems corresponding to the state table based on the identifier data of the data to be queried. When data persistence is required, multiple state storage systems corresponding to the state table can be jointly determined as the target state storage system.
[0105] In the above embodiments of this application, multiple state storage systems are restarted when the stream computing system restarts.
[0106] In one alternative embodiment, after the stream computing system starts, multiple state storage systems can be restarted. That is, after the stream computing system crashes and restarts, it is only necessary to reopen the state storage system to restore the state to the previous checkpoint without loading data.
[0107] In the above embodiments of this application, determining a target state storage system from multiple state storage systems based on the identifier data of the data to be queried carried in the state query instruction includes: determining the target identifier data range to which the identifier data of the data to be queried belongs from the identifier data ranges corresponding to the multiple state storage systems, wherein different identifier data ranges correspond to different state storage systems; and determining the state storage system corresponding to the target identifier data range in the multiple state storage systems as the target state storage system.
[0108] The target identifier range mentioned above can be the identifier range to which the identifier data to be written belongs, that is, the range of concurrent keys. Optionally, different identifier data ranges correspond to different state storage systems. For multiple concurrency, a state storage system can be created for each concurrency of each state table. The range of keys for each concurrency can be divided according to the situation of each state storage system. It is only necessary to ensure that the range of keys between different concurrencies does not overlap. For example, the first state storage system stores data with keys 1-100, and the second state storage system stores data with keys 101-200.
[0109] In one optional embodiment, the target identifier data interval to which the identifier data of the data to be written belongs can be determined from the identifier data intervals corresponding to multiple state storage systems based on the identifier data of the data to be written. For example, the key of the data to be written is 10, and identifier data interval 1 is used to store data with keys 1-100. Therefore, the target identifier data interval corresponding to the data to be written can be determined as identifier data interval 1. Furthermore, the state storage system corresponding to the target identifier data interval can be determined as the target state storage system, that is, the state storage system corresponding to identifier data interval 1 is determined as the target state storage system.
[0110] In the above embodiments of this application, determining a target state storage system from multiple state storage systems based on the identifier data of the data to be written carried in the state write instruction includes: determining the target identifier data interval to which the identifier data of the data to be written belongs from the identifier data intervals corresponding to the multiple state storage systems based on the identifier data of the data to be written; and determining the state storage system corresponding to the target identifier data interval in the multiple state storage systems as the target state storage system.
[0111] In one optional embodiment, based on the identifier data of the data to be queried, the target identifier data interval to which the identifier data of the data to be queried belongs is determined from the identifier data intervals corresponding to multiple state storage systems. For example, the key of the data to be queried is 20, and identifier data interval 1 is used to store data with keys 1-100. Therefore, the target identifier data interval corresponding to the data to be queried can be determined to be identifier data interval 1. Furthermore, the state storage system corresponding to the target identifier data interval can be determined to be the target state storage system, that is, the state storage system corresponding to identifier data interval 1 is determined to be the target state storage system.
[0112] In the above embodiments of this application, the method further includes: obtaining an adjustment instruction to adjust the identifier data range corresponding to multiple state storage systems; shutting down multiple state storage systems; determining at least one new identifier data range based on the adjustment instruction; and creating a state storage system corresponding to at least one new identifier data range.
[0113] In an optional embodiment, after receiving an adjustment instruction to adjust the identifier data range corresponding to multiple state storage systems, it can be determined that the concurrency needs to be adjusted. Therefore, multiple state storage systems can be shut down, and at least one new identifier data range can be determined based on the adjustment instruction. At least one new identifier data range corresponding to the state storage system can be created. That is, a new state storage system is created according to the new concurrency, thereby reducing the impact on the overall stream computing task.
[0114] In the above embodiments of this application, the method further includes: obtaining the memory usage of multiple state storage systems; when the sum of the memory usages is greater than a preset amount, determining cache items of the multiple state storage systems, wherein the cache items are used to represent data persistently stored in the key-value storage system; and evicting cache items based on the access information of the cache items until the sum of the memory usages is less than the preset amount.
[0115] In one alternative embodiment, a stream computing system may have multiple state storage systems, thus requiring a centralized system (such as...). Figure 5 The central memory control module (shown) ensures that the total memory usage of the state storage systems does not exceed a certain limit. Optionally, each state storage system needs to record its own cache (i.e., the data already persistently stored in the key-value storage system mentioned above) usage and cache access frequency information, and report this information to the centralized system. The centralized system periodically determines how much memory each state storage system should evict based on this information, eliminating as few frequently accessed items as possible, and then notifies each state storage system to perform the corresponding eviction. That is, when the total memory usage exceeds the preset limit, the system identifies cache items from multiple state storage systems and evicts them based on access information such as the access frequency of the cache items, until the total memory usage is less than the preset limit.
[0116] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0117] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0118] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, they can also be implemented by hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0119] Example 2
[0120] According to an embodiment of this application, a state storage system 60 is also provided, such as... Figure 6 Shown, including:
[0121] Multiple interfaces 61 are connected to the stream computing system 62. Different interfaces are used to receive different instructions sent by the stream computing system 62. These different instructions include: state write instructions, state query instructions, and persistence instructions. A first linked list 63 is used to store data that has not been persisted to the key-value storage system. A hash table 64 contains at least one second linked list 6401 (only one is shown in the figure). Different second linked lists 6401 correspond to different identification data. The different second linked lists 6401 are configured to store different versions of data, version information of different versions of data, and persistence information of whether different versions of data are persisted to the key-value storage system.
[0122] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0123] Example 3
[0124] According to an embodiment of this application, a storage system 70 for a stream computing system is also provided, such as... Figure 7As shown, it includes: a key-value storage system 71, which stores multiple state tables, with different state tables used to store data of different data types generated by the stream computing system; and a set of multiple state storage systems 72, which corresponds one-to-one with the multiple state tables, and the set of state storage systems 72 includes at least one state storage system 73 (only one is shown in the figure).
[0125] This application does not impose specific restrictions on the data in the aforementioned multiple status tables, for example, Figure 7 State table 1, state table 2, up to state table n, optionally, each state table corresponds one-to-one with state storage system set 72, optionally, each state storage system set 72 includes at least one state storage system 73. Figure 7 In this example, we will only use a state storage system set including a state storage system 73 as an example.
[0126] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0127] Example 4
[0128] According to an embodiment of this application, an apparatus for implementing the state storage method of the above-described stream computing system is also provided. Figure 8 This is a schematic diagram of a state storage device for a stream computing system according to Embodiment 4 of this application, as shown below. Figure 8 As shown, the device includes: a first determining module 802 and a data management module 804.
[0129] The first determining module 802 is used to respond to the state processing instructions sent by the stream computing system and determine the target state storage system corresponding to the state processing instructions from multiple state storage systems pre-created for the state tables stored in the key-value storage system. The state processing instructions include: state query instructions, state write instructions, and persistence instructions. The data management module 804 is used to perform data management on the hash table and / or the first linked list in the target state storage system based on the state processing instructions. The hash table stores at least one second linked list corresponding to the identifier data. The second linked list is configured to store different versions of data. The first linked list is used to store data that has not been persisted to the key-value storage system.
[0130] In the above embodiments of this application, the data management module 804 includes: a first determining unit, used to determine the identifier data of the data to be queried carried in the status query instruction; a second determining unit, used to determine whether there is target identifier data in the target state storage system that is the same as the identifier data of the data to be queried; a third determining unit, used to determine the second linked list corresponding to the target identifier data from the target state storage system when the target identifier data exists in the hash table of the target state storage system, wherein the second linked list also stores version information of different versions of data; and a first obtaining unit, used to obtain the data corresponding to the largest version information in the second linked list corresponding to the target identifier data as the data to be queried, and send the data to be queried to the stream computing system.
[0131] In the above embodiments of this application, the device further includes: a first sending module, configured to send a query instruction to a key-value storage system, wherein the query instruction is used to query whether the key-value storage system stores the data to be queried; a second sending module, configured to send the data to be queried to a stream computing system when the key-value storage system sends the data to be queried; and a third sending module, configured to send a prompt message to the stream computing system when the key-value storage system does not send the data to be queried, wherein the prompt message is used to indicate that the key-value storage system does not store the data to be queried.
[0132] In the above embodiments of this application, the device further includes: a first storage module, used to store target identification data and data to be queried data into a hash table of the target state storage system, and to mark the persistent information of the data to be queried as already persisted.
[0133] In the above embodiments of this application, the data management module 804 further includes: a fourth determining unit, used to determine the data to be written carried in the state write instruction, and the identification data of the data to be written; a fifth determining unit, used to determine the second linked list corresponding to the identification data of the data to be written from the hash table of the target state storage system, wherein the second linked list also stores version information of different versions of data; a sixth determining unit, used to obtain the data corresponding to the maximum version information from the second linked list corresponding to the identification data of the data to be written, and determine the target state of the data corresponding to the maximum version information, wherein the target state includes: not triggered persistence, not completed persistence, and already persisted; and a writing unit, used to write the data to be written into the target state storage system based on the target state.
[0134] In the above embodiments of this application, the sixth determining unit includes: an acquisition subunit, configured to acquire target persistence information and target version information of the data corresponding to the maximum version information from the second linked list corresponding to the identifier data of the data to be written; a first determining subunit, configured to determine that the target state is not triggered persistence when the target version information is the same as the preset version information of the target state storage system; a second determining subunit, configured to determine that the target state is not persisted when the target version information is different from the preset version information and the target persistence information is that the data corresponding to the maximum version information is not persisted; and a third determining subunit, configured to determine that the target state is persisted when the target version information is different from the preset version information and the target persistence information is that the data corresponding to the maximum version information is persisted.
[0135] In the above embodiments of this application, the writing unit includes: a first storage subunit, used to overwrite the data corresponding to the maximum version information with the data to be written when the target state is either not triggered or already persisted; and a second storage subunit, used to store the data to be written into a second linked list corresponding to the identifier data of the data to be written when the target state is not yet persisted, and to store the preset version information as the version information of the data to be written into the second linked list corresponding to the identifier data of the data to be written.
[0136] In the above embodiments of this application, the device further includes: a second storage module, used to store the data to be written to the first linked list of the target state storage system when the target state is not yet persisted or has been persisted.
[0137] In the above embodiments of this application, the data management module 804 further includes: an update unit for updating the preset version information of the target state storage system; a second acquisition unit for acquiring at least one piece of data stored in the first linked list of the target state storage system; and a clearing unit for storing at least one piece of data in the key-value storage system and clearing the first linked list.
[0138] In the above embodiments of this application, the device further includes: a modification module, configured to modify the persistence information of at least one piece of data stored in the target state storage system to indicate that at least one piece of data has been persistently stored in the key-value storage system.
[0139] In the above embodiments of this application, the device further includes: a second determining module, configured to determine a second linked list corresponding to the identifier data of at least one piece of data from the hash table of the target state storage system, wherein the second linked list also stores version information of different versions of data; a third storage module, configured to use at least one piece of data as a cache item of the target state storage system when at least one piece of data is the data corresponding to the maximum version information in the second linked list corresponding to the identifier data of at least one piece of data; and a release module, configured to release at least one piece of data when at least one piece of data is not the data corresponding to the maximum version information in the second linked list corresponding to the identifier data of at least one piece of data.
[0140] In the above embodiments of this application, the first determining module 802 includes: a seventh determining unit, configured to determine a target state storage system from multiple state storage systems based on the identifier data of the data to be queried carried in the state query instruction when the state processing instruction includes a state query instruction; an eighth determining unit, configured to determine a target state storage system from multiple state storage systems based on the identifier data of the data to be written carried in the state write instruction when the state processing instruction includes a state query instruction; and a ninth determining unit, configured to determine multiple state storage systems as the target state storage system when the state processing instruction includes a persistence instruction.
[0141] In the above embodiments of this application, the device further includes: a restart module, used to restart multiple state storage systems in the event of a restart of the stream computing system.
[0142] In the above embodiments of this application, the seventh determining unit includes: a fourth determining subunit, used to determine the target identifier data interval to which the identifier data of the data to be queried belongs from the identifier data intervals corresponding to multiple state storage systems based on the identifier data of the data to be queried, wherein different identifier data intervals correspond to different state storage systems; and a sixth determining subunit, used to determine the state storage system corresponding to the target identifier data interval in the multiple state storage systems as the target state storage system.
[0143] In the above embodiments of this application, the eighth determining unit includes: a seventh determining subunit, used to determine the target identifier data interval to which the identifier data of the data to be written belongs from the identifier data intervals corresponding to multiple state storage systems based on the identifier data of the data to be written; and a ninth determining subunit, used to determine the state storage system corresponding to the target identifier data interval in the multiple state storage systems as the target state storage system.
[0144] In the above embodiments of this application, the device further includes: a second acquisition module, configured to acquire an adjustment instruction for adjusting the identifier data ranges corresponding to multiple state storage systems; a shutdown module, configured to shut down multiple state storage systems; a third determination module, configured to determine at least one new identifier data range based on the adjustment instruction; and create a state storage system corresponding to at least one new identifier data range.
[0145] In the above embodiments of this application, the device further includes: a third acquisition module, used to acquire the memory usage of multiple state storage systems; a fourth determination module, used to determine cache items of multiple state storage systems when the sum of memory usage is greater than a preset usage, wherein the cache item is used to represent data that has been persistently stored in the key-value storage system; and an eviction module, used to evict cache items based on the access information of the cache items until the sum of memory usage is less than the preset usage.
[0146] It should be noted that the first determining module 802 and the data management module 804 mentioned above correspond to steps S402 to S404 in Embodiment 1. The two modules and the corresponding steps implement the same instances and application scenarios, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above modules can also be part of the device and run in the computer terminal 10 provided in Embodiment 1.
[0147] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0148] Example 5
[0149] Embodiments of this application may provide an electronic device, which may be any one of a group of electronic devices. Optionally, in this embodiment, the aforementioned electronic device may also be replaced by a terminal device such as a mobile terminal.
[0150] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0151] In this embodiment, the computer terminal described above can execute the program code in the method.
[0152] Optionally, Figure 9This is a structural block diagram of an electronic device according to an embodiment of this application. As shown in the figure, the electronic device A may include: one or more (only one is shown in the figure) processors 902, memory 904, memory controller, and peripheral interfaces, wherein the peripheral interfaces are connected to a radio frequency module, an audio module, and a display.
[0153] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and apparatus in the embodiments of this application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to terminal A via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0154] The processor can invoke information and applications stored in memory via a transmission device to perform the following steps: In response to a state processing instruction sent by the stream computing system, determine the target state storage system corresponding to the state processing instruction from multiple state storage systems pre-created for the state tables stored in the key-value storage system, wherein the state processing instruction includes: a state query instruction, a state write instruction, and a persistence instruction; Based on the state processing instruction, perform data management on the hash table and / or the first linked list in the target state storage system, wherein the hash table stores at least one second linked list corresponding to the identifier data, the second linked list is configured to store different versions of data, and the first linked list is used to store data that has not been persistently stored in the key-value storage system.
[0155] Optionally, the processor may also execute program code that performs the following steps: determining the identifier data of the data to be queried carried in the state query instruction; determining whether there is target identifier data in the target state storage system that is the same as the identifier data of the data to be queried; if the target identifier data exists in the hash table of the target state storage system, determining the second linked list corresponding to the target identifier data from the target state storage system, wherein the second linked list also stores version information of different versions of data; obtaining the data corresponding to the largest version information in the second linked list corresponding to the target identifier data as the data to be queried, and sending the data to be queried to the stream computing system.
[0156] Optionally, the processor may also execute program code that performs the following steps: sending a query instruction to the key-value storage system, wherein the query instruction is used to query whether the key-value storage system stores the data to be queried; if the key-value storage system sends the data to be queried, sending the data to be queried to the stream computing system; if the key-value storage system does not send the data to be queried, sending a prompt message to the stream computing system, wherein the prompt message is used to indicate that the key-value storage system does not store the data to be queried.
[0157] Optionally, the processor may also execute program code that performs the following steps: storing the target identifier data and the data to be queried in a hash table of the target state storage system, and marking the persistent information of the data to be queried as already persisted.
[0158] Optionally, the processor may also execute program code that performs the following steps: determining the data to be written carried in the state write instruction, and the identifier data of the data to be written; determining the second linked list corresponding to the identifier data of the data to be written from the hash table of the target state storage system, wherein the second linked list also stores version information of different versions of the data; obtaining the data corresponding to the maximum version information from the second linked list corresponding to the identifier data of the data to be written, and determining the target state of the data corresponding to the maximum version information, wherein the target state includes the following: persistence not triggered, persistence not completed, and persistence completed; and writing the data to be written to the target state storage system based on the target state.
[0159] Optionally, the processor may also execute program code with the following steps: obtaining the target persistence information and target version information of the data corresponding to the maximum version information from the second linked list corresponding to the identifier data of the data to be written; determining the target state as not triggered persistence if the target version information is the same as the preset version information of the target state storage system; determining the target state as not completed persistence if the target version information is different from the preset version information and the target persistence information is that the data corresponding to the maximum version information is not persisted; and determining the target state as persisted if the target version information is different from the preset version information and the target persistence information is that the data corresponding to the maximum version information is persisted.
[0160] Optionally, the processor may also execute program code with the following steps: if the target state is that persistence has not been triggered or persistence has been completed, overwrite the data corresponding to the maximum version information with the data to be written; if the target state is that persistence has not been completed, store the data to be written in the second linked list corresponding to the identifier data of the data to be written, and store the preset version information as the version information of the data to be written in the second linked list corresponding to the identifier data of the data to be written.
[0161] Optionally, the processor may also execute program code that performs the following steps: when the target state is incomplete or has been persisted, the data to be written is stored in the first linked list of the target state storage system.
[0162] Optionally, the processor may also execute program code that performs the following steps: updates the preset version information of the target state storage system; obtains at least one piece of data stored in the first linked list of the target state storage system; stores at least one piece of data in the key-value storage system and clears the first linked list.
[0163] Optionally, the processor may also execute program code that modifies the persistence information of at least one piece of data stored in the target state storage system to indicate that at least one piece of data has been persistently stored in the key-value storage system.
[0164] Optionally, the processor may also execute program code that performs the following steps: determining a second linked list corresponding to the identifier data of at least one piece of data from the hash table of the target state storage system, wherein the second linked list also stores version information of different versions of data; if at least one piece of data is the data corresponding to the maximum version information in the second linked list corresponding to the identifier data of at least one piece of data, then using at least one piece of data as a cache entry of the target state storage system; if at least one piece of data is not the data corresponding to the maximum version information in the second linked list corresponding to the identifier data of at least one piece of data, then releasing at least one piece of data.
[0165] Optionally, the processor may also execute program code that performs the following steps: when the state processing instruction includes a state query instruction, a target state storage system is determined from multiple state storage systems based on the identifier data of the data to be queried carried in the state query instruction; when the state processing instruction includes a state query instruction, a target state storage system is determined from multiple state storage systems based on the identifier data of the data to be written carried in the state write instruction; when the state processing instruction includes a persistence instruction, multiple state storage systems are determined as the target state storage system.
[0166] Optionally, the processor may also execute program code that restarts multiple state storage systems in the event of a restart of the stream computing system.
[0167] Optionally, the processor may also execute program code that performs the following steps: determining the target identifier data range to which the identifier data of the data to be queried belongs from the identifier data ranges corresponding to multiple state storage systems, wherein different identifier data ranges correspond to different state storage systems; and determining the state storage system corresponding to the target identifier data range in the multiple state storage systems as the target state storage system.
[0168] Optionally, the processor may also execute program code that performs the following steps: determining the target identifier data range to which the identifier data to be written belongs from the identifier data ranges corresponding to multiple state storage systems; and determining the state storage system corresponding to the target identifier data range in the multiple state storage systems as the target state storage system.
[0169] Optionally, the processor may also execute program code that performs the following steps: obtaining adjustment instructions for adjusting the identifier data ranges corresponding to multiple state storage systems; shutting down multiple state storage systems; determining at least one new identifier data range based on the adjustment instructions; and creating a state storage system corresponding to at least one new identifier data range.
[0170] Optionally, the processor may also execute program code that performs the following steps: obtains the memory usage of multiple state storage systems; if the sum of the memory usages is greater than a preset amount, determines the cache entries of the multiple state storage systems, wherein the cache entries are used to represent data that has been persistently stored in the key-value storage system; and evicts the cache entries based on their access information until the sum of the memory usages is less than the preset amount.
[0171] In this embodiment, in response to a state processing instruction sent by the stream computing system, a target state storage system corresponding to the state processing instruction is determined from multiple state storage systems pre-created for the state tables stored in the key-value storage system. The state processing instruction includes a state query instruction, a state write instruction, and a persistence instruction. Based on the state processing instruction, data management is performed on the hash table and / or the first linked list in the target state storage system. The hash table stores at least one second linked list corresponding to the identifier data. The second linked list is configured to store different versions of data, and the first linked list is used to store data that has not been persistently stored in the key-value storage system. It is noteworthy that multiple state storage systems can be pre-created for the state tables stored in the key-value storage system. This allows for data management using a multi-concurrency approach, improving computation speed. Furthermore, different versions of data are stored using hash tables in the target state storage system. Upon receiving a state processing instruction, the target state storage system corresponding to the instruction can be determined promptly. Based on the instruction, data management is performed on the hash table and / or the first linked list in the target state storage system. Storing at least one identifier using a hash table not only saves memory but also acts as a cache, reducing disk reads and lookups of the target state. This improves the performance of the target state storage system and solves the technical problem of low performance in state storage in related technologies.
[0172] It will be understood by those skilled in the art that the structure shown in the figure is merely illustrative, and the electronic device may also be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. This figure does not limit the structure of the aforementioned electronic device. For example, electronic device A may include more or fewer components (such as a network interface, a display device, etc.) than shown in the figure, or may have a different configuration than shown in the figure.
[0173] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0174] Example 6
[0175] Embodiments of this application also provide a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store program code executed by the method provided in the above embodiments.
[0176] Optionally, in this embodiment, the storage medium may be located in any one of the electronic devices in the group of electronic devices in the computer network, or in any one of the mobile terminals in the group of mobile terminals.
[0177] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: in response to a state processing instruction sent by the stream computing system, determining a target state storage system corresponding to the state processing instruction from multiple state storage systems pre-created for the state table stored in the key-value storage system, wherein the state processing instruction includes: a state query instruction, a state write instruction, and a persistence instruction; and based on the state processing instruction, performing data management on a hash table and / or a first linked list in the target state storage system, wherein the hash table stores at least one second linked list corresponding to the identifier data, the second linked list is configured to store different versions of data, and the first linked list is used to store data that has not been persistently stored in the key-value storage system.
[0178] Example 7
[0179] The embodiments of this application also provide a computer program product. Optionally, in this embodiment, the computer program product may include a computer program that, when executed by a processor, implements the method provided in the above embodiments.
[0180] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0181] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0182] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0183] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0184] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0185] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0186] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A state storage method for a stream computing system, characterized in that, include: In response to a state processing instruction sent by a stream computing system, a target state storage system corresponding to the state processing instruction is determined from multiple state storage systems pre-created for state tables stored in a key-value storage system. The state processing instruction includes: a state query instruction, a state write instruction, and a persistence instruction. Based on the state processing instructions, data management is performed on the hash table and / or the first linked list in the target state storage system. The hash table stores at least one second linked list corresponding to the identifier data. The second linked list is configured to store different versions of data. The first linked list is used to store data that has not been persistently stored in the key-value storage system.
2. The method according to claim 1, characterized in that, When the state processing instruction includes the state query instruction, the step of managing data on the hash table and / or the first linked list in the target state storage system based on the state processing instruction includes: Determine the identifier data of the data to be queried carried in the status query instruction; Determine whether there is target identifier data in the target state storage system that is identical to the identifier data of the data to be queried; If the target identifier data exists in the hash table of the target state storage system, a second linked list corresponding to the target identifier data is determined from the target state storage system, wherein the second linked list also stores version information of the different versions of the data; The data corresponding to the maximum version information in the second linked list corresponding to the target identifier data is obtained as the data to be queried, and the data to be queried is sent to the stream computing system.
3. The method according to claim 2, characterized in that, If the target identifier data does not exist in the hash table of the target state storage system, the method further includes: Send a query command to the key-value storage system, wherein the query command is used to query whether the key-value storage system stores the data to be queried; Upon receiving the data to be queried from the key-value storage system, the data to be queried is sent to the stream computing system; If the data to be queried is not received from the key-value storage system, a prompt message is sent to the stream computing system, wherein the prompt message indicates that the data to be queried is not stored in the key-value storage system.
4. The method according to claim 3, characterized in that, Upon receiving the data to be queried from the key-value storage system, the method further includes: The target identifier data and the data to be queried are stored in the hash table of the target state storage system, and the persistence information of the data to be queried is marked as persistent.
5. The method according to claim 1, characterized in that, When the state processing instruction includes the state write instruction, the step of managing data on the hash table and / or the first linked list in the target state storage system based on the state processing instruction includes: Determine the data to be written carried in the status write instruction, and the identification data of the data to be written; From the hash table of the target state storage system, determine the second linked list corresponding to the identifier data of the data to be written, wherein the second linked list also stores version information of the different versions of the data; Retrieve the data corresponding to the maximum version information from the second linked list corresponding to the identifier data of the data to be written, and determine the target state of the data corresponding to the maximum version information, wherein the target state includes: Persistence not triggered, persistence not completed, and persistence already completed; Based on the target state, the data to be written is written to the key-value storage system.
6. The method according to claim 5, characterized in that, The second linked list also stores persistence information on whether the different versions of the data are persistently stored in the key-value storage system; Determining the target state of the data corresponding to the maximum version information includes: Obtain the target persistence information and target version information of the data corresponding to the maximum version information from the second linked list corresponding to the identifier data of the data to be written; If the target version information is the same as the preset version information of the target state storage system, the target state is determined to be the non-triggered persistence. If the target version information is different from the preset version information, and the target persistent information is that the data corresponding to the maximum version information has not been persisted, the target state is determined to be "incomplete persistence". If the target version information is different from the preset version information, and the target persistent information is the data persistent corresponding to the maximum version information, then the target state is determined to be "already persisted".
7. The method according to claim 6, characterized in that, The step of writing the data to be written to the key-value storage system based on the target state includes: If the target state is either not persisted or already persisted, the data to be written will overwrite the data corresponding to the maximum version information. If the target state is not yet persistent, the data to be written is stored in the second linked list corresponding to the identifier data of the data to be written, and the preset version information is used as the version information of the data to be written and stored in the second linked list corresponding to the identifier data of the data to be written.
8. The method according to claim 6, characterized in that, The method further includes: If the target state is either not persisted or has been persisted, the data to be written is stored in the first linked list of the target state storage system.
9. The method according to claim 1, characterized in that, When the state processing instruction includes the persistence instruction, the step of managing data on the hash table and / or the first linked list in the target state storage system based on the state processing instruction includes: Update the preset version information of the target state storage system; Obtain at least one piece of data stored in the first linked list of the target state storage system; The at least one data is stored in the key-value storage system, and the first linked list is cleared.
10. The method according to claim 9, characterized in that, After storing the at least one data in the key-value storage system, the method further includes: The persistence information of at least one piece of data stored in the target state storage system is modified to indicate that at least one piece of data has been persistently stored in the key-value storage system.
11. The method according to claim 9, characterized in that, After storing the at least one data in the key-value storage system, the method further includes: The second linked list corresponding to the identifier data of the at least one piece of data is determined from the hash table of the target state storage system, wherein the second linked list also stores version information of the different versions of the data; If the at least one piece of data is the data corresponding to the maximum version information in the second linked list corresponding to the identifier data of the at least one piece of data, the at least one piece of data shall be used as a cache item of the target state storage system; If the at least one piece of data is not the data corresponding to the maximum version information in the second linked list corresponding to the identifier data of the at least one piece of data, the at least one piece of data will be released.
12. The method according to claim 1, characterized in that, The step of determining the target state storage system corresponding to the state processing instruction from multiple state storage systems pre-created for the state tables stored in the key-value storage system includes: When the status processing instruction includes the status query instruction, the target status storage system is determined from the plurality of status storage systems based on the identifier data of the data to be queried carried in the status query instruction. When the state processing instruction includes the state query instruction, the target state storage system is determined from the plurality of state storage systems based on the identifier data of the data to be written carried in the state write instruction. If the state processing instruction includes the persistence instruction, the plurality of state storage systems are determined as the target state storage system.
13. The method according to claim 12, characterized in that, If the stream computing system restarts, the plurality of state storage systems will be restarted.
14. The method according to claim 12, characterized in that, The step of determining the target state storage system from the plurality of state storage systems based on the identifier data of the data to be queried carried in the state query instruction includes: The target identifier data range to which the identifier data of the data to be queried belongs is determined from the identifier data ranges corresponding to the multiple state storage systems, wherein different identifier data ranges correspond to different state storage systems; The state storage system corresponding to the target identifier data range in the plurality of state storage systems is determined as the target state storage system.
15. The method according to claim 12, characterized in that, The step of determining the target state storage system from the plurality of state storage systems based on the identifier data of the data to be written carried in the state write instruction includes: Determine the target identifier data range to which the identifier data of the data to be written belongs from the identifier data ranges corresponding to the multiple state storage systems; The state storage system corresponding to the target identifier data range in the plurality of state storage systems is determined as the target state storage system.
16. The method according to claim 14 or 15, characterized in that, The method further includes: Obtain adjustment instructions to adjust the identifier data ranges corresponding to the multiple state storage systems; Shut down the multiple state storage systems; Based on the adjustment instruction, at least one new identification data range is determined; Create a state storage system corresponding to the at least one new identifier data range.
17. The method according to claim 12, characterized in that, The method further includes: Get the memory usage of multiple state storage systems; If the sum of the memory usages exceeds a preset amount, cache entries of the plurality of state storage systems are determined, wherein the cache entries are used to represent data that has been persistently stored in the key-value storage system; Based on the access information of the cached items, the cached items are evicted until the sum of the memory usage is less than the preset usage.
18. A state storage system, characterized in that, include: Multiple interfaces are connected to the stream computing system. Different interfaces are used to receive different instructions sent by the stream computing system, including: state write instructions, state query instructions, and persistence instructions. The first linked list is used to store data that has not been persistently stored in the key-value storage system; A hash table contains at least one second linked list, with different second linked lists corresponding to different identifier data. The different second linked lists are configured to store different versions of data, version information of the different versions of data, and persistence information of whether the different versions of data are persistently stored in the key-value storage system.
19. A storage system for a stream computing system, characterized in that, include: The key-value storage system stores multiple state tables, with different state tables used to store data of different data types generated by the stream computing system; A set of multiple state storage systems, each corresponding to one of the multiple state tables, wherein the set of state storage systems includes at least one state storage system as described in claim 18.
20. An electronic device, characterized in that, include: Memory, which stores executable programs; A processor for running the program, wherein the program, when running, performs the method according to any one of claims 1 to 17.
21. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored executable program, wherein, when the executable program is executed, it controls the device on which the storage medium is located to perform the method according to any one of claims 1 to 17.
22. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 17.