A cache data processing method, device, electronic device and storage medium

By writing cached data to the off-heap cache area in Flink's distributed system, the lag and performance degradation caused by frequent garbage collection in Flink applications is solved, and the stability and performance of the system are improved.

CN113742095BActive Publication Date: 2025-05-16BEIJING WODONG TIANJUN INFORMATION TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110048902.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-14
Publication Date
2025-05-16
Estimated Expiration
2041-01-14

AI Technical Summary

Technical Problem

The cached data generated by the big data computing engine Flink during real-time calculations leads to frequent garbage collection of local memory, resulting in application lag and performance degradation.

Method used

In distributed systems, cached data is written to the off-heap cache area by calling microservices, which is not affected by the JVM memory management mechanism, avoiding frequent garbage collection.

Benefits of technology

It effectively avoids application lag and performance degradation caused by frequent garbage collection, and improves Flink's robustness and stability in real-time calculation of massive data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113742095B_ABST
    Figure CN113742095B_ABST
Patent Text Reader

Abstract

The embodiments of the present invention are applicable to the field of computer technology, and provide a cache data processing method, device, electronic device and storage medium, wherein the cache data processing method includes: when a first working node in a distributed system generates cache data during the execution of a task, calling a microservice corresponding to the first working node; based on the microservice corresponding to the first working node, writing the cache data into a corresponding cache area; the cache area is an off-heap cache area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a cache data processing method, device, electronic device and storage medium. Background Art

[0002] The big data computing engine Flink generates cache data when performing real-time computing on streaming data. Related technologies store cache data in local memory, which is an in-heap cache. Local memory will perform garbage collection (GC) as cache data is written. In some cases, such as e-commerce promotions, Flink will receive massive amounts of streaming data, which will generate a large amount of cache data. Due to the continuous writing of cache data, local memory will frequently perform GC, causing Flink applications to freeze and performance to degrade. Summary of the invention

[0003] In order to solve the above problems, an embodiment of the present invention provides a cache data processing method, device, electronic device and storage medium, so as to at least solve the problem of frequent garbage collection of local memory in related technologies causing application freezes and performance degradation.

[0004] The technical solution of the present invention is achieved in this way:

[0005] In a first aspect, an embodiment of the present invention provides a cache data processing method, the method comprising:

[0006] When a first working node in a distributed system generates cache data during task execution, a microservice corresponding to the first working node is called;

[0007] Based on the microservice corresponding to the first working node, the cache data is written into a corresponding cache area; the cache area is an off-heap cache area.

[0008] In the above solution, before the microservice corresponding to the first working node writes the serialized data into the corresponding cache area, the method further includes:

[0009] Determine the memory usage of the cache area corresponding to the first working node;

[0010] When the memory usage is greater than the set value, the first cache data in the corresponding cache area is deleted; the first cache data represents the cache data with the largest timestamp in the corresponding cache area; the timestamp represents the time interval between the time when the corresponding cache data was last accessed and the current time.

[0011] In the above solution, the microservice corresponding to the first working node writes the serialized data into the corresponding cache area, including:

[0012] Serializing the cached data based on the microservice corresponding to the first working node to obtain corresponding serialized data;

[0013] Based on the microservice corresponding to the first working node, the serialized data is written into the corresponding cache area.

[0014] In the above scheme, the method further comprises:

[0015] Create a microservice corresponding to the first working node;

[0016] Create a corresponding cache area based on the microservice corresponding to the first working node.

[0017] In the above solution, the step of creating a corresponding cache area based on the microservice corresponding to the first working node includes:

[0018] Create a serializer and a deserializer corresponding to the cache area; the serializer is used to serialize the cache data; the deserializer is used to convert the serialized data in the cache area into the corresponding cache data;

[0019] Create a buffer based on the set memory size, set memory structure, corresponding serializer and deserializer.

[0020] In the above scheme, the method further comprises:

[0021] Obtain at least one first cache data;

[0022] The set memory size is determined based on the memory size of the cache area occupied by each first cache data of the at least one first cache data and the cache data amount of the set cache area.

[0023] In the above scheme, the method further comprises:

[0024] Upon receiving a read request from the first working node to the corresponding cache area, reading serialized data corresponding to the read request from the corresponding cache area based on the setting microservice;

[0025] The serialized data is deserialized based on the setting microservice to obtain cache data corresponding to the read request.

[0026] In a second aspect, an embodiment of the present invention provides a cache data processing device, the device comprising:

[0027] A calling module, used for calling a microservice corresponding to a set first working node in a distributed system when the first working node generates cache data in the process of executing a task;

[0028] A writing module is used to write the cache data into a corresponding cache area based on the microservice corresponding to the first working node; the cache area is an off-heap cache area.

[0029] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a processor and a memory, wherein the processor and the memory are interconnected, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the steps of the cache data processing method provided in the first aspect of the embodiment of the present invention.

[0030] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, including: the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the cache data processing method provided in the first aspect of the embodiment of the present invention are implemented.

[0031] In the embodiment of the present invention, when the first working node in the distributed system generates cache data during the execution of a task, the microservice corresponding to the set first working node is called, and the cache data is written into the corresponding cache area based on the microservice corresponding to the first working node, wherein the cache area is an off-heap cache area. In the embodiment of the present invention, the cache data is stored in the off-heap cache area, and the off-heap cache area does not need to comply with the memory management mechanism of the JVM to perform GC, thereby solving the problem of frequent GC in related technologies causing application freezes and performance degradation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic diagram of an implementation flow of a cache data processing method provided by an embodiment of the present invention;

[0033] Figure 2 It is a schematic diagram of a cache data processing flow provided by an embodiment of the present invention;

[0034] Figure 3 It is a schematic diagram of a storage structure of cache data provided by an embodiment of the present invention;

[0035] Figure 4 is a schematic diagram of a state storage format provided by an embodiment of the present invention;

[0036] Figure 5 It is a schematic diagram of an implementation flow of another cache data processing method provided by an embodiment of the present invention;

[0037] Figure 6It is a schematic diagram of an implementation flow of another cache data processing method provided by an embodiment of the present invention;

[0038] Figure 7 It is a schematic diagram of an implementation flow of another cache data processing method provided by an embodiment of the present invention;

[0039] Figure 8 It is a schematic diagram of an implementation flow of another cache data processing method provided by an embodiment of the present invention;

[0040] Fig. 9 It is a schematic diagram of an implementation flow of another cache data processing method provided by an embodiment of the present invention;

[0041] Fig.10 It is a schematic diagram of an implementation flow of another cache data processing method provided by an embodiment of the present invention;

[0042] Fig.11 It is a schematic diagram of a cache data processing flow provided by an application embodiment of the present invention;

[0043] Fig.12 It is a schematic diagram of a cache data processing flow provided by an application embodiment of the present invention;

[0044] Fig.13 is a schematic diagram of a cache data processing device provided by an embodiment of the present invention;

[0045] Fig.14 It is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0047] Flink is a distributed streaming data flow engine written in Java and Scala. Flink executes any streaming data program in a data-parallel and pipeline manner. The data of streaming computing is often fleeting. Of course, in real business scenarios, it is impossible to say that all data will go away after coming in, and nothing will be left. Then the things left are called state, which can be translated into state data in Chinese, that is, cache data in the embodiment of the present invention. Related technologies will save state data as objects in Java heap memory (taskManager). The heap memory complies with the memory management mechanism of the Java Virtual Machine (JVM). When the remaining storage space of the heap memory is insufficient, the garbage collector will scan the heap memory space, identify objects that are no longer used by the application, and release their space. This process is called GC. Since the garbage collector needs to scan the heap and needs to suspend the application thread (Stop-The-World, STW) during the scan, too much cached data will cause the GC overhead to increase, thereby affecting the performance of the application and causing the application to freeze.

[0048] In view of the shortcomings of the above-mentioned related technologies, an embodiment of the present invention provides a cache data processing method, which can at least solve the problem of application freeze caused by frequent GC. In order to illustrate the technical solution of the present invention, a specific embodiment is used for description below.

[0049] Figure 1 The embodiment of the present invention provides a schematic diagram of a method for processing cache data. The execution subject of the method for processing cache data is an electronic device, which includes a desktop computer, a notebook computer, a server, etc. Figure 1 , the cache data processing method includes:

[0050] S101, when a first working node in a distributed system generates cache data during task execution, a microservice corresponding to the first working node is called.

[0051] Here, the distributed system is explained by taking the Flink distributed cluster as an example. The Flink distributed cluster includes a control node (JobManager) and multiple working nodes (TaskManager). The TaskManager receives the tasks to be deployed from the JobManager, then uses the Slot resources to start the Task, establishes a network connection for data access, receives data and starts data processing. At the same time, data interaction between TaskManagers is carried out in the form of data streams. Each TaskManager in Flink is a JVM process. A TaskManager executes multiple tasks concurrently in multiple threads. Task is the smallest unit of Flink task scheduling. Among them, the first working node refers to any TaskManager in the Flink distributed cluster.

[0052] Each Task corresponds to a task. When TaskManager executes a task, each Task will generate cache data. Related technologies will store the cache data in the heap memory corresponding to the Task. Figure 2 , Figure 2 It is a schematic diagram of a cache data processing flow provided by an embodiment of the present invention. All original data enters the user code and then is output to the downstream. If the reading and writing of state is involved in the middle, these state data will be stored in the local statebackend, which is a storage backend used to save the state.

[0053] In Flink, State is mainly divided into Operator State and Keyed State. The main differences between the two are as follows: 1. Operator State has no concept of current key, while the value of Keyed State always corresponds to a current key. 2. The Operator State backend has only one on-heap (in-heap cache) implementation, while the Keyed State backend has two implementation methods: on-heap and off-heap (RocksDB). 3. Operator State requires manual implementation of snapshot and recovery methods, while Keyed State is implemented by the backend itself and is transparent to the user. 4. The data scale of Operator State is usually relatively small, while the scale of Keyed State is relatively large.

[0054] Figure 3This is a schematic diagram of a cache data storage structure provided by an embodiment of the present invention. Flink comes with the following out-of-the-box state backends: MemoryStateBackend, FsStateBackend, and RocksDBStateBackend. Figure 3 As shown in the figure, Operator State can be stored in any of the above storage methods, and the Operator State backend is on-heap. When Keyed State is stored in MemoryStateBackend and FsStateBackend, it is on-heap, and when it is stored in RocksDBStateBackend, it is off-heap. Among them, GC will occur in all except rocksDB.

[0055] RocksDB is an open-source LSM key-value storage database developed by Facebook and is widely used in stand-alone components of big data systems. Flink's Keyed State is essentially a key-value pair, so it is consistent with the data model of RocksDB. Figure 4 , Figure 4 is a schematic diagram of a state storage format provided by an embodiment of the present invention. Figure 4 These are the storage formats of “window state” and “value state” in RocksDB. All stored keys and values ​​are serialized into bytes for storage.

[0056] In an embodiment of the present invention, when a first working node in a distributed system generates cache data during the execution of a task, a microservice corresponding to the first working node is called.

[0057] Among them, each first working node can correspond to a microservice individually, or several first working nodes can correspond to the same microservice together, and can be actually configured according to the distributed environment.

[0058] Here, the microservice needs to be created in advance, and the corresponding microservice is automatically called when the working node generates cache data. In the embodiment of the present invention, the function of the microservice is to create a cache area and read and write cache data to the cache area.

[0059] refer to Figure 5 , the cache data processing method also includes:

[0060] S501: Create a microservice corresponding to the first working node.

[0061] Here, microservices can be created based on microservice frameworks such as Springboot and Spring cloud, and microservices can be created based on the resources of distributed systems.

[0062] S502: Create a corresponding cache area based on the microservice corresponding to the first working node.

[0063] Here, the corresponding cache area created based on the microservice corresponding to the first working node is an in-heap cache area, that is, the cache area is an OHC cache. OHC stands for Off-Heap-Cache, which is a Java-based key-value off-heap cache framework. Off-heap memory does not need to comply with the memory management mechanism of the JVM, and it is managed by the operating system. Among them, when the Java program is running, the memory area managed by the JVM is called the heap.

[0064] refer to Figure 6 In one embodiment, the creating a corresponding cache area based on the microservice corresponding to the first working node includes:

[0065] S601, creating a serializer and a deserializer corresponding to the cache area; the serializer is used to serialize the cache data; the deserializer is used to convert the serialized data in the cache area into the corresponding cache data.

[0066] Since the off-heap cache is used in the form of byte arrays, the cache data needs to be serialized and deserialized when writing and reading the cache data.

[0067] Here, the serializer is used to serialize the cached data. The serializer implements the serialization method serialize() to serialize the cached data into a byte array for easy storage in the OHC storage.

[0068] The deserializer is used to convert the serialized data in the cache area into the corresponding cache data. The deserializer is used to implement the deserialize method deserialize(), which is the reverse process of the serialization method and is used to deserialize the byte array taken out from the OHC and restore it to the cache data.

[0069] S602, creating a cache area based on the set memory size, the set memory structure, the corresponding serializer and the deserializer.

[0070] In addition to setting the serializer and deserializer corresponding to the cache area, you also need to set the cache area's memory size, memory structure, and data replacement strategy.

[0071] Here, in actual applications, you need to select the appropriate memory structure according to business needs, such as String type key, value, etc.

[0072] Since the cache does not use JVM for garbage collection, when the cache memory reaches a certain size, data replacement is required. For example, the data replacement strategy can be the Least Recently Used (LRU) algorithm, which selects the data that has not been used the longest to be eliminated. The algorithm assigns an access field to each data to record the time t since the data was last accessed. When a data needs to be eliminated, the data with the largest t value, that is, the data that has been used the least recently, is selected for elimination.

[0073] Before creating a cache, you also need to pre-set the cache memory size. Figure 7 In one embodiment, the cache data processing method further includes:

[0074] S701, obtaining at least one piece of first cache data.

[0075] Here, at least one first cache data is online real data, that is, cache data generated by a Task, and at least one first cache data is randomly extracted from the online real data.

[0076] S702: Determine the set memory size based on the memory size of the cache area occupied by each first cache data of the at least one first cache data and the set cache data amount of the cache area.

[0077] Here, the set memory size can be obtained by multiplying the average value of the memory size of the cache area occupied by each first cache data in the at least one first cache data by the cache data amount of the set cache area. For example, if the average value of the memory size of the cache area occupied by the at least one first cache data is 1MB, and the cache data amount of the set cache area is 1000, then the set memory size is 1MB×1000=1000MB.

[0078] Then, based on the set memory size, the set memory structure, the data replacement strategy, the corresponding serializer and deserializer, a cache area is created through the microservice.

[0079] S102: Write the cache data into a corresponding cache area based on the microservice corresponding to the first working node; the cache area is an off-heap cache area.

[0080] When the first working node generates cache data, the microservice corresponding to the first working node is called, and the cache data is written into the corresponding cache area through the microservice. Here, the cache area corresponding to the first working node is created based on the corresponding microservice.

[0081] Here, since the cache area is an off-heap cache area (OHC), unlike the on-heap cache, the off-heap cache is not controlled by the JVM. The application itself is responsible for allocating and releasing memory, so it is not affected by GC. Therefore, there is no application lag problem caused by GC.

[0082] In the embodiment of the present invention, when the first working node in the distributed system generates cache data during the execution of a task, the microservice corresponding to the set first working node is called, and the cache data is written into the corresponding cache area based on the microservice corresponding to the first working node, wherein the cache area is an off-heap cache area. In the embodiment of the present invention, the cache data is stored in the off-heap cache area, and the off-heap cache area does not need to comply with the memory management mechanism of the JVM to perform GC, thereby solving the problem of frequent GC in related technologies causing application freezes and performance degradation.

[0083] refer to Figure 8 In one embodiment, the microservice corresponding to the first working node writes the serialized data into the corresponding cache area, including:

[0084] S801, serialize the cached data based on the microservice corresponding to the first working node to obtain corresponding serialized data.

[0085] Since the off-heap cache is used in the form of byte arrays, the cache data needs to be serialized and deserialized.

[0086] When the first working node generates cache data in the process of executing a task, the cache data is serialized based on the microservice corresponding to the first working node to obtain corresponding serialized data.

[0087] S802: Write the serialized data into a corresponding cache area based on the microservice corresponding to the first working node.

[0088] Based on the microservice, the serialized data is written into the cache area corresponding to the first working node.

[0089] refer to Fig. 9 In one embodiment, the method further comprises:

[0090] S901, when receiving a read request from the first working node to the corresponding cache area, read serialized data corresponding to the read request from the corresponding cache area based on the setting microservice.

[0091] S902: Deserialize the serialized data based on the setting microservice to obtain cache data corresponding to the read request.

[0092] Similarly, when reading cached data from the cache area, read the corresponding byte array according to the key value in the read request, and then call the deserialize() method of the deserializer to deserialize the serialized data and restore it to cached data.

[0093] Since the cache area does not use the memory management mechanism of the JVM for garbage collection, when the memory of the cache area is used to a certain size, memory replacement is required. The present invention can use the LRU algorithm for memory replacement.

[0094] refer to Fig.10 In one embodiment, before the microservice corresponding to the first working node writes the serialized data into the corresponding cache area, the method further includes:

[0095] S1001, determining the memory usage of the cache area corresponding to the first working node.

[0096] For example, if the maximum memory of the cache area corresponding to the first working node is 8G, the current memory usage is 7G.

[0097] S1002, when the memory usage is greater than the set value, delete the first cache data in the corresponding cache area; the first cache data represents the cache data with the largest timestamp in the corresponding cache area; the timestamp represents the time interval between the time when the corresponding cache data was last accessed and the current time.

[0098] When the memory usage is greater than the set value, for example, the set value is 6G and the current memory usage is 7G, ​​the first cached data in the corresponding cache area is deleted. Here, when writing cached data to the cache area, a timestamp is set for the cached data. The timestamp indicates the time interval between the last time the cached data was accessed and the current time. The larger the timestamp, the longer the corresponding cached data has not been used. Deleting the first cached data means deleting the cached data in the cache area that has not been used for the longest time.

[0099] In actual applications, when the microservice needs to write cache data into the cache area and the memory usage is greater than the set value, the first cache data in the cache area is deleted, and then it is determined whether the cache data can be written into the cache area. If it cannot be written, the first cache data in the cache area is deleted again until the cache data can be written into the cache area.

[0100] The embodiment of the present invention optimizes the GC of the Flink real-time computing program by placing cached data in the microservice, so that the cached data will not be garbage collected by the JVM, but memory management is performed through the self-configured strategy, thereby solving the problem of application freeze caused by GC. The present invention effectively enhances the robustness, stability and fault tolerance of the Flink distributed cluster in real-time computing of massive data, and provides automation and intelligent convenience for the operation and maintenance of massive real-time computing.

[0101] refer to Fig.11 , Fig.11 It is a schematic diagram of a cache data processing flow provided by an application embodiment of the present invention. First, the input data stream input stream enters each TaskManager node of the Flink distributed cluster. When the Task in each TaskManager node generates cache data during the execution of the task, the microservice (Micro service) corresponding to the TaskManager node is called, and the microservice serializes the cache data and writes it to the corresponding cache area (Cache). When the TaskManager node needs to read the cache data in the cache area, the corresponding microservice reads the data from the cache area and restores the data after deserialization.

[0102] refer to Fig.12 , Fig.12 It is a schematic diagram of a cache data processing flow provided by an application embodiment of the present invention. First, a microservice is created, and a cache area is created based on the microservice. When creating a cache area, the cache size and the replacement strategy are first set, and then the memory structure of the cache area is set, and a serializer and a deserializer are created. Finally, a cache area (off-heap cache) is created based on the cache size, replacement strategy, memory structure, serializer and deserializer. The microservice is also used to provide read and write services for cache data. When it is necessary to write cache data to the cache area, first determine whether the capacity of the cache area is a threshold. The capacity refers to the memory size used by the cache area. When the capacity is greater than the threshold, the cache area is cached and replaced, and the most recently unused data in the cache area is deleted, and then the cache data is serialized and written to the cache area. If the capacity is less than the threshold, the cache data is directly serialized and written to the cache area. When the microservice reads the cache data in the cache area, it reads the corresponding byte array in the cache area according to the key value in the read request, and then calls the deserializer deserialize() method to deserialize and restore the data.

[0103] The application embodiment of the present invention stores cache data in an off-heap cache area through microservices. The off-heap cache area does not perform garbage collection of the JVM, but instead performs memory management through its own configured strategy, thereby solving the problem of application freezes caused by GC.

[0104] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0105] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0106] It should be noted that the technical solutions described in the embodiments of the present invention can be arbitrarily combined without conflict.

[0107] In addition, in the embodiments of the present invention, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0108] refer to Fig.13 , Fig.13 is a schematic diagram of a cache data processing device provided by an embodiment of the present invention, such as Fig.13 As shown, the device includes: a calling module and a writing module.

[0109] A calling module, used for calling a microservice corresponding to a set first working node in a distributed system when the first working node generates cache data in the process of executing a task;

[0110] A writing module is used to write the cache data into a corresponding cache area based on the microservice corresponding to the first working node; the cache area is an off-heap cache area.

[0111] The device also includes:

[0112] A determination module, used to determine the memory usage of the cache area corresponding to the first working node;

[0113] A deletion module is used to delete the first cache data in the corresponding cache area when the memory usage is greater than a set value; the first cache data represents the cache data with the largest timestamp in the corresponding cache area; the timestamp represents the time interval between the time when the corresponding cache data was last accessed and the current time.

[0114] The writing module is specifically used for:

[0115] Serializing the cached data based on the microservice corresponding to the first working node to obtain corresponding serialized data;

[0116] Based on the microservice corresponding to the first working node, the serialized data is written into the corresponding cache area.

[0117] The device also includes:

[0118] A first creation module, used to create a microservice corresponding to the first working node;

[0119] The second creation module is used to create a corresponding cache area based on the microservice corresponding to the first working node.

[0120] The second creation module is specifically used for:

[0121] Create a serializer and a deserializer corresponding to the cache area; the serializer is used to serialize the cache data; the deserializer is used to convert the serialized data in the cache area into the corresponding cache data;

[0122] Create a buffer based on the set memory size, set memory structure, corresponding serializer and deserializer.

[0123] The device also includes:

[0124] An acquisition module, used to acquire at least one first cache data;

[0125] The memory determination module is used to determine the set memory size based on the memory size of the cache area occupied by each first cache data in the at least one first cache data and the cache data amount of the set cache area.

[0126] The device also includes:

[0127] A reading module, configured to read serialized data corresponding to the read request from the corresponding cache area based on the setting microservice when receiving a read request from the first working node to the corresponding cache area;

[0128] A deserialization module is used to deserialize the serialized data based on the setting microservice to obtain cache data corresponding to the read request.

[0129] In actual application, the calling module and the writing module can be implemented by a processor in an electronic device, such as a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU) or a programmable gate array (FPGA).

[0130] It should be noted that: the cache data processing device provided in the above embodiment only uses the division of the above modules as an example when performing cache data processing. In actual applications, the above processing can be assigned to different modules as needed, that is, the internal structure of the device is divided into different modules to complete all or part of the processing described above. In addition, the cache data processing device provided in the above embodiment and the cache data processing method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0131] Based on the hardware implementation of the above program modules and in order to implement the method of the embodiment of the present application, the embodiment of the present application also provides an electronic device. Fig.14 Schematic diagram of the hardware structure of the electronic device of the present application embodiment. Fig.14 As shown, the electronic equipment includes:

[0132] Communication interface, which can exchange information with other devices such as network equipment;

[0133] The processor is connected to the communication interface to realize information exchange with other devices, and is used to execute the method provided by one or more technical solutions of the electronic device side when running the computer program. The computer program is stored in the memory.

[0134] Of course, in actual applications, the various components in the electronic device are coupled together through the bus system. It is understandable that the bus system is used to achieve connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig.14 In the specification, various buses are labeled as bus systems.

[0135] The memory in the embodiment of the present application is used to store various types of data to support the operation of the electronic device. Examples of such data include: any computer program used to operate on the electronic device.

[0136] It can be understood that the memory can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAM bus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 130 described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memories.

[0137] The method disclosed in the above embodiment of the present application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the program in the memory and completes the steps of the above method in combination with its hardware.

[0138] Optionally, when the processor executes the program, the corresponding processes implemented by the electronic device in each method of the embodiments of the present application are implemented, which will not be described in detail here for the sake of brevity.

[0139] In an exemplary embodiment, the present application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a first memory storing a computer program, and the computer program can be executed by a processor of an electronic device to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.

[0140] In the several embodiments provided in the present application, it should be understood that the disclosed devices, electronic devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0141] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0142] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0143] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium, which, when executed, executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, disks or optical disks.

[0144] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0145] It should be noted that the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.

[0146] In addition, in the examples of the present application, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0147] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A cache data processing method, characterized in that: The method comprises: When a first working node in a distributed system generates cache data during task execution, a microservice corresponding to the first working node is called; The cache data is written into a corresponding cache area based on the microservice corresponding to the first working node; the cache area is an off-heap cache area; the memory size of the cache area is pre-set before creating the cache area according to the memory size of the cache area occupied by each first cache data in the at least one first cache data obtained and the cache data amount of the set cache area; Wherein, the method further comprises: Create a serializer and a deserializer corresponding to the cache area; the serializer is used to serialize the cache data; the deserializer is used to convert the serialized data in the cache area into the corresponding cache data; Based on the set memory size, the set memory structure, the corresponding serializer and the deserializer, a cache area is created through the microservice.

2. The method according to claim 1, characterized in that Before the microservice corresponding to the first working node writes the serialized data into the corresponding cache area, the method further includes: Determine the memory usage of the cache area corresponding to the first working node; When the memory usage is greater than the set value, the first cache data in the corresponding cache area is deleted; the first cache data represents the cache data with the largest timestamp in the corresponding cache area; the timestamp represents the time interval between the time when the corresponding cache data was last accessed and the current time.

3. The method according to claim 1, characterized in that: The microservice corresponding to the first working node writes the serialized data into a corresponding cache area, including: Serializing the cached data based on the microservice corresponding to the first working node to obtain corresponding serialized data; Based on the microservice corresponding to the first working node, the serialized data is written into the corresponding cache area.

4. The method according to claim 1, characterized in that: The method further comprises: Create a microservice corresponding to the first working node.

5. The method according to claim 3, characterized in that: The method further comprises: In case of receiving a read request from the first working node to the corresponding cache area, reading serialized data corresponding to the read request from the corresponding cache area based on the setting microservice; The serialized data is deserialized based on the setting microservice to obtain cache data corresponding to the read request.

6. A cache data processing device, characterized in that: include: A calling module, used for calling a microservice corresponding to a set first working node in a distributed system when the first working node generates cache data in the process of executing a task; A writing module, used to write the cached data into a corresponding cache area based on the microservice corresponding to the first working node; The cache area is an off-heap cache area, and is created by the microservice based on a set memory size, a set memory structure, and a serializer and a deserializer corresponding to the created cache area; The serializer is used to serialize the cache data; the deserializer is used to convert the serialized data in the cache area into the corresponding cache data; The memory size of the cache area is preset before creating the cache area according to the memory size of the cache area occupied by each first cache data of the at least one first cache data acquired and the cache data amount of the set cache area.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the cache data processing method according to any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the cache data processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Operation method, device and electronic equipment for multi-stage cache

    CN107102896A