Data processing method and device
By using multiple memory file descriptors to represent the address space of the data object in a distributed system, the problem of low memory utilization caused by the overall cache of data objects is solved, and more efficient memory usage is achieved.
Patent Information
- Application Number
- CN202410178037.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-02-07
- Publication Date
- 2025-05-20
AI Technical Summary
In a distributed system, the cache unit of the data object is the whole data object, resulting in a low memory utilization rate when only part of the data in the data object is needed to be accessed.
By determining multiple memory file descriptors corresponding to the address space of the application data of the target data object in the memory in the storage server, and sending the memory addresses of these memory file descriptors to the storage client, data object access with the memory file descriptor as granularity is realized.
This method allows data objects to be split and accessed according to the granularity of memory file descriptors, improving memory utilization and allowing memory to cache more hot data of data objects.
Smart Images

Figure CN120020730A_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 202311541648.5 and the application title "A Data Processing Method, Apparatus and Other Devices" submitted to the National Intellectual Property Administration on November 17, 2023. The entire content of which is incorporated herein by reference. Technical Field
[0002] Embodiments of this application relate to the storage field, and in particular, to a data processing method and apparatus. Background Art
[0003] Data caching technology is a storage unit between traditional memory and a processor, used to temporarily store accessed data and instructions to improve the access speed of a computer processor. Data caching usually uses memory to temporarily store and quickly access data. When a computer application needs to read a certain data, it first checks whether the data already exists in the cache. If so, it can directly read from the cache, thereby reducing the time and overhead required to read data from a disk or network.
[0004] A Memory File Descriptor (MEMFD) is a system call in the Linux kernel used to create an anonymous, temporary, memory-based file that can be used in scenarios such as temporary files, shared memory, and inter-process communication. In a distributed system, it is mainly used for data sharing between processes to achieve the purpose of data passing between processes without copying.
[0005] However, when a distributed system creates corresponding memory file descriptors for data objects to indicate the data of the data objects, the data caching unit is the entire data object. One memory file descriptor is required to correspond to the data of one data object. When a large data operation only needs to access part of the data in a data object, the entire data object will also be cached, resulting in low memory utilization. Summary of the Invention
[0006] This application provides a data processing method and apparatus, thereby improving memory utilization in a large data scenario.
[0007] In a first aspect, the present application provides a data processing method, which is applied to a storage server in a distributed system. The storage server is connected to a storage client in the distributed system. In the process of the data processing method, first, the storage server receives a read request for a target data object sent by the storage client. The read request carries parameters of the target data object, and the parameters of the target data object are used to indicate the address space in memory of the application data included in the target data object. Then, the storage server determines a plurality of memory file descriptors corresponding to the address space, and the plurality of memory file descriptors indicate the application data included in the target data object. Finally, the storage server sends the memory addresses of the plurality of memory file descriptors to the storage client.
[0008] Among them, the parameters of the target data object include the identifier of the target data object, the offset address of the target data object, and the length of the target data object. In this way, when the data object is stored in a plurality of memory file descriptors, the storage server can determine the data range of the address space of the data object according to the object identifier, offset address, and length, so as to determine the plurality of memory file descriptors corresponding to the data range.
[0009] In the existing memory sharing scheme in a distributed system, a storage server or a storage server service is deployed on each node of the distributed system. An application can create a data object through the storage server, create a memory file for storing the data object, and at the same time map the memory address of the memory file to the process of the application through the memory file descriptor through the storage client for sharing. The application writes application data to the memory address and returns the object identifier (object identify) of the data object to the application, so that the application can read the data in the object at any node in the distributed system through the object identifier. However, the caching unit of the application data is the entire data object, and one memory file descriptor is required to correspond to the data of one data object. When a large data operation only needs to access a part of the data in the data object, the entire data object indicated by the memory file descriptor will also be cached, resulting in low memory utilization.
[0010] In the data processing method of the present application, compared with the above memory sharing scheme, the data range of the address space in memory of the application data included in a data object corresponds to multiple memory file descriptors. That is, when the data object is cached, it is split into multiple segments of data at a certain granularity, and the address space in memory of each segment of data corresponds to a memory file descriptor. When the data object is accessed, the storage server returns to the storage client the memory addresses of the multiple memory file descriptors corresponding to the data range of the address space in memory of the application data of the data object, thereby completing the access to the data object at the granularity of the memory file descriptor. In this way, the data object can be split and stored and accessed at the granularity of the memory file descriptor, enabling hot data memory storage at the granularity of the memory file descriptor. Thus, the hot data of a data object is stored in memory, and the cold data of the data object is persistently stored in the system file, enabling memory to cache more hot data of data objects and improving memory utilization.
[0011] As a possible implementation, when the storage server determines the multiple memory file descriptors corresponding to the data range of the address space in memory of the application data of the target data object, if there are multiple memory file descriptors corresponding to the data range of the address space, it determines that the multiple memory file descriptors to be sent to the storage client are the multiple memory file descriptors corresponding to the data range of the address space. If there are no multiple memory file descriptors corresponding to the data range of the address space, the storage server applies for new memory file descriptors to load the application data included in the target data object, and determines that the multiple memory file descriptors to be sent to the storage client are the new memory file descriptors. In this way, the storage server can read the data object from the memory file descriptor or read the data object after the memory file descriptor loads the data according to the data object reading requirements of the storage client, making the memory file descriptor more flexible when loading or storing the data object and improving memory utilization.
[0012] As a possible implementation, before the storage server responds to the storage client for data object access, it also needs to create the target data object and cache the application data of the target data object. Optionally, the storage server receives an application request sent by the storage client. The application request includes the application data included in the target data object and is used to request to write the application data included in the target data object. Then, it determines multiple memory file descriptors in the memory file descriptor pool according to the application request. The memory file descriptor pool includes at least one unreleased memory file descriptor, and then writes the application data into the memory files indicated by the multiple memory file descriptors according to the application request. In this way, the storage server uses the unreleased memory file descriptors in the memory file descriptor pool to write the application data, recycling the memory file descriptors, avoiding memory page faults when creating new memory file descriptors every time application data needs to be cached, and reducing data write latency.
[0013] Optionally, when there are idle memory file descriptors in the memory file descriptor pool, the storage server determines multiple memory file descriptors in the memory file descriptor pool; when there are no idle memory file descriptors in the memory file descriptor pool, the storage server creates multiple memory file descriptors from the operating system. Among them, when the storage server creates memory file descriptors, if there are n missing memory file descriptors in the memory file descriptor pool, the storage server creates n memory file descriptors from the operating system. The value of n can be 1, or the total number of memory file descriptors required to store application data, or any value between 1 and the total number of memory file descriptors required to store application data.
[0014] Optionally, this application does not limit the size of the memory file descriptor.
[0015] For example, the size of each memory file descriptor among the multiple memory file descriptors is the same, that is, the size of the memory file indicated by each memory file descriptor is the same. In this way, the storage server can quickly determine the number of memory file descriptors required according to the storage capacity requirement of the application data, improving the data processing efficiency.
[0016] Another example is that the size of each memory file descriptor among the multiple memory file descriptors is partially or completely different. In this way, the storage server can flexibly select memory file descriptors of corresponding sizes according to the storage capacity requirement of the application data, avoiding memory waste caused by size mismatch and improving the memory utilization rate.
[0017] Optionally, when the storage server writes application data into multiple memory file descriptors, it splits the application data and writes it into multiple memory file descriptors in units of memory file descriptors.
[0018] As a possible implementation, the storage server persists the cold data in the application data in units of memory file descriptors, and releases the memory file descriptors after persistent storage back to the memory file descriptor pool. In this way, for the hot data and cold data of the same data object, the hot data can be cached in the memory, and the cold data can be written to the system file for persistent storage, enabling the memory to cache the hot data of more data objects and improving the memory utilization rate.
[0019] Second aspect, the present application provides a data processing method, which is applied to a storage client in a distributed system, and the storage client is connected to a storage server in the distributed system. In the process of the data processing method, first, the storage client sends a read request for a target data object to the storage server to instruct the storage server to determine a data range corresponding to the address space of the application data included in the target data object in the memory, and send the memory addresses of multiple memory file descriptors to the storage client. Then, the storage client maps the memory addresses of the multiple memory file descriptors to a continuous address space. Finally, the storage client returns the starting address of the continuous address space to the application.
[0020] As a possible implementation, the storage client may also send an application request to the storage server. The application request includes the application data included in the data object and is used to request to write the application data included in the target data object. The storage server determines multiple memory file descriptors in the memory file descriptor pool according to the application request, and writes the application data to the memory files indicated by the multiple memory file descriptors according to the application request. The memory file descriptor pool includes at least one unreleased memory file descriptor. Then, the storage client maps the memory addresses of the multiple memory file descriptors to a continuous address space, writes the application data to the continuous address space, and then returns the object identifier of the data object to the application. The object identifier is used to indicate the target data object.
[0021] The difference between the second aspect and the first aspect lies in the different execution entities. The second aspect can be combined with any possible implementation of the first aspect, which will not be elaborated here.
[0022] Third aspect, the present application provides a data processing system, including a storage client and a storage server. The storage server is used to receive a read request for a target data object sent by the storage client; the read request carries parameters of the target data object, and the parameters of the target data object are used to indicate the address space of the application data included in the target data object in the memory, where the parameters of the target data object include the identifier of the target data object, the offset address of the target data object, and the length of the target data object; determine multiple memory file descriptors corresponding to the address space; the multiple memory file descriptors indicate the application data included in the target data object; send the memory addresses of the multiple memory file descriptors to the storage client. The storage client is used to send a read request for a target data object to the storage server, map the memory addresses of the multiple memory file descriptors to a continuous address space, and return the starting address of the continuous address space to the application.
[0023] Fourthly, the present application provides a data processing device, including a transceiver module and a processing module. The transceiver module is configured to receive a read request for a target data object sent by a storage client; the read request carries parameters of the target data object, and the parameters of the target data object are used to indicate the address space of the application data included in the target data object in the memory, wherein the parameters of the target data object include an identifier of the target data object, an offset address of the target data object, and a length of the target data object. The processing module is configured to determine a plurality of memory file descriptors corresponding to the address space; the plurality of memory file descriptors indicate the application data included in the target data object. The transceiver module is further configured to send the memory addresses of the plurality of memory file descriptors to the storage client.
[0024] As a possible implementation manner, the data processing device may further include other modules that perform the operation steps of the data processing method described in the first aspect.
[0025] Fifthly, the present application provides a data processing device, including a transceiver module and a processing module. The transceiver module is configured to send a read request for a target data object to a storage server; the read request carries parameters of the target data object, and the parameters of the target data object are used to indicate the address space of the application data included in the target data object in the memory, wherein the parameters of the target data object include an identifier of the target data object, an offset address of the target data object, and a length of the target data object, so as to instruct the storage server to determine a plurality of memory file descriptors corresponding to the address space and send the memory addresses of the plurality of memory file descriptors to the storage client; the plurality of memory file descriptors indicate the application data included in the target data object. The processing module is configured to map the memory addresses of the plurality of memory file descriptors to a continuous address space. The transceiver module is further configured to send the starting address of the continuous address space to the application.
[0026] As a possible implementation manner, the data processing device may further include other modules that perform the operation steps of the data processing method described in the second aspect.
[0027] Regarding the technical principles and beneficial effects of the second aspect, the third aspect, the fourth aspect, and the fifth aspect, reference may be made to the relevant descriptions of the first aspect above, and details are not described herein again.
[0028] Sixthly, a computing device cluster is provided, including at least one computing device, and each computing device includes a processor and a memory. The processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, so that the computing device cluster executes the data processing method described in any possible implementation manner of the first aspect above.
[0029] In a seventh aspect, there is provided a cluster of computing devices, including at least one computing device, and each computing device includes a processor and a memory. The processor of at least one computing device is configured to execute instructions stored in the memory of at least one computing device, so that the cluster of computing devices performs the data processing method described in any possible implementation manner of the second aspect above.
[0030] In an eighth aspect, there is provided a computer program product. The computer program product includes a computer program or instructions, and when the computer program or instructions are run on a computer, the computer is caused to execute the data processing method described in any possible implementation manner of the first aspect above.
[0031] In a ninth aspect, there is provided a computer program product. The computer program product includes a computer program or instructions, and when the computer program or instructions are run on a computer, the computer is caused to execute the data processing method described in any possible implementation manner of the second aspect above.
[0032] In a tenth aspect, there is provided a computer-readable storage medium. The readable storage medium includes: a computer program or instructions; when the computer program or instructions are run on a computer, the computer is caused to execute the data processing method described in any possible implementation manner of the first aspect above.
[0033] In an eleventh aspect, there is provided a computer-readable storage medium. The readable storage medium includes: a computer program or instructions; when the computer program or instructions are run on a computer, the computer is caused to execute the data processing method described in any possible implementation manner of the second aspect above. Description of the Drawings
[0034] Figure 1 It is a schematic diagram of the architecture of a distributed system provided by this application;
[0035] Figure 2 It is a schematic diagram of the structure of a virtualization node provided by this application;
[0036] Figure 3 It is a schematic diagram of the structure of a virtualization layer provided by this application;
[0037] Figure 4 It is a schematic diagram of the flowchart of a data processing method provided by this application;
[0038] Figure 5 It is a schematic diagram of the flowchart of the management steps of a memory file descriptor provided by this application;
[0039] Figure 6 It is a schematic diagram of the structure of a data processing device provided by this application;
[0040] Figure 7 Structural schematic diagram of another data processing device provided by this application;
[0041] Figure 8 Structural schematic diagram of a computing device provided by this application;
[0042] Figure 9 Structural schematic diagram of a computing device cluster provided by this application;
[0043] Figure 10 Structural schematic diagram of the connection between computing devices through a network provided by this application. Specific implementation manners
[0044] The data processing method provided by the embodiments of this application can be applied to distributed scenarios in the storage field. The following briefly introduces the technologies that this application may involve.
[0045] (1) Distributed system
[0046] A distributed system, also known as a distributed computer system, refers to a system formed by connecting multiple dispersed computing devices through communication lines. Different functions of the system (such as processing, control, storage, etc.) are distributed in each computing device. According to the system functions, distributed systems include distributed systems, distributed computing systems, distributed message queue systems, distributed machine learning systems, etc.
[0047] Among them, a distributed system refers to a system that dispersedly stores data on multiple independent storage nodes. The distributed network storage system adopts an extensible system structure and uses multiple storage nodes to share the storage load. It not only improves the reliability, availability, and access efficiency of the storage system but is also easy to expand.
[0048] (2) Distributed Shared Memory (DSM)
[0049] The shared memory technology provides an abstraction of a unified address space for upper-layer applications, enabling computing tasks running on different hardware units to access the content of local memory and the memory content on remote hardware units in a unified addressing manner. The single-machine shared memory solution has been widely applied in multi-core processors. Distributed shared memory is an important technology that emerged in the development of parallel processing and has also been applied to application scenarios such as distributed key-value storage systems and distributed transaction processing systems. Distributed shared memory provides a logically unified address space, and any node can directly perform read and write operations on this address space. It has the advantages of the scalability of the distributed memory structure and also has the advantages of good generality, portability, and easy programming of the shared memory structure.
[0050] (3) Memory File Descriptor (MEMFD)
[0051] A file descriptor (FD) is a tool in the system call interface for accessing and operating on files. It is a non - negative integer used to identify a record of an open file maintained by the kernel for a specific process. Each file descriptor corresponds to an open file, and these file descriptors are required to indicate the file to be operated on when performing file read and write operations. In fact, it is an index value pointing to the table maintained by the kernel for each process's open files. When a program opens an existing file or creates a new file, the kernel returns a file descriptor to the process. A memory file descriptor is a file descriptor that points to a memory file.
[0052] This application provides a data processing method, especially a "data processing method for splitting and caching data objects in memory files indicated by multiple memory file descriptors". First, the storage server receives a read request for a target data object sent by the storage client. The read request carries parameters of the target data object, and the parameters of the target data object are used to indicate the address space in memory of the application data included in the target data object. Among them, the parameters of the target data object include the identifier of the target data object, the offset address of the target data object, and the length of the target data object. Then, the storage server determines multiple memory file descriptors corresponding to the address space. Finally, the storage server sends the memory addresses of the multiple memory file descriptors to the storage client, so that the storage client maps the memory addresses of the multiple memory file descriptors to a continuous address space and returns the starting address of the continuous address space to the application.
[0053] Based on this data processing method, the data range of the address space in memory of the application data included in a data object corresponds to multiple memory file descriptors, that is, when the data object is cached, it is split into multiple segments of data at a certain granularity, and the address space in memory of each segment of data corresponds to a memory file descriptor. When the data object is accessed, the storage server sends the memory addresses of the multiple memory file descriptors corresponding to the data range of the address space in memory of the application data of the data object to the storage client, thus completing the access to the data object at the granularity of the memory file descriptor. In this way, the data object can be split and stored and accessed at the granularity of the memory file descriptor, enabling hot data memory storage at the granularity of the memory file descriptor, so as to store the hot data of a data object in memory and persistently store the cold data of the data object in the system file, enabling the memory to cache the hot data of more data objects and improving the memory utilization rate.
[0054] The following will describe in detail the implementation manner of the embodiments of this application with reference to the accompanying drawings.
[0055] Figure 1 This is a schematic diagram of the architecture of a distributed system provided for this application. As Figure 1 shown, the distributed system 100 includes a computing server cluster 110, a storage server cluster 120, a management server cluster 130, a network device cluster 140, and a user terminal 150. Among them, the computing server cluster 110, the storage server cluster 120, and the management server cluster 130 are respectively in communication with the user terminal 150 through the network device cluster 140.
[0056] The computing server cluster 110 includes one or more computing servers ( Figure 1 two computing servers, namely computing server 111 and computing server 112, are shown in the figure, but it is not limited to two computing servers).
[0057] As computing resources in the distributed system 100, such as servers, desktop computers, etc., the computing servers are used to generate and allocate computing resources based on virtualization technology according to user needs. At the hardware level, a processor and a memory are provided in the computing server ( Figure 1 not shown in the figure), and the computing function of the computing server is implemented by the processor running the program in the memory. The computing server can also read / write data in each storage server in the storage server cluster 120 according to user needs.
[0058] The storage server cluster 120 includes one or more storage servers ( Figure 1 two storage servers, namely storage server 121 and storage server 122, are shown in the figure, but it is not limited to two storage servers).
[0059] As storage resources in the distributed system 100, such as servers, desktop computers, or controllers of storage arrays, hard disk enclosures, etc., the storage servers are used to provide services such as logical disk storage, semi-structured data storage, and integrated backup for cloud virtual machines in the distributed system 100. Hardware-wise, a network card, a processor, and a memory are provided in the storage server. The processor in the storage server is used to process data from outside the storage server. The network card is used to control the access process to the memory, such as controlling address signals, data signals, and various command signals, so that the storage server can provide the memory as a storage resource to users. The memory is used to store data and can include memory and / or a hard disk. Memory refers to the internal memory that directly exchanges data with the processor. Memory can read and write data quickly at any time and serves as a temporary data storage for the operating system or other running programs. Different from memory, the hard disk reads and writes data more slowly than memory and is usually used to persistently store data.
[0060] The management server cluster 130 includes one or more management servers ( Figure 1 as shown in Figure 1 , two management servers, namely management server 131 and management server 132, are shown, but it is not limited to two management servers).
[0061] The management server is used to manage all computing services, shared storage, and networks of the entire distributed system 100, and at the same time provide an application program interface (API) for managing the entire node to users or administrators. In this application, the distributed system 100 can provide the program product of this application to users by providing an accessible application program interface.
[0062] The network device cluster 140 includes one or more switches and routers. As Figure 1 shown, in this embodiment, the network device cluster 140 includes router 141, switch 142, switch 143, switch 144, and switch 145. Among them, the user terminal 150 is connected to router 141 through the Internet, router 141 is then connected to switch 143 through switch 142, and switch 143 is connected to each computing server in the computing server cluster 110. Switch 144 is respectively connected to each computing server in the computing server cluster 110 and each storage server in the storage server cluster 120. Switch 145 is respectively connected to each computing server in the computing server cluster 110, each storage server in the storage server cluster 120, and each management server in the management server cluster 130.
[0063] Optionally, the number and type of switches included in the network device cluster 140 can be adjusted according to the requirements of the distributed system 100. Switches 142, 143, 144, and 145 can be switches with different functions. For example, switch 142 is a core switch, and switches 143, 144, and 145 are switches for managing specific network segments. Exemplarily, switch 142 can be a core switch, switch 143 can be an internal and external switching network segment switch, switch 144 can be a storage network segment switch, and switch 145 can be a management network segment switch.
[0064] The user terminal 150 includes one or more user terminals ( Figure 1 as shown in Figure 1 , two user terminals, namely user terminal 151 and user terminal 152, are shown, but it is not limited to two user terminals). The user terminal includes interfaces and applications required to access the distributed system 100.
[0065] It should be noted that Figure 1It is only a schematic diagram and should not be construed as a limitation to this application. Other devices or other architectures may also be included in the distributed system 100, which are not drawn in Figure 1 For example Figure 1 The computing server cluster 110, the storage server cluster 120, and the management server cluster 140 in
[0066] The computing servers, storage servers, etc. in the distributed system 100 provided by this application may be nodes in a cloud platform. Based on the devices of the distributed system 100 shown in Figure 1 The distributed system 100 realizes the functions of nodes such as computing nodes and storage nodes based on software as a service (SaaS), platform as a service (PaaS), and infrastructure as a service (IaaS), and provides services (such as computing services, storage services, and network services, etc.) for the user terminal 150 through the nodes. The node may be a cloud service node (such as a computing node and a storage node, etc.) virtualized by using the resources (such as computing resources and storage resources, etc.) of the distributed system 100.
[0067] For example Figure 2 As shown in Figure 2 The computing server cluster 110 in the distributed system 100 is virtualized into storage clients ( Figure 2 Two storage clients, namely storage client 211 and storage client 212, are shown in
[0068] The storage client is used to run an application. The application applies for multiple memory file descriptors from the storage server. The storage client is used to map the memory spaces corresponding to the memory files indicated by the multiple memory file descriptors into a continuous address space and share it with the application, write application data into the address space, and return the object identifier to the application.
[0069] When the application on the storage client reads the target data object, it reads the application data in the memory files indicated by multiple memory file descriptors corresponding to the data range of the target data object from the storage server through a read request for the target data object. The storage server is used to map the memory addresses of the memory files indicated by the multiple memory file descriptors to a continuous address space, and return the starting address of the continuous address space to the storage client, so that the application can access the data according to the starting address.
[0070] The storage server is used to reserve at least one memory file descriptor according to the configured cache space size at startup, and store application data in units of memory file descriptors.
[0071] The storage server circularly uses the memory file descriptors in the memory file descriptor pool. For example, it applies for memory file descriptors from the memory file descriptor pool according to the storage requirements of data objects, releases the memory file descriptors back to the memory file descriptor pool, or eliminates the memory file descriptors from the memory file descriptor pool to the file system. Among them, the memory file descriptor pool includes at least one unreleased memory file descriptor.
[0072] Next, in combination with Figure 3 the virtualization hierarchical structure of the distributed system 100 will be described.
[0073] As Figure 3 shown, the IaaS platform 310 is used to perform virtualization of all infrastructure resources in the distributed system 100, and provide virtual resources (such as computing resources, network resources, and storage resources) for users through a software-defined method. Among them, the infrastructure resources refer to the resources provided by the computing server cluster 110, storage server cluster 120, and / or network device cluster 130 in the distributed system 100.
[0074] The PaaS platform 320 is used to implement the runtime environment and application support functions of the distributed system 100, so that users can apply for computing units within the quota instead of virtual resources to run their own services. Optionally, the computing unit can be a container, and the distributed system 100 deploys and runs the user's code by scheduling containers. It should be noted that the number of containers in the PaaS platform 320 can be one or more, Figure 3 and only one container is taken as an example here.
[0075] As a possible implementation, the distributed system 100 can inject one or more components into the container to implement the deployment and running of the code. Optionally, the resources (computing resources or storage resources) used by multiple components in the same container can belong to the same hardware device (such as a computing server or a storage server) in the distributed system 100, or can belong to different hardware devices.
[0076] The SaaS application 330 is used to form service nodes from the applications deployed by users in the form of API responses based on the IaaS platform 310 and the PaaS platform 320, and provide them to users. The application of the SaaS application 330 and the container can communicate through a web server.
[0077] It should be noted that Figure 3 This is only a schematic diagram and should not be construed as a limitation of this application. Other modules may also be included in the virtualized hierarchical structure of the distributed system 100, which are not drawn in Figure 3 the figure.
[0078] Next, the data processing method provided in this embodiment will be specifically described with reference to the accompanying drawings.
[0079] The steps of the data processing method provided in this application are executed by the nodes virtualized by the distributed system 100, such as the client, the server, etc. Next, in combination with Figure 4 , taking the storage client 211 and the storage server 221 as examples, the data processing method provided in the embodiments of this application will be described.
[0080] Please refer to Figure 4 , Figure 4 FIG. is a schematic flowchart of a data processing method provided in this application. The data processing method may include the following steps 410-step 480.
[0081] Step 410, the storage client 211 obtains a data processing request.
[0082] The data processing request may be generated by an application in response to a user operation, and the user operation may be sent by the terminal 151. The data processing request may be used to indicate any type of data operation on the target data object, such as a data write operation, a data read operation, a data deletion operation, a data overwrite operation, a data query operation, etc.
[0083] Step 420, the storage client 211 generates a read request for the target data object according to the data processing request.
[0084] The read request includes the parameters of the target data object. The parameters of the target data object are used to indicate the address space in the memory of the application data included in the target data object.
[0085] As a possible implementation, the parameters of the target data object are used to indicate the address space in the memory of the application data of the data object that the application indicated by the data processing request needs to access. The application data of the data object that the application needs to access may be part or all of the application data of the target data object.
[0086] As a possible implementation, the parameters of the target data object include an object identifier, an offset address, and a length. The object identifier is used to indicate the data object. The offset address is used to indicate the distance between the starting position of the data block of the data object and the base address, where the data block is the data block that the application wants to read, and the base address is the starting address of the data object. The length is used to indicate the data length of the data object.
[0087] Step 430: The storage client 211 sends a read request to the storage server 221.
[0088] The storage client 211 sends a read request to the storage server 221, and the read request includes the parameters of the target data object.
[0089] Step 440: The storage server 221 receives the read request sent by the storage client 211.
[0090] Step 450: The storage server 221 determines multiple memory file descriptors corresponding to the data range of the address space.
[0091] The storage server 221 determines multiple memory file descriptors corresponding to the data range of the address space indicated by the parameters of the target data object.
[0092] As a possible implementation, the object identifier, the offset address, and the length specify a virtual address space. The starting position to the ending position of this address space is called the data range corresponding to this address space. A memory file descriptor indicates a created memory file. The memory file has a corresponding address space in the memory, and this address space can be called the address space corresponding to the memory file descriptor. The address space corresponding to each memory file descriptor is used to store application data. The storage server 221 determines the address space in the memory indicated by the application data of the target data object according to the mapping relationship between the address space indicated by the parameters of the target data object and the address space in the memory, and uses the multiple memory file descriptors to which the address space in this segment of memory belongs as the multiple memory file descriptors corresponding to the data range of the address space indicated by the parameters of the target data object.
[0093] Among them, for the multiple memory file descriptors corresponding to the data range of the address space indicated by the parameters of the target data object, each memory file descriptor is used to indicate a memory file. The multiple memory files corresponding to the multiple memory file descriptors are used to cache part or all of the application data of the target data object indicated by the object identifier of the parameters of the target data object.
[0094] As a possible implementation, the target data object is created by the application on the storage server 221, the memory file descriptor is applied for by the application from the storage server 221, and the application data is written by the storage client 211 into the memory files indicated by multiple memory file descriptors. For the specific steps, please refer to Figure 5 Steps 510 - 550 shown, which will not be elaborated here.
[0095] As a possible implementation, if part or all of the application data of the target data object indicated by the parameters of the target data object is not in the memory files indicated by the memory file descriptors, that is, part or all of the application data of the target data object indicated by the parameters of the target data object is not in the memory, the storage server 221 applies to create a new memory file descriptor, and loads the application data not in the memory from the system file through the memory file indicated by the new memory file descriptor.
[0096] Step 460: The storage server 221 sends the memory addresses of multiple memory file descriptors to the storage client 211.
[0097] The memory addresses of multiple memory file descriptors are the physical addresses in the memory of the memory files indicated by multiple memory file descriptors. The corresponding relationship between any memory file descriptor and the memory address is determined when the memory file descriptor is created.
[0098] As a possible implementation, if the application data of the target data object indicated by the parameters of the target data object is in the memory file descriptors, that is, in the memory files indicated by the memory file descriptors, the storage server 221 sends the memory addresses of multiple memory file descriptors storing the application data to the storage client 211. Among them, the multiple memory file descriptors storing the application data are the multiple memory file descriptors corresponding to the data range of the address space indicated by the parameters of the target data object. In this embodiment, the memory file descriptors storing the application data can be understood as the memory file descriptors of the corresponding memory files that have stored the application data.
[0099] As a possible implementation, if all of the application data of the target data object indicated by the parameters of the target data object is not in the memory file descriptors, the storage server 221 sends the memory address of the newly created memory file descriptor applied for to the storage client 211.
[0100] As a possible implementation, if a part of the application data of the target data object indicated by the parameter of the target data object is not in the memory file descriptor, the storage server 221 sends the memory address of the newly created memory file descriptor to be applied for to the storage client 211, and the memory addresses of at least one memory file descriptor storing the application data. Among them, at least one memory file descriptor storing the application data is at least one memory file descriptor containing the application data among the multiple memory file descriptors corresponding to the data range of the address space indicated by the parameter of the target data object. In this embodiment, the multiple memory file descriptors corresponding to the data range of the address space indicated by the parameter of the target data object can also be understood as the multiple memory file descriptors of multiple memory files corresponding to the data range of the address space indicated by the parameter of the target data object.
[0101] Step 470: The storage client 211 maps the memory addresses of the multiple memory file descriptors to a continuous address space.
[0102] The storage client 211 maps the memory addresses of the multiple memory file descriptors to the address space of the application process, and converts the memory addresses of the multiple memory file descriptors into a continuous address space during the mapping process. In this way, the storage client 211 converts the possibly discontinuous address spaces of the multiple memory file descriptors into a continuous virtual address space during the memory address mapping, which is convenient for the application to call the application data in the memory based on the continuous virtual address space, ensuring that the application can also normally call the application data when the application data is split and stored in units of memory file descriptors. In this embodiment, the application data is split and stored in units of memory file descriptors, which can also be understood as the application data is split and stored in units of the memory file size indicated by the memory file descriptor.
[0103] Step 480: The storage client 211 returns the starting address of the continuous address space to the application.
[0104] The storage client 211 can perform data processing operations on the application data according to the starting address of the continuous address space.
[0105] The application of the storage client 211 performs data processing operations on the application data according to the starting address of the continuous address space.
[0106] The above step 480 is a common operation of shared memory and will not be elaborated here.
[0107] In a possible embodiment of the present application, after the above step 480, the storage server 211 may also record the usage information of the memory file descriptor (such as the access frequency of the memory file descriptor, the access times of the memory file descriptor, the most recent access time of the memory file descriptor, etc.), spill the cold data to the system file for persistent storage, and then release the memory file descriptor corresponding to the cold data back to the memory file descriptor pool. The memory space corresponding to the memory file indicated by the memory file descriptor in the memory file descriptor pool is not released.
[0108] In a possible embodiment of the present application, the target data object accessed by the storage client 211 from the storage server 221 may be on a storage server other than the storage server 211, such as the storage server 222. Then, the storage server 221 reads the application data of the target data object from the storage server 222 according to the parameters of the target data object, and then executes the above steps 460 - 480.
[0109] Based on the above data processing method, the data range of the address space in the memory corresponding to the application data included in a data object corresponds to multiple memory file descriptors, that is, when the data object is cached, it is split into multiple segments of data according to a certain granularity, and the address space of each segment of data in the memory corresponds to a memory file descriptor. When the data object is accessed, the storage server returns the memory addresses corresponding to multiple memory file descriptors of the data range of the address space in the memory of the application data of the data object to the storage client, thereby completing the access to the data object with the memory file descriptor as the granularity. In this way, the data object can be split and stored and accessed with the memory file descriptor as the granularity, and the in-memory storage of hot data with the memory file descriptor as the granularity can be realized, so that the hot data of a data object is stored in the memory, and the cold data of the data object is persistently stored in the system file, enabling the memory to cache more hot data of data objects and improving the memory utilization rate.
[0110] As described above in conjunction with Figure 4 the processing flow of the data processing method provided by the present application has been described. Next, in conjunction with Figure 5 the management steps of the memory file descriptor will be described in detail.
[0111] Please refer to Figure 5 , Figure 5 which is a schematic flow diagram of the management steps of a memory file descriptor provided by the present application. The management steps of the memory file descriptor may include the following steps 510 - 570.
[0112] Step 510: The storage client 211 sends an application request to the storage server 221.
[0113] The application request is used to instruct the storage server 211 to create a memory file descriptor for storing the application data of the target data object. The application request includes the application data of the target data object and is used to request to write the application data included in the target data object.
[0114] Step 520: The storage server 211 determines multiple memory file descriptors in the memory file descriptor pool according to the application request.
[0115] The storage server 211 determines multiple memory file descriptors in the memory file descriptor pool according to the size of the application data of the target data object. The total data capacity of the multiple memory file descriptors is greater than or equal to the size of the application data of the target data object. In this embodiment, the total data capacity of the multiple memory file descriptors can also be understood as the total data capacity of the memory files indicated by the multiple memory file descriptors.
[0116] As a possible implementation, when there are idle memory file descriptors in the memory file descriptor pool, the memory file descriptors for storing the application data are preferentially determined in the memory file descriptor pool. The memory file descriptor pool includes one or more unreleased memory file descriptors. In this way, when the storage server 211 uses the unreleased memory file descriptors, the memory file descriptors have corresponding memory address spaces in the memory, and there will be no page fault exception, reducing the overall operation latency of the distributed system 100.
[0117] For example, when the total data capacity of one or more idle memory file descriptors in the memory file descriptor pool is greater than or equal to the size of the application data, all the multiple memory file descriptors for storing the application data are obtained from the memory file descriptor pool.
[0118] For another example, when the total data capacity of one or more idle memory file descriptors in the memory file descriptor pool is less than the size of the application data, the storage server 211 obtains all the current idle memory file descriptors from the memory file descriptor pool to store part of the application data, and applies for one or more new memory file descriptors from the operating system (OS) to store the remaining part of the application data.
[0119] As a possible implementation, when there are no idle memory file descriptors in the memory file descriptor pool, multiple memory file descriptors are applied for and created from the operating system to store the application data of the data object.
[0120] Among them, at least one memory file descriptor in the memory file descriptor pool is n memory file descriptors reserved according to the configured cache space size when the storage server 211 is started. A memory space with the same size as the memory files indicated by the n memory file descriptors is reserved in the memory. n is a preset quantity. The capacity sizes of each memory file descriptor can be the same or different. In this embodiment, the capacity size of the memory file descriptor can also be understood as the capacity size of the memory file indicated by the memory file descriptor.
[0121] Step 530: The storage client 221 maps the memory addresses of multiple memory file descriptors to a continuous address space.
[0122] The storage client 221 maps the memory addresses of multiple memory file descriptors to the address space of the application process, and converts the memory addresses of the multiple memory file descriptors into a continuous address space during the mapping process.
[0123] Step 540: The storage client 221 writes application data to the continuous address space.
[0124] The application of the storage client 221 writes the application data of the data object to be cached to the continuous address space. Since there is a mapping relationship between the continuous address space and the memory addresses of multiple memory file descriptors, that is, the application of the storage client 221 writes the application data of the target data object to be cached to the memory addresses of multiple memory file descriptors.
[0125] As a possible implementation, the storage client 221 splits the application data and writes it to multiple memory file descriptors in units of memory file descriptors. In this embodiment, splitting the application data and writing it to multiple memory file descriptors can also be understood as splitting the application data and writing it to the memory files indicated by multiple memory file descriptors.
[0126] Step 550: The storage client 221 returns the object identifier of the target data object to the application.
[0127] After the storage client 221 returns the object identifier of the target data object to the application, the application can initiate a data processing request according to the object identifier.
[0128] The management steps of the memory file descriptor, in addition to the process of creating the management steps of the memory file descriptor involved in the above steps 510 - 550, may also involve the release process of the memory file descriptor. The release process of the memory file descriptor can be as shown in the following steps 560 - 570.
[0129] Step 560: When the system memory is sufficient, the storage server 211 releases the memory file descriptor to the memory file descriptor pool.
[0130] As a possible implementation, sufficient system memory means that the system memory is greater than or equal to a preset memory threshold.
[0131] As a possible implementation, after the storage server 211 releases the memory file descriptor to the memory file descriptor pool, the memory file descriptor in the memory file descriptor pool is not released and there is a corresponding address space in the memory.
[0132] Step 570: When the system memory is insufficient, the storage server 211 eliminates the memory file descriptor from the memory file descriptor pool.
[0133] As a possible implementation, insufficient system memory means that the system memory is less than the preset memory threshold.
[0134] As a possible implementation, after the storage server 211 eliminates the memory file descriptor from the memory file descriptor pool, the memory file descriptor is released and there is no corresponding address space in the memory.
[0135] Based on the above steps 510 - 570, through the memory file descriptor pool to manage processes such as application creation and elimination of memory description files, the recycling of memory file descriptors in an unreleased state is achieved, avoiding high latency caused by page faults when creating memory file descriptors for the first time.
[0136] To cooperate with the above Figure 4 or Figure 5 shown data processing method, the present application also provides a data processing device 600, which can be used to implement the functions of the storage server in the above Figure 4 or Figure 5 shown data processing method. As Figure 6 shown, the data processing device 600 includes a transceiver module 610 and a processing module 620.
[0137] The transceiver module 610 is used to receive a read request for a target data object sent by the storage client; the read request carries parameters of the target data object, and the parameters of the target data object are used to indicate the address space in the memory of the application data included in the target data object. Among them, the parameters of the target data object include the identifier of the target data object, the offset address of the target data object, and the length of the target data object. For example, the transceiver module 610 is used to execute step 440 as Figure 4 shown.
[0138] The processing module 620 is used to determine multiple memory file descriptors corresponding to the address space; the multiple memory file descriptors indicate the application data included in the target data object. For example, the processing module 620 is used to execute step as Figure 4Step 450 shown above.
[0139] The transceiver module 610 is further configured to send the memory addresses of multiple memory file descriptors to the storage client. For example, the transceiver module 610 is configured to execute steps such as Figure 4 Step 460 shown above.
[0140] As a possible implementation, the transceiver module 610 is further configured to receive an application request sent by the storage client; the application request is used to request to write the application data included in the target data object. The processing module 620 is further configured to, in response to the application request, determine multiple memory file descriptors in the memory file descriptor pool, and write the application data into the multiple memory file descriptors; wherein, the memory file descriptor pool includes at least one unreleased memory file descriptor.
[0141] As a possible implementation, the processing module 620 is specifically configured to: when there are idle memory file descriptors in the memory file descriptor pool, determine multiple memory file descriptors in the memory file descriptor pool.
[0142] As a possible implementation, the processing module 620 is specifically configured to: when there are no idle memory file descriptors in the memory file descriptor pool, create multiple memory file descriptors from the operating system.
[0143] As a possible implementation, the processing module 620 is specifically configured to: split the application data and write it into multiple memory file descriptors in units of memory file descriptors.
[0144] As a possible implementation, the processing module 620 is further configured to: persistently store the cold data in the application data in units of memory file descriptors; release the memory file descriptors after persistent storage back to the memory file descriptor pool.
[0145] Wherein, both the transceiver module 610 and the processing module 620 can be implemented by software or by hardware. Exemplarily, next, taking the transceiver module 610 as an example, the implementation manner of the transceiver module 610 will be introduced. Similarly, the implementation manner of the processing module 620 can refer to the implementation manner of the transceiver module 610.
[0146] As an example of a software functional unit, the transceiver module 610 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the transceiver module 610 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers for running the code may be distributed in the same availability zone (AZ) or in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.
[0147] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is achieved through the communication gateway.
[0148] As an example of a hardware functional unit, the transceiver module 610 may include at least one computing device, such as a server. Alternatively, the transceiver module 610 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0149] The multiple computing devices included in the transceiver module 610 can be distributed in the same region or in different regions. The multiple computing devices included in the transceiver module 610 can be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the transceiver module 610 can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), and generic array logic (GALs).
[0150] It should be noted that in other embodiments, either the transceiver module 610 or the processing module 620 can be used to execute any step in the data processing method. The steps to be implemented by the transceiver module 610 and the processing module 620 can be specified as needed. By implementing different steps in the data processing method through the transceiver module 610 and the processing module 620 respectively, all functions of the data processing apparatus 600 can be realized.
[0151] To cooperate with the above-mentioned Figure 4 or Figure 5 data processing method shown, the present application also provides a data processing apparatus 700, which can be used to implement the function of storing the client in the above-mentioned Figure 4 or Figure 5 data processing method shown. As Figure 7 shown, the data processing apparatus 700 includes:
[0152] A transceiver module 710, configured to send a read request for a target data object to a storage server; the read request carries parameters of the target data object, and the parameters of the target data object are used to indicate the address space of the application data included in the target data object in the memory. Among them, the parameters of the target data object include the identifier of the target data object, the offset address of the target data object, and the length of the target data object, so as to instruct the storage server to determine multiple memory file descriptors corresponding to the address space and send the memory addresses of the multiple memory file descriptors to the storage client; the multiple memory file descriptors indicate the application data included in the target data object. For example, the transceiver module 710 is used to execute step 430 as Figure 4 shown.
[0153] A processing module 720, configured to map the memory addresses of the multiple memory file descriptors to a continuous address space. For example, the processing module 720 is used to execute step 470 as Figure 4 shown.
[0154] The transceiver module 710 is further configured to return the starting address of the continuous address space to the application. For example, the transceiver module 710 is used to execute step 480 as Figure 4 shown.
[0155] As a possible implementation, the transceiver module 710 is further configured to send an application request to the storage server; the application request is used to request to write the application data included in the target data object, so as to instruct the storage server to determine a plurality of memory file descriptors in the memory file descriptor pool in response to the application request, and write the application data into the plurality of memory file descriptors; wherein, the memory file descriptor pool includes at least one unreleased memory file descriptor. The processing module 720 is further configured to map the memory addresses of the plurality of memory file descriptors to a continuous address space. The transceiver module 710 is further configured to write the application data into the continuous address space, and return the target object identifier of the data object to the application; the object identifier is used to indicate the target data object.
[0156] Both the transceiver module 710 and the processing module 720 can be implemented by software or by hardware. Exemplarily, the implementation manner of the transceiver module 710 will be introduced next. Similarly, the implementation manner of the processing module 720 can refer to the implementation manner of the transceiver module 710.
[0157] As an example of a software functional unit, the transceiver module 710 may include code running on a computing instance. Wherein, the computing instance may be at least one of computing devices such as a physical host (computing device), a virtual machine, a container, etc. Further, the above computing devices may be one or more. For example, the transceiver module 710 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the application program may be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers for running the code may be distributed in the same AZ or in different AZs, and each AZ includes one data center or multiple geographically close data centers. Wherein, generally one region may include multiple AZs.
[0158] Similarly, the multiple hosts / virtual machines / containers for running the code may be distributed in the same VPC or in multiple VPCs. Wherein, generally one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.
[0159] As an example of a hardware functional unit, the transceiver module 710 may include at least one computing device, such as a server, etc. Alternatively, the transceiver module 710 may also be a device implemented by using ASIC or PLD. Wherein, the above PLD may be implemented by CPLD, FPGA, GAL or any combination thereof.
[0160] The multiple computing devices included in the transceiver module 710 can be distributed in the same region or in different regions. The multiple computing devices included in the transceiver module 710 can be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the transceiver module 710 can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices can be any combination of computing devices such as servers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), and generic array logic (GALs).
[0161] This application also provides a computing device 800. As Figure 8 shown, the computing device 800 includes: a bus 802, a processor 804, a memory 806, and a communication interface 808. The processor 804, the memory 806, and the communication interface 808 communicate with each other through the bus 802. The computing device 800 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 800.
[0162] The bus 802 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 8 only one line is shown herein, but it does not mean that there is only one bus or one type of bus. The bus 802 can include a path for transmitting information between various components of the computing device 800 (for example, the memory 806, the processor 804, and the communication interface 808).
[0163] The processor 804 can include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0164] The memory 806 may include volatile memory, such as random access memory (RAM). The processor 804 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0165] The executable program code is stored in the memory 806, and the processor 804 executes the executable program code to implement the functions of the respective modules included in the foregoing data processing device 600 or data processing device 700, thereby implementing the data processing method. That is, instructions for executing the data processing method are stored on the memory 806.
[0166] Alternatively, executable code is stored in the memory 806, and the processor 804 executes the executable code to implement the functions of the foregoing storage client or storage server, thereby implementing the data processing method. That is, instructions for executing the data processing method are stored on the memory 806.
[0167] The communication interface 808 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 800 and other devices or a communication network.
[0168] Considering that the data processing method provided in this application is applied to the distributed system 100, the respective infrastructures of the computing server cluster 110, the storage server cluster 120, etc. of the distributed system 100 usually include multiple computing devices. Therefore, this application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0169] As Figure 9 shown, the computing device cluster includes at least one computing device 800. Instructions for executing the data processing method may be stored in the same manner in the memory 806 of one or more of the computing devices 800 in the computing device cluster.
[0170] In some possible implementations, parts of the instructions for executing the data processing method may also be stored separately in the memories 806 of one or more of the computing devices 800 in the computing device cluster. In other words, the combination of one or more computing devices 800 can jointly execute the instructions for executing the data processing method.
[0171] It should be noted that the memories 806 in different computing devices 800 in the computing device cluster can store different instructions, respectively for executing some functions of the data processing device 600 or the data processing device 700. That is to say, the instructions stored in the memories 806 of different computing devices 800 can implement the functions of one or more modules included in the data processing device 600 or the data processing device 700.
[0172] In some possible implementations, one or more computing devices in the computing device cluster can be connected through a network. Among them, the network can be a wide area network or a local area network, etc. Figure 10 A possible implementation is shown. As Figure 10 shown, two computing devices 800A and 800B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the memory 806 in the computing device 800A stores instructions for executing the functions of one or more of the transceiver module 610 and the processing module 620. Figure 10 Taking the example that the memory 806 in the computing device 800A stores instructions for executing the function of the transceiver module 610. At the same time, the memory 806 in the computing device 800B stores instructions for executing the functions of one or more of the transceiver module 610 and the processing module 620. Figure 10 Taking the example that the memory 806 in the computing device 800B stores instructions for executing the function of the processing module 620.
[0173] It should be understood that Figure 10 the functions shown in the computing device 800A can also be completed by multiple computing devices 800. Similarly, the functions of the computing device 800B can also be completed by multiple computing devices 800.
[0174] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the data processing method as Figure 4 or Figure 5 shown, or as Figure 4 or Figure 5Store the steps executed by the client or the storage server in the data processing method shown.
[0175] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that can be stored by a computing device or a data storage device such as a data center including one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to perform Figure 4 or Figure 5 the data processing method shown, or to perform the steps executed by the client or the storage server in the data processing method shown as Figure 4 or Figure 5 shown.
[0176] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center including one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium may be a solid-state drive.
[0177] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician may use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0178] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0179] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0180] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0181] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0182] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.
[0183] In this application, "at least one" means one or more, and "a plurality of" means two or more. "And / or" describes the relationship between associated objects and indicates that there can be three relationships. For example, A and / or B can represent the cases of A existing alone, A and B existing simultaneously, and B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item)" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.
[0184] It should be noted that in this application, words such as "exemplary" or "for example" are used to give examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0185] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the protection scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data processing method, characterized in that: Applied to a storage service end, the method includes: Receive a read request for a target data object sent by a storage client; the read request carries parameters of the target data object, the parameters of the target data object are used to indicate an address space in the memory of application data included in the target data object, wherein the parameters of the target data object include an identifier of the target data object, an offset address of the target data object, and a length of the target data object; Determine a plurality of memory file descriptors corresponding to the address space; the plurality of memory file descriptors indicate application data included in the target data object; The memory addresses of the multiple memory file descriptors are sent to the storage client.
2. The method according to claim 1, characterized in that Before determining the multiple memory file descriptors corresponding to the address space, the method further includes: receiving an application request sent by the storage client; the application request is used to request writing application data included in the target data object; In response to the application request, the plurality of memory file descriptors are determined in a memory file descriptor pool, and the application data is written to the plurality of memory file descriptors; wherein the memory file descriptor pool includes at least one unreleased memory file descriptor.
3. The method according to claim 2, characterized in that The applying for the multiple memory file descriptors in the memory file descriptor pool includes: When there are free memory file descriptors in the memory file descriptor pool, the multiple memory file descriptors are determined in the memory file descriptor pool.
4. The method according to claim 2, characterized in that: The applying for the multiple memory file descriptors in the memory file descriptor pool includes: When there are no free memory file descriptors in the memory file descriptor pool, the multiple memory file descriptors are created from the operating system.
5. The method according to claim 2, characterized in that: Writing the application data to the multiple memory file descriptors comprises: The application data is split and written into the multiple memory file descriptors based on the granularity of the memory file descriptors.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Using memory file descriptors as granularity, persistently storing cold data in the application data; The persistently stored memory file descriptor is released back to the memory file descriptor pool.
7. A data processing method, characterized in that: Applied to a storage client, the method comprises: Sending a read request for a target data object to a storage service end; the read request carries parameters of the target data object, the parameters of the target data object are used to indicate an address space in a memory of application data included in the target data object, wherein the parameters of the target data object include an identifier of the target data object, an offset address of the target data object, and a length of the target data object, so as to instruct the storage service end to determine a plurality of memory file descriptors corresponding to the address space, and send memory addresses of the plurality of memory file descriptors to the storage client; the plurality of memory file descriptors indicate application data included in the target data object; Mapping the memory addresses of the multiple memory file descriptors into a continuous address space; The starting address of the continuous address space is returned to the application.
8. The method according to claim 7, characterized in that The method further comprises: Sending an application request to the storage service end; the application request is used to request writing the application data included in the target data object, so as to instruct the storage service end to determine the multiple memory file descriptors in the memory file descriptor pool in response to the application request, and write the application data to the multiple memory file descriptors; wherein the memory file descriptor pool includes at least one unreleased memory file descriptor; Mapping the memory addresses of the multiple memory file descriptors into a continuous address space; Writing the application data into the continuous address space; The target object identifier of the data object is returned to the application; the object identifier is used to indicate the target data object.
9. A distributed system, characterized in that: The system includes a storage client and a storage service client; The storage service end is used to receive a read request for a target data object sent by a storage client; the read request carries parameters of the target data object, and the parameters of the target data object are used to indicate an address space in the memory of application data included in the target data object, wherein the parameters of the target data object include an identifier of the target data object, an offset address of the target data object, and a length of the target data object; determine multiple memory file descriptors corresponding to the address space; the multiple memory file descriptors indicate the application data included in the target data object; and send memory addresses of the multiple memory file descriptors to the storage client; The storage client is used to send a read request for a target data object to a storage service end; map the memory addresses of the multiple memory file descriptors into a continuous address space; and return a starting address of the continuous address space to an application.
10. A data processing device, characterized in that: The device comprises: a transceiver module, configured to receive a read request for a target data object sent by a storage client; the read request carries parameters of the target data object, the parameters of the target data object are used to indicate an address space in the memory of application data included in the target data object, wherein the parameters of the target data object include an identifier of the target data object, an offset address of the target data object, and a length of the target data object; A processing module, configured to determine a plurality of memory file descriptors corresponding to the address space; the plurality of memory file descriptors indicating application data included in the target data object; The transceiver module is also used to send the memory addresses of the multiple memory file descriptors to the storage client.
11. A data processing device, characterized in that: The device comprises: A transceiver module, used for sending a read request for a target data object to a storage service end; the read request carries parameters of the target data object, and the parameters of the target data object are used to indicate an address space in a memory of application data included in the target data object, wherein the parameters of the target data object include an identifier of the target data object, an offset address of the target data object, and a length of the target data object, so as to instruct the storage service end to determine a plurality of memory file descriptors corresponding to the address space, and send memory addresses of the plurality of memory file descriptors to the storage client; the plurality of memory file descriptors indicate application data included in the target data object; A processing module, used for mapping the memory addresses of the multiple memory file descriptors into a continuous address space; The transceiver module is further used to return the starting address of the continuous address space to the application.
12. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 6.
13. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to claim 7 or 8.
14. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 6.
15. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the method according to claim 7 or 8.
16. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 6.
17. A computer-readable storage medium, characterized in that: The method comprises computer program instructions, and when the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to claim 7 or 8.
Citation Information
Cited By
Memory allocation method, network equipment and memory allocation device
CN121187803A