Data processing method and apparatus
By implementing the prefetch mechanism on the object storage server, the object data that the client may access is cached in advance, solving the problem of high delay in the object storage server when processing multiple object operation requests, and improving resource utilization and client acquisition speed.
Patent Information
- Application Number
- PCT/CN2024/120858
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2024-09-24
- Publication Date
- 2025-06-19
AI Technical Summary
When the existing object storage server handles multiple object operation requests, it needs to frequently obtain and cache object data, resulting in a high delay.
By implementing the prefetch mechanism on the server, the data of the first target object is obtained and cached from the target object storage device in advance according to the prefetch object identification carried in the second object access request of the client, and the cached data is directly returned when the first object access request is received.
It reduces the latency in object storage, improves resource utilization when the server is idle, and the speed at which the client acquires objects.
Smart Images

Figure CN2024120858_19062025_PF_FP_ABST
Abstract
Description
Data processing method and device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 13, 2023, with application number 202311715740.9 and application name “A method, device and other equipment for data processing”, and the Chinese patent application filed with the State Intellectual Property Office on March 27, 2024, with application number 202410361777.4 and application name “Data processing method and device”, all contents of which are incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of storage, and more particularly to a data processing method and device. Background Art
[0003] In the digital age, data has become a core asset across all industries and requires efficient, scalable, and highly reliable storage and management. Object storage is an object-centric storage method that stores data as objects rather than files. It offers advantages such as high reliability, scalability, and high performance.
[0004] User operations on objects are often performed multiple times over a single persistent connection. However, current object storage servers retrieve the corresponding object data upon receiving a client request. When the number of operations is high, each operation requires the corresponding object data to be retrieved based on the request, resulting in high latency.
[0005] Summary of the Invention
[0006] The present application provides a data processing method and device, thereby reducing the latency problem in object storage.
[0007] In a first aspect, the present application provides a data processing method, which is applied to the server of an object storage service. In the process of the data processing method, the server first receives a first object access request sent by a client of the object storage service, and the first object access request includes a first target object. Then, the server determines that the data of the first target object has been cached, and the data of the first target object is obtained and cached from the target object storage device based on the pre-fetch object identifier carried by the second object access request, and the pre-fetch object identifier is used to indicate that the first target object is a pre-fetch object. Finally, the server sends the data of the first target object to the client of the object storage service.
[0008] Based on the above data processing method, the server determines that the prefetched object to be accessed by the client next is the first target object based on the second object access request that precedes the first object access request. This allows the server to complete the acquisition and caching of the first target object before receiving the first object access request from the client. Upon receiving the first object access request, the cached data of the first target object is returned to the client. This improves resource utilization during idle time on the server, increases the speed at which clients retrieve objects, and reduces latency in object storage.
[0009] As a possible implementation method, before receiving the first object access request, the server needs to obtain and cache the data of the first target object based on the prefetch object identifier carried in the object access request received before the first object access. The server receives the second object access request sent by the client, and the second object access request includes the second target object and the prefetch object identifier. Then, based on the prefetch object identifier, the server obtains and caches the data of the first target object from the target object storage device, and sends the data of the second target object to the client. In this way, when the client sends the second object access request to the server, in addition to carrying the second target object, it also includes the prefetch object identifier of the prefetch object that the client may access next time. The server can obtain and cache the data of the first target object based on the prefetch object identifier, realize the pre-storage of the prefetch object, save the time of obtaining and caching the data of the first target object in the subsequent operation process of the first target object, and improve the speed of the client obtaining the object.
[0010] As a possible implementation method, the server also includes a pre-storage queue and a communication queue. The server determines that the received object access request includes a pre-fetch object identifier, adds the access request for the pre-fetch object in the received object access request to the pre-storage queue, and adds the access request for the target object in the received object access request to the communication queue. The server determines that the received object access request does not include a pre-fetch object identifier, and adds the received object access request to the communication queue. In this way, the server uses the pre-storage queue and the communication queue to process the access request for the pre-fetch object and the access request for the target object respectively, and processes the two access requests asynchronously, avoiding the acquisition of the pre-fetch object and the cache blocking the return of the target object, thereby improving the utilization rate of the server during idle period and further improving the speed of the client acquiring objects.
[0011] As a possible implementation method, the prefetch object identifier includes a whole object prefetch identifier and a range object prefetch identifier. The whole object prefetch identifier is used to indicate the object name of the prefetch object. The range object prefetch identifier is used to indicate the data range of the prefetch object. In this way, the server can cache all the data of the prefetch object according to the whole object prefetch identifier, and can also cache part of the data of the prefetch object according to the range object prefetch identifier, thereby improving the accuracy of the operation of the prefetch object. When the client only needs to prefetch part of the data of the object, it only obtains and caches part of the data of the prefetch object, without obtaining all the data of the prefetch object, thereby avoiding the waste of cache space and improving the utilization rate of cache resources.
[0012] As one possible implementation, the server retrieves the data of the first target object from the target object storage device and caches the first byte and the main body of the data in that first target object in sequence. This way, before inserting a task into the pre-fetched queue, the first byte of the pre-fetched object is cached before processing the main body of the data. This reduces the time required to send the first byte, allowing the client to quickly retrieve the first byte and improve the first byte experience.
[0013] As a possible implementation, the pre-fetched queue is a priority queue, with the first byte of uncached data in the pre-fetched queue being prioritized first, the smallest data body being prioritized second, and the most recently accessed data being prioritized third. This ensures that pre-fetched objects are more aligned with the tenant's next request, improving object access performance.
[0014] As a possible implementation, the server determines that the cache time of the first target object is greater than or equal to a preset time, and deletes the cached data of the first target object, thereby further improving the utilization of the cache space.
[0015] In a second aspect, the present application provides a data processing device, comprising a transceiver module and a processing module. The transceiver module is used to receive a first object access request sent by a client of an object storage service; the first object access request includes a first target object. The processing module is used to determine the cached data of the first target object in response to the first object access request; the data of the first target object is obtained and cached from the target object storage device based on a prefetch object identifier carried by a second object access request, and the prefetch object identifier is used to indicate that the first target object is a prefetch object. The transceiver module is also used to send the data of the first target object to the client.
[0016] As a possible implementation, the transceiver module is further configured to receive a second object access request from a client; the second object access request includes a second target object and a prefetch object identifier. The processing module is further configured to retrieve and cache data of the first target object from the target object storage device based on the prefetch object identifier. The transceiver module is further configured to send the data of the second target object to the client.
[0017] As a possible implementation, the server also includes a pre-storage queue and a communication queue. The processing module is further configured to: determine if a received object access request includes a pre-fetch object identifier, add the access request for the pre-fetch object in the received object access request to the pre-storage queue, and add the access request for the target object in the received object access request to the communication queue; and determine if the received object access request does not include a pre-fetch object identifier, add the received object access request to the communication queue.
[0018] As a possible implementation, the prefetch object identifier includes a whole object prefetch identifier and a range object prefetch identifier. The whole object prefetch identifier is used to indicate the object name of the prefetch object, and the range object prefetch identifier is used to indicate the data range of the prefetch object.
[0019] As a possible implementation manner, the processing module is specifically configured to: obtain data of the first target object from the target object storage device, and sequentially cache the first byte and the data body of the data of the first target object.
[0020] As a possible implementation, the pre-stored queue is a priority queue, the data whose first byte is not cached in the pre-stored queue has the first priority, the data with the smallest data size has the second priority, and the data with the most recent access time has the third priority.
[0021] As a possible implementation manner, the processing module is further configured to: determine that the cache time of the first target object is greater than or equal to a preset time, and delete the cached data of the first target object.
[0022] As a possible implementation manner, the data processing device may further include other modules that execute the operation steps of the data processing method of the first aspect.
[0023] Regarding the technical principles and beneficial effects of the second aspect, please refer to the relevant description of the first aspect mentioned above, and no further details will be given here.
[0024] In a third aspect, an object storage service system is provided, including a client and a server. The client is configured to execute the data processing method of any possible implementation of the first aspect, and the server is configured to send an object access request to the server and receive data of a target object from the server.
[0025] In a fourth aspect, a computing device cluster is provided, comprising at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the data processing method described in any possible implementation of the first aspect.
[0026] In a fifth aspect, a computer program product is provided, which includes a computer program or instructions, and when the computer program or instructions are executed on a computer, causes the computer to execute the data processing method described in any possible implementation of the first aspect.
[0027] In a sixth aspect, a computer-readable storage medium is provided. The computer-readable storage medium includes a computer program or instructions that, when executed on a computer, causes the computer to execute the data processing method described in any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] FIG1 is a schematic diagram of a scenario of an object storage service provided by this application;
[0029] FIG2 is a schematic diagram of the architecture of an object storage system provided by this application;
[0030] FIG3 is a flow chart of a data processing method provided by the present application;
[0031] FIG4 is a flow chart of a pre-fetch object processing step provided by the present application;
[0032] FIG5 is a flow chart of the steps of processing a pre-stored queue and a communication queue provided by the present application;
[0033] FIG6 is a schematic diagram of a pre-fetched object identifier provided by the present application;
[0034] FIG7 is a schematic diagram of another pre-fetch object identification provided by the present application;
[0035] FIG8 is a schematic diagram of a first byte and a data body provided by this application;
[0036] FIG9 is a schematic structural diagram of a data processing device provided by the present application;
[0037] FIG10 is a schematic diagram of the structure of a computing device provided by the present application;
[0038] FIG11 is a schematic diagram of the structure of a computing device cluster provided by this application;
[0039] FIG12 is a schematic diagram of a structure of a network connection between computing devices provided by the present application. DETAILED DESCRIPTION
[0040] The data processing method provided in the embodiment of the present application can be applied to object storage service scenarios in the storage field. The following is a brief introduction to the technologies that may be involved in this application.
[0041] (1) Object Storage Service (OBS): OBS is an object-based storage service that offers the advantages of massive capacity, security, high reliability, and low cost. OBS is an Internet-accessible service. Tenants can establish a connection with the object storage service layer (a computing node that supports OBS) through an object storage client, create buckets in the storage nodes managed by the object storage service layer, and then access and manage objects in the buckets. By default, the object storage service layer uses a sequential distribution method to store objects in buckets. The sequential distribution method can also be called a lexicographic distribution or a range distribution.
[0042] (2) Object: An object is the basic unit of data storage in an object storage system. An object is actually a collection of a file's data and its related attribute information (metadata). Data uploaded by tenants to OBS is stored in buckets as objects. Objects consist of three parts: a key, metadata, and data. The key, or object name, is a UTF-8-encoded character sequence with a length greater than 0 and no more than 1024. Each object in a bucket has a unique object key.
[0043] (3) Bucket: A bucket is a container for storing objects in OBS. Object storage provides a flat storage method based on buckets and objects. All objects in a bucket are at the same logical level, eliminating the multi-level tree directory structure in the file system. Each bucket has its own storage category, access rights, region, and other attributes. Tenants can create buckets with different storage categories and access rights and configure more advanced attributes to meet the storage requirements of different scenarios.
[0044] (4) Object storage device (OSD): OSD is the basic storage unit of the object storage system, which is set on the physical disk. Specifically, it is a fixed-size storage space of the physical disk. The object storage system manages the physical disks of multiple computing devices in the form of OSD.
[0045] Tenants can log in to the cloud management platform on the public cloud access page using a pre-registered account and password, and after successful login, select and purchase corresponding public cloud services on the public cloud access page, such as OBS, virtual machine services, container services, etc. For OBS, tenants can further create multiple buckets through the configuration interface or application programming interface (API) on the public cloud access page provided by the cloud platform. There is no limit on the number and total size of objects stored in each bucket, and tenants do not need to consider the scalability of data. OBS is a service based on the Representational State Transfer API (REST API) and the Hypertext Transport Protocol (HTTP) and Hypertext Transfer Protocol Secure (HTTPS). Tenants can locate bucket resources through the Uniform Resource Locator (URL), also referred to as the access domain name in this application, or simply the domain name. Among them, in OBS, the bucket name is globally unique and cannot be modified, that is, the bucket created by a tenant cannot have the same name as other buckets created by the tenant, nor can it have the same name as buckets created by other tenants.
[0046] Figure 1 shows a schematic diagram of at least one application scenario in an embodiment of the present application. As shown in Figure 1, each bucket can include multiple objects, and the objects between buckets are isolated from each other. The tenant logs in to the cloud management platform through the object storage client, selects and purchases the cloud service of the object storage service on the cloud management platform. After the purchase, the tenant can store objects based on the functions provided by the object storage service. Among them, the cloud management platform is mainly used to manage the infrastructure for running the object storage service. Exemplarily, the infrastructure for running the object storage service may include multiple data centers set up in different regions, and each data center includes multiple servers. The data center can provide basic resources for the object storage service, such as computing resources, storage resources, etc. Therefore, when purchasing and using the object storage service, the tenant mainly pays for the resources used. Specifically, the object storage service provides the domain name of the bucket, and the tenant can access the domain name of the bucket through the object storage client to upload data to the bucket or download data from the bucket, wherein the uploaded data is stored in the bucket in the form of objects.
[0047] Figure 2 shows a schematic diagram of the architecture of an object storage system in an embodiment of the present application. The object storage system includes an object storage client, an object storage service layer, and an object storage device.
[0048] The object storage client can be installed on a terminal device, such as a mobile phone, laptop, tablet, PDA, wireless terminal in a smart city, or wireless device in a smart home. Tenants log in to the cloud management platform through the client and select and purchase object storage cloud services on the cloud management platform.
[0049] The object storage service layer manages one or more object storage devices and includes one or more servers, which can be located in the service devices of the service layer. Exemplarily, the service devices connect to the storage devices via a switch. In another exemplary embodiment, the service devices connect to the terminal devices via a remote connection gateway.
[0050] Tenants log in to the object storage service layer using their accounts on the client. In the service layer, they select and purchase object storage cloud services, such as creating buckets and configuring bucket names. After the service layer detects the tenant's operation (such as creating a bucket), it can send a creation command to the object storage device. The creation command includes information such as the bucket name and domain name. The creation command is used to notify the storage device to create a bucket and save the bucket name, bucket domain name, and other information. After the bucket is created, the tenant can access the bucket domain name through the object storage client, locate the bucket on the storage device, and upload, download, and delete objects from the bucket.
[0051] The service layer includes a control unit and an operating system. The control unit runs on the operating system, which includes a disk driver and a physical network card driver. The control unit controls the disk controller through the disk driver to set the physical disk as multiple persistent storage units (persistence log, PLOG).
[0052] The infrastructure for running the object storage service can be configured with multiple data centers in different regions. For example, object storage devices are configured in multiple data centers in different regions, with each data center including multiple object storage devices. Each object storage device includes at least one physical disk. For example, the disk controller configures physical disks 11 and 12 in Region 1 as four PLOGs.
[0053] After receiving the creation instruction of bucket 1, the service layer notifies the control unit to create bucket 1. The control unit creates bucket 1 on physical disk 11 and physical disk 12 in the object storage device through the operating system. The PLOG of bucket 1 is distributed on physical disk 11 and physical disk 12, and the bucket name, bucket domain name and other information of bucket 1 are saved.
[0054] When a tenant needs to upload an object to bucket 1 through a client, the tenant can trigger the object storage device through the client to write the object to bucket 1. For example, the control unit in the object storage service layer can write the object to bucket 1.
[0055] When a tenant needs to download an object from bucket 1 through a client, the tenant can trigger the object storage device through the client to read the object from bucket 1. For example, a control unit in the service layer can obtain the object from bucket 1 and feed it back to the client.
[0056] It is worth noting that Figures 1 and 2 are merely schematic diagrams and should not be construed as limiting the present application. Any object storage server or object storage device with object storage capabilities can apply the data processing method provided herein. Furthermore, the object storage application scenario illustrated in Figure 1 and the object storage system architecture illustrated in Figure 2 may also include other devices or architectures not shown in Figures 1 or 2.
[0057] Next, the data processing method provided by this embodiment will be described in detail with reference to the accompanying drawings.
[0058] The steps of the data processing method provided by the present application are executed by any client of the object storage client and any server of the object storage service layer shown in Figure 2. Next, the data processing method provided by the present application is described in conjunction with Figure 3.
[0059] Step 301: The client sends a first object access request to the server.
[0060] The client sends a first object access request to the server in response to the tenant's operation.
[0061] As a possible implementation manner, the first object access request includes a first target object and a data operation.
[0062] Optionally, the first target object in the first object access request may be a unique identification name (identity) of the first target object. The data operation may be data reading, data writing, and the like.
[0063] In a possible embodiment, the client may call an application programming interface (API) corresponding to functions such as upload, download, and delete to implement the interaction of object access requests.
[0064] Step 302: The server receives a first object access request sent by the client.
[0065] Step 303: The server determines that the data of the first target object has been cached.
[0066] The server determines that a pre-fetch request for the data of the first target object has been executed before receiving the first object access request, and the data of the first target object has been cached in the pre-storage queue.
[0067] As a possible implementation, the data of the first target object is acquired and cached based on the pre-fetched object identifier of the first target object carried in the second object access request, wherein the server receives the second object access request earlier than the server receives the first object access request.
[0068] For specific steps of the server acquiring and caching the data of the first target object according to the pre-fetched object identifier of the first target object carried in the second object access request, please refer to steps 401 to 403 shown in FIG. 4 , which will not be repeated here.
[0069] Step 304: The server sends the data of the first target object to the client of the object storage service.
[0070] The server sends the cached data of the first target object to the client in the form of a first object access response, where the first object access response includes the data of the first target object.
[0071] Step 305: The client receives data of the first target object.
[0072] The client receives data of the first target object sent by the server in the form of a first object access response.
[0073] Based on the above data processing method, the server determines that the first target object is the prefetched object that the client will next access based on the second object access request that precedes the first object access request. This allows the server to complete the acquisition and caching of the first target object before receiving the first object access request from the client. Upon receiving the first object access request, the cached data of the first target object is sent to the client. This improves resource utilization during idle time on the server, increases the speed at which clients retrieve objects, and reduces latency in object storage.
[0074] The above, combined with Figure 3, illustrates the overall process by which a server sends cached prefetched object data to a client. Prior to this, the server must acquire and cache the prefetched object data, so that upon receiving a first object access request, it can send the cached second target object to the client that sent the first object access request. Next, combined with Figure 4, the server's steps for acquiring and caching prefetched object data are explained.
[0075] Step 401: The client sends a second object access request to the server.
[0076] The client sends a second object access request to the server in response to the tenant's operation.
[0077] As a possible implementation manner, the second object access request includes the second target object and a pre-fetched object identifier.
[0078] Optionally, the second target object in the second object access request may be a unique identification name (identification) of the second target object. The second object access request may also include a data operation, which may be data reading, data writing, and the like.
[0079] Optionally, the prefetch object identifier is used to indicate the data of the prefetch object. For example, if the data operation corresponding to the second object access request and the operation corresponding to the first object access request in Figure 3 are multiple operations in a long connection, and the data of the first target object is prefetched data after the tenant obtains the data of the second target object, then the prefetch object identifier is used to indicate that the data of the first target object is the data of the prefetch object.
[0080] In a possible embodiment of the present application, the first object access request and the second object access request may be messages. The prefetch object identifier may be carried in a message header, referred to as a prefetch object header field. The prefetch object identifier may also be carried in the message body. The specific form of the prefetch object identifier can be found in FIG6 and the related description below and will not be repeated here.
[0081] Step 402: The server receives a second object access request.
[0082] The server receives the second object access request and extracts the pre-fetched object identifier from the second object access request.
[0083] Step 403: The server obtains and caches the data of the first target object according to the pre-fetched object identifier.
[0084] The server obtains and caches data of the first target object from the object storage device storing the first target object according to the pre-fetched object identifier.
[0085] As a possible implementation, after the server extracts the prefetch object identifier from the second object access request, the second object access request without the prefetch object identifier is an access request for the target object. The prefetch object identifier can be regarded as an access request for the prefetch object, also known as a prefetch request.
[0086] In a possible embodiment of the present application, access requests for target objects and access requests for prefetched objects can be categorized and added to a pre-storage queue and a communication queue. The pre-storage queue is used to maintain pre-fetched requests, i.e., access requests for pre-fetched objects, while the communication queue is used to maintain access requests for target objects, i.e., access requests for objects that do not carry a pre-fetched object identifier. The specific implementation of the pre-storage queue and the communication queue is shown in Figure 5 and the related description below, and will not be repeated here.
[0087] Step 404: The server sends the data of the second target object to the client.
[0088] The server obtains the data of the second target object from the object storage device storing the data of the second target object, and sends the data of the second target object to the client.
[0089] Based on steps 401-404 above, the client sends an access request with a Prefetch Object header field to the server, notifying the server of the object it intends to access in its next request. This allows the server to asynchronously retrieve and cache the prefetched object in advance using a pre-storage queue. This provides a client-server collaborative method for prefetching and caching prefetched objects, thereby reducing object access latency.
[0090] The above, combined with Figure 4, illustrates how the server obtains and caches data for the first target object that the client may subsequently read. Based on this, the process for the server to obtain and cache data for the first target object can be optimized using a pre-stored queue and a communication queue. Next, combined with Figure 5, the data processing steps based on the pre-stored queue and the communication queue are explained.
[0091] Step 501: The client sends a second object access request to the server.
[0092] As a possible implementation manner, the specific manner in which the client sends the second object access request to the server is shown in step 401 of FIG. 4 , and will not be described in detail here.
[0093] Step 502: The server receives a second object access request.
[0094] As a possible implementation manner, the specific manner in which the client sends the second object access request to the server is shown in step 402 of FIG. 4 and will not be described in detail here.
[0095] Step 503: The server splits the second object access request into an access request for the pre-fetched object and an access request for the target object.
[0096] The server extracts the prefetch object identifier from the second object access request as an access request to the prefetch object, and uses the second object access request without the prefetch object identifier after extracting the prefetch object identifier as an access request to the target object.
[0097] Step 504: The server adds the access request for the pre-fetched object to the pre-stored queue.
[0098] As a possible implementation, a pre-storage queue, also known as a PreSave Queue, is used to maintain pre-fetch requests.
[0099] Step 505: The server executes an access request to the pre-fetched object according to the pre-stored queue, and obtains and caches the data of the first target object.
[0100] As a possible implementation method, the server determines that the pre-fetched object is the first target object based on the access request for the pre-fetched object, obtains the data of the first target object from the object storage device storing the data of the first target object, and caches the data of the first target object to the pre-storage queue.
[0101] As a possible implementation, the pre-fetch queue is a priority queue. Data with the first uncached byte in the pre-fetch queue is prioritized first, data with the smallest body size is prioritized second, and data with the most recent access time is prioritized third. The server asynchronously caches pre-fetched objects based on the priority of the pre-fetch queue, ensuring that the pre-fetched objects are more suitable for the tenant's next request and improving object access performance.
[0102] As a possible implementation method, for the data of objects cached in the pre-stored queue, the server can be configured to periodically scan the data of pre-fetched objects downloaded in the pipeline of the pre-stored queue according to a preset period. If the data of the pre-fetched object has exceeded the preset time, the data of the pre-fetched object is deleted.
[0103] In a possible embodiment, if the data of the pre-fetched object in the pre-stored queue has been released when the server receives the object access request from the client, the server re-acquires the data of the pre-fetched object and sends the data of the pre-fetched object to the client.
[0104] Step 506: The server adds the access request to the target object to the communication queue.
[0105] As a possible implementation method, the communication queue, also known as the Connect Queue, and the pre-stored queue can be two different threads.
[0106] Step 507: The server executes the access request to the target object according to the communication queue, and obtains the data of the second target object.
[0107] As a possible implementation manner, the server determines that the target object is the second target object according to the access request to the target object, and obtains the data of the second target object from the object storage device storing the data of the second target object.
[0108] Step 508: The server sends the data of the second target object to the client.
[0109] As a possible implementation manner, the specific manner in which the server sends the data of the second target object to the client is the same as the manner in which the server sends the data of the first target object to the client in step 304 shown in FIG3 , and is not repeated here.
[0110] Step 509: The client receives data of the second target object.
[0111] As a possible implementation manner, the specific manner in which the client receives the data of the second target object is the same as the manner in which the client receives the data of the first target object in step 305 shown in FIG3 , and is not described again here.
[0112] Based on steps 501-509, since the communication queue and the pre-storage queue can be two different threads, the process of retrieving and caching the pre-fetched object data in the pre-storage queue is an asynchronous background task that does not block the communication queue thread, thereby ensuring normal object data access and the performance of the object storage system. Furthermore, the server can retrieve and cache the pre-fetched object data according to the pre-storage queue during idle periods (e.g., when the communication queue is empty), improving resource utilization on the server during idle periods and increasing the speed at which clients can retrieve large quantities of objects.
[0113] In each step shown in FIG3-FIG5 above, the server needs to identify, obtain or cache the prefetched object based on the prefetched object identifier carried in the object access request. The specific form of the prefetched object identifier is explained below in conjunction with FIG6 and FIG7.
[0114] As shown in Figure 6, the object access request may be a message, including a header and a body. The pre-fetched object identifier may be carried in the header of the object access request.
[0115] As shown in Figure 7, the object access request may be a message, including a header and a body. The pre-fetched object identifier may also be carried in the body of the object access request.
[0116] For example, the object access request header carries x-obs-next-payload:streaming-next-key-payload, indicating that the prefetched object identifier is located in the body of the object access request. The body data carries x-obs-next-key as a segmentation identifier at the end. The portion before the segmentation identifier is the object data, and the portion after the segmentation identifier is the prefetched object identifier.
[0117] As a possible implementation manner, the prefetch object identifier includes an entire object prefetch identifier (X-obs-next-key) and a range object prefetch identifier (X-obs-next-range).
[0118] The whole object prefetch identifier is used to indicate the object name of the prefetch object. For example, the whole object prefetch identifier is X-obs-next-key:objectname1 / X-End, indicating that the object name of the prefetch object is objectname1. In this embodiment, objectname1 may be the object name of the first target object.
[0119] The range object prefetch flag indicates the data range of the prefetch object. For example, if the range object prefetch flag is X-obs-next-range:bytes 0-5000 / X-End, it means that the prefetched data (the data to be obtained and cached) corresponds to the data range of 0-5000 bytes of the prefetch object.
[0120] On the basis of the data processing method shown in Figures 3-5 provided in this application, in order to improve the user's first byte experience, as shown in Figure 8, when the task in the pre-stored queue is executed, the server distinguishes between the first byte of the object (after entering the pre-stored queue, the metadata of the object is obtained in a first-in-first-out order) and the data body. Before inserting the task into the pre-stored queue, the first byte of the pre-fetched object will be cached first, and then the data body will be processed to reduce the time required to send the first byte, so that the client can quickly obtain the first byte and improve the first byte experience.
[0121] As a possible implementation manner, the first byte may be the first byte of the data of the object, or may be the first batch of bytes (such as 512 bytes) of the data of the object.
[0122] To support the data processing method shown in Figures 3-5 above, this application also provides a data processing device 900, which can be used to implement the server-side functions in the data processing method shown in Figures 3-5 above. As shown in Figure 9, the data processing device 900 includes a transceiver module 910 and a processing module 920.
[0123] The transceiver module 910 is configured to receive a first object access request sent by a client of the object storage service; the first object access request includes a first target object.
[0124] Processing module 920 is used to determine the cached data of the first target object in response to the first object access request; the data of the first target object is obtained and cached from the target object storage device according to the prefetch object identifier carried by the second object access request, and the prefetch object identifier is used to indicate that the first target object is a prefetch object.
[0125] The transceiver module 910 is further configured to send the data of the first target object to the client.
[0126] As a possible implementation, the transceiver module 910 is further configured to receive a second object access request from a client; the second object access request includes a second target object and a prefetch object identifier. The processing module 920 is further configured to retrieve and cache data of the first target object from the target object storage device based on the prefetch object identifier. The transceiver module 910 is further configured to send the data of the second target object to the client.
[0127] As a possible implementation, the server also includes a pre-storage queue and a communication queue. Processing module 920 is further configured to: determine if a received object access request includes a pre-fetch object identifier, add the access request for the pre-fetch object in the received object access request to the pre-storage queue, and add the access request for the target object in the received object access request to the communication queue; and determine if the received object access request does not include a pre-fetch object identifier, add the received object access request to the communication queue.
[0128] As a possible implementation, the prefetch object identifier includes a whole object prefetch identifier and a range object prefetch identifier. The whole object prefetch identifier is used to indicate the object name of the prefetch object, and the range object prefetch identifier is used to indicate the data range of the prefetch object.
[0129] As a possible implementation manner, the processing module 920 is specifically configured to obtain data of the first target object from the first target bucket, and sequentially cache the first byte and the data body of the data of the first target object.
[0130] As a possible implementation, the pre-stored queue is a priority queue, the data whose first byte is not cached in the pre-stored queue has the first priority, the data with the smallest data size has the second priority, and the data with the most recent access time has the third priority.
[0131] As a possible implementation manner, the processing module 920 is specifically configured to: determine that the cache time of the first target object is greater than or equal to a preset time, and delete the cached data of the first target object.
[0132] The transceiver module 910 and the processing module 920 can be implemented in software or hardware. For example, the implementation of the transceiver module 910 will be described below using the transceiver module 910 as an example. Similarly, the implementation of the processing module 920 can refer to the implementation of the transceiver module 910.
[0133] As an example of a software functional unit, the transceiver module 910 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Furthermore, the computing instance may be one or more. For example, the transceiver module 910 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple geographically close data centers. Typically, a region may include multiple AZs.
[0134] Similarly, multiple hosts / virtual machines / containers running the code can be distributed within the same virtual private cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Cross-region communication between two VPCs within the same region, or between VPCs in different regions, requires a communication gateway within each VPC to interconnect the VPCs.
[0135] As an example of a hardware functional unit, the transceiver module 910 may include at least one computing device, such as a server. Alternatively, the transceiver module 910 may be implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0136] The multiple computing devices included in the transceiver module 910 can be distributed in the same region or in different regions. The multiple computing devices included in the transceiver module 910 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the transceiver module 910 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, GALs, and other computing devices.
[0137] It should be noted that, in other embodiments, any module in the transceiver module 910 or the processing module 920 can be used to execute any step in the data processing method, and the steps that the transceiver module 910 and the processing module 920 are responsible for implementing can be specified as needed. The full functions of the data processing device 900 are realized by implementing different steps in the data processing method through the transceiver module 910 and the processing module 920 respectively.
[0138] This application also provides a computing device 1000. As shown in Figure 10, computing device 1000 includes a bus 1002, a processor 1004, a memory 1006, and a communication interface 1008. Processor 1004, memory 1006, and communication interface 1008 communicate with each other via bus 1002. Computing device 1000 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in computing device 1000.
[0139] Bus 1002 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG10 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 1002 may include a path for transmitting information between various components of computing device 1000 (e.g., memory 1006, processor 1004, and communication interface 1008).
[0140] The processor 1004 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0141] The memory 1006 may include volatile memory, such as random access memory (RAM). The processor 1004 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0142] The memory 1006 stores executable program codes, and the processor 1004 executes the executable program codes to respectively implement the functions of the various modules included in the aforementioned data processing device 900, thereby implementing the data processing method. In other words, the memory 1006 stores instructions for executing the data processing method.
[0143] Alternatively, the memory 1006 stores executable code, and the processor 1004 executes the executable code to implement the aforementioned client or server functions, thereby implementing the data processing method. In other words, the memory 1006 stores instructions for executing the data processing method.
[0144] The communication interface 1008 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1000 and other devices or a communication network.
[0145] Considering that the data processing method provided in this application is applied to an object storage system, which may include multiple clients or multiple servers, this application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0146] As shown in Figure 11, the computing device cluster includes at least one computing device 1000. The memory 1006 in one or more computing devices 1000 in the computing device cluster may store the same instructions for executing the data processing method.
[0147] In some possible implementations, the memory 1006 of one or more computing devices 1000 in the computing device cluster may also store some instructions for executing the data processing method. In other words, the combination of one or more computing devices 1000 can jointly execute the instructions for executing the data processing method.
[0148] It should be noted that the memory 1006 in different computing devices 1000 in the computing device cluster may store different instructions, each for executing part of the functions of the data processing apparatus 1000. In other words, the instructions stored in the memory 1006 in different computing devices 1000 may implement the functions of one or more modules included in the data processing apparatus 1000.
[0149] In some possible implementations, one or more computing devices in a computing device cluster may be connected via a network. The network may be a wide area network (WAN) or a local area network (LAN), etc. FIG12 illustrates a possible implementation. As shown in FIG12 , two computing devices 1000A and 1000B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1006 in the computing device 1000A stores instructions for executing the functions of one or more of the transceiver module 910 and the processing module 920. FIG12 takes the example of the memory 1006 in the computing device 1000A storing instructions for executing the functions of the transceiver module 910. At the same time, the memory 1006 in the computing device 1000B stores instructions for executing the functions of one or more of the transceiver module 910 and the processing module 920. FIG12 takes the example of the memory 1006 in the computing device 1000B storing instructions for executing the functions of the processing module 920.
[0150] It should be understood that the functionality of the computing device 1000A shown in FIG12 may also be accomplished by multiple computing devices 1000. Similarly, the functionality of the computing device 1000B may also be accomplished by multiple computing devices 1000.
[0151] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes the data processing method shown in FIG3 , or the steps executed by the client or server in the data processing method shown in FIG3 .
[0152] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform the data processing method shown in Figure 3, or the steps performed by the client or server in the data processing method shown in Figure 3.
[0153] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (such as infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0154] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0155] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0156] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0157] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0158] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0159] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk, or an optical disk.
[0160] In this application, "at least one" means one or more, and "more" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b and c can be single or multiple.
[0161] It should be noted that, in this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0162] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A data processing method, characterized in that: A server end applied to an object storage service, the server end is used to access at least one object storage device managed by a cloud platform, the at least one object storage device is located in multiple data centers in different regions, each data center includes multiple servers, and the method includes: Receive a first object access request sent by a client of the object storage service; the first object access request includes a first target object; In response to the first object access request, determine the cached data of the first target object; the data of the first target object is obtained and cached from the target object storage device according to the pre-fetch object identifier carried in the second object access request, and the pre-fetch object identifier is used to indicate that the first target object is a pre-fetch object; Sending data of the first target object to the client.
2. The method according to claim 1, characterized in that Before receiving the first object access request sent by the client of the object storage service, the method further includes: receiving the second object access request sent by the client; the second object access request includes a second target object and the pre-fetched object identifier; Acquire and cache the data of the first target object from the target object storage device according to the pre-fetch object identifier; The data of the second target object is sent to the client.
3. The method according to claim 1 or 2, characterized in that: The server also includes a pre-stored queue and a communication queue, and the method also includes: Determining that the received object access request includes a pre-fetch object identifier, adding an access request for the pre-fetch object in the received object access request to the pre-storage queue, and adding an access request for the target object in the received object access request to the communication queue; It is determined that the received object access request does not include a pre-fetch object identifier, and the received object access request is added to the communication queue.
4. The method according to any one of claims 1 to 3, characterized in that The pre-fetch object identifier includes an entire object pre-fetch identifier and a range object pre-fetch identifier. The entire object pre-fetch identifier is used to indicate the object name of the pre-fetch object, and the range object pre-fetch identifier is used to indicate the data range of the pre-fetch object.
5. The method according to claim 3 or 4, characterized in that: The acquiring and caching the data of the first target object from the target object storage device according to the pre-fetch object identifier includes: The data of the first target object is obtained from the target object storage device, and the first byte and the data body of the data of the first target object are cached in sequence.
6. The method according to claim 5, characterized in that The pre-stored queue is a priority queue, the data whose first byte is not cached in the pre-stored queue has a first priority, the data whose data body has the smallest data size has a second priority, and the data with the most recent access time has a third priority.
7. The method according to any one of claims 1 to 6, characterized in that After sending the data of the first target object to the client, the method further includes: It is determined that the cache time of the first target object is greater than or equal to a preset time, and the cached data of the first target object is deleted.
8. A data processing device, characterized in that: The device is applied to a server of an object storage service, the server is used to access at least one object storage device managed by a cloud platform, the at least one object storage device is located in multiple data centers in different regions, each data center includes multiple servers, and the device includes: A transceiver module, configured to receive a first object access request sent by a client of an object storage service; the first object access request includes a first target object; a processing module, configured to determine cached data of the first target object in response to the first object access request; the data of the first target object is obtained and cached from a target object storage device according to a pre-fetch object identifier carried in the second object access request, the pre-fetch object identifier being used to indicate that the first target object is a pre-fetch object; The transceiver module is further used to send the data of the first target object to the client.
9. An object storage service system, characterized in that: include: The client is used to send a second object access request to the server of the object storage service; the object access request includes a second target object and a pre-fetch object identifier, and the pre-fetch object identifier is used to indicate that the first target object is a pre-fetch object, so as to instruct the server to The pre-fetch object identifier obtains and caches the data of the first target object from the target storage device, and sends the data of the second target object to the client; Also used for sending a first object access request to the server, and receiving data of the first target sent by the client; A server, used for receiving a first object access request sent by the client; the first object access request includes a first target object; in response to the first object access request, determining cached data of the first target object; the data of the first target object is obtained and cached from a target object storage device according to a prefetch object identifier carried in a second object access request, and the prefetch object identifier is used to indicate that the first target object is a prefetch object; sending the data of the first target object to the client;.
10. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 7.
11. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster is caused to perform the method according to any one of claims 1 to 7.
12. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data prefetching method and device
CN104063330A
Object access method and device
CN109871181A
Ceph-based network target range rear-end storage system design method
CN110750334A
Data consistency storage method and system for object storage device
CN111124301A
Metadata prefetching system and method for distributed file system
CN113688113A