Method, system and equipment for optimizing object enumeration performance in object storage based on large model and medium

By training access pattern recognition large models to predict hotspot objects and optimize query strategies, the performance bottleneck of traditional object storage systems under high concurrency and massive data is solved, efficient execution of object listing operations is achieved, and system performance and resource utilization are improved.

CN120256465APending Publication Date: 2025-07-04SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375546.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional object storage systems have problems such as performance bottlenecks, high index maintenance costs, low query efficiency and inability to adapt to rapidly changing data access patterns when performing object enumeration operations, which are particularly obvious in high concurrency scenarios.

Method used

By collecting access logs and object metadata, training access pattern recognition large models, predicting the hot objects that users need to access, and dynamically adjusting the query strategy and cache methods, loading hot objects into cache or memory in advance, and optimizing query task allocation.

Benefits of technology

Effectively reduce query latency, improve response speed and throughput, optimize resource utilization, reduce storage costs, and enhance user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256465A_ABST
    Figure CN120256465A_ABST
Patent Text Reader

Abstract

The invention provides a method, a system, equipment and a medium for optimizing object enumeration performance in object storage based on a large model, and belongs to the technical field of data storage, the method comprises the following steps: collecting access logs and object metadata in a storage system to construct a data set, and using the data set to train an access mode to identify the large model; responding to the object enumeration request, analyzing the context of the object enumeration request by using an access mode identification large model, and predicting a hot object needing to be accessed by a user in advance; dynamically adjusting a query strategy and an access mode of a storage node based on a prediction result, caching and pre-loading a hotspot object in advance, and then executing an object enumeration operation; and collecting an enumeration result and performance data of the object enumeration operation, and optimizing and updating the access mode recognition large model in combination with the new access log. In the object enumeration operation, the query delay can be effectively reduced, the response speed and throughput can be improved, the resource utilization rate can be optimized, the storage cost can be reduced, and the user experience can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of data storage, and particularly relates to a method, system, device, and medium for optimizing object listing performance in object storage based on a large model. Background Art

[0002] With the development of cloud computing, big data, and Internet of Things technologies, the demand for the storage and management of massive data is increasing. As a flexible and highly scalable data storage solution, the object storage system has become a widely used storage system in Internet services, enterprise-level storage, and data backup. A typical object storage system uses objects as the basic storage unit to store large-scale unstructured data.

[0003] However, when faced with massive data storage, traditional object storage systems will encounter performance bottlenecks during object listing operations. The object listing operation is to list the list of objects that meet the conditions in the storage system according to specific conditions, such as prefixes, tags, etc. In a storage system, object listing operations usually involve the following two types of query operations: full listing and conditional listing. Full listing means traversing all objects, while conditional listing requires filtering data according to conditions specified by users, such as prefixes, timestamps, or tags. In the case of extremely large amounts of data, traditional object listing operations have the following problems: First, performance is limited. Since the number of objects in the object storage system is huge, traditional listing methods need to traverse the data set, which not only increases query latency but also increases the load on the storage system, affecting the processing efficiency of concurrent requests. Second, the commonly used index-based object lookup method in object storage systems leads to a significant decrease in the maintenance cost and query efficiency of indexes when the number of stored objects is large. Third, in high-concurrency usage scenarios, multiple requests may simultaneously require listing a large number of objects, and traditional listing methods cannot handle this, resulting in a decline in the performance of the storage system. Fourth, the data access patterns of storage systems change rapidly, and traditional optimization methods are usually based on fixed rules and static indexes and cannot be adjusted in real time to adapt to these changes. Summary of the Invention

[0004] In a first aspect, an embodiment of this application provides a method for optimizing object listing performance in object storage based on a large model, including the following steps: S1. Collect access logs and object metadata in the storage system to construct a data set, and use the data set to train an access pattern recognition large model; S2. In response to an object listing request, use the access pattern recognition large model to analyze the context of the object listing request and predict in advance the hot objects that the user needs to access; S3. Dynamically adjust the query strategy and the access method of storage nodes based on the prediction results, and cache and preload hot objects in advance, and then perform the object listing operation; S4. Collect the enumeration results and performance data of the object enumeration operation, and optimize and update the access pattern recognition large model in combination with the new access logs.

[0005] Further, the access logs include the access frequency, access time, visitor identity, and request type of the object; the object metadata includes the object ID, label, size, and creation time. The specific steps for constructing the dataset are as follows: Analyze the access logs to obtain access behavior data; the access behavior data includes access frequency, time distribution, and access relationships. Clean and process the access logs and access behavior data to obtain a structured dataset.

[0006] Further, the specific steps for training the access pattern recognition model using the dataset are as follows: Select a base model to construct the access pattern recognition model, use the structured dataset as input, and use predicting hot data and cold data within a set future time period as the training objective, and start model training. During the training process, train the model using a combination of supervised learning and unsupervised learning, and optimize it through cross-validation and reinforcement learning until the model performance meets the requirements to obtain the access pattern recognition model.

[0007] Further, the specific steps of step S2 are as follows: S21. When an object enumeration request is detected, the storage system collects the context of the object enumeration request and inputs it into the access pattern recognition large model; the context of the object enumeration request includes request parameters and user behavior history. S22. The access pattern recognition large model recognizes the user's historical access pattern, combines object relevance and labels, and predicts in advance the hot objects that need to be queried currently.

[0008] Further, the specific steps of step S3 are as follows: S31. Determine the storage node to which the hot object belongs as the routing object and set it as the target storage node. S32. Determine the target storage node as the assignment node for the query task. S33. Pre-load the hot object into the cache or memory in advance, and pre-load the hot object into different target storage nodes in advance. S34. Execute the object enumeration operation, split the object enumeration operation into several query tasks, and assign each query task to the target storage node. S35. Each target storage node executes the assigned query task and queries the required objects in turn according to the order of the cache or memory, the hot objects pre-loaded by itself, and the original query strategy. S36. Merge the query results of the query tasks of each target storage node and return them to the user to complete the object listing operation.

[0009] Further, the specific steps of step S35 are as follows: S351. Each target storage node starts and executes the assigned query task; S352. The query task first looks for the required object from the hot objects in the cache or memory; If the requirement is met, go to step S36; If the requirement is not met, go to step S353; S353. The query task takes its own target storage node as the routing target and queries the required object from its own hot objects or the hot objects preloaded by itself; If the requirement is met, go to step S36; If the requirement is not met, go to step S354; S354. The query task executes the query of the required object according to the original query strategy until the requirement is met.

[0010] Further, the specific steps of step S4 are as follows: S41. Collect the listing result and performance data after the object listing operation is completed; the listing result includes listing success or failure; the performance data includes the listing response time; S42. Analyze the listing result to obtain the listing success rate; S43. Optimize the access pattern recognition large model according to the listing success rate and performance data, and adjust the query strategy and cache strategy; S44. Regularly collect new access logs in the storage system and add them to the dataset to retrain the access pattern recognition large model to complete model update.

[0011] In a second aspect, an embodiment of the present application further provides a system for optimizing the object listing performance in object storage based on a large model, including: A data collection and model training module, configured to collect access logs and object metadata in the storage system to construct a dataset, and use the dataset to train an access pattern recognition large model; A hot object prediction module, configured to respond to an object listing request, analyze the context of the object listing request using the access pattern recognition large model, and predict in advance the hot objects that the user needs to access; An object listing optimization module, configured to dynamically adjust the query strategy and the access method of the storage node based on the prediction result, and cache and preload the hot objects in advance, and then perform the object listing operation; A model optimization and update module, which is used to collect the enumeration results and performance data of object enumeration operations, and optimize and update the access pattern recognition large model by combining new access logs.

[0012] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for optimizing the object enumeration performance in object storage based on a large model as described in the first aspect are implemented.

[0013] In a fourth aspect, an embodiment of the present application further provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method for optimizing the object enumeration performance in object storage based on a large model as described in the first aspect are implemented.

[0014] From the above technical solutions, it can be seen that the present application has the following advantages: In the method, system, device, and medium for optimizing the object enumeration performance in object storage based on a large model provided by the present application, by preloading hot objects into the cache or memory, the number of disk I / Os is reduced, and the response time is optimized from the second level to the millisecond level; by the dynamic routing strategy, the cross-node query overhead is reduced, saving bandwidth and computing resources, and it is applicable to the EB-level massive data scenario. The present invention improves the performance of object enumeration operations in the object storage system. Especially in the high-concurrency and massive data environment, it can effectively reduce the query latency, improve the response speed and throughput of the system. At the same time, through the intelligent cache and preloading mechanism, the resource utilization rate is optimized, the storage cost is reduced, and the user experience is enhanced. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions of the present application, the drawings required to be used in the description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic flowchart of the method for optimizing the object enumeration performance in object storage based on a large model of the present invention.

[0017] Figure 2 It is a schematic flowchart of the system for optimizing the object enumeration performance in object storage based on a large model of the present invention. Detailed Embodiments

[0018] In the specific steps of the method for optimizing the object enumeration performance in the object storage based on the large model, which will be described in detail below, various embodiments of the present disclosure will be described more comprehensively. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but rather the present disclosure should be understood to cover all adjustments, equivalents and / or alternatives that fall within the spirit and scope of the various embodiments of the present disclosure.

[0019] For example, with the rapid development of cloud computing, big data and Internet of Things technologies, the demand for storage and management of massive data has shown explosive growth. Object storage systems have become a widely used data storage solution in the fields of Internet services, enterprise storage, and data backup due to their flexible and powerful scalability. In these systems, objects, as the basic storage unit, carry the task of storing large-scale unstructured data.

[0020] However, when processing massive amounts of data, traditional object storage systems face many challenges when performing object enumeration operations. Object enumeration operations are designed to filter out a list of objects that meet the conditions from the storage system based on specific conditions, such as prefixes, tags, etc. Specifically, object enumeration operations mainly include two types: full enumeration and conditional enumeration. Full enumeration requires checking all objects in the storage system one by one, while conditional enumeration accurately filters data based on user-set conditions, such as prefixes, timestamps, or tags.

[0021] When the amount of data is extremely large, traditional object enumeration operations expose a series of problems. First, performance bottlenecks are prominent. Due to the extremely large number of objects stored in object storage systems, traditional enumeration methods need to traverse the entire data set, which not only significantly increases query latency, but also greatly increases the load on the storage system, thereby affecting the processing efficiency of other concurrent requests. Second, there are issues with index maintenance costs and query efficiency. Object storage systems usually rely on indexes to speed up object lookups, but when the number of stored objects is extremely large, the maintenance cost of indexes rises sharply and query efficiency also drops significantly. In addition, in high-concurrency usage scenarios, multiple requests may require the enumeration of a large number of objects at the same time, and traditional enumeration methods often cannot effectively cope with this, resulting in a significant drop in storage system performance. Finally, the data access patterns of storage systems change rapidly. Traditional optimization methods are usually based on fixed rules and static indexes, which are difficult to adjust in real time to adapt to these rapidly changing access patterns.

[0022] To address the above problems, this embodiment provides a method for optimizing the object listing performance in object storage based on a large model, which improves the performance of object listing operations in the object storage system. Especially in a high-concurrency and massive data environment, it can effectively reduce query latency, improve the system's response speed and throughput. At the same time, through an intelligent caching and preloading mechanism, it optimizes resource utilization, reduces storage costs, and enhances the user experience.

[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Please refer to Figure 1 The following is a flowchart of a method for optimizing the object listing performance in object storage based on a large model in a specific embodiment. The method includes the following steps: S1. Collect access logs and object metadata in the storage system to build a dataset, and use the dataset to train a large model for access pattern recognition; It should be noted that by collecting detailed access logs and object metadata, a training dataset is constructed to provide data input for model training and improve the prediction accuracy of the model. By cleaning and processing the data, invalid or incorrect data is removed to ensure the quality of the dataset and improve the efficiency and effect of model training; S2. In response to an object listing request, use the large model for access pattern recognition to analyze the context of the object listing request and predict in advance the hot objects that the user needs to access; It should be noted that by analyzing the request context, the object listing request of the user is responded to in real time, providing personalized prediction results to improve the user experience. Combining the user's historical access patterns and object relevance, the hot objects that the user may access are predicted in advance to improve the accuracy of the prediction; S3. Dynamically adjust the query strategy and the access method of the storage node based on the prediction results, and cache and preload the hot objects in advance, and then perform the object listing operation; It should be noted that by determining the storage node to which the hot object belongs, the query strategy is dynamically adjusted, and the query task is assigned to the target storage node to improve the query efficiency. By preloading the hot object into the cache or memory in advance and preloading it into the target storage node, the query latency is reduced and the response speed is improved. By executing the query task in steps and preferentially querying from the cached and preloaded hot objects, unnecessary data retrieval is reduced and the query efficiency is improved; S4. Collect the enumeration results and performance data of the object enumeration operation, and optimize and update the access pattern recognition large model by combining the new access logs; It should be noted that by collecting the enumeration results and performance data, combining the new access logs, continuously optimizing the access pattern recognition large model to ensure the performance of the model; according to the enumeration success rate and performance data, dynamically adjusting the query strategy and cache strategy to improve the performance rate of the storage system.

[0025] In this embodiment, the large model is trained by collecting access logs and object metadata to realize the acquisition of user access patterns; the trained model is used to predict hot objects to provide a basis for optimization operations; the query strategy, caching, and preloading are adjusted based on the prediction results to greatly improve the query efficiency; the model is optimized and updated by collecting the enumeration results and new logs, enabling the system to continuously adapt to changing access patterns.

[0026] Furthermore, as a refinement and extension of the specific implementation manner of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, another method for optimizing the object enumeration performance in object storage based on a large model is provided. This method includes the following steps: S1. Collect the access logs and object metadata in the storage system to build a dataset, and use the dataset to train the access pattern recognition large model; the access logs include the access frequency, access time, visitor identity, and request type of the object; the object metadata includes the object ID, label, size, and creation time; Taking the object storage system of an e-commerce enterprise as an example, continuously collect access logs. For example, a certain user frequently accesses pictures of women's clothing products between 3 pm and 5 pm every day in the past week, with an average daily access frequency of 20 times, and the request type is mostly to view details; at the same time, collect object metadata, such as the object ID of the product picture is "P001", the label is "Women's Summer Dress", the size is 2MB, and the creation time is "May 1, 2024"; The specific steps for building the dataset are as follows: Analyze the access logs to obtain access behavior data; the access behavior data includes access frequency, time distribution, and access relationship; It should be noted that the access relationship is obtained by constructing an object co-occurrence graph, and the edge weight in the graph represents the frequency of objects being called by the same access sequence; Clean and process the access logs and access behavior data to obtain a structured dataset; Exemplarily, conduct in-depth analysis on the collected access logs, count the access frequencies, time distributions of different commodity objects, and the access relationships between different users accessing the same commodity or related commodities; for example, if it is found that users who have purchased "summer dresses for women" often access "sandals for women" subsequently; clean the access logs and access behavior data, remove incorrect or incomplete data, and process them in a specific format to form a structured data set; Convert the structured data set into a vector form, denoted as X = [x1, x2, …, x n , where x i represents the i th input feature vector; The specific steps for training an access pattern recognition model using the data set are as follows: Select a basic model to construct an access pattern recognition model, use the structured data set as the input, and use predicting hot data and cold data within a set future time period as the training objective, and start model training; Exemplarily, select Transformer as the basic model to construct an access pattern recognition model, use the structured data set as the input, and use predicting which commodity data will be hot data (frequently accessed) and which will be cold data (rarely accessed) within the next week as the training objective, and start model training; During the training process, adopt supervised learning to optimize the model prediction accuracy using the labeled historical access results; adopt unsupervised learning to mine potential access patterns in the data; Specifically, select the Transformer model as the basic model to construct an access pattern recognition model, and use the multi-head attention mechanism of Transformer to calculate the query parameter , key parameter and value parameter respectively for each input feature vector i of each head;

[0027] Among them, is the query parameter a learnable weight matrix, is the key parameter a learnable weight matrix, is the value parameter a learnable weight matrix; Calculate the attention score between the query parameter and the key parameter through dot product to obtain the attention matrix ;

[0028] Among them, dk is the key parameter In terms of dimensions, the softmax function can normalize the scores; Calculate the attention matrix Multiply with the value parameter to obtain the output of the self-attention mechanism ;

[0029] Combine Simplify to ; Concatenate the input feature vectors i for each head and obtain the final multi-head attention output Z through a linear transformation;

[0030] where is the learnable weight matrix; Process the multi-head attention output Z using a feed-forward neural network:

[0031] where and are the two weight matrices of the feed-forward neural network, and are the bias vectors of the feed-forward neural network, and max() is the ReLU activation function; The output of the feed-forward neural network is the final output of the access pattern recognition model; Exemplarily, the access pattern recognition model needs to predict the access heat of commodity data within the next week, that is, which commodity data belongs to hot data and which belongs to cold data. The output Y of the model can be expressed as:

[0032] where represents the predicted access heat value of the j-th commodity data; During the training process, the model is trained by combining supervised learning and unsupervised learning, and optimized through cross-validation and reinforcement learning until the model performance meets the requirements to obtain the access pattern recognition model; Specifically, the difference between the predicted access heat value and the actual access heat value is used as the loss function for supervised learning. Exemplarily, the mean squared error function is selected as the loss function Ls:

[0033] where is the actual access heat value of the j-th commodity data; By mining the potential patterns in the data, such as the access relationship between commodities, a loss function for unsupervised learning is constructed. For example, a contrastive learning loss function Lu is selected so that data with similar access patterns are closer in the feature space and data with different access patterns are farther away. The supervised learning loss and the unsupervised learning loss are weighted summed to obtain the total loss function L:

[0034] in, It is the weight coefficient that balances the role of supervised learning and unsupervised learning in training; Using the stochastic gradient descent algorithm, according to the total loss function L The parameters of the model (such as the weight matrix W and bias vector b) are updated to minimize the loss function and improve model performance; At the same time, through cross-validation, the data set is divided into multiple subsets, and training and validation are performed in turn to adjust model parameters. Reinforcement learning is used to dynamically optimize the model strategy based on the feedback of model predictions and actual access conditions until the model performance meets the requirements and a large model for access pattern recognition is obtained. It should be noted that you can also choose deep learning models such as RNN and LSTM models as the basic model, or other large pre-trained models as the basic model; S2. In response to the object enumeration request, the access pattern recognition model is used to analyze the context of the object enumeration request, and the hot objects that the user needs to access are predicted in advance; the specific steps of step S2 are as follows: S21. When an object enumeration request is detected, the storage system collects the context of the object enumeration request and inputs the access pattern recognition model; the context of the object enumeration request includes request parameters and user behavior history; For example, when a user initiates an object enumeration request to query "Summer women's clothing discount products", the storage system immediately collects the request parameters ("Summer women's clothing discount products") and the user's behavior history, such as whether the user has browsed women's dresses and purchased discounted products many times in the past month, and inputs this information into the access pattern recognition model; S22. The access pattern recognition model recognizes the user's historical access pattern, combines object relevance and tags, and predicts the hot objects that need to be queried in advance; For example, the access pattern recognition model identifies the user's historical access pattern, and based on the relevance and labels of the product objects, predicts that the user may also be interested in summer women's clothing products such as "women's sun protection clothing" and "women's sun hats", and identifies these predicted product data as hot objects; S3. Dynamically adjust the query strategy and the access method of storage nodes based on the prediction results, cache and preload hot objects in advance, and then perform the object listing operation; the specific steps of step S3 are as follows: S31. Determine the storage node to which the hot object belongs as the routing object and set it as the target storage node; Exemplarily, by analyzing the storage location of the hot object, determine the storage node storing hot objects such as "ladies' sunscreen clothes" and "ladies' sun hats" as the routing object and set it as the target storage node; Specifically, the target storage node is selected based on the weighted score of the real-time load rate of the storage node and the object heat; It should be noted that the hot objects predicted in step S2 also carry object heat; Scores of each node ; Among them is the weight coefficient that balances the roles of object heat and node load rate in the selection of the target storage node; The target storage node is selected in descending order of the scores of each node; S32. Determine the target storage node as the allocation node for the query task; S33. Preload the hot object into the cache or memory in advance, and preload the hot object into different target storage nodes in advance; For example, preload the data of "ladies' sunscreen clothes" at node 1 and node 2 in data center A and node 3 in data center B; It should be noted that the hot objects are classified according to the object heat; For hot objects with object heat greater than the first threshold, they can be preloaded into the memory cache in advance; For hot objects with object heat less than the first threshold and greater than the second threshold, they can be preloaded into the SSD cache; For hot objects with object heat less than the second threshold and greater than the third threshold, they can be replicated among the target storage nodes in advance; S34. Perform the object listing operation, split the object listing operation into several query tasks, and allocate each query task to the target storage node; S35. Each target storage node executes the assigned query task and queries the required objects in turn according to the order of the cache or memory, the hot objects preloaded by itself, and the original query strategy; S36. Merge the query results of the query tasks of each target storage node and return them to the user to complete the object listing operation; S4. Collect the listing results and performance data of the object listing operation, and optimize and update the access pattern recognition large model in combination with the new access log; the specific steps of step S4 are as follows: S41. Collect the listing results and performance data after the object listing operation is completed; the listing results include listing success or failure; the performance data includes the listing response time; For example, successfully list 10 discounted summer women's clothing items and performance data (listing response time is 200ms); S42. Analyze the listing results to obtain the listing success rate; S43. Optimize the access pattern recognition large model according to the listing success rate and performance data, and adjust the query strategy and cache strategy; Exemplarily, according to the listing success rate and performance data, if it is found that the response time is long, optimize the access pattern recognition large model and adjust the query strategy, such as optimizing the query task allocation method, or adjust the cache strategy, such as increasing the cache duration of popular items; S44. Regularly collect new access logs in the storage system and add them to the dataset to retrain the access pattern recognition large model to complete model update; It should be noted that the access pattern recognition large model adopts an incremental learning method, and the weight of new data in the dataset is calculated according to the data freshness w;

[0035] Among them, is the interval for data collection; It should be noted that the higher the freshness, the greater the weight of the data, and the lower the freshness, the smaller the weight of the data, so as to capture the latest access pattern changes.

[0036] In an embodiment of the present invention, based on step S35, a possible embodiment will be given below to non - restrictively elaborate on its specific implementation.

[0037] The specific steps of step S35 are as follows: S351. Each target storage node starts and executes the assigned query task; S352. The query task first looks for the required object from the hot objects in the cache or memory; If the requirement is met, go to step S36; If the requirement is not met, go to step S353; S353. The query task uses its own target storage node as the routing target and queries the required object from its own hot objects or the hot objects pre - loaded by itself; If the requirement is met, go to step S36; If the requirement is not met, go to step S354; S354. The query task performs the query of the required objects according to the original query strategy until the requirements are met.

[0038] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0039] As Figure 2 shown, the following are embodiments of a system for optimizing the object listing performance in object storage based on a large model provided by the embodiments of the present disclosure. This system and the method for optimizing the object listing performance in object storage based on a large model in the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiments of the system for optimizing the object listing performance in object storage based on a large model, reference can be made to the embodiments of the method for optimizing the object listing performance in object storage based on a large model.

[0040] The system includes: A data collection and model training module, configured to collect access logs and object metadata in the storage system to construct a data set, and use the data set to train a large model for access pattern recognition; A hot object prediction module, configured to respond to an object listing request, analyze the context of the object listing request using the large model for access pattern recognition, and predict in advance the hot objects that the user needs to access; An object listing optimization module, configured to dynamically adjust the query strategy and the access method of the storage nodes based on the prediction result, and cache and preload the hot objects in advance, and then perform the object listing operation; A model optimization and update module, configured to collect the listing results and performance data of the object listing operation, and optimize and update the large model for access pattern recognition in combination with the new access logs.

[0041] Through the cooperation of the data collection and model training module, the hot object prediction module, the object listing optimization module, and the model optimization and update module in this embodiment, the performance of the object listing operation is optimized.

[0042] The method for optimizing the object listing performance in object storage based on a large model provided by the embodiments of the present application can be applied to an electronic device. Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described and / or claimed herein.

[0043] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, keys, a camera, a display screen, and a SIM card interface, etc.

[0044] It can be understood that the structure schematically shown in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0045] The processor may include one or more processing units. For example, the processor may include a central processing unit (CPU), etc., an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0046] Among them, the processor can be the nerve center and command center of the electronic device. The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching instructions and executing instructions.

[0047] A memory can also be set in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory can save the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the system efficiency.

[0048] The above-mentioned electronic device implements the technical solution of the method for optimizing the object listing performance in object storage based on a large model in the present application, that is, constructing a data set by collecting access logs and object metadata in the collection and storage system, and using the data set to train an access pattern recognition large model; in response to an object listing request, using the access pattern recognition large model to analyze the context of the object listing request, and predicting in advance the hot objects that the user needs to access; dynamically adjusting the query strategy and the access method of the storage node based on the prediction result, and caching and preloading the hot objects in advance, and then performing the object listing operation; collecting the listing result and performance data of the object listing operation, and optimizing and updating the access pattern recognition large model in combination with the new access log, achieving the beneficial effects of improving the performance of the object listing operation in the object storage system, effectively reducing the query latency, improving the response speed and throughput of the system; at the same time, through the intelligent caching and preloading mechanism, optimizing the resource utilization rate, reducing the storage cost, and enhancing the user experience.

[0049] In the storage medium provided by the present application, there is a program product capable of implementing the method for optimizing the object listing performance in object storage based on a large model.

[0050] The method for optimizing the object listing performance in object storage based on a large model includes: collecting access logs and object metadata in the storage system to construct a data set, and using the data set to train an access pattern recognition large model; in response to an object listing request, using the access pattern recognition large model to analyze the context of the object listing request, and predicting in advance the hot objects that the user needs to access; dynamically adjusting the query strategy and the access method of the storage node based on the prediction result, and caching and preloading the hot objects in advance, and then performing the object listing operation; collecting the listing result and performance data of the object listing operation, and optimizing and updating the access pattern recognition large model in combination with the new access log.

[0051] In some possible embodiments, the high-reusability method for data exchange between contract management and heterogeneous systems of the present disclosure may be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.

[0052] The storage medium of the present disclosure may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0053] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for optimizing the object listing performance in object storage based on a large model, characterized in that It includes the following steps: S1. Collect access logs and object metadata in the storage system to build a dataset, and use the dataset to train a large access pattern recognition model; S2. In response to an object listing request, use the large access pattern recognition model to analyze the context of the object listing request and predict in advance the hot objects that the user needs to access; S3. Dynamically adjust the query strategy and the access method of the storage node based on the prediction result, and cache and preload the hot objects in advance, and then perform the object listing operation; S4. Collect the listing results and performance data of the object listing operation, and optimize and update the large access pattern recognition model in combination with the new access logs.

2. The method for optimizing the object listing performance in object storage based on a large model according to claim 1, wherein The access logs include the access frequency, access time, visitor identity, and request type of the object; the object metadata includes the object ID, label, size, and creation time; The specific steps for building the dataset are as follows: Analyze the access logs to obtain access behavior data; the access behavior data includes access frequency, time distribution, and access relationship; Clean and process the access logs and access behavior data to obtain a structured dataset.

3. The method for optimizing the object listing performance in object storage based on a large model according to claim 2, wherein The specific steps for training the access pattern recognition model using the dataset are as follows: Select a basic model to build an access pattern recognition model, use the structured dataset as input, and use predicting hot data and cold data within a future set time period as the training target, and start model training; During the training process, the model is trained by combining supervised learning and unsupervised learning, and optimized through cross-validation and reinforcement learning methods until the model performance meets the requirements to obtain an access pattern recognition model.

4. The method for optimizing the object listing performance in object storage based on a large model according to claim 2, wherein The specific steps of step S2 are as follows: S21. When an object listing request is detected, the storage system collects the context of the object listing request and inputs it into the large access pattern recognition model; the context of the object listing request includes request parameters and user behavior history; S22. The large access pattern recognition model recognizes the user's historical access pattern, combines object relevance and labels, and predicts in advance the hot objects that need to be queried currently.

5. The method for optimizing the object listing performance in object storage based on a large model according to claim 4, wherein The specific steps of step S3 are as follows: S31. Determine the storage node to which the hot object belongs as the routing object and set it as the target storage node; S32. Determine the target storage node as the allocation node for the query task; S33. Preload the hot object into the cache or memory in advance, and preload the hot object into different target storage nodes in advance; S34. Perform the object listing operation, split the object listing operation into several query tasks, and allocate each query task to the target storage node; S35. Each target storage node executes the allocated query task and queries the required objects in the order of the cache or memory, the hot objects preloaded by itself, and the original query strategy; S36. Merge the query results of the query tasks of each target storage node and return them to the user to complete the object listing operation.

6. The method for optimizing the object listing performance in object storage based on a large model according to claim 5, wherein The specific steps of step S35 are as follows: S351. Each target storage node starts and executes the allocated query task; S352. The query task first looks for the required object from the hot object in the cache or memory; If the requirement is met, go to step S36; If the requirements are not met, go to step S353; S353. The query task queries the required object from its own hot objects or pre-loaded hot objects with its own target storage node as the routing target; If the requirements are met, go to step S36; If the requirements are not met, go to step S354; S354. The query task executes the query of the required object according to the original query strategy until the requirements are met.

7. The method for optimizing the object listing performance in object storage based on a large model according to claim 5, wherein The specific steps of step S4 are as follows: S41. Collect the listing result and performance data after the object listing operation is completed; the listing result includes listing success or failure; the performance data includes the listing response time; S42. Analyze the listing result to obtain the listing success rate; S43. Optimize the access pattern recognition large model according to the listing success rate and performance data, and adjust the query strategy and cache strategy; S44. Regularly collect new access logs in the storage system and add them to the dataset to retrain the access pattern recognition large model to complete model update.

8. A system for optimizing the object listing performance in object storage based on a large model, characterized in that, Including: A data collection and model training module, configured to collect access logs and object metadata in the storage system to construct a dataset, and use the dataset to train an access pattern recognition large model; A hot object prediction module, configured to respond to an object listing request, analyze the context of the object listing request using the access pattern recognition large model, and predict in advance the hot objects that the user needs to access; An object listing optimization module, configured to dynamically adjust the query strategy and the access method of the storage node based on the prediction result, and cache and pre-load the hot objects in advance, and then perform the object listing operation; A model optimization and update module, configured to collect the listing result and performance data of the object listing operation, and optimize and update the access pattern recognition large model in combination with new access logs.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps of the method for optimizing the object listing performance in object storage based on a large model according to any one of claims 1 to 7 when executing the program.

10. A storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by the processor, implements the steps of the method for optimizing the object listing performance in object storage based on a large model according to any one of claims 1 to 7.