Data storage method and system for AI capability fusion application platform

Through real-time monitoring and dynamic adjustment of the storage mapping of AI tasks and data, the problem of uneven allocation of storage resources in the existing technology is solved, efficient and flexible storage resource management is achieved, and the storage performance of the AI ​​application platform is improved.

CN120066411APending Publication Date: 2025-05-30SHANGHAI SHENGTONG ZHIMING TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510147164.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing technology is difficult to dynamically adjust the allocation of storage resources while meeting the access needs of artificial intelligence models, resulting in lagging storage efficiency and being unable to respond to complex and changing data access needs in a timely manner.

Method used

By monitoring the changes in AI task access load and data access patterns in real time, extracting feature vectors of tasks and data, performing grouping and matching score calculations, and dynamically adjusting the storage mapping of task grouping and data packets to ensure flexibility and efficiency of resource allocation.

Benefits of technology

Real-time optimization of the storage resources of the AI ​​capability convergence application platform is achieved, storage efficiency and access performance are improved, and it can respond to changes in task load and data access mode in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066411A_ABST
    Figure CN120066411A_ABST
Patent Text Reader

Abstract

The invention provides a data storage method and system for an AI capability fusion application platform, and relates to the technical field of artificial intelligence. The objective of the invention is to dynamically optimize the storage allocation efficiency of the AI task and the multi-type data. The method comprises the following steps: extracting an access mode, a time delay demand and a throughput demand of an AI task, and generating a task feature vector; extracting a data access frequency, a data size and a read-write mode, and generating a data feature vector; the tasks and the data are grouped through the feature vectors, and matching scores between task groups and data groups are calculated. And splitting and recombining the low-matching-degree grouping combination, and optimizing storage allocation. The method further comprises the steps that when the task load or the data access mode changes remarkably, dynamic adjustment is triggered, the matching score is recalculated, and only the affected groups are locally updated. And when the access conflicts or the concurrency exceeds the limit, re-grouping or merging is executed, so that the access bottleneck is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly, to a data storage method and system for an AI capability fusion application platform. Background Art

[0002] In the current field of artificial intelligence technology, a large amount of data storage requirements need to be processed during the training and inference processes of large models. However, due to the huge amount of data and high storage costs, how to reduce storage costs while meeting the model access requirements has become an urgent problem for technicians to solve.

[0003] To solve the above problems, for example, Chinese Patent Application No. CN117931077A discloses a data storage optimization method for large models. This method performs reference counting on the data involved in the model inference process, classifies the model data according to the reference count values, stores the data with high reference frequencies in storage resources with higher performance and lower latency, and stores the data with low reference frequencies in storage resources with higher latency and lower costs. This way reduces storage costs while meeting the data access requirements.

[0004] Again, for example, Chinese Patent Application No. CN118228821A discloses a data storage method for large models in multiple scenarios. This method counts according to the corresponding relationship between the data involved in the inference process and the scenario attributes during the inference process, classifies the model data according to the reference count values of the scenario attributes, and stores the model data corresponding to different scenario attributes in different storage resources. This method can improve the professionalism of model data in specific scenarios and enhance the service quality of large models in different scenarios under the condition of limited storage resources.

[0005] However, the above methods are mainly based on static or periodic data reference counting, lack a fast response mechanism for real-time changing task loads or data access patterns, cannot adjust the storage resource allocation in time, and there is a problem of lagging storage efficiency. Therefore, a more flexible dynamic data storage method is needed to cope with the complex and changeable data access requirements in the AI capability fusion application platform. Summary of the Invention

[0006] In view of the deficiencies of the prior art, embodiments of the present disclosure provide a data storage method and system for an AI capability fusion application platform.

[0007] In a first aspect, embodiments of the present disclosure provide a data storage method for an AI capability fusion application platform, including:

[0008] Extract the access patterns, latency requirements, and throughput requirements of multiple AI tasks from the target AI platform, respectively forming multiple task feature vectors; extract the access frequencies, data sizes, and read / write patterns from multiple types of data in the target AI platform, respectively forming multiple data feature vectors;

[0009] Group multiple AI tasks according to the task feature vectors to form several task groups; group multiple types of data according to the data feature vectors to form several data groups;

[0010] Calculate the matching scores between each pair of the task groups and the data groups respectively to generate a first matching result; wherein, the matching score is weighted and calculated by the feature similarity between the task group and the data group;

[0011] In the first matching result, mark the group combinations with matching scores lower than the preset threshold as target group combinations, and perform local fine-tuning on the target group combinations to generate a second matching result; wherein, the local fine-tuning includes: splitting and reorganizing the target group combinations, and keeping the group combinations with matching scores higher than the preset threshold unchanged;

[0012] Based on the second matching result and the storage requirement information corresponding to each group combination, map each task group and data group to different storage levels, where the storage requirement information includes: latency sensitivity requirement, tolerance requirement, and capacity requirement.

[0013] As an optional implementation, it further includes:

[0014] When the real-time eigenvalue of the access load of the AI task or the data access mode changes relative to the initial eigenvalue collected during system initialization, and the change amplitude exceeds the predetermined percentage threshold of dynamic adjustment, trigger the dynamic adjustment operation;

[0015] Recalculate the matching scores of the task groups and data groups that have changed, and only update the matching results of the task groups and data groups affected by the change;

[0016] For the target groups with matching scores lower than the preset threshold in the updated matching results, perform the following steps on the task units or data units in the target groups:

[0017] Reallocate the task units or data units with low matching degrees to other groups with higher feature similarity;

[0018] In response to the existing groups being unable to meet the similarity requirements, create new groups according to the feature similarity;

[0019] Recalculate the matching score for the reallocated task unit or data unit and update the grouping structure.

[0020] As an alternative implementation, it further includes:

[0021] In response to detecting that the access conflict probability or concurrency exceeds a second preset threshold, perform regrouping operations on the target groups that have been mapped to the corresponding storage levels;

[0022] If there are still target groups with scores lower than the preset threshold after regrouping, split the target groups or merge them with similar groups, recalculate the matching scores, and then update the storage mapping.

[0023] As an alternative implementation, the value of the second preset threshold is determined based on at least one of the following methods:

[0024] The access conflict probability distribution calculated based on historical statistical methods;

[0025] The upper limit of concurrency configured by the system administrator;

[0026] Dynamically adjusted through reinforcement learning or adaptive algorithms to meet the tolerance requirements in different target scenarios.

[0027] As an alternative implementation, the regrouping operations include:

[0028] Identify the task units and data units of the target groups that have been mapped to the corresponding storage levels;

[0029] For the units in the target groups whose similarity is insufficient to maintain the current mapping relationship, relocate them to existing groups that are more similar to their feature vectors;

[0030] If the similarity between all existing groups and the target groups is lower than a first threshold, create a new group to accommodate the task units and data units of the target groups that have been mapped to the corresponding storage levels, and update the matching scores of the units.

[0031] As an alternative implementation, the triggering conditions for performing splitting or merging in response to there still being target groups with scores lower than the preset threshold after regrouping include at least one of the following:

[0032] The similarity between the task units or data units within the target group is lower than a predetermined difference threshold;

[0033] The similarity of the feature vectors between the target group and other groups exceeds a predetermined merging threshold;

[0034] The access latency or conflict of the target group in the current storage level fails to be reduced to within the preset safe range.

[0035] As an alternative implementation, the updating of the storage mapping after recalculating the matching score includes:

[0036] Determine whether the newly added sub - groups due to splitting or the newly formed groups due to merging still match the latency sensitivity requirements and capacity requirements of the corresponding storage levels;

[0037] If there is a mismatch, perform local re - allocation between this group and the remaining groups;

[0038] Finally, generate a new matching result and register it in the monitoring module.

[0039] In a second aspect, an embodiment of the present disclosure also provides a data storage system for an AI capability fusion application platform, including:

[0040] An acquisition module, a grouping module, a first matching module, a second matching module, and a mapping module.

[0041] The acquisition module is configured to extract access patterns, latency requirements, and throughput requirements of multiple AI tasks from a target AI platform, respectively forming multiple task feature vectors; extract access frequencies, data sizes, and read - write patterns from multiple types of data in the target AI platform, respectively forming multiple data feature vectors;

[0042] The grouping module is configured to group multiple AI tasks according to the task feature vectors, forming several task groups; group multiple types of data according to the data feature vectors, forming several data groups;

[0043] The first matching module is configured to calculate the matching score between each pair of the task groups and the data groups respectively, generating a first matching result; wherein, the matching score is weighted and calculated by the feature similarity between the task group and the data group;

[0044] The second matching module is configured to mark the group combinations with matching scores lower than a preset threshold in the first matching result as target group combinations, and perform local fine - tuning on the target group combinations, generating a second matching result; wherein, the local fine - tuning includes: splitting and reorganizing the target group combinations, and keeping the group combinations with matching scores higher than the preset threshold unchanged;

[0045] The mapping module is configured to map each task group and data group to different storage levels based on the second matching result and the storage requirement information corresponding to each group combination, wherein the storage requirement information includes: latency sensitivity requirements, tolerance requirements, and capacity requirements.

[0046] Compared with the prior art, the present invention monitors the changes in the access load and data access patterns of AI tasks in real time, and re-matches and stores the mapping of task groups and data groups in a timely manner, avoiding uneven allocation of storage resources, and significantly improving storage efficiency and access performance. For the grouping combinations with low matching degrees, split or reorganize them, and dynamically create new groups to ensure that the data or tasks within the groups have high similarity, thereby improving the accuracy and speed of data access and storage. According to the differences in task characteristics and data characteristics, map different groups to different levels of storage resources to ensure that high-real-time tasks and high-frequency accessed data are allocated to low-latency storage devices, reducing access latency. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 is a flowchart of a data storage method for an AI capability fusion application platform provided by an embodiment of the present disclosure;

[0048] Figure 2 is a flowchart of a method for updating storage mapping after recalculating matching scores provided by an embodiment of the present disclosure;

[0049] Figure 3 is a schematic diagram of a data storage system for an AI capability fusion application platform provided by an embodiment of the present disclosure.

[0050] Reference numerals: 10, acquisition module; 20, grouping module; 30, first matching module; 40, second matching module; 50, mapping module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0052] See Figure 1 as shown, which is a flowchart of a data storage method for an AI capability fusion application platform provided by an embodiment of the present disclosure. The method includes steps S101 to S105, where:

[0053] S101: Extract the access patterns, latency requirements, and throughput requirements of multiple AI tasks from the target AI platform, and respectively form multiple task feature vectors; extract the access frequency, data size, and read / write mode from multiple types of data in the target AI platform, and respectively form multiple data feature vectors;

[0054] S102: Group multiple AI tasks according to the task feature vectors to form several task groups; group multiple types of data according to the data feature vectors to form several data groups;

[0055] S103: Calculate the matching scores between each pair of the task groups and the data groups respectively to generate a first matching result; wherein, the matching score is calculated by weighting the feature similarity between the task group and the data group.

[0056] S104: In the first matching result, mark the group combinations with matching scores lower than the preset threshold as target group combinations, and perform local fine-tuning on the target group combinations to generate a second matching result; wherein, the local fine-tuning includes: splitting and reorganizing the target group combinations, and keeping the group combinations with matching scores higher than the preset threshold unchanged.

[0057] S105: Based on the second matching result and the storage requirement information corresponding to each group combination, map each task group and data group to different storage levels, wherein the storage requirement information includes: delay sensitivity requirement, tolerance requirement, and capacity requirement.

[0058] Among them, the target AI platform, that is, the AI capability fusion application platform. The target AI platform includes several AI tasks with different functional requirements and performance requirements, such as image recognition tasks, natural language processing tasks, and data mining tasks, etc. The data types involved in the platform include video data, text data, and structured table data.

[0059] Regarding the above S101:

[0060] First, during system initialization, the access patterns, delay requirements, and throughput requirements of each AI task are sequentially collected from the target AI platform to form corresponding multiple task feature vectors. The subscript i used in the following description represents the i-th AI task, and m represents the total number of AI tasks. Therefore, the feature vectors of all tasks can be expressed as {T 1 , T 2 , …, T m}. Specifically, the system performs the following operations for each AI task (i.e., for all i ∈ {1, …, m}):

[0061] For obtaining the access pattern:

[0062] In specific implementation, the call frequency and average response time requirement of the task can be monitored. For example, the image recognition task has a relatively concentrated call frequency per day and a high requirement for real-time response; the natural language processing task is called relatively frequently, but has a high requirement for the peak throughput carrying capacity; the data mining task is called relatively concentrated and executed in batches at night, has a high requirement for throughput, but the delay requirement is relatively loose.

[0063] For evaluating the delay requirement:

[0064] In specific implementations, the maximum delay threshold that each task can tolerate under different loads can be obtained by continuously sampling the response time of task requests. For example, image recognition tasks allow a maximum response delay of 200 milliseconds in application scenarios, natural language processing tasks allow 300 milliseconds, and data mining tasks allow 1 second.

[0065] Assessment of throughput requirements:

[0066] In specific implementations, the processing volume of each task within a period of time (such as the number of requests processed per minute or per second) can be counted, and the throughput standards that need to be guaranteed can be determined in combination with business objectives. For example, image recognition tasks need to process 100 images per second, natural language processing tasks need to process 1,000 paragraphs of text per second, and data mining tasks can process massive amounts of data at one time but require high parallelism.

[0067] After normalizing the above key indicators such as access mode, latency requirement, and throughput requirement, we construct the corresponding task feature vector, such as T i =(v i ,l i ,th i ), where v i represents the feature value related to the access mode of the i-th task, l i Indicates the latency requirement of the task, th i Indicates the throughput requirement of the task.

[0068] Furthermore, the access frequency, data size, and read / write mode are extracted from the multi-type data of the target AI platform to form multiple data feature vectors. The subscript j used in the following description represents the jth type of data, and n represents the total number of data categories or data groups. Therefore, the feature vector of all data can be expressed as {D 1 ,D 2 ,…,D n}.

[0069] In specific implementations, the number of times each type of data is read or updated in a given time window (such as every day or every hour) can be counted through the monitoring log of the database or storage system. For example, access to video data is concentrated during peak hours, access to text data is relatively even, and structured table data will be frequently updated during business peaks. Statistics are performed on the average size and maximum / minimum values ​​of each type of data. For example, the average size of video files is in the hundreds of megabytes, the average size of text data is smaller but the quantity is huge, and structured data may be a table of several GB. Determine whether each type of data is mainly read-only or read-write, or whether there are a large number of random writes / batch writes. For example, video data is mainly read-only, text data has both batch write operations and real-time updates, and structured data is frequently updated during peak hours.

[0070] Similarly, after uniformly scaling or normalizing the access frequency, data size, and read-write mode related metrics, a data feature vector D j =(f j , s j , rw j ), where f j represents the access frequency of the j-th type of data, s j represents the data size magnitude, and rw j represents the read-write mode feature.

[0071] Regarding the above S102:

[0072] In a specific implementation, based on the above task feature vector, multiple AI tasks are grouped to form several task groups. Exemplarily, the system uses the Hierarchical Clustering or K-means algorithm to perform aggregation analysis on all AI tasks.

[0073] Exemplarily, for all task feature vector sets {T 1 , T 2 , …, T m}, similarity calculation is performed. Cosine similarity or Euclidean distance can be used as the measurement criterion; according to the similarity matrix or distance matrix, at a preset number of clusters or similarity threshold, tasks with similar features are automatically grouped into the same group. For example, image recognition and speech recognition tasks with high real-time response requirements are classified into the "real-time high-sensitivity task group", natural language processing, text analysis, etc. are classified into the "medium real-time task group", and data mining and batch statistical analysis tasks are classified into the "offline processing task group", etc.

[0074] Similarly, based on the data feature vector, various types of data are clustered and grouped to form several data groups. Preliminary clustering can be performed according to type features such as video data, text data, and structured table data, and then more refined grouping is carried out according to the similarity of access patterns and read-write patterns. For example, read-only video data and large files with low-frequency access are grouped into one group, and text log data with frequent continuous write operations is separately grouped into one group, etc.

[0075] Regarding the above S103:

[0076] After completing the preliminary grouping, the matching score between each pair of task groups and data groups is calculated to generate a first matching result. The calculation method of the matching score can be:

[0077] score(G T , G D )

[0078] = w1 ·sim(Delay requirement, Access frequency) + w 2 ·sim(Throughput requirement, Data size)

[0079] + w 3 ·sim(Access pattern, Read / write pattern)

[0080] Among them, G T represents a certain task group, and G D represents a certain data group. sim represents the quantitative calculation of feature similarity, and w 1 , w 2 , w 3 are respectively weighted coefficients, used to reflect the influence proportion of different delays, throughputs, and access patterns on the overall matching degree. In specific implementation, the weights can be configured in advance by operation and maintenance personnel or system administrators to meet the business requirements of the platform.

[0081] Regarding the above S104:

[0082] In specific implementation, in the first matching result, the group combinations with matching scores lower than the preset threshold are marked as target group combinations. For these target group combinations, perform local fine-tuning operations to generate the second matching result.

[0083] Exemplarily, the local fine-tuning operations may include:

[0084] Conduct a more refined analysis of the task and data characteristics within the target group combination to confirm the main reasons for the low matching degree. For example, the delay requirement does not match the access characteristics of the current data, or the throughput requirement cannot be matched due to limited data group capacity;

[0085] Split and reorganize the target group combination: If a task group contains tasks with extremely high and medium-low delay requirements at the same time, it can be further split; If the data group contains data with mixed read / write and extremely different access patterns, the subclass data can also be separately split into new groups;

[0086] During the reorganization process, keep the group combinations with matching scores higher than the preset threshold unchanged to ensure that the existing stable matching relationships are not disrupted; For the combinations that need to be adjusted, recombine them according to more refined features or merge them with similar groups to form new group combinations.

[0087] Regarding the above S105:

[0088] In specific implementation, after completing the local fine-tuning, generate the second matching result. Based on this result, and the storage requirement information corresponding to each group combination (such as delay sensitivity requirement, tolerance requirement, capacity requirement, etc.), map each task group and data group to different storage levels.

[0089] Exemplarily, task groups with high time-delay sensitivity and high throughput requirements and their supporting data groups are mapped to the SSD cache layer or the high-speed memory computing layer; task groups and data groups with medium real-time requirements are allocated to performance-type disks or main storages; and offline batch processing task groups and corresponding massive data groups are allocated to capacity-type hard disk storages or cold storages according to capacity requirements.

[0090] It should be noted that during the implementation process, the storage layer mapping scheme can be flexibly set in combination with the actual business requirements and the status of hardware resources. For example, after comprehensively analyzing the time delay, throughput, and access frequency of each packet combination, the specific placement location is determined. Finally, after completing the entire mapping, the system generates the latest storage allocation scheme according to the second matching result and records the mapping information between each packet combination and the corresponding storage layer in the background, providing data support for subsequent dynamic scheduling or monitoring.

[0091] In this way, the initial grouping is first completed based on the AI tasks and data feature vectors, then the matching degree calculation and local fine-tuning are performed, and finally the matched task groups and data groups are mapped to the appropriate storage layer, so as to meet the time delay and throughput requirements of the platform under various AI application scenarios and take into account the efficient allocation and utilization of storage resources.

[0092] As an alternative implementation method, it further includes:

[0093] When the access load of the AI task or the real-time characteristic value of the data access mode changes relative to the initial characteristic value collected during system initialization, and the change amplitude exceeds the predetermined percentage threshold for dynamic adjustment, a dynamic adjustment operation is triggered;

[0094] Recalculate the matching scores of the task groups and data groups that have changed, and only update the matching results of the task groups and data groups affected by the change;

[0095] For the target groups with matching scores lower than the preset threshold in the updated matching results, the following steps are performed on the task units or data units in the target groups:

[0096] Reallocate the task units or data units with low matching degrees to other groups with higher feature similarity;

[0097] In response to the inability of the existing groups to meet the similarity requirements, new groups are created according to the feature similarity;

[0098] For the reallocated task units or data units, recalculate their matching scores and update the group structure.

[0099] In a specific implementation, during the operation phase, the present disclosure continuously monitors metrics such as the real-time access load and data read / write patterns of AI tasks, and compares them with the initial eigenvalue collected during system initialization. To this end, the system maintains two sets of eigenvalues:

[0100] 1) Initial eigenvalue: The task feature vector recorded during system initialization or after the previous grouping is stable and the data feature vector where i represents the i-th AI task, j represents the j-th type of data, and the superscript 0 represents the feature vector of a certain task i and data type j at the "initial" time;

[0101] 2) Real-time eigenvalue: The system collects the current task feature vector at fixed time intervals (such as every 5 minutes or every 30 minutes) or when an event is triggered (such as when the access volume suddenly increases) and the data feature vector This real-time eigenvalue may include: the call frequency of the task, the actual throughput, the latency observation value, as well as the access frequency, size change, read / write pattern, etc. of the data.

[0102] where the superscript t represents the feature vector of task i and the j-th type of data at the "real-time monitoring" moment (for example, dynamically collected after the system has been running for a period of time).

[0103] When the system detects that the gap between the real-time eigenvalue of any task unit or data unit and the corresponding initial eigenvalue exceeds a predefined dynamic adjustment percentage threshold (such as 20% or 30%), a dynamic adjustment operation is triggered. This percentage threshold can be configured by the system administrator according to business stability requirements or determined through historical analysis.

[0104] After the dynamic adjustment operation is triggered, the system does not perform a full recalculation for all task groups and data groups, but only updates the matching scores for the groups to which the task units or data units affected by the change belong, as well as the corresponding data groups or task groups in the matching table. The reason is that other groups that have not changed significantly still maintain a relatively stable matching degree in the short term and do not require a full recalculation.

[0105] In a specific implementation, if a certain task has a 40% increase in its actual throughput or access pattern compared to its initial value, it indicates that there is a significant change in the task load and its task group needs to be re-evaluated.

[0106] Find the task group G T to which this task unit belongs, and the set of data groups {G D1 ,G D2, …}, if it is a certain data unit that changes, then search for the data group G where it is located D and the corresponding task group.

[0107] Furthermore, based on the new real-time eigenvalue, similar to the matching score calculation method described above:

[0108] new_score(G T , G D )

[0109] = w 1 ·sim(new latency requirement, access frequency) + w 2 ·sim(new throughput requirement, data size)

[0110] + w 3 ·sim(new access pattern, read / write pattern)

[0111] Compare the newly obtained matching score with the original score, and update the matching score of this combination in the system. If the score increases significantly, maintain the original group mapping; if the score drops below a preset threshold, mark this combination as the "target group combination" and further execute the subsequent steps.

[0112] In a specific implementation, if there is a target group with a matching score lower than the preset threshold in the updated matching result, it means that the current task group or data group no longer matches the new feature situation. At this time, the system needs to adjust the task units or data units in the target group.

[0113] Exemplarily, other existing groups with higher feature similarity can be searched in the matching structure. For example, if a certain task unit was originally in the "medium real-time task group", but now its latency requirement has increased significantly and is closer to the "high real-time" feature, it can be transferred to the "real-time high-sensitivity task group".

[0114] The reallocation of data units is the same. If a certain type of data was originally in the "large file low access frequency" group, but actual monitoring finds that its access frequency continues to climb, it can be migrated to the "high access frequency data group".

[0115] In addition, if the features of any existing group are not close enough to the changed features of this task or data unit (for example, the similarity values are all lower than a certain first threshold), then a new group is automatically created according to the new feature vector, and this task unit or data unit is placed in it.

[0116] For example, if the running mode of a certain task unit becomes extremely hybrid (requiring both high concurrency and medium latency), and there is no group with similar features in the existing groups, the system creates a "hybrid requirement group" to accommodate this task unit.

[0117] After completing the reallocation or creating a new group, it is necessary to recalculate the affected matching relationships, such as scoring the new task group where the transferred unit is located and its corresponding data group; or matching and scoring the newly created group with all possible data groups.

[0118] According to the scoring results, update the final mapping relationship of the unit or the group. If the matching score meets the preset threshold, the adjustment is completed; otherwise, continue with the next step of local fine-tuning.

[0119] After all affected units are adjusted, update and record the group structure and matching table in the background to ensure that the latest mapping information can be used for subsequent access scheduling.

[0120] In this way, when the present disclosure detects that the task or data characteristics deviate significantly from the initial settings, it can timely perform dynamic adjustment on the groups with low matching scores. Compared with the static group management method, when the task load or data access pattern fluctuates with the business requirements, the system can quickly reallocate resources and groups, avoiding performance bottlenecks caused by the decrease in matching degree. Only locally update the affected groups instead of recalculating all group matches in full volume, which can save computing resources and reduce the impact on the normal operation of the system. In a dynamic environment, new groups and new matching relationships can be continuously created, merged or split, so that the group strategy of the system always matches the actual business requirements.

[0121] In this way, by real-time monitoring, local recalculation, and reallocation, the high matching degree and rationality between the task groups and data groups are maintained. This can not only ensure the storage efficiency and access performance of the AI capability integration application platform under the condition of changing business load, but also take into account the overall resource utilization rate of the system.

[0122] As an alternative implementation manner, it further includes:

[0123] In response to detecting that the access conflict probability or concurrency exceeds the second preset threshold, perform a regrouping operation on the target group that has been mapped to the corresponding storage level;

[0124] If there is still a target group with a score lower than the preset threshold after regrouping, split the target group or merge it with a similar group, recalculate the matching score, and then update the storage mapping.

[0125] As an alternative implementation manner, the triggering conditions for performing splitting or merging in response to there still being a target group with a score lower than the preset threshold after regrouping include at least one of the following:

[0126] The similarity between the task units or data units within the target group is lower than the predetermined difference threshold;

[0127] The similarity of the eigenvectors of the target group and other groups exceeds a predetermined merging threshold;

[0128] The access latency or conflict of the target group in the current storage level fails to be reduced to within the preset safe range.

[0129] Please refer to Figure 2 , as an alternative implementation, Figure 2 is a flowchart of a method for updating a storage mapping after recalculating a matching score provided by an embodiment of the present disclosure, including steps S201 to S203, where:

[0130] S201: Determine whether the newly added sub - groups due to splitting or the newly formed groups due to merging still match the latency sensitivity requirements and capacity requirements of the corresponding storage level;

[0131] S202: If there is a mismatch, perform local re - allocation between this group and the remaining groups;

[0132] S203: Finally, generate a new matching result and register it in the monitoring module.

[0133] Among them, β 2 is the second preset threshold, which is a new parameter specifically used in this implementation to measure whether the access conflict probability or concurrency exceeds the limit. If it is detected that the access conflict probability or concurrency exceeds this threshold, the re - grouping operation is triggered.

[0134] In addition, if the concurrency needs to be recorded, then P j can represent the concurrency of the j - th group.

[0135] In a specific implementation, to avoid access blocking or performance bottlenecks in some storage levels, the platform continuously monitors the access conflict probability and concurrency of each group in the actual access scenario.

[0136] Exemplarily, the number of resource contentions, lock waits, or conflict retries of a certain group within a unit time can be counted through the logs of the storage layer or the monitoring module, and divided by the total number of accesses to obtain the conflict probability.

[0137] If the conflict probability is greater than the safe range allowed by the system for a long time, it indicates that this group faces a high risk of access conflict at the storage level.

[0138] For concurrency, it can be estimated or monitored through indicators such as the number of connections, the number of parallel processing tasks, or the request queue length.

[0139] If the concurrency corresponding to a certain group continues to rise and exceeds the value preset by the administrator (such as β 2), indicating that it is difficult for this group to carry additional load in the current storage layer and there is a potential bottleneck.

[0140] When it is detected that the access conflict probability or concurrency volume of any packet exceeds the second preset threshold β 2 the system triggers a regrouping operation, attempting to reduce conflicts or share the concurrent load by adjusting the packet structure.

[0141] After triggering the regrouping operation, locate the "target packet" that has been mapped to a specific storage level (such as the cache layer, performance disk layer, or capacity hard disk layer), that is, the packet with conflicts or super concurrency. The regrouping operation may include the following steps:

[0142] Obtain the detailed composition of this packet and confirm which task units {T i} and data units {D j} are included.

[0143] At the same time, retrieve the feature vectors of these units, such as latency requirements, throughput requirements, read / write modes, etc., to prepare for subsequent splitting or merging.

[0144] Furthermore, determine the main reasons for high conflicts or high concurrency. For example, it is because tasks with extremely high access frequencies are mixed into low-frequency tasks in the task units, or because the read / write modes of the data units are incompatible.

[0145] According to the analysis results, decide which units to split or merge. For example, if the access patterns of certain sub-class tasks are too different from other units, they can be separately split into new packets.

[0146] Furthermore, the matching scoring mechanism described above can be used to recalculate the matching degree of the units that need to be migrated or reorganized with each alternative packet.

[0147] For example, if a high-concurrency task has a too low matching degree with the current data packet but is more compatible with another high-concurrency data packet in terms of throughput requirements, the task can be transferred to a more suitable packet.

[0148] After determining the new packet scheme, perform "splitting" or "partial transfer" on the target packet, and reallocate the units in it to packets with higher similarity in the same layer or other layers. Finally, generate a new packet structure and mapping table to reduce the access conflicts or concurrent pressure in the current layer.

[0149] In addition, after completing a regrouping operation, it is necessary to check the matching scores of the newly generated or adjusted packets one by one (refer to the scoring formula in the previous text) and determine whether there are still "target packets" with scores lower than the preset threshold. If it is found that some packets are still not good in the new structure, the following measures can be further taken:

[0150] If the task units or data units within the target group have too large differences themselves and the internal similarity is not sufficient to maintain a unified group (for example, some tasks require extremely low latency while others mainly focus on throughput), the target group can be split into multiple sub - groups.

[0151] For each sub - group formed after splitting, calculate its matching score with the corresponding data group or task group again. If the score reaches the threshold, the splitting is completed; otherwise, continue to fine - tune.

[0152] If the similarity of the feature vectors between the target group and other groups is high enough, even exceeding the predetermined merging threshold, these two groups can be merged into a new group to integrate resources or improve access efficiency.

[0153] After the merging is completed, it is necessary to recalculate the matching score of the newly merged group to confirm that it still matches the latency, throughput, and capacity requirements of the storage hierarchy.

[0154] For the new group or sub - group after the above splitting or merging, confirm again whether it meets the latency sensitivity requirements, capacity requirements, etc. in the storage hierarchy.

[0155] If the new group still cannot meet the requirements at the current level, it can be considered to be migrated to a storage layer with higher performance or larger capacity; and finally, update the storage mapping information of the system.

[0156] In this way, by triggering regrouping when the access conflict probability or concurrency exceeds the limit, and splitting or merging the target group with unqualified scores after regrouping, the present disclosure can split or merge high - conflict and high - concurrency groups in a timely manner through dynamic monitoring, effectively reducing the risk of access congestion in the storage hierarchy and making the system run more stably. By reasonably diverting or merging groups with severe access conflicts, over - occupation or repeated waste of storage resources by a single group can be avoided, thereby improving the overall resource utilization rate. In the iteration of multiple regroupings and subsequent splitting / merging operations, the system gradually evolves into a more reasonable grouping relationship, making the mapping between AI tasks and data features between different storage layers more scientific and accurate.

[0157] In this way, by performing regrouping operations on the target group whose access conflict probability or concurrency exceeds the second preset threshold, and splitting or merging when necessary, the storage system can actively cope with the performance challenges in the high - concurrency and high - conflict scenarios of AI tasks, maintain dynamic balance and continuously optimize the matching degree between the grouping and the storage hierarchy.

[0158] As an optional implementation manner, the value of the second preset threshold is determined according to at least one of the following methods:

[0159] The access conflict probability distribution calculated based on the historical statistical method;

[0160] Based on the upper limit of concurrency configured by the system administrator;

[0161] Dynamically adjusted through reinforcement learning or adaptive algorithms to meet the tolerance requirements in different target scenarios.

[0162] In specific implementations, after running for a period of time, data on the access conflict probability of each group at different load levels will be accumulated. For example, the conflict probability statistics of the same group during peak and off-peak periods are collected every day to form a time series or distribution function. The relationship between the concurrency and the conflict probability can also be recorded simultaneously. For example, when the concurrency is 100, 200, and 500 respectively, how the corresponding conflict probability changes. The collected conflict probability data is subjected to distribution fitting, such as normal distribution, Poisson distribution, or custom probability distribution, and key characteristic values (such as the 95th percentile, 99th percentile, or the mean plus several times the standard deviation) are calculated from it. Taking the 95th percentile as an example, if the conflict probability of a certain group is greater than this percentile, it is considered to be "higher than most historical situations" and can be used as a reference for judging whether the threshold is exceeded. The administrator can select one of the above characteristic values (such as the 99th percentile) as the second preset threshold β 2 That is, when the conflict probability of the group is monitored in real time and exceeds this percentile, the regrouping operation is triggered. In this way, based on the real operation history, the threshold can be ensured to take into account a certain degree of business security redundancy through statistical methods.

[0163] In addition, in some scenarios, the system administrator or the operation and maintenance team may directly know the concurrency limit of a single server or a certain storage pool. For example, a storage server can support a maximum of 1000 simultaneous connections, and obvious performance degradation will occur if it goes beyond this limit. Similarly, for different types of storage levels (such as cache layer, performance disk, capacity hard disk), their concurrency upper limits are also different. The administrator can directly set the second preset threshold β 2 to a certain value, such as 800, 900, or 1000, and when it exceeds this value, it is determined that the group exceeds the bearable range of the current level.

[0164] When the system detects that the concurrency of a certain group approaches or exceeds the second preset threshold β 2 , regrouping will be triggered to avoid continuous load increase leading to failures or performance crashes.

[0165] In addition, the system can introduce a controller based on reinforcement learning (RL), or use an adaptive algorithm to perform feedback learning on the real-time monitored conflict probability, concurrency, and group scoring.

[0166] Under this framework, the system will continuously try different betas 2 values, observe the impact on overall performance (such as throughput, average response time, resource utilization, etc.), and use the results as rewards or penalties for learning.

[0167] Initially, the system can set an empirical value of beta 2 as a starting point; subsequently, based on the real-time observed conflict situations and concurrent loads, fine-tune the current beta 2 such as increasing it by 5% or decreasing it by 5%.

[0168] If the system performance improves after adjustment, the algorithm increases the exploration in the current direction; if the performance deteriorates, it retreats or changes the direction. After multiple rounds of iteration, the system gradually converges to a beta 2 value that can adapt to the current business scenario, thus meeting the tolerance requirements in different scenarios.

[0169] In this way, the method for determining the "second preset threshold" based on multiple methods in the present disclosure can enable the system to flexibly adapt to the access conflict and concurrent control requirements in different AI capability fusion application scenarios, ensuring better storage performance and resource utilization efficiency in the face of load fluctuations.

[0170] Based on the same inventive concept, the present disclosure embodiments also provide a data storage system for an AI capability fusion application platform corresponding to the data storage method for an AI capability fusion application platform. Since the principle of problem-solving of the system in the present disclosure embodiments is similar to the above-mentioned data storage method for an AI capability fusion application platform in the present disclosure embodiments, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0171] Refer to Figure 3 As shown, it is a schematic diagram of the data storage system for an AI capability fusion application platform provided by the present disclosure embodiments. The system includes: a collection module 10, a grouping module 20, a first matching module 30, a second matching module 40, and a mapping module 50.

[0172] The collection module 10 is used to extract the access patterns, latency requirements, and throughput requirements of multiple AI tasks from the target AI platform, respectively forming multiple task feature vectors; extract the access frequencies, data sizes, and read / write patterns from multiple types of data in the target AI platform, respectively forming multiple data feature vectors;

[0173] The grouping module 20 is used to group multiple AI tasks according to the task feature vectors, forming several task groups; group multiple types of data according to the data feature vectors, forming several data groups;

[0174] The first matching module 30 is configured to calculate the matching scores between each pair of the task groups and the data groups respectively, and generate a first matching result; wherein, the matching score is calculated by weighting the feature similarity between the task group and the data group.

[0175] The second matching module 40 is configured to mark the group combinations with matching scores lower than a preset threshold as target group combinations in the first matching result, and perform local fine-tuning on the target group combinations to generate a second matching result; wherein, the local fine-tuning includes: splitting and reorganizing the target group combinations, and keeping the group combinations with matching scores higher than the preset threshold unchanged.

[0176] The mapping module 50 is configured to map each task group and data group to different storage levels based on the second matching result and the storage requirement information corresponding to each group combination, wherein the storage requirement information includes: latency sensitivity requirement, tolerance requirement, and capacity requirement.

[0177] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic. It should be understood that determining B based on A does not mean determining B only based on A, but also B can be determined based on A and / or other information.

[0178] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0179] In the description of this specification, the descriptions referring to terms such as "an embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

Claims

1. A data storage method for an AI capability fusion application platform, characterized in that: include: Extract access patterns, latency requirements, and throughput requirements of multiple AI tasks from the target AI platform to form multiple task feature vectors. Extract access frequency, data size, and read / write mode from multiple types of data on the target AI platform to form multiple data feature vectors; According to the task feature vector, multiple AI tasks are grouped to form a number of task groups; according to the data feature vector, multiple types of data are grouped to form a number of data groups; Calculating the matching score between each pair of the task group and the data group respectively to generate a first matching result; wherein the matching score is weightedly calculated by the feature similarity between the task group and the data group; In the first matching result, the grouping combination with a matching score lower than the preset threshold is marked as a target grouping combination, and local fine-tuning is performed on the target grouping combination to generate a second matching result; wherein the local fine-tuning includes: splitting and reorganizing the target grouping combination, and maintaining the grouping combination with a matching score higher than the preset threshold unchanged; Based on the second matching result and the storage requirement information corresponding to each group combination, each task group and data group are mapped to different storage levels, wherein the storage requirement information includes: delay sensitivity requirement, tolerance requirement, and capacity requirement.

2. The data storage method for the AI ​​capability fusion application platform according to claim 1 is characterized in that: Also includes: When the real-time characteristic value of the access load or data access mode of the AI ​​task changes relative to the initial characteristic value collected during system initialization, and the change exceeds the predetermined percentage threshold of dynamic adjustment, a dynamic adjustment operation is triggered; Recalculate the matching scores of the changed task groups and data groups, and only update the matching results of the task groups and data groups affected by the changes; For the target groups whose matching scores in the updated matching results are lower than the preset threshold, the following steps are performed on the task units or data units in the target groups: Reassign low-matching task units or data units to other groups with higher feature similarity; In response to existing groups failing to meet similarity requirements, creating new groups based on feature similarity; For the reallocated task units or data units, the matching scores are recalculated and the grouping structure is updated.

3. The data storage method for the AI ​​capability fusion application platform according to claim 2 is characterized in that: Also includes: In response to detecting that the access conflict probability or the concurrency exceeds a second preset threshold, performing a regrouping operation on the target groups mapped to the corresponding storage tier; In response to the target group still having a score lower than the preset threshold after the regrouping, the target group is split or merged with similar groups, and the storage mapping is updated after the matching score is recalculated.

4. The data storage method for the AI ​​capability fusion application platform according to claim 3 is characterized in that: The value of the second preset threshold is determined according to at least one of the following methods: The access conflict probability distribution calculated based on historical statistical methods; Based on the upper limit of concurrency configured by the system administrator; Dynamically adjust through reinforcement learning or adaptive algorithms to meet the tolerance requirements in different target scenarios.

5. The data storage method for the AI ​​capability fusion application platform according to claim 4 is characterized in that: The regrouping operation includes: Identify task units and data units mapped to target groups of corresponding storage tiers; For the units in the target group whose similarity is not enough to maintain the current mapping relationship, they are relocated to an existing group that is closer to their feature vectors; If the similarities between all existing groups and the target group are lower than a first threshold, a new group is created to accommodate the task units and data units of the target group mapped to the corresponding storage level, and the matching scores of the units are updated.

6. The data storage method for the AI ​​capability fusion application platform according to claim 5 is characterized in that: In response to the target group still having a score lower than the preset threshold after the regrouping, the triggering condition for executing the splitting or merging includes at least one of the following: The similarity between task units or data units in the target group is lower than a predetermined difference threshold; The similarity between the feature vectors of the target group and other groups exceeds a predetermined merging threshold; The access delay or conflict of the target group in the current storage layer cannot be reduced to the preset safety range.

7. The data storage method for the AI ​​capability fusion application platform according to claim 6 is characterized in that: The updating of the storage mapping after recalculating the matching score comprises: Determine whether the newly added sub-groups due to splitting or the new groups formed due to merging still match the latency sensitivity requirements and capacity requirements of the corresponding storage tier; If there is a mismatch, a local redistribution is performed between the group and the remaining groups; Finally, new matching results are generated and registered in the monitoring module.

8. A data storage system for an AI capability fusion application platform, characterized in that: include: A collection module, a grouping module, a first matching module, a second matching module, and a mapping module; The acquisition module is used to extract access modes, latency requirements, and throughput requirements of multiple AI tasks from the target AI platform to form multiple task feature vectors respectively; Extract access frequency, data size, and read / write mode from multiple types of data on the target AI platform to form multiple data feature vectors; The grouping module is used to group multiple AI tasks according to the task feature vector to form a plurality of task groups; and to group multiple types of data according to the data feature vector to form a plurality of data groups; The first matching module is used to calculate the matching score between each pair of the task group and the data group, and generate a first matching result; wherein the matching score is weightedly calculated by the feature similarity between the task group and the data group; The second matching module is used to mark the grouping combinations with matching scores lower than the preset threshold in the first matching result as target grouping combinations, and perform local fine-tuning on the target grouping combinations to generate a second matching result; wherein the local fine-tuning includes: splitting and reorganizing the target grouping combinations, and maintaining the grouping combinations with matching scores higher than the preset threshold unchanged; The mapping module is used to map each task group and data group to different storage levels based on the second matching result and the storage requirement information corresponding to each group combination, wherein the storage requirement information includes: delay sensitivity requirement, tolerance requirement, and capacity requirement.

Citation Information

Patent Citations

  • Large model-oriented data storage optimization method

    CN117931077A

  • Large model data storage method for multiple scenes

    CN118228821A

Cited By

  • Heterogeneous storage acceleration system and method for edge AI calculation

    CN122131988A

  • Heterogeneous storage acceleration system and method for edge ai computing

    CN122131988B