An Offline Data Loading Method Based on Dynamic Priority Adjustment
By applying DQN algorithm in data loading, real-time monitoring and adjustment of data block loading priority, the problem of difficult to flexibly adjust data loading priority in traditional methods is solved, and efficient data loading and system performance improvement in offline environments is achieved.
Patent Information
- Application Number
- CN202510215701.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Traditional data loading and management methods are difficult to flexibly adjust the priority of data loading according to changes in actual use, resulting in waste of resources and low loading efficiency. Especially in complex offline environments, it is impossible to effectively respond to dynamic changes, affecting the system's response speed and overall performance.
Using a method based on DQN algorithm, the access frequency of each data block in the device is monitored in real time, the comprehensive value of the data block is determined, and the order and weight of the data blocks are determined based on the comprehensive value of the data blocks in real time, so that different data blocks have different priorities.
It realizes automatic adjustment of loading priorities according to the real-time needs of the device and users, maximizes the efficiency of limited resources, ensures that the device can quickly obtain the most important data in offline scenarios, and significantly improves data loading efficiency and system performance.
Smart Images

Figure CN119690543B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital data processing, and particularly to an offline data loading method based on dynamic priority adjustment. Background Art
[0002] In recent years, with the rapid development of technologies such as the Internet of Things and artificial intelligence, various devices continuously generate a large amount of data. Especially in offline scenarios with limited resources, unstable networks, or batch processing requirements, how to efficiently and timely obtain and manage this data has become an urgent problem to be solved. In fields such as forest resource monitoring, geospatial analysis, and disaster management, the scale of remote sensing datasets is increasing day by day, bringing huge storage and processing challenges. How to reduce the data volume while ensuring data quality and improve data processing efficiency has become a key issue in current data management.
[0003] Traditional data loading and management methods often rely on preset static strategies and are difficult to flexibly adjust the priority of data loading according to changes in actual use, resulting in resource waste, low loading efficiency, and even affecting device performance and user experience. Especially in complex offline environments, traditional methods cannot effectively cope with dynamically changing requirements, thus affecting the system's response speed and overall performance. For example, the invention patent application with the application number 202411388666.9 uses a fixed grid model that cannot be dynamically updated, affecting accuracy. Especially when processing high-resolution remote sensing data, the data processing speed is slow.
[0004] DQN (Deep Q-Network) is an algorithm that combines deep learning and reinforcement learning and is used to solve control problems in Markov Decision Process (MDP). It is based on the idea of value iteration and guides the agent to select the best action in each state by estimating the value function (Q-value function) of each state-action pair. Two key networks are used in the DQN algorithm: the main network and the target network. These two networks are essentially deep neural networks, but they have different functions: the main network is used to predict the expected reward, that is, the Q-value, given a state (s) and an action (a). The Q-value represents the cumulative reward expected by the agent after taking an action. The main network learns through interaction with the environment and continuously updates its parameters to better predict the Q-value. The target network is an auxiliary network, and its parameters (θ') are copies of the main network's parameters (θ). The purpose of the target network is to provide a stable training target for the DQN algorithm. It uses the same structure as the main network, but the weights are updated more slowly. The target network usually directly copies the parameters of the main network after a certain number of training steps. Summary of the Invention
[0005] The object of the present invention is to provide a method for solving the above problems, breaking the traditional data loading mechanism based on static rules, and being able to automatically adjust the loading priority according to the real-time requirements of devices and users, namely, an offline data loading method based on dynamic priority adjustment.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows: An offline data loading method based on dynamic priority adjustment, comprising the following steps;
[0007] S1, the user accesses the overall data through the device , where the data space occupied by the i-th data block is , the available space of the device to access the overall data D is C, 1 ≤ i ≤ N;
[0008] S2, define the comprehensive value of ;
[0009] ,
[0010] ,
[0011] ,
[0012] wherein, is a weight parameter, and , is the expected access probability of is the importance probability of is the importance score of is the unaccessed status score of , are respectively the mean and standard deviation of N importance scores;
[0013] S3, create the data block loading process as a Markov decision process, and define the state, action, and reward;
[0014] The state includes the set of indexes of the loaded data blocks and the remaining available space of the device;
[0015] The action is to select the next loadable data block, and the loadable data block is a data block that has not been loaded and the occupied data space is less than the remaining available space;
[0016] The reward is the comprehensive value of the loadable data block selected in the action;
[0017] S4. Based on the DQN algorithm, a main network and a target network are established. The network structures of the two are the same. The input of the main network is the state, and the output is the Q - function value of the action executed in this state.
[0018] S5. At each time step, an action is executed to generate an experience. The experience at time step t is , and are the states at time steps t and t + 1 respectively, is the state the action that maximizes the Q - function value, is the comprehensive value of the loadable data block corresponding to it;
[0019] The state and the action The corresponding Q - function value is obtained according to the following formula;
[0020] ,
[0021] In the formula, T is the total number of preset time steps, represents the expected calculation in the state and the action , is the discount factor, k is the time step k, t ≤ k ≤ T, the comprehensive value corresponding to the time step k;
[0022] S6. Preset the number of iteration rounds and the update time interval of the target network , train the main network based on experience replay, update the parameters of the main network by minimizing the loss function L, and when reaching , update the parameters of the target network with the parameters of the main network until the number of iteration rounds is reached, and obtain the trained main network and target network;
[0023] S7. When the user accesses the overall data, use the trained main network to obtain the current state, calculate the Q - function values of each action in this state, and select the action with the largest Q - function value to execute.
[0024] As an optimization: and are obtained according to the following method;
[0025] S21. Preset the initial importance score and the initial unvisited state score , and for all data blocks, the and are the same;
[0026] S22. When the user accesses the overall data D multiple times, count The number of times accessed by the user and the number of times not accessed and update according to the following formula the importance score of and the non - access status score ;
[0027] ,
[0028] .
[0029] Preferably: In step S4, generating the experience at time step t includes steps S41 - S43;
[0030] S41, obtain the state at time step t , , where is the index set of the data blocks already loaded at time step t, is the remaining available space of the device at time step t;
[0031] S42, execute multiple actions when in the state , mark the action with the maximum Q - function value as , and mark the comprehensive value of the data blocks that can be loaded in as ;
[0032] S43, generate the experience at time step t , where is the state at time step t + 1 .
[0033] Preferably: In S6, the loss function L is obtained according to the following formula;
[0034] ,
[0035] ,
[0036] In the formula, y is the target value of the state and the action , is the maximum value of the Q - function value in the target network when the state is and the action is .
[0037] Preferably: The overall data is an image set composed of forestry remote - sensing images.
[0038] Compared with the prior art, the advantages of the present invention are as follows: Based on the DQN algorithm, the present invention determines the comprehensive value of data blocks by real-time monitoring the access frequencies of various data blocks in the device, so as to evaluate the importance of data blocks; and determines the loading order and weight of data blocks according to the real-time comprehensive value of data blocks, so that different data blocks have different priorities; finally, through the dynamic allocation of priorities for different data blocks, the use efficiency of limited resources is maximized, ensuring that the device can quickly obtain the most important data in the offline scenario.
[0039] Therefore, the present invention breaks the traditional data loading mechanism based on static rules and adopts an adaptive optimization strategy, which can automatically adjust the loading priority according to the real-time needs of the device and users.
[0040] Through this method, data management no longer depends on fixed loading rules, but can flexibly cope with various complex and dynamic application scenarios. Especially for large-scale data sets such as forestry remote sensing data, this method can significantly improve the data loading efficiency, reduce the system response time, improve the overall performance of the device, and reduce the waste of resources caused by overloading unimportant data in the offline environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0042] The present invention will be further described below in conjunction with embodiments.
[0043] Embodiment 1: Refer to Figure 1 , an offline data loading method based on dynamic priority adjustment, comprising the following steps;
[0044] S1, the user accesses the overall data through the device , where the i-th data block occupies a data space of 、the available space of the device to access the overall data D is C, 1 ≤ i ≤ N;
[0045] S2, define the comprehensive value of ;
[0046] ,
[0047] ,
[0048] ,
[0049] where, is a weight parameter, and , is The expected access probability of is The importance probability of is The importance score of is The unvisited status score of , where e is the natural constant. and are the mean and standard deviation of N importance scores respectively;
[0050] S3. Create the data block loading process as a Markov decision process, and define the state, action, and reward;
[0051] The state includes the set of loaded data block indices and the remaining available space of the device;
[0052] The action is to select the next loadable data block, and the loadable data block is a data block that has not been loaded and whose occupied data space is less than the remaining available space;
[0053] The reward is the comprehensive value of the loadable data block selected in the action;
[0054] S4. Based on the DQN algorithm, establish a main network and a target network. The network structures of the two are the same. The input of the main network is the state, and the output is the Q function value of the action executed in this state;
[0055] S5. Execute actions at each time step to generate experiences. The experience at time step t is , and are the states at time steps t and t + 1 respectively, is the state The action that maximizes the Q function value, is The comprehensive value of the corresponding loadable data block;
[0056] The state The action The corresponding Q function value Is obtained according to the following formula;
[0057] ,
[0058] In the formula, T is the total number of preset time steps, Indicates the expected calculation in the state and the action , is the discount factor, k is the time step k, t ≤ k ≤ T, The comprehensive value corresponding to the time step k;
[0059] S6. Preset the number of iteration rounds and the update time interval of the target network , train the main network based on experience replay to update the parameters of the main network by minimizing the loss function L, and when reaching , update the parameters of the target network with the parameters of the main network until the number of iteration rounds is reached, obtaining the trained main network and target network;
[0060] S7. When the user accesses the overall data, use the trained main network to obtain the current state, calculate the Q-function values of each action in this state, and select the action with the largest Q-function value to execute. Since the action in the present invention is to select the next loadable data block, executing the action with the largest Q-function value means loading the loadable data block with the largest Q-function value, thus realizing the function of dynamically selecting data blocks based on dynamic priorities in the present invention.
[0061] In this embodiment, 、 is obtained according to the following method;
[0062] S21. Preset the initial importance score 、initial unvisited state score , and the 、 of all data blocks are the same;
[0063] S22. When the user accesses the overall data D multiple times, count the number of times the is accessed by the user 、the number of times not accessed , and update the importance score 、unvisited state score according to the following formula;
[0064] ,
[0065] .
[0066] In step S4, generating the experience at time step t includes steps S41~S43;
[0067] S41. Obtain the state at time step t , , where is the index set of the data blocks already loaded at time step t, is the remaining available space of the device at time step t;
[0068] S42. Execute multiple actions in the state , mark the action with the largest Q-function value as , and The comprehensive value of the loadable data block in is marked as ;
[0069] S43, generate the experience at time step t , where is the state at time step t+1 .
[0070] In S6, the loss function L is obtained according to the following formula;
[0071] ,
[0072] ,
[0073] In the formula, y is the target value of the state , action . is in the target network, the state , action when the maximum value of the Q function value.
[0074] The overall data is an image set composed of forestry remote sensing images.
[0075] Based on the method of the present invention, when a user accesses the overall data, the loading priority and strategy can be automatically adjusted according to real-time feedback and monitoring requirements, improving the flexibility of the model. By preferentially loading key data blocks, redundant calculations are reduced and the processing speed is improved.
[0076] Embodiment 2: In the DQN algorithm, the deep Q network includes a main network and a target network. A specific process for training the deep Q network given in this embodiment is as follows:
[0077] (1) Initialization:
[0078] Environment: Select and initialize the environment to be trained, such as OpenAI Gym;
[0079] Neural network: Initialize two neural network models, the main network and the target network, which have the same structure;
[0080] Experience replay buffer: Used to store experiences;
[0081] Parameters: Set The initial importance score of , the initial unvisited state score , the discount factor , the learning rate a, the exploration rate , other hyperparameters, etc.; in this embodiment , , =0.9.
[0082] (2) Training is performed at each time step. The training at time step t includes the following (2.1) to (2.6);
[0083] (2.1) At each time step t, select an action according to the current state, using - ε-greedy policy: with probability randomly select an action, and with probability select the action with the largest Q-function value as ;
[0084] (2.2) Execute the action and observe the reward returned by the action and the next state ;
[0085] (2.3) Store the experience in the experience replay buffer;
[0086] (2.4) Experience replay: randomly sample a small batch of experiences from the experience replay buffer;
[0087] (2.5) For each experience, calculate the Q-function value;
[0088] (2.6) Use this small batch of experiences to update the main network parameters, and when reaching (such as 500, 1000 time steps), copy the main network to the target network.
[0089] (3) Update the exploration rate;
[0090] (4) Check the termination conditions, such as: reaching the preset time step, or the remaining available space is not enough to load any unloaded data blocks, then end the training.
[0091] (5) The data blocks selected at each time step form the final policy in sequence, and according to this final policy, load the selected set of data blocks.
[0092] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. An offline data loading method based on dynamic priority adjustment, characterized in that: The steps include: S1, users access overall data through devices , where the i-th data block The data space occupied is , the available space for the device to access the overall data D is C, 1≤i≤N; S2, Definition The comprehensive value ; , , , in, is the weight parameter, and , for The expected access probability of for The probability of importance, for Importance rating, for The unvisited state score of , e is a natural constant, , are the mean and standard deviation of N importance scores respectively; S3, creates the data block loading process as a Markov decision process, defining states, actions, and rewards; The status includes a set of loaded data block indexes and remaining available space on the device; The action is to select the next loadable data block, wherein the loadable data block is a data block that has not been loaded and occupies a data space smaller than the remaining available space; The reward is the comprehensive value of the loadable data block selected in the action; S4, based on the DQN algorithm, a main network and a target network are established, and the network structures of the two are the same. The input of the main network is the state, and the output is the Q function value of the action executed in the state; S5, performs actions at each time step to generate experience, where the experience at time step t is , , are the states at time steps t and t+1 respectively, Status The action that maximizes the Q function value, for The comprehensive value of the corresponding loadable data block; state ,action The corresponding Q function value According to the following formula: , Where T is the total number of preset time steps, Indicates in status and actions The expected calculation is performed below. is the discount factor, k is the time step k, t≤k≤T, The comprehensive value corresponding to time step k; S6, preset iteration rounds, target network update time interval , based on experience replay training main network, to minimize the loss function L to update the main network parameters, and when reaching When , the parameters of the target network are updated with the parameters of the main network until the iteration round is reached, and the trained main network and target network are obtained; S7, the user accesses the overall data, uses the trained main network to obtain the current state, calculates the Q function value of each action in this state, and selects the action with the largest Q function value to execute.
2. The offline data loading method based on dynamic priority adjustment according to claim 1, characterized in that: , Obtained according to the following method; S21, preset The initial importance score of , Initial unvisited status score , and all data blocks , same; S22, when the user accesses the overall data D multiple times, statistics Number of times visited by users , the number of times not visited , and update according to the following formula Importance Rating , Unvisited Status Score ; , 。 3. The offline data loading method based on dynamic priority adjustment according to claim 1, characterized in that: In step S4, generating experience at time step t includes steps S41 to S43; S41, get the state of time step t , ,in is the index set of the loaded data blocks at time step t, is the remaining available space of the device at time step t; S42, in state Execute multiple actions at the same time, and mark the action with the largest Q function value as ,Will The comprehensive value of the loadable data block is marked as ; S43, generate experience at time step t ,in is the state at time step t+1 .
4. The offline data loading method based on dynamic priority adjustment according to claim 1, characterized in that: In S6, the loss function L is obtained according to the following formula: , , In the formula, y is the state ,action The target value of In the target network, the state ,action The maximum value of the Q function when .
5. The offline data loading method based on dynamic priority adjustment according to claim 1, characterized in that: The overall data is an image set consisting of forestry remote sensing images.
Citation Information
Patent Citations
A method for mapping forestry resources based on remote sensing data
CN118941743B
H5 page loading method and device based on artificial intelligence, equipment and medium
CN115080147A
AI-based environment resource dynamic loading method in real-time cloud rendering
CN115460254A