Data acquisition efficiency optimization method based on artificial intelligence and reinforcement learning

By acquiring data in real time in a dynamic environment and generating new action strategies and adjusting data acquisition parameters, the problem of poor adaptability of deep reinforcement learning models in a dynamic environment is solved, and efficient data acquisition and system performance improvement is achieved.

CN120197527AActive Publication Date: 2025-06-24SHENZHEN HANLEY TECH CO LTD

Patent Information

Application Number
CN202510683165.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

In dynamic environments, existing deep reinforcement learning models have reduced data acquisition efficiency due to poor adaptability, especially in scenarios where multitasking goals or resource constraints, it is difficult to quickly respond to environmental changes.

Method used

The initial data acquisition model is constructed through a deep reinforcement learning algorithm, and the environment state data is obtained in real time in the dynamic environment, new action strategies are generated based on current and historical data, data acquisition parameters are adjusted, and model parameters are continuously iteratively updated to adapt to environmental changes.

Benefits of technology

This method can respond to environmental changes more quickly, maintain efficient data acquisition capabilities, reduce the possibility of policy failure, improve system stability and reliability, adapt to the needs of complex scenarios, and improve overall system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197527A_ABST
    Figure CN120197527A_ABST
Patent Text Reader

Abstract

The invention discloses a data acquisition efficiency optimization method based on artificial intelligence and reinforcement learning, and relates to the technical field of artificial intelligence and data processing, and the method comprises the steps: S1, constructing an initial data acquisition model through a deep reinforcement learning algorithm; s2, acquiring environment state data in real time in a dynamic environment; s3, generating a new action strategy based on the current environment state data and the historical data; S4, adjusting data acquisition parameters according to the generated action strategy; and S5, continuously iterating and updating model parameters to adapt to a new environment state, the data acquisition efficiency optimization method based on artificial intelligence and reinforcement learning can significantly improve the performance of the model in long-term operation, reduces performance reduction caused by neglecting long-term dynamic characteristics in a traditional method, and improves the efficiency of data acquisition. The problem of poor model adaptability caused by rapid and frequent state change in a dynamic environment in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and data processing, and particularly relates to a method for optimizing data collection efficiency based on artificial intelligence and reinforcement learning. Background Art

[0002] Data collection in a dynamic environment is an important part of many intelligent systems, and its main purpose is to support the training, optimization, and decision-making of the system by collecting high-quality data. In this process, deep reinforcement learning methods are widely used because they can automatically adapt to the changes in complex environments through trial-and-error learning. However, due to the rapid and frequent state changes in the dynamic environment, traditional deep reinforcement learning models often have poor adaptability.

[0003] Specifically, when existing deep reinforcement learning algorithms perform data collection in a dynamic environment, they mainly rely on fixed policy update mechanisms and static model parameters. Facing the frequently changing environmental states, this method is prone to situations such as lagging model updates or ineffective policies, resulting in a decrease in data collection efficiency. Especially in scenarios involving multiple task objectives or resource constraints, existing methods are difficult to quickly respond to changes, thereby affecting the data collection effect and system performance. In addition, in a dynamic environment, traditional reinforcement learning models usually ignore the long-term dynamic characteristics of the environment and only learn based on short-term observation results, which makes the models perform poorly in long-term adaptability. For example, in fields such as intelligent monitoring and real-time optimization, the rapidly changing environmental states pose higher requirements for real-time performance and robustness of the models, and existing technologies are difficult to meet these needs. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for optimizing data collection efficiency based on artificial intelligence and reinforcement learning, so as to solve the problem of poor model adaptability in the existing technology due to rapid and frequent state changes in a dynamic environment.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for optimizing data collection efficiency based on artificial intelligence and reinforcement learning, the method comprising: S1. Construct an initial data collection model through a deep reinforcement learning algorithm; S2. Real-time obtain environmental state data in a dynamic environment; S3. Generate a new action policy based on the current environmental state data and historical data. Specifically, map the current environmental state data and historical data to a policy combination, generate a logical next action policy in the policy set, and combine with the reward function in deep reinforcement learning to evaluate and select the generated policy. The formula for generating the next action policy is: ; Wherein, Indicates the new round of action strategy set generated after the transposition operation After that, Indicates at the th iteration, the data collection action strategy set Indicates the transposition logic operation function, which is used to perform logical adjustment on the current action strategy set For logical adjustment, Indicates the current time step of the policy iteration; S4. Adjust the data collection parameters according to the generated action strategy; S5. Continuously iterate and update the model parameters to adapt to the new environmental state. Specifically, represent the parameters of the reinforcement learning model as complex numbers, iterate to generate the distribution of the model parameters, simulate the complex process of the model adapting to the rapidly changing environment, analyze whether the parameter distribution converges or diverges, adjust the iteration step size, so that the model gradually approaches the optimal parameter configuration. The specific formula for iteratively generating the distribution of the model parameters is: ; Among them, Indicates the complex representation of the current model parameters, Indicates the model parameters generated in the next iteration, Indicates the complex parameters corresponding to the environmental state, Indicates the number of iterations.

[0006] Preferably, the S1 includes: Convert the data collection task characteristics into an image form, construct an input tensor, design a multi-layer convolutional neural network model, extract features and biases with convolutional kernels, and enhance the non-linear expression ability through activation functions. The specific formula is: ; Among them, Indicates the output of the th layer of neurons, Indicates the output of the -1th layer of neurons, Indicates the index of the network layer, Indicates the th layer of weight matrix, Indicates the th layer of bias vector, Indicates the activation function;

[0007] Use the data collection simulation data to supervise and train the model, optimize the initial strategy, and generate an action strategy prediction model for data collection.

[0008] Preferably, the S2 includes: Model the state variables of the dynamic environment, real-time simulate the change trend of the environmental state over time, and dynamically adjust the timing and frequency of data collection according to the simulation results. The specific formula for simulating the change trend of the environmental state over time is: , ; Among them, represents the amount of resources in a dynamic environment, represents the activity intensity of the acquisition device, represents time, represents the natural growth rate of resources, represents the consumption rate of the acquisition device on the amount of resources, represents the improvement coefficient of the acquisition efficiency, represents the resource consumption of the acquisition device.

[0009] Preferably, the S4 includes: Optimally model the optimization of data acquisition parameters as a constrained optimization problem, transform the constrained problem into an unconstrained optimization, solve the optimal solution of parameter adjustment, and adjust the acquisition parameters in real time according to the optimization results to ensure that the acquisition efficiency remains optimal when the environment changes. The specific formula for solving the optimal solution of parameter adjustment is: ; Among them, represents the set of data acquisition parameters, represents the weight coefficient of the constraint condition on the objective function, represents the th weight coefficient of the resource constraint condition on the objective function, represents the objective function of data acquisition efficiency, represents the th constraint condition in data acquisition, represents the number of the constraint condition, represents the total number of resource constraint conditions involved in the data acquisition system, represents the objective function constructed in the Lagrangian dual optimization.

[0010] Preferably, the S3 further includes: Collect the environmental state data and the environmental state data of the previous round; Input the current environmental state data and historical data into the initial data acquisition model; Generate a new action strategy based on the predicted value output by the model; Select the final action through the action probability output by the model. If the probability P of the current action is greater than or equal to the set threshold T, then select the current action, otherwise select a random action, where P represents the action probability and T represents the threshold.

[0011] Preferably, the generating a new action strategy based on the predicted value output by the model includes: Obtain the model predicted value V(s) under the current environmental state; Calculate the weighted average prediction value W = α × V(s) + (1 - α) × V according to historical data prev , where α represents the weighting coefficient, V(s) represents the prediction value in the current environmental state, and V prev represents the prediction value of the previous state; Set a preset value T W , and generate an action strategy based on the weighted average prediction value. If the weighted average prediction value W is greater than or equal to the preset value T W , then adopt this action strategy.

[0012] Preferably, the S1 further includes: In the convolutional neural network model, use convolutional kernels of different sizes in the convolutional layer to extract multi-scale features of the target distribution. Add a pooling layer after the convolutional layer to reduce the feature dimension through max pooling or average pooling. Integrate the local features extracted by the convolutional and pooling layers in the fully connected layer to generate global acquisition strategy parameters.

[0013] Preferably, the S2 further includes: Initialization of the resource amount in the dynamic environment. Obtain the initial resource distribution through historical data and environmental sensors, and update the parameters in real time according to the environmental state during the modeling process 、 、 and .

[0014] Preferably, the objective function of the data acquisition efficiency in the S4 is: = c - d - e; where c represents the acquisition rate, d represents the energy consumption, and e represents the data loss rate.

[0015] Preferably, the convolutional layer of the convolutional neural network model adopts dilated convolution.

[0016] From the above technical solutions, it can be seen that the present invention has the following beneficial effects: The method for optimizing data acquisition efficiency based on artificial intelligence and reinforcement learning constructs an initial data acquisition model through a deep reinforcement learning algorithm, obtains environmental state data in real time in a dynamic environment, generates a new action policy based on the current environmental state data and historical data, adjusts the data acquisition parameters according to the generated action policy, and continuously iteratively updates the model parameters to adapt to the new environmental state. It can respond more quickly to the frequent changes in the environment, thereby maintaining high-efficiency data acquisition capabilities in a dynamic environment, reducing the possibility of policy failure, improving the stability and reliability of the system in complex scenarios, maximizing data acquisition efficiency under limited resource conditions, enhancing the overall system performance, adapting to the requirements of various complex scenarios such as intelligent monitoring and real-time optimization, significantly improving the performance of the model in long-term operation, and reducing the performance degradation caused by traditional methods ignoring long-term dynamic characteristics. It solves the problem of poor model adaptability in the existing technology due to rapid and frequent state changes in a dynamic environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] As Figure 1 shown, the present invention provides a technical solution: a method for optimizing data acquisition efficiency based on artificial intelligence and reinforcement learning, the method comprising: S1. Construct an initial data acquisition model through a deep reinforcement learning algorithm; S2. Obtain environmental state data in real time in a dynamic environment; S3. Generate a new action policy based on the current environmental state data and historical data. Specifically, map the current environmental state data and historical data to a policy combination, generate a logical next action policy in the policy set, and evaluate and select the generated policy in combination with the reward function in deep reinforcement learning. The formula for generating the next action policy is: ; wherein, represents a new round of action policy set generated after the transposition operation , represents the data acquisition action policy set at the th iteration, Represents a transposition logic operation function for the current action policy set to perform logical adjustment Represents the current time step of policy iteration; S4. Adjust the data acquisition parameters according to the generated action policy; S5. Continuously iterate and update the model parameters to adapt to the new environmental state. Specifically, represent the parameters of the reinforcement learning model in complex number form, iterate to generate the distribution of model parameters, simulate the complex process of the model adapting to a rapidly changing environment, analyze whether the parameter distribution converges or diverges, adjust the iteration step size, so that the model gradually approaches the optimal parameter configuration. The specific formula for iteratively generating the distribution of model parameters is: ; Among them, represents the complex number representation of the current model parameters, represents the model parameters generated in the next iteration, represents the complex number parameters corresponding to the environmental state, represents the number of iterations.

[0020] The present invention optimizes the data acquisition efficiency through a deep reinforcement learning algorithm, dynamically generates action policies based on the environmental state and historical data, and combines a reward function to realize the selection and optimization of policies. During the data acquisition process, the data acquisition efficiency is continuously improved by adjusting the data acquisition parameters in real time. By describing the model state in complex number parameter form and performing iterative updates, it can more quickly adapt to complex and changing environmental states, and gradually optimize the model in a dynamic environment. In a specific implementation, the environmental state and historical data are mapped to a policy set, and the next action policy in the policy set is generated according to the logical operation function and the pros and cons of the policy are evaluated using the reward function of deep reinforcement learning. In order to adapt to dynamic environmental changes, the present invention further describes the model parameters in complex number form, generates new parameters through complex number mapping relationships and iterative formulas, thereby dynamically optimizing the model performance. Through the real-time policy optimization of the deep reinforcement learning algorithm, the data acquisition efficiency is significantly improved, which is applicable to complex and changing dynamic environments. Based on the parameter representation and iterative update method in complex number form, it can quickly respond to environmental state changes and adjust model parameters to ensure an efficient data acquisition process. The complex number parameter iterative optimization method can simulate the convergence or divergence state of the model in a complex environment, improve the robustness and adaptability of the model. The present invention reduces the need for manual intervention and the operation complexity by dynamically adjusting the data acquisition parameters, and for the first time combines the complex number parameter form with deep reinforcement learning, providing an innovative technical solution for data acquisition optimization in a dynamic environment.

[0021] S1 includes converting the data acquisition task features into an image form, constructing an input tensor, designing a multi-layer convolutional neural network model, extracting features and biases with convolutional kernels, and enhancing the non-linear expression ability through activation functions. The specific formula is as follows: ; Among them, represents the output of the -th layer of neurons, represents the output of the -th layer of neurons, represents the index of the network layer, represents the -th layer weight matrix, represents the -th layer bias vector, represents the activation function; Use the data acquisition simulation data to supervise and train the model, optimize the initial strategy, and generate an action strategy prediction model for data acquisition.

[0022] When constructing the data acquisition model in the present invention, the task features are efficiently extracted and modeled through the convolutional neural network in deep learning. First, the data acquisition task features are converted into an image form suitable for neural network processing, and an input tensor is constructed for network training. By designing a multi-layer convolutional neural network and combining the characteristics of convolutional kernels, the local patterns of data features can be extracted and the spatial relationships can be captured. The addition of the bias term is used to enhance the sensitivity of the model to the input data, and the introduction of the activation function enhances the non-linear expression ability of the model, enabling it to handle complex feature mapping relationships. In specific implementation, the output of the convolutional layer is represented by the formula , where the weight matrix and bias vector are continuously optimized through training and adjusted according to specific task requirements. Through the supervised learning method, the simulation data is used to optimize the parameter configuration and initialization strategy of the network. Finally, a model that can predict the data acquisition action strategy is generated to guide the actual data acquisition task. Through the local receptive field mechanism of the convolutional kernel, the spatial patterns of the data acquisition task features are effectively captured, improving the efficiency and accuracy of feature extraction. The activation function is used to enhance the non-linear expression ability of the network, enabling it to better handle complex task features and optimize the prediction performance of the action strategy. Through the supervised training of the simulation data, the optimization of the model parameters is achieved, enabling it to efficiently handle various data acquisition task scenarios. The weight sharing mechanism of the convolutional network reduces the number of model parameters, lowers the training complexity, and improves the calculation efficiency at the same time. The optimized model can generate more accurate action strategy predictions, thereby improving the data acquisition efficiency and task completion rate.

[0023] S2 includes modeling the state variables of the dynamic environment, simulating the changing trend of the environmental state over time in real time, and dynamically adjusting the timing and frequency of data collection according to the simulation results. The specific formula for simulating the changing trend of the environmental state over time is as follows: , ; Among them, represents the amount of resources in the dynamic environment, represents the activity intensity of the collection device, represents time, represents the natural growth rate of resources, represents the consumption rate of the amount of resources by the collection device, represents the promotion coefficient of the collection efficiency, represents the resource consumption of the collection device.

[0024] The present invention models the state variables of the dynamic environment, simulates the dynamic interaction between resources and devices, so as to optimize the timing and frequency of data collection. In the formula , the natural growth of the amount of resources over time is represented by αx, while the consumption rate of resources by the collection device is described by −βxy. At the same time, the activity intensity y of the collection device is driven by the amount of resources x, and its promotion rate is represented by δxy, and the consumption rate is represented by −γy. The formula is . Through the above system of equations, the coupling relationship between resource growth and the behavior of the collection device can be dynamically simulated. When the amount of resources is sufficient, the activity intensity y of the collection device will increase with the increase of resources x, thereby improving the collection efficiency. When the amount of resources gradually decreases, the device activity intensity will decrease due to insufficient resources, thus avoiding resource exhaustion caused by over-collection. In specific implementation, by adjusting the collection frequency y of the collection device in real time, the efficiency of data collection can be optimized to ensure that the collection process is both efficient and sustainable. In addition, this method dynamically predicts the best time window for data collection by simulating the changing trend of the environmental state, thereby avoiding collection delay or resource waste. By dynamically simulating the interaction relationship between resources and collection devices, the collection timing and frequency can be determined more accurately, significantly improving the data collection efficiency, avoiding resource exhaustion caused by over-collection, realizing the efficient utilization and sustainability of resources through real-time adjustment of the device activity intensity, and being able to quickly respond to the changing trend of the dynamic environment and adapt to complex and changeable scenarios by using state variable modeling and real-time simulation technology. The present invention describes the state change of the dynamic environment through a clear mathematical model, provides a scientific basis for the optimization of the collection process, improves the credibility and universality of the solution, reduces unnecessary collection operations by optimizing the collection frequency and intensity, and reduces resource waste and energy consumption.

[0025] S4 includes optimizing the data acquisition parameters by modeling them as a constrained optimization problem, transforming the constrained problem into an unconstrained optimization, solving for the optimal solution of parameter adjustment, and adjusting the acquisition parameters in real time according to the optimization results to ensure that the acquisition efficiency remains optimal when the environment changes. The specific formula for solving the optimal solution of parameter adjustment is as follows: ; Wherein, represents the set of data acquisition parameters, represents the weight coefficient of the constraint condition on the objective function, represents the th weight coefficient of the resource constraint condition on the objective function, represents the objective function of data acquisition efficiency, represents the th constraint condition in data acquisition, represents the number of the constraint condition, represents the total number of resource constraint conditions involved in the data acquisition system, represents the objective function constructed in the Lagrangian dual optimization.

[0026] In the present invention, the optimization problem of data acquisition parameters is formalized as a constrained optimization problem, and the Lagrangian dual optimization method is used to solve the optimal parameter configuration. During the acquisition process, the goal is to optimize the data acquisition efficiency while satisfying various constraint conditions (such as device resource limitations, power consumption limitations, etc.). Formula expresses the core idea of the optimization, where the objective function represents the maximization of the acquisition efficiency, and the constraint condition represents the resource limitations or performance requirements that need to be satisfied during the operation of the system. By introducing the Lagrangian multiplier , the original constrained problem is transformed into an unconstrained optimization problem, and the gradient descent method or other optimization algorithms are used to solve it. During the optimization process, the parameter set is adjusted in real time according to the dynamic changes of the environment. For example, when some resource constraint conditions become more stringent, the corresponding weight coefficient It will be dynamically adjusted to preferentially satisfy more critical constraint conditions while maintaining relative optimization of data acquisition efficiency. The optimization results are directly used to adjust the parameter configuration of the acquisition device to ensure that the system operates at the highest efficiency in a dynamic environment. By solving the Lagrangian optimization problem, the data acquisition parameter configuration can be dynamically optimized to improve the system acquisition efficiency. By formalizing various constraint conditions and dynamically adjusting the weights, it ensures the efficient execution of data acquisition tasks under complex resource constraints. When the environment changes, it can quickly adjust the acquisition parameters to ensure that the system continues to operate in the optimal state. The use of scientific solution methods such as Lagrangian optimization and gradient descent provides a theoretical support for the adjustment of data acquisition parameters. The optimization process is efficient and reliable, balancing the relationship between the objective function and resource constraint conditions, achieving the optimal matching of acquisition efficiency and resource consumption, and avoiding resource waste or efficiency decline.

[0027] S3 also includes collecting the current environmental status data and the environmental status data of the previous round; inputting the current environmental status data and historical data into the initial data acquisition model; generating a new action strategy based on the predicted value output by the model; and selecting the final action through the action probability output by the model. If the probability P of the current action is greater than or equal to the set threshold T, the current action is selected; otherwise, a random action is selected, where P represents the action probability and T represents the threshold.

[0028] The present invention optimizes the decision-making process of action policy generation by introducing the action probability mechanism in reinforcement learning. First, the current environmental state data and the environmental state data of the previous round are collected. By inputting these data into the initial data collection model, the model analyzes the current environment and outputs corresponding predicted values. The predicted values are used to generate a new action policy, and a probability value P is assigned to each possible action. In the selection of the final action, the action probability value P output by the model is used to determine whether to execute the current action. If the probability P of an action is greater than or equal to the set threshold T, it indicates that the model has a high confidence in this action, and at this time, this action is directly selected for execution; conversely, if P < T, a random action is selected, thereby providing the possibility of exploration for the model and avoiding the local optimum problem caused by premature convergence. This decision-making method combines the "exploration-exploitation" strategy in reinforcement learning, which not only ensures the execution efficiency of high-probability actions but also increases the diversity of decisions through random actions, improving the overall model performance and data collection efficiency. The selection mechanism based on action probability can execute high-probability actions more reasonably, significantly improving the data collection efficiency. By adding random selection to low-probability actions, the exploration ability of the model is increased, avoiding the model from falling into the local optimum solution prematurely. Using the predicted values of the model to generate the probability distribution of the action policy makes the decision-making process more intelligent and has a stronger ability to adapt to dynamic environmental changes. By adjusting the threshold T, the balance point between exploration and exploitation can be flexibly controlled to meet the requirements of different data collection scenarios. Combining the current environmental state and historical data enhances the adaptability of the model to complex environments and improves the stability of the collection efficiency.

[0029] Generating a new action policy based on the predicted values output by the model includes obtaining the model predicted value V(s) in the current environmental state; calculating the weighted average predicted value W = α×V(s)+(1−α)×V prev , where α represents the weighting coefficient, V(s) represents the predicted value in the current environmental state, and V prev represents the predicted value of the previous state; setting a preset value T W , and generating an action policy based on the weighted average predicted value, where if the weighted average predicted value W is greater than or equal to the preset value T W , then this action policy is adopted.

[0030] The present invention dynamically optimizes the generation process of the action policy by combining the predicted values of the current environmental state and historical data. First, the system calculates the predicted value V(s) of the current environmental state through a deep reinforcement learning model, and this predicted value reflects the potential benefits of taking specific actions in the current state. To improve the robustness of decision-making, the influence of historical data is further introduced, and through the weighted average formula W = α×V(s)+(1−α)×V prev for the current predicted value V(s) and the predicted value V of the previous state prevPerform fusion. The weighting coefficient α is used to balance the influence of the current state and historical data. When α is large, the decision-making is more inclined to rely on the predicted value of the current state; conversely, when α is small, the influence of historical data on the decision-making is more significant. Then, according to the set preset value T W Perform policy selection. If the weighted average prediction value W is greater than or equal to the preset value T W , it indicates that the current action policy is favorable under the comprehensive evaluation of historical and current states, and the system directly adopts this action policy; otherwise, the system selects other actions according to the preset mechanism. In this way, the relationship between real-time response and long-term policy optimization is effectively balanced. By introducing the weighted average prediction value, combining the influence of the current state and historical data, the impact of single-state data fluctuations on decision-making is reduced, and the stability and reliability of action policy generation are improved. According to the comparison between the weighted prediction value and the preset value, the optimal action policy can be selected more accurately, the data acquisition efficiency can be improved, and the predicted value V of the previous state can be fully utilized prev , providing additional context information for the decision-making process, which helps to cope with complex dynamic environments. By adjusting the weighting coefficient α and the preset value T W , the weights of historical data and current predicted values can be flexibly changed to adapt to the requirements of different application scenarios. Making decisions by combining historical and current data can reduce the impact of model prediction uncertainty on action selection and enhance the robustness and anti-interference ability of the system.

[0031] S1 also includes using different-sized convolutional kernels in the convolutional layer of the convolutional neural network model to extract multi-scale features of the target distribution, adding a pooling layer after the convolutional layer to reduce the feature dimension through max pooling or average pooling, and integrating the local features extracted by the convolutional and pooling layers in the fully connected layer to generate global acquisition policy parameters.

[0032] Through an improved convolutional neural network model, the present invention further optimizes the feature extraction and policy generation processes in data acquisition tasks. First, in the convolutional layer, convolutional kernels of different sizes (such as 3×3, 5×5, 7×7, etc.) are used to capture multi-scale features of the target distribution. Smaller convolutional kernels can extract detailed features, while larger convolutional kernels can capture broader context information. By combining multiple convolutional kernels, the model can simultaneously focus on local details and global patterns, thereby more comprehensively characterizing the characteristics of the target distribution. Then, a pooling layer is added after the convolutional layer, and max pooling or average pooling operations are used to reduce the dimensionality of the features. Max pooling retains significant features by selecting the maximum value in the local window, while average pooling calculates the mean of the local window to smooth the feature representation. This operation can effectively reduce the data dimension, lower the computational complexity, while retaining key information and preventing the model from overfitting. Finally, in the fully connected layer, the local features extracted by the convolutional and pooling layers are integrated to generate global acquisition policy parameters. The fully connected layer realizes the mapping from local features to global features through weighted summation of all features, providing support for the generation of subsequent data acquisition action strategies. Through the above methods, the finally formed global acquisition policy parameters can accurately reflect the current data distribution and its characteristics, providing a reliable basis for efficient data acquisition. Capturing multi-scale features of the target distribution through convolutional kernels of different sizes improves the comprehensiveness and accuracy of the model's characterization of the target distribution characteristics. The introduction of the pooling layer effectively reduces the feature dimension, reduces the computational burden, while retaining key feature information through max pooling or average pooling. The fully connected layer integrates the local features extracted by the convolutional and pooling layers into global acquisition policy parameters, making the acquisition policy more comprehensive and accurate. Through the improved convolutional neural network structure, the adaptability of the model to complex data distributions is enhanced, and the efficiency and accuracy of data acquisition tasks are improved. The pooling operation and feature dimensionality reduction effectively prevent the model from overfitting on small sample datasets and enhance the generalization ability of the model.

[0033] S2 also includes the initialization of the resource quantity in the dynamic environment. The initial resource distribution is obtained through historical data and environmental sensors, and the parameters are updated in real time according to the environmental state during the modeling process. 、 、 and 。

[0034] In the initial stage of dynamic environment modeling, obtaining the initial distribution of resources is a key step. This process constructs the initial state of the resource quantity by combining historical data and information collected in real time by environmental sensors. These initial data can provide a reference basis for subsequent modeling and parameter updates, ensuring that the model starts running from an accurate starting point. During the modeling process, the parameters are dynamically adjusted according to the changes in the environmental state: represents the natural growth rate of resources and can be adjusted according to the change in the resource replenishment rate; Represents the consumption rate of resources by the acquisition device, and its value can be updated according to the device performance and the acquisition pressure in the environment; Represents the improvement coefficient of the acquisition efficiency, which is related to the device performance and the characteristics of resource distribution, and is dynamically adjusted to optimize the efficiency; Represents the resource consumption rate of the acquisition device, and its value can be adjusted with the changes in the device usage intensity and environmental constraint conditions. By combining historical data and real-time sensor information for the initialization of the resource quantity, the accuracy in the initial stage of modeling is significantly improved, providing a reliable basis for subsequent optimization. By adjusting the parameters α, β, δ, and γ in real time, the model can quickly respond to the changes in the dynamic environment, improving the data acquisition efficiency. The dynamically updated parameters ensure the reasonable allocation and efficient utilization of resources, avoiding the problems of over-acquisition or resource waste. The real-time parameter update makes the acquisition strategy more accurate and reliable, thus reducing the strategy deviation caused by environmental changes and improving the decision-making efficiency. The model parameters can be dynamically adjusted according to different environments and resource distribution characteristics, expanding the applicable scope of the method.

[0035] S4 includes = c - d - e; where, c represents the acquisition rate, d represents the energy consumption, and e represents the data loss rate. The invention optimizes the data acquisition efficiency by introducing the objective function , comprehensively considering the acquisition rate, energy consumption, and data loss rate. In this objective function, the acquisition rate c is the index to be maximized, representing the amount of data acquired per unit time. The higher its value, the higher the acquisition efficiency of the system; the energy consumption d is the index to be minimized, representing the energy consumed to complete the data acquisition task. The lower its value, the more energy-efficient the system; the data loss rate e is the index to be minimized, representing the proportion of data loss caused by various reasons (such as network delay or device failure) during the acquisition process. The lower its value, the higher the reliability of data acquisition. The construction of the objective function reflects the comprehensive balance of the three indexes. By maximizing during the optimization process, the system can simultaneously improve the acquisition efficiency, reduce the energy consumption, and reduce the data loss, thus achieving the overall optimization of the data acquisition efficiency. In the specific implementation, optimization algorithms such as Lagrangian dual optimization or gradient descent can be combined, and the data acquisition parameter k can be dynamically adjusted according to the objective function to maintain efficient operation in a complex dynamic environment. The objective function The acquisition rate, energy consumption, and data loss rate are simultaneously considered, enabling comprehensive optimization of data acquisition efficiency, enhancing the overall performance of the system, effectively controlling energy consumption while maximizing the acquisition rate, providing an efficient and energy-saving solution for the system, ensuring more complete collected data by minimizing the data loss rate, enhancing the reliability and stability of the system. The simple form of the objective function is convenient for extension and adjustment. The weights or constraint conditions can be adjusted according to the requirements of different application scenarios to achieve flexible optimization. The objective function can be combined with dynamic parameter adjustment methods to adapt to changing environmental states and improve the adaptability and robustness of the optimization results.

[0036] In the convolutional layer of the convolutional neural network model, dilated convolution is adopted to expand the receptive field and improve the ability to extract the distribution characteristics of targets at different scales, thereby enhancing the generalization performance of the data acquisition strategy. Specifically, in the design of the convolutional neural network, due to the fixed receptive field of traditional convolutional operations, large-scale spatial features may not be effectively captured. By introducing dilated convolution, the receptive field can be expanded without increasing the computational complexity, enabling the model to extract deep features of the data distribution at different scales, thus improving the strategy optimization ability of data acquisition. In this embodiment, the settings of dilated convolution are as follows: in the convolutional layer, dilated convolutional kernels with different dilation rates are introduced, such as 3×3 convolutional kernels with dilation rates of 1, 2, and 4 respectively, so that the receptive field grows exponentially, thereby capturing richer spatial information. In the deep convolutional layer, residual connections are adopted to retain the original data features, and at the same time, the deep features of dilated convolution are introduced to avoid the problem of gradient disappearance and improve the model stability. Combining the max-pooling or average-pooling of the pooling layer further reduces the feature dimension, reduces redundant information, and at the same time retains the most important target distribution features. In the fully connected layer, multi-scale features extracted by dilated convolution are integrated, and through the Softmax layer or Sigmoid activation function, the final global data acquisition strategy parameters are generated to improve the acquisition efficiency. Through the above method, the convolutional neural network model of the present invention can effectively utilize dilated convolution to more accurately model the distribution of target data in the reinforcement learning environment, improve the acquisition efficiency, and at the same time keep the computational complexity controllable, providing an optimized strategy for efficient data acquisition tasks.

[0037] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An optimization method for data acquisition efficiency based on artificial intelligence and reinforcement learning, characterized in that, The method includes: S1. Construct an initial data acquisition model through a deep reinforcement learning algorithm; S2. Obtain environmental state data in real time in a dynamic environment; S3. Generate a new action policy based on the current environmental state data and historical data. Specifically, map the current environmental state data and historical data into a policy combination, generate a logical next action policy in the policy set, and combine the reward function in deep reinforcement learning to evaluate and select the generated policy. The formula for generating the next action policy is: ; Among them, represents the new round of action strategy set generated after the transposition operation ; represents the data acquisition action strategy set at the -th iteration; represents the transposition logic operation function, which is used to perform logical adjustment on the current action strategy set ; represents the current time step of the policy iteration. S4. Adjust data acquisition parameters according to the generated action policy; S5. Continuously iterate and update the model parameters to adapt to the new environmental state. Specifically, represent the parameters of the reinforcement learning model in complex number form, iteratively generate the distribution of the model parameters, simulate the complex process of the model adapting to a rapidly changing environment, analyze whether the parameter distribution converges or diverges, adjust the iteration step size, and make the model gradually approach the optimal parameter configuration. The specific formula for iteratively generating the distribution of the model parameters is as follows: ; Among them, represents the complex representation of the current model parameters, represents the model parameters generated in the next iteration, represents the complex parameters corresponding to the environmental state, represents the number of iterations.

2. The data acquisition efficiency optimization method based on artificial intelligence and reinforcement learning according to claim 1, characterized in that: The S1 includes: Convert the data acquisition task features into an image form, construct an input tensor, design a multi-layer convolutional neural network model, extract features and biases with convolutional kernels, and enhance the non-linear expression ability through activation functions. The specific formula is as follows: ; Among them, represents the output of the -th layer of neurons, represents the output of the -1-th layer of neurons, represents the index of the number of network layers, represents the -th layer of weight matrix, represents the -th layer of bias vector, represents the activation function; Use data acquisition simulation data to supervise and train the model, optimize the initial policy, and generate an action policy prediction model for data acquisition.

3. A method for optimizing data acquisition efficiency based on artificial intelligence and reinforcement learning according to claim 1, characterized in that: The S2 includes: Model the state variables of the dynamic environment, simulate the changing trend of the environmental state over time in real time, dynamically adjust the timing and frequency of data acquisition according to the simulation results. The specific formula for simulating the changing trend of the environmental state over time is: , ; Among them, represents the amount of resources in a dynamic environment, represents the activity intensity of the acquisition device, represents time, represents the natural growth rate of resources, represents the consumption rate of the acquisition device on the amount of resources, represents the improvement coefficient of the acquisition efficiency, represents the resource consumption of the acquisition device.

4. An optimization method for data acquisition efficiency based on artificial intelligence and reinforcement learning according to claim 1, characterized in that: The S4 includes: Model the optimization of data acquisition parameters as a constrained optimization problem, transform the constrained problem into an unconstrained optimization, solve the optimal solution for parameter adjustment, and adjust the acquisition parameters in real time according to the optimization results to ensure that the acquisition efficiency remains optimal when the environment changes. The specific formula for solving the optimal solution for parameter adjustment is: ; Among them, represents the set of data acquisition parameters, represents the weight coefficient of the constraint condition on the objective function, represents the weight coefficient of the th resource constraint condition on the objective function, represents the objective function of data acquisition efficiency, represents the th constraint condition in data acquisition, represents the number of the constraint condition, represents the objective function constructed in the Lagrangian dual optimization.

5. A method for optimizing data acquisition efficiency based on artificial intelligence and reinforcement learning according to claim 1, characterized in that: The S3 further includes: Collect environmental state data and the environmental state data of the previous round; Input the current environmental state data and historical data into the initial data acquisition model; Generate a new action policy based on the predicted value output by the model; Select the final action through the action probability output by the model. If the probability P of the current action is greater than or equal to the set threshold T, then select the current action, otherwise select a random action, where P represents the action probability and T represents the threshold.

6. The data acquisition efficiency optimization method based on artificial intelligence and reinforcement learning according to claim 5, characterized in that: The generating a new action policy based on the predicted value output by the model includes: Obtain the model predicted value V(s) in the current environmental state; Calculate the weighted average prediction value W = α × V(s) + (1 - α) × V according to historical data prev , where α represents the weighting coefficient, V(s) represents the prediction value in the current environmental state, and V prev represents the prediction value of the previous state; Set a preset value T W , generate an action policy based on the weighted average predicted value, where if the weighted average predicted value W is greater than or equal to the preset value T W , then adopt this action policy.

7. An optimization method for data acquisition efficiency based on artificial intelligence and reinforcement learning according to claim 2, characterized in that: The S1 further includes: In the convolutional neural network model, use convolutional kernels of different sizes in the convolutional layer to extract multi-scale features of the target distribution, add a pooling layer after the convolutional layer, reduce the feature dimension through max pooling or average pooling, and integrate the local features extracted by the convolutional and pooling layers in the fully connected layer to generate global acquisition policy parameters.

8. An optimization method for data acquisition efficiency based on artificial intelligence and reinforcement learning according to claim 3, characterized in that: The S2 further includes: Initialization of the resource quantity in a dynamic environment, obtaining the initial resource distribution through historical data and environmental sensors, and updating the parameters in real time according to the environmental state during the modeling process , , and .

9. The data acquisition efficiency optimization method based on artificial intelligence and reinforcement learning according to claim 4, wherein: The objective function of the data acquisition efficiency in S4 is as follows: = c - d - e; where c represents the acquisition rate, d represents the energy consumption, and e represents the data loss rate.

10. A method for optimizing data collection efficiency based on artificial intelligence and reinforcement learning according to claim 7, characterized in that: The convolutional layer of the convolutional neural network model uses dilated convolution.

Citation Information

Patent Citations

  • Vehicle data acquisition frequency dynamic adjustment method based on deep reinforcement learning

    CN110213827A

  • Network data acquisition efficiency optimization method and system based on deep reinforcement learning

    CN114710410A

  • Metadata acquisition method based on user-defined strategy optimization algorithm

    CN119719084A

  • Database adaptive data flow acquisition optimization method and system based on reinforcement learning

    CN119719783A

  • Data acquisition method and system for multiple data sources

    CN119884670A

Cited By

  • Flexible job shop dynamic scheduling method and system considering machine aging

    CN121303637A

  • AI-based enterprise data acquisition method, equipment and medium

    CN121524243A