A marine fishery resource big data dynamic allocation method based on reinforcement learning
By constructing a slowly varying reinforcement learning model for multi-objective collaborative scheduling, the problems of unstable water exchange and accurate feeding system in offshore aquaculture cages were solved, efficient resource allocation and energy consumption optimization were achieved, and the environmental adaptability and strategy stability of the system were improved.
Patent Information
- Application Number
- CN202511041377.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-28
AI Technical Summary
The layout of offshore aquaculture cages is affected by factors such as ocean currents, wind and waves, and pollution diffusion, resulting in unstable water exchange capacity, reduced feeding system accuracy, and increased maintenance and operating costs. In addition, the reinforcement learning model has ambiguous strategies in multi-objective conflicts and mutation environments, making it difficult to achieve effective resource allocation.
A slowly varying reinforcement learning model for multi-objective collaborative scheduling is constructed. By introducing a target conflict recognition function and dynamic reward weight adjustment, combined with adaptive representation values and a transient transition model, the dynamic allocation of marine fishery resources is achieved, including initializing data collection, building a model, generating adaptive representation values, and making adaptive decisions.
It improves resource utilization efficiency, reduces energy consumption, enhances the system's ability to respond to environmental changes, avoids policy failure, and provides an efficient and controllable resource regulation solution.
Smart Images

Figure CN120562994B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing and adaptive processing technology, and specifically relates to a method for dynamic allocation of marine fishery resource big data based on reinforcement learning. Background Art
[0002] Pelagic aquaculture, conducted in waters far from the coastline, effectively addresses the current situation of intense competition for nearshore resources. It also offers advantages such as greater environmental protection and less impact on the natural ecosystem, as well as enhanced water exchange and pollution diffusion capabilities for the aquatic products. Despite this, the layout of cages in pelagic aquaculture still suffers from insufficient stability in water exchange and pollution diffusion. This instability is related to the layout of the cages. Cages floating in the open ocean are significantly affected by currents, wind and waves, and the spread of pollution. Cages that are too close together can restrict water flow, leading to localized hypoxia and pollution accumulation. However, cages that are too far apart can complicate control of the cages. In particular, excessively long feeding system pipes or feeding paths can reduce the accuracy of feeding amounts and timing, further leading to feed waste and increased energy consumption. The core goal of reinforcement learning is to use big data to train a strategy that maximizes cumulative rewards in a given environment. Using reinforcement learning to control cage placement can significantly improve energy efficiency and aquaculture performance. However, policy ambiguity is prone to occur in practical applications. This can lead to decreased control stability in offshore aquaculture systems, primarily due to frequent cage relocations and increased wear and tear in automated feeding systems, resulting in high maintenance and operating costs. This policy ambiguity arises from the inherent multi-objective nature of reinforcement learning strategies. In multi-objective reinforcement learning models, inherent conflicts between these objectives and the inconsistent orientation of these multiple objectives create a vast range of choices for real-time policy deployment. This increases the randomness of policy actions and slows down training convergence. For example, when cage spacing is small, water exchange capacity is poor. Consequently, the reinforcement learning strategy tends to increase the distance between cages, while increasing the feeding path tends to decrease the distance between cages. These two strategies counteract each other in their allocation mechanism, making policy ambiguity in this range impossible to establish a fast-converging decision output.
[0003] Such pelagic environments often exhibit a significant coexistence of trend evolution and sudden disturbances. Specifically, under conventional ocean currents, water parameters such as dissolved oxygen, salinity, and temperature exhibit a relatively stable, slowly varying mechanism. Simultaneously, sudden events, including typhoons and red tides, can cause dramatic, short-term changes in ocean conditions, leading to abrupt shifts in these slowly varying parameters. These sudden shifts disrupt the system's existing resource utilization balance, exposing reinforcement learning to failures in learned strategies. Consequently, such pelagic systems exhibit significant spatiotemporal stratification and coupled feedback. The resource planning approach developed is primarily driven by the coordinated scheduling of various resources. On the one hand, it must consider trade-offs between multiple objectives and maintain stability. On the other hand, when sudden disturbances occur, risk aversion and survival become the dominant objectives, prioritizing previously secondary indicators, such as emergency power supply and localized sewage disposal, for temporary strategies. Such sudden shifts in the dominant objective can significantly bias reinforcement learning models trained on a coordinated equilibrium, leading to policy failures and even policy collapse. Therefore, these two types of problems are not isolated from each other, but rather alternately dominate the strategy generation process of the reinforcement learning system during the natural regulation cycle. Especially in the transition phase before and after a sudden change, the system must be able to flexibly and promptly switch from multi-objective balanced scheduling to risk-driven control, and achieve self-recovery and return to a coordinated track after the sudden change. Therefore, a dynamic allocation method for marine fishery resources big data based on reinforcement learning is urgently needed. Summary of the Invention
[0004] The purpose of the present invention is to propose a method for dynamic allocation of marine fishery resource big data based on reinforcement learning to solve one or more technical problems existing in the prior art and at least provide a beneficial option or create conditions.
[0005] To achieve the above objectives, according to one aspect of the present invention, a method for dynamic allocation of marine fishery resources big data based on reinforcement learning is provided, the method comprising the following steps:
[0006] S100, initializing the marine fishery resource data scene and collecting fishery data;
[0007] S200, a slowly varying reinforcement learning model for multi-objective collaborative scheduling using fishery data;
[0008] S201, obtaining control strategies for fishery resource allocation from a slowly varying reinforcement learning model;
[0009] S300, generating an adaptive representation value in real time during the running process of the slowly varying reinforcement learning model;
[0010] S400, constructing an imminent transition model according to the adaptation characterization value and obtaining an imminent transition order value;
[0011] S500, making an adaptive decision by using the critical transition order value.
[0012] Furthermore, in step S100, the method for initializing the marine fishery resource data scene and collecting fishery data is: a sensor system and a control device for marine environment monitoring and resource usage perception are deployed in the offshore aquaculture farm area, and the data collected by the sensor system and the control device are used as fishery data; wherein the sensor system includes at least one water quality sensor, a dissolved oxygen sensor, a salinity meter, a thermometer, a wind and wave meter, an energy consumption sensor and a position positioning device; the sensor system is used to collect the spatiotemporal state information, power usage status and aquaculture unit distribution of the fishery aquaculture environment in real time; the control device includes a cage position control unit, an automatic feeding unit, an energy distribution execution module and a data communication module, and the control device is used to receive scheduling instructions and execute cage layout adjustment, energy supply and feeding operations.
[0013] Furthermore, in step S200, the method for constructing a multi-objective collaborative scheduling slowly varying reinforcement learning model using fishery data is as follows: constructing a reinforcement learning model using fishery data, wherein the model uses a state compression mechanism to perform dimensionality reduction processing on the original marine environment data;
[0014] The optimization objectives of the reinforcement learning model were set as power consumption, cage spacing, water exchange efficiency, and feeding path length. A reinforcement learning loss function was constructed using a multi-objective reward structure. The multi-objective reward structure included the feeding accuracy per unit resource usage, the improvement in water exchange capacity due to cage spacing, and the coverage of the feeding system per unit resource. The resulting reinforcement learning model was denoted as a slowly varying reinforcement learning model.
[0015] Furthermore, in step S200, the method of constructing a slowly-varying reinforcement learning model for multi-objective collaborative scheduling using fishery data also includes: introducing a target conflict identification function into the slowly-varying reinforcement learning model to detect the real-time directional deviation between each optimization target, the directional deviation being obtained by calculating the gradient vector angle of each target sub-reward item, and when it is detected that the directional deviation between any two target sub-reward items exceeds a preset conflict threshold, the reward weight coefficient corresponding to each target in the slowly-varying reinforcement learning model is dynamically adjusted according to the state of the marine environment.
[0016] Although the goal conflict identification and weight adjustment mechanism of this method does not directly affect the final action output of the policy network, it substantially changes the policy update path by adjusting the proportion and gradient direction of each goal sub-reward during training, indirectly optimizing the action selection tendency and output stability formed by the reinforcement learning model under multi-goal conditions.
[0017] Beneficial effects: Taking into account the complex characteristics of multi-objective and multi-disturbance in resource regulation tasks in the distant ocean fishery system, the key optimization control of the reinforcement learning model under the policy fuzzy state is achieved by introducing gradient directional evaluation and environmental state-driven reward weight adaptive adjustment strategy: In traditional multi-objective reinforcement learning, the target direction inconsistency caused by multiple disturbances in the marine environment characteristics leads to nonlinear changes, and the policy update path will oscillate repeatedly in the gradient space and it is difficult to converge to a stable control strategy. Therefore, this method uses a dynamic weight coordination mechanism to enable the strategy to obtain clear directional dominance in the conflict interval, thereby eliminating the ambiguity and uncertainty of the policy space and constructing a dynamic policy scheduling capability with resource priority adaptability.
[0018] Furthermore, in step S201, the method for obtaining a control strategy from a slowly varying reinforcement learning model for allocating fishery resources is: using the control strategy output by the slowly varying reinforcement learning model to dynamically adjust key control units in the marine fishery scenario, wherein the key control units include but are not limited to a feeding system, a power distribution device, and a cage spacing adjustment mechanism.
[0019] The slowly varying reinforcement learning model generates a specific control strategy based on the current environmental state and the results of a multi-objective reward function. This control strategy, as the model's output signal, is used to control the regulatory behavior of multiple key devices in offshore fisheries scenarios, such as precise automatic feeding control, spatial priority sorting of power distribution, and adjustment of cage floating positions. This process, based on the optimal action sequence output by the reinforcement learning policy network, achieves a balance between resource efficiency and aquaculture benefits, avoiding redundant adjustments and energy waste.
[0020] Furthermore, in step S300, the method for generating an adaptive representation value for the running process of the slowly varying reinforcement learning model in real time is: obtaining the reward value, strategy output result and strategy entropy state of the current reinforcement learning model, and performing difference measurement with the corresponding values at the previous moment: obtaining the change amplitude of the reward value, the degree of difference in the strategy output probability distribution, and the fluctuation of the strategy entropy, and using them as three dimensions for measuring adaptability; performing a weighted combination of the above three dimensions to form an adaptive representation value that can be continuously updated over time, which is used to represent the current strategy's ability to respond to changes in environmental states.
[0021] The current reinforcement learning model refers to the last or latest execution of the slowly varying reinforcement learning model, and the previous moment refers to the timing or moment of the last execution of the slowly varying reinforcement learning model in the reverse time direction.
[0022] Beneficial effects: The adaptive representation value is a state perception indicator dynamically generated during the operation of the reinforcement learning strategy, and plays a key role in connecting the previous and the next. Through the collaborative reinforcement learning model, it can capture multi-dimensional feedback such as reward changes, strategy output stability, and strategy entropy fluctuations, and reflect in real time the strategy's ability to respond to changes in the current marine environment in slowly changing scenarios. At the same time, as input data for further identification of impending transitions, it provides mathematical support for fluctuation trends to determine whether the system is in an unstable transition stage before and after a sudden change. Since the environmental state of distant-water fisheries often exhibits the coupling characteristics of alternating slow changes and sudden changes, a single strategy indicator is usually difficult to cope with complex dynamic scenarios. Therefore, the adaptive representation value, as a comprehensive feedback signal of the reinforcement learning model's reward value, strategy output results, and strategy entropy state, can effectively perceive the critical change point of the strategy robustness, thereby initiating a response mechanism before the strategy fuzzy state evolves into strategy failure. Therefore, the adaptive representation value is the core feature that enables the dynamic and flexible conversion between collaborative optimization and risk defense.
[0023] Furthermore, in step S400, a transient transition model is constructed based on the adaptive characterization value. The method for obtaining the transient transition order value is as follows: any moment of obtaining the control strategy is used as a modeling point; an integer value is preset as the backtest value ST, and its value range is ST∈[3,10]; the exponential average of the adaptive characterization values corresponding to any modeling point to the previous ST and previous 2×ST modeling points is recorded as the short-term benchmark value STEma and the long-term benchmark value LTEma respectively; the absolute value of the difference between the short-term benchmark value and the long-term benchmark value is recorded as the absolute deviation; the ratio of the absolute deviation to the long-term benchmark value is calculated to obtain the fluctuation response D t ;
[0024] The critical transition order value is obtained by constructing a critical transition model based on the Sigmoid function through nonlinear mapping of the fluctuation response.
[0025] In order to effectively process the fluctuation response Dt and maximize its influence within a specific range, this step designs a transient transition model based on the Sigmoid function to perform nonlinear mapping on it, thereby obtaining the transient transition order value L at the current moment. t ;
[0026] ATCV t Indicates the adaptive representation value corresponding to the current modeling point, which is used to perform nonlinear transformation on the fluctuation response; t Performing a hyperbolic tangent transformation to convert the current state of the marine fishery environment into a range more suitable for the transient model, thus making the quantification of current conditions more flexible and effective;
[0027] The transient transition model is constructed by adapting the characterization value ATCV tApplying the hyperbolic tangent function for nonlinear compression makes its numerical distribution closer to the actual perception boundary of resource response in the marine fishery environment. This processing mechanism simulates the gradual perception process of abnormal signals by the resource scheduling system in real sea areas, avoiding the risk of excessive response of the policy system by the extreme edge value. By introducing the value into the form of exponential function, it is combined with the dynamic factor D t The difference forms a combination, which strengthens the model's ability to amplify slight disturbances. This sensitive amplification mechanism conforms to the physical characteristics of the ocean system that the initial change is the first perturbation, and can capture the signal that the environment enters an unstable area in advance, realizing the pre-perception of potential mutations. The transition order value L output by the Sigmoid function mapping is t It is controlled between 0 and 1 to achieve direct numerical connection with the policy control system and improve the interpretability and stability of the model. This greatly enhances the resilience and strategic foresight of the resource allocation system under severe environmental disturbances.
[0028] Because existing methods perform nonlinear mapping of adaptive characterization values based on short-term and long-term baseline values, and then directly generate a processing path for the transient transition order value, they suffer from a single response mechanism and uneven discrimination sensitivity when faced with the dynamic environmental background of slow-changing trends and short-term disturbances in the open ocean. This makes it difficult to effectively characterize the transition evolution characteristics of the system during the transition from a stable state to a sudden state. Especially in areas sensitive to sea condition regulation, such as the initial stage of wind and wave amplification or near the critical point of pollution diffusion, this method is prone to over-response or response lag, thereby reducing the timeliness and adaptability of resource control strategies. Existing technologies lack an indicator system that can decompose and analyze disturbance structures at multiple time scales, and are unable to accurately identify disturbance driving sources and quantitatively express evolutionary trends. This makes it difficult to meet the practical needs of the dual goals of risk warning and robust resource scheduling in the open ocean aquaculture system.
[0029] To this end, this paper proposes an optimization approach: by introducing three types of dynamic characteristic factors differentiated by time dimension, representing disturbance intensity, trend change rate, and multi-round abnormal accumulation behavior, and constructing a dynamic weighted fusion mechanism with hierarchical nested logic, the model can proactively identify key disturbance characteristics at different scales, enabling adaptive identification of mutation risks and adjustment of response levels. This strategy enhances the selective response capability while preserving the model's sensitivity, significantly improving the system's insight into mutation precursors and the strategic fault tolerance margin, providing intelligent support with ecological resilience and risk aversion for distant-water fishery resource management.
[0030] Preferably, in step S400, the method for obtaining the transition value of the transition model according to the adaptive characteristic value is: taking any time point when the control strategy is obtained as a modeling point, taking any modeling point and its previous Num modeling points as its dynamic modeling window, and constructing a window characteristic value sequence corresponding to the adaptive characteristic value from the dynamic modeling window; Num is the number of modeling points in the current 12-48 hours;
[0031] The sea area disturbance factor a is calculated through the dynamic modeling window t , the environmental sudden change factor β t , and the risk accumulation factor γ t , and the above factors are brought into the transition model to obtain the transition value Lt: L t = max (|a t |, |β t |) x exp (γ t ); wherein the sea area disturbance factor is used to judge the degree of non-stability of the parameters obtained by the sensing system in the current environment, and its mathematical expression is the standard deviation of the window characteristic value sequence.
[0032] However, since the transition model is sensitive to extreme values and cannot effectively capture the dynamic changes of data in actual application, the sea area disturbance factor often has the problem of insufficient sensitivity to extreme values, so another solution is provided, and the implementation process is: defining the window amplitude of the current modeling point as the range of the corresponding window characteristic value sequence, and then the amplitude stability coefficient PSI is the ratio of the minimum value of the window amplitude of each modeling point in the dynamic modeling window to the window amplitude.
[0033] The difference between each adaptive characteristic value in the window characteristic value sequence and ATCV1 is calculated one by one, all the differences are squared and weighted, the weighted result is multiplied by PSI, and then accumulated; the accumulated result is divided by the number of time steps Num, and then the square root is taken, and the final output value is the sea area disturbance factor. The factor functions to quantify the shock amplitude of the adaptive characteristic value deviating from the average level, and reflects the abnormal fluctuation intensity of the environmental parameters.
[0034] The environmental sudden change factor is a risk gradient for the model to identify trend instability in advance: taking the absolute value of the difference between the adaptive characteristic value step difference of the current modeling point and the previous modeling point as the difference change, and the ratio of the difference change to the modeling time distance is the environmental sudden change factor.
[0035] Wherein the adaptive characteristic value step difference refers to the difference between the adaptive characteristic value of any modeling point and the adaptive characteristic value of the previous modeling point; the modeling time distance is the time distance between the current modeling point and the previous modeling point, and the unit is minute.
[0036] The role of the environmental sudden change factor is to capture the acceleration or deceleration of the rate of change of the adaptive characteristic value and identify the evolutionary kinetic energy of the environmental deterioration trend;
[0037] The cumulative amount of adaptive representation values exceeding the threshold within the dynamic modeling window is recorded as the risk backlog factor to quantify the degree of long-term existence of local abnormal states.
[0038] The specific calculation method is as follows: obtain the 90th percentile of the adaptive characterization value in the limited historical data as the abnormal threshold τ, traverse each adaptive characterization value in the dynamic modeling window, and if the adaptive characterization value is greater than the threshold τ, mark it as an abnormal point and calculate its threshold excess, that is, the difference between the adaptive characterization value and τ; if the value is less than or equal to τ, ignore the point; accumulate the excess values of all abnormal points to obtain the risk backlog factor. This factor quantifies the cumulative effect of the continuous deviation of the environmental state from the normal range and represents the energy reserve of systemic risk. Its mathematical expression is:
[0039] The time constraint for historical data is to limit the time length to at least one month. Because ocean data usually needs to adapt to tidal phenomena, 30 natural days is a relatively complete backtesting period. It can be adjusted to a longer time length, such as one quarter or one year, according to the actual modeling application direction.
[0040] This model uses a sliding window mechanism to capture the temporal evolution of adaptive representation values, effectively characterizing the degree of environmental instability. The marine disturbance factor reflects the magnitude of abnormal fluctuations in environmental parameters, the environmental sudden change factor detects the acceleration of deterioration, and the risk backlog factor quantifies the degree of sustained deviation from normal conditions. These three factors are combined in a nonlinear model to form a highly sensitive threshold transition value, providing a quantitative basis for subsequent decision-making regarding sudden change risks.
[0041] Beneficial Effects: By introducing a modeling mechanism for the critical transition order, a hierarchical quantitative framework for the coupled response to trend evolution and risk mutations was established in the context of offshore fishery resource regulation, enabling structured perception and semantic interpretation of multidimensional disturbance signals in the dynamic environment of fishery resources. The critical transition order not only serves as the core judgment basis for strategy switching, but its deeper quantitative significance lies in the use of adaptive representation values as dynamic driving variables in the continuous time domain to perform interval projection and function mapping on the response trajectories of the three stages of slow change, criticality, and mutation in the environmental state evolution path. This enables the system to quantify the sensitivity of the current state to future mutation trends in real time based on the coupling results of multiple implicit change factors such as disturbance source intensity, frequency, and acceleration, once the disturbance is initially present. This builds an intermediary judgment bridge between the reinforcement learning policy network and real-world environmental feedback, enhancing the system's ability to discern and adapt to environmental non-stationary processes while ensuring policy stability. This enables the reinforcement learning control model to have the ability to self-perceive changes in environmental risks and actively adjust decision-making paths, promoting the paradigm shift of AI-based fishery resource management methods from passive response to predictive regulation, and providing more efficient, controllable and ecologically coordinated solutions for power allocation, feeding path regulation, cage layout, etc. in offshore aquaculture scenarios.
[0042] Furthermore, in step S500, the method for making adaptive decisions through the critical transition order value is: setting 5-15 natural days as a reference time period, obtaining all critical transition order values within the reference time period as an order value set; if the current critical transition order value is less than the lower quartile of the order value set, it is defined that a transition mark appears; when the proportion of transition marks appearing in a natural day exceeds half, it is defined that the current environment is in a mutation risk state, suspending the model control output and triggering an early warning, or executing a preset safety control strategy.
[0043] Under sudden risk conditions, the system automatically suspends the control output of the slowly varying reinforcement learning model to prevent resource allocation imbalances caused by policy mismatch. This triggers one of two intervention measures: a remote manual alert prompting operations and maintenance personnel to take over immediately; and the system executes a pre-set safety control strategy, prioritizing key survival actions such as oxygen supply, flow control, and maintaining basic feeding, while limiting energy-intensive and precision-critical operations to maintain the system's minimum stable operating state. This mechanism effectively prevents uncontrolled decision-making by the reinforcement learning strategy under sudden risk conditions, improving the resilience and safety of the system's operations.
[0044] Preferably, all undefined variables in the present invention, if not clearly defined, can be manually set thresholds.
[0045] The present invention also provides a system for dynamically allocating marine fishery resource big data based on reinforcement learning. The system comprises: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for dynamically allocating marine fishery resource big data based on reinforcement learning are implemented. The system can be run on computing devices such as desktop computers, laptop computers, PDAs, and cloud data centers. The executable system may include, but is not limited to, a processor, a memory, and a server cluster. The processor executes the computer program to run in the following system units:
[0046] The ocean scene initialization unit is used to initialize the marine fishery resource data scene and collect fishery data;
[0047] A reinforcement learning model building unit, used to build a slowly varying reinforcement learning model for multi-objective collaborative scheduling using fishery data;
[0048] An adaptation representation calculation unit, used to generate an adaptation representation value in real time during the running process of the slowly varying reinforcement learning model;
[0049] The imminent transition identification unit is used to construct an imminent transition model according to the adaptation characterization value and obtain an imminent transition order value.
[0050] The model execution decision unit is used to make adaptive decisions through the critical transition order value.
[0051] The beneficial effects of the present invention are as follows: the present invention provides a method for dynamic allocation of marine fishery resource big data based on reinforcement learning. By introducing a modeling mechanism of the transient transition order value, a hierarchical quantitative framework for the coupled response of trend evolution and risk mutation is established in the scenario of offshore fishery resource regulation to realize structured perception and semantic interpretation of multidimensional disturbance signals in the dynamic environment of fishery resources. In the continuous time domain, the adaptive characterization value is used as the dynamic driving variable, and the response trajectories of the three stages of slow change, criticality and mutation in the environmental state evolution path are interval projected and function mapped. The system can quantify the sensitivity of the current state to future mutation trends in real time based on the coupling results of multiple implicit change factors such as the intensity, frequency and acceleration of the disturbance source when the disturbance is initially presented, thereby constructing an intermediary judgment bridge between the reinforcement learning strategy network and the real environment feedback, while ensuring the stability of the strategy, enhancing the system's ability to distinguish and adapt to environmental non-stationary processes. This enables the reinforcement learning control model to have the ability to self-perceive changes in environmental risks and actively adjust decision-making paths, promoting the paradigm shift of AI-based fishery resource management methods from passive response to predictive regulation, and providing more efficient, controllable and ecologically coordinated solutions for power allocation, feeding path regulation, cage layout, etc. in offshore aquaculture scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The above and other features of the present invention will become more apparent through a detailed description of the embodiments shown in conjunction with the accompanying drawings. In the drawings of the present invention, the same reference numerals represent the same or similar elements. Obviously, the drawings described below are only some embodiments of the present invention. It is possible for a person skilled in the art to derive other drawings based on these drawings without inventive effort. In the drawings:
[0053] Figure 1 Shown is a flow chart of a method for dynamic allocation of marine fishery resources big data based on reinforcement learning;
[0054] Figure 2 Shown is a structural diagram of a marine fishery resource big data dynamic allocation system based on reinforcement learning. DETAILED DESCRIPTION
[0055] The following will be combined with the embodiments and drawings to clearly and completely describe the concept, specific structure and technical effects of the present invention so as to fully understand the purpose, scheme and effect of the present invention. It should be noted that the embodiments and features in the embodiments of this application can be combined with each other unless there is a conflict.
[0056] like Figure 1 The figure shows a flow chart of a dynamic allocation method of marine fishery resources big data based on reinforcement learning. Figure 1To illustrate a method for dynamic allocation of marine fishery resources big data based on reinforcement learning according to an embodiment of the present invention, the method includes the following steps:
[0057] S100, initializing the marine fishery resource data scene and collecting fishery data;
[0058] S200, a slowly varying reinforcement learning model for multi-objective collaborative scheduling using fishery data;
[0059] S201, obtaining control strategies for fishery resource allocation from a slowly varying reinforcement learning model;
[0060] S300, generating an adaptive representation value in real time during the running process of the slowly varying reinforcement learning model;
[0061] S400, constructing an imminent transition model according to the adaptation characterization value and obtaining an imminent transition order value;
[0062] S500, making an adaptive decision by using the critical transition order value.
[0063] Furthermore, in step S100, the method for initializing the marine fishery resource data scene and collecting fishery data is: a sensor system and a control device for marine environment monitoring and resource usage perception are deployed in the offshore aquaculture farm area, and the data collected by the sensor system and the control device are used as fishery data; wherein the sensor system includes at least one water quality sensor, a dissolved oxygen sensor, a salinity meter, a thermometer, a wind and wave meter, an energy consumption sensor and a position positioning device; the sensor system is used to collect the spatiotemporal state information, power usage status and aquaculture unit distribution of the fishery aquaculture environment in real time; the control device includes a cage position control unit, an automatic feeding unit, an energy distribution execution module and a data communication module, and the control device is used to receive scheduling instructions and execute cage layout adjustment, energy supply and feeding operations.
[0064] The purpose of this step is to construct the multidimensional environmental input space required by the reinforcement learning system, providing high-quality raw data support for subsequent policy learning and scheduling models. The sensors in the sensing system are used to capture the marine environmental conditions of the aquaculture waters. A GNSS global navigation satellite positioning module and time synchronization mechanism are used to ensure that the collected data has sufficient spatial accuracy and temporal continuity. The control device is used for the actual deployment of the subsequent reinforcement learning strategy and feedback on the operating status. This operational status feedback includes energy consumption, position response, execution error, etc. after the strategy is executed, and is used to design the reward function of the reinforcement learning model and evaluate the system stability.
[0065] Fishery data includes measurements from the water quality sensors, dissolved oxygen sensors, salinity meters, thermometers, wind and wave meters, energy consumption sensors, and position positioning devices in the sensor system. It also includes operational status data fed back by the cage position control unit, automatic feeding unit, and energy allocation execution module in the control device. This operational status data includes the deviation between cage position adjustment instructions and actual response displacement, the difference between feeding amount and expected ratio, and the execution time and energy consumption feedback of power allocation tasks. This data is uniformly transmitted back through the communication module, forming a multi-source, heterogeneous fishery data set with time, spatial, and task tags, which is referred to as fishery data. The energy consumption sensor is the energy consumption data collector. In some embodiments, the system can optionally be equipped with a local pollutant monitoring module. This module includes pollutant sensors located below specific cages or at the intersection of aquaculture areas to monitor the concentrations of representative indicators such as ammonia nitrogen, nitrite, and suspended solids. These sensors, combined with spatial location, are used to construct pollutant distribution heat maps to support strategy training for environmental carrying capacity targets.
[0066] The sensor system's position location device should be set for each cage. This is because state representation in reinforcement learning models relies on spatial perception. The relative distance between cages is a key decision-making factor in multi-objective conflicts, such as water exchange and feeding routes. It also underlies marine environmental parameters with significant spatial gradients, such as currents, salinity, and dissolved oxygen. Therefore, the sensor data associated with the position location device is a key basis for measuring the deterioration of aquaculture conditions. Ultimately, control behaviors such as cage movement and feeding route planning also require strong spatial coupling with cage position data to allocate or fine-tune strategies. Each water quality sensor, dissolved oxygen sensor, salinity meter, thermometer, wind and wave meter, and energy consumption sensor should have one and only one corresponding position location device for simultaneous operation.
[0067] Furthermore, in step S200, the method for constructing a multi-objective collaborative scheduling slowly varying reinforcement learning model using fishery data is as follows: constructing a reinforcement learning model using fishery data, wherein the model uses a state compression mechanism to perform dimensionality reduction processing on the original marine environment data; extracting core environmental state variables for strategy modeling;
[0068] The optimization objectives of the reinforcement learning model were set as power consumption, cage spacing, water exchange efficiency, and feeding path length. A reinforcement learning loss function was constructed using a multi-objective reward structure. The multi-objective reward structure included the feeding accuracy per unit resource usage, the improvement in water exchange capacity due to cage spacing, and the coverage of the feeding system per unit resource. The resulting reinforcement learning model was denoted as a slowly varying reinforcement learning model.
[0069] The state compression mechanism is implemented by a VAE variational autoencoder or an autoencoder, and the VAE variational autoencoder is preferentially selected because the sea state data itself is disturbed by many factors, the VAE variational autoencoder has the property of modeling the probability distribution of the state, can balance the modeling accuracy and the uncertainty expression capability, is more suitable for reflecting these interference factors than the ordinary autoencoder, is more friendly to the abnormal sensitivity of the marine environment monitoring, and is more compatible with the strategy uncertainty in the open-sea culture scenario from another angle, so that the strategy output constructed under the premise is more robust.
[0070] The multi-objective reward structure includes but is not limited to: the first item is the feeding precision of unit resource use, that is, the actual feeding success rate under unit energy consumption, which is used to measure the number of effective feeding behaviors generated per 1 kilojoule of electric energy.
[0071] The second item is the promotion effect of net cage spacing on the water exchange capacity, that is, the influence intensity R_Exch of the distance change between adjacent net cages on the local flow rate and the dissolved oxygen concentration, which is used to measure the promotion degree of the structure adjustment on the water ventilation and ecological recovery capacity, and the mathematical expression is R_Exch=(△VS+△CO) / △DS, wherein △VS, △CO and △DS respectively represent the promotion value of the average flow rate after adjustment, the promotion value of the average dissolved oxygen concentration after adjustment and the distance change amplitude value of the net cage.
[0072] The third item is the influence of power use on the coverage range of the feeding system, that is, the number of feeding net cages that can be completed per unit path energy consumption, which is used to measure the coverage efficiency of the feeding operation under the condition of limited energy consumption; the mathematical expression is the ratio of the feeding net cage to the energy consumption, and the energy consumption includes the moving path energy consumption and the feeding energy consumption.
[0073] When the local pollutant monitoring module is set, the fourth item of the multi-objective reward structure is the influence of the local pollutant concentration on the environmental carrying threshold, that is, the margin ratio between the unit area pollutant concentration and the corresponding ecological safety threshold, which is used to measure whether the current breeding intensity is close to the ecological overload risk, and the mathematical expression is the ratio of any pollutant concentration to the corresponding ecological threshold. The ecological threshold is a preset value, and when the value is exceeded, it is considered that the corresponding pollutant concentration is too large.
[0074] The reinforcement learning strategy also includes establishing an action change penalty term and a strategy entropy regularization term, which are used to improve the convergence speed and strategy stability of the model in the target conflict interval, the action change penalty term is used to inhibit the strategy output jump to avoid frequent adjustment of the net cage layout or the feeding path, and the strategy entropy regularization term is used to maintain the stability and explorability of the strategy distribution to prevent the strategy from converging to an extreme solution too early in the target conflict; the combination of the penalty term and the regularization term can ensure that the model still has good training stability and decision continuity when facing the conflict of multiple target guides.
[0075] Furthermore, in step S200, the method for constructing a slowly varying reinforcement learning model for multi-objective collaborative scheduling using fishery data further includes: introducing a target conflict identification function into the slowly varying reinforcement learning model to detect real-time directional deviations between optimization targets, where the directional deviations are obtained by calculating the gradient vector angles of each target sub-reward item; and dynamically adjusting the reward weight coefficients corresponding to each target in the slowly varying reinforcement learning model based on the state of the ocean environment when the directional deviation between any two target sub-reward items exceeds a preset conflict threshold. The strategy weights can be iteratively updated based on the strategy feedback loss of the reinforcement learning strategy network to achieve adaptive balance between multiple targets and accelerate the convergence of the ocean resource scheduling strategy.
[0076] The target sub-reward item refers to the reward item value of the reward function constructed separately for each optimization target that can be obtained at any time in the reinforcement learning model. In this method, the sub-reward items include the corresponding feeding efficiency reward, water exchange capacity reward, power consumption minimization reward and pollution control index reward in the multi-target reward structure. The target conflict identification function takes the gradient of each sub-reward item to the current strategy as input, outputs whether there is an optimization conflict for the current goal, and outputs the recommended value of the strategy weight adjustment. The gradient vector angle refers to the angle between the gradient directions of any two sub-reward functions. It is a real-time directional deviation of the instantaneous gradient, and the corresponding conflict threshold is 60°; its mathematical expression is: θ k1,k2 = arccos ((∇ θ R k1 ∇ θ R k2 ) / || ∇ θ R k1 ||·||∇ θ R k2 ||); where ∇ θ R k1 and ∇ θ R k2 The sub-reward functions representing any two optimization objectives are respectively the gradient vectors of the policy parameters; the reward weight coefficient is the weight factor of each sub-reward item in the total reward function; the reinforcement learning policy network is a neural network structure used to output control actions, including adjusting the position of the cage, controlling the running path of the feeding machine, etc.; the policy feedback loss is the loss term that the policy network relies on when updating, including policy gradient error, Q value error, entropy loss term or penalty term.
[0077] In addition, the gradient vector angle can be replaced by a trend inconsistency indicator. The trend inconsistency indicator measures the consistency of the slope direction of the sub-reward value changes. The slope angle is used to determine whether there is a continuous reverse fluctuation between targets, which is a continuous gradient. The corresponding conflict threshold is 60°.
[0078] When the conflict threshold is exceeded, the weight ratio of the sub-reward items is temporarily adjusted based on the state of the ocean environment. This means that environmental indicators such as remaining power and cage density are obtained to increase the weight of the most critical objective item in the current task, while moderately reducing the strategic influence of secondary objectives. For example, when the gradient angle between the two sub-objectives of feeding accuracy per unit resource use and the improvement of water exchange capacity due to cage spacing is greater than the set conflict angle threshold, and the system detects that the current feeding power is in a low surplus state, the reward weight of the feeding accuracy sub-objective per unit energy consumption is increased, and the weight of the water exchange efficiency improvement goal due to the expansion of cage spacing is correspondingly reduced, guiding the strategy to prioritize energy-saving feeding. This adjustment mechanism can be executed periodically during each round of training or online deployment.
[0079] Furthermore, in step S201, the method for obtaining a control strategy from a slowly varying reinforcement learning model for allocating fishery resources is: using the control strategy output by the slowly varying reinforcement learning model to dynamically adjust key control units in the marine fishery scenario, wherein the key control units include but are not limited to a feeding system, a power distribution device, and a cage spacing adjustment mechanism.
[0080] The slowly varying reinforcement learning model generates a specific control strategy based on the current environmental state and the results of a multi-objective reward function. This control strategy, as the model's output signal, is used to control the regulatory behavior of multiple key devices in offshore fisheries scenarios, such as precise automatic feeding control, spatial priority sorting of power distribution, and adjustment of cage floating positions. This process, based on the optimal action sequence output by the reinforcement learning policy network, achieves a balance between resource efficiency and aquaculture benefits, avoiding redundant adjustments and energy waste.
[0081] Furthermore, in step S300, the method for generating an adaptive representation value for the running process of the slowly varying reinforcement learning model in real time is: obtaining the reward value, strategy output result and strategy entropy state of the current reinforcement learning model, and performing difference measurement with the corresponding values at the previous moment: obtaining the change amplitude of the reward value, the degree of difference in the strategy output probability distribution, and the fluctuation of the strategy entropy, and using them as three dimensions for measuring adaptability; performing a weighted combination of the above three dimensions to form an adaptive representation value that can be continuously updated over time, which is used to represent the current strategy's ability to respond to changes in environmental states.
[0082] The current reinforcement learning model refers to the last or latest execution of the slowly varying reinforcement learning model, and the previous moment refers to the timing or moment of the last execution of the slowly varying reinforcement learning model in the reverse time direction.
[0083] The mathematical expression of the adaptive representation value is: Let the actual reward value RR obtained from the current reinforcement learning model t , the strategy output distribution πt , and strategy entropy STH t ; Then the change range of the reward value is: △RR t =R t -R t-1 , used to measure the volatility of the returns brought by the current strategy. The degree of difference in the probability distribution of the strategy output is D KL (π t ||π t-1 ), which is used to characterize whether the strategy output state is stable. The fluctuation of strategy entropy is △STH t =STH t -STH t-1 , used to reflect the certainty of the model action selection. Adaptive characterization value ATCV = α·|△RR t |+β·D KL (π t ||π t-1 )+γ·|△STH t |; where α, β and γ are preset weighting coefficients respectively.
[0084] Furthermore, in step S400, a transient transition model is constructed based on the adaptive characterization value. The method for obtaining the transient transition order value is as follows: any moment of obtaining the control strategy is used as a modeling point; an integer value is preset as the backtest value ST, and its value range is ST∈[3,10]; the exponential average of the adaptive characterization values corresponding to any modeling point to the previous ST and previous 2×ST modeling points is recorded as the short-term benchmark value STEma and the long-term benchmark value LTEma respectively; the absolute value of the difference between the short-term benchmark value and the long-term benchmark value is recorded as the absolute deviation; the ratio of the absolute deviation to the long-term benchmark value is calculated to obtain the fluctuation response D t ; ;
[0085] The critical transition order value is obtained by constructing a critical transition model based on the Sigmoid function through nonlinear mapping of the fluctuation response.
[0086] In order to effectively process the fluctuation response Dt and maximize its influence within a specific range, this step designs a transient transition model based on the Sigmoid function to perform nonlinear mapping on it, thereby obtaining the transient transition order value L at the current moment. t , the specific transient transition model is: ;
[0087] ATCV t Indicates the adaptive representation value corresponding to the current modeling point, which is used to perform nonlinear transformation on the fluctuation response; tPerforming a hyperbolic tangent transformation to convert the current state of the marine fishery environment into a range more suitable for the transient model, thus making the quantification of current conditions more flexible and effective;
[0088] Preferably, in step S400, a transient transition model is constructed based on the adaptive characterization value, and a method for obtaining the transient transition order value is as follows: any moment of obtaining a control strategy is used as a modeling point, any modeling point and its preceding Num modeling points are used as its dynamic modeling window, and the dynamic modeling window is constructed into a window eigenvalue sequence corresponding to the adaptive characterization value; Num is the number of modeling points in the current 12-48 hours;
[0089] Calculate the sea disturbance factor α through the dynamic modeling window t , environmental sudden change factor β t and the risk backlog factor γ t And bring the above factors into the critical transition model to calculate the critical transition order value Lt: L t =max(|α t |,|β t |)×exp(γ t ); The sea disturbance factor is used to determine the degree of instability and drastic changes in the parameters obtained by the sensor system in the current environment. Its mathematical expression is the standard deviation of the window eigenvalue sequence;
[0090] However, due to the problem that the transient transition model is sensitive to extreme values and cannot effectively capture the dynamic changes of data in actual applications, this sea disturbance factor often has the problem of insufficient sensitivity to extreme values. Therefore, another solution is provided. The implementation process is: define the window amplitude of the current modeling point as the range of the corresponding window eigenvalue sequence, and the amplitude stability coefficient PSI is the ratio of the minimum value of the window amplitude corresponding to each modeling point in the dynamic modeling window to the window amplitude;
[0091] The difference between each adaptive representation value and ATCV1 within the window eigenvalue sequence is calculated one by one. All differences are squared and weighted, and the weighted results are multiplied by the PSI and accumulated. The accumulated result is divided by the number of time steps (Num) and the square root is taken. The final output value is the sea disturbance factor. This factor quantifies the amplitude of the adaptive representation value's deviation from the average level and reflects the intensity of abnormal fluctuations in environmental parameters. ; where i1 is the cumulative variable, ATCV i1 and PSI i1 Represent the adaptive representation value and amplitude stability coefficient of the i1th modeling point in the reverse time direction at the current moment, respectively. e is a natural constant. The weighted processing aims to give higher weights to smaller i1 values, so that the adaptive representation value of the most recent time step has a greater impact on the final result.
[0092] The environmental sudden change factor is the risk gradient for the model to identify trend instability in advance: the absolute value of the difference between the adaptation characterization value step difference between the current modeling point and its previous modeling point is taken as the differential change, and the ratio of the differential change to the modeling time distance is the environmental sudden change factor;
[0093] The adaptive characterization value step difference refers to the difference in adaptive characterization value between any modeling point and its previous modeling point; the modeling time distance refers to the time distance between the current modeling point and its previous modeling point, in minutes.
[0094] The role of the environmental sudden change factor is to capture the acceleration or deceleration of the rate of change of the adaptive characteristic value and identify the evolutionary kinetic energy of the environmental deterioration trend. Its mathematical expression is: ; Among them, △ATCV t and ΔATCV t-1 They represent the adaptive representation value step difference between the current modeling point and its previous modeling point, and △t represents the modeling time distance corresponding to the current modeling point;
[0095] The cumulative amount of adaptive representation values exceeding the threshold within the dynamic modeling window is recorded as the risk backlog factor to quantify the degree of long-term existence of local abnormal states.
[0096] The specific calculation method is as follows: obtain the 90th percentile of the adaptive characterization value in the limited historical data as the abnormal threshold τ, traverse each adaptive characterization value in the dynamic modeling window, and if the adaptive characterization value is greater than the threshold τ, mark it as an abnormal point and calculate its threshold excess, that is, the difference between the adaptive characterization value and τ; if the value is less than or equal to τ, ignore the point; accumulate the excess values of all abnormal points to obtain the risk backlog factor. This factor quantifies the cumulative effect of the continuous deviation of the environmental state from the normal range and represents the energy reserve of systemic risk. Its mathematical expression is:
[0097] The time constraint for historical data is to limit the time length to at least one month. Because ocean data usually needs to adapt to tidal phenomena, 30 natural days is a relatively complete backtesting period. It can be adjusted to a longer time length, such as one quarter or one year, according to the actual modeling application direction.
[0098] Furthermore, in step S500, the method for making adaptive decisions through the critical transition order value is: setting 5-15 natural days as a reference time period, obtaining all critical transition order values within the reference time period as an order value set; if the current critical transition order value is less than the lower quartile of the order value set, it is defined that a transition mark appears; when the proportion of transition marks appearing in a natural day exceeds half, it is defined that the current environment is in a mutation risk state, suspending the model control output and triggering an early warning, or executing a preset safety control strategy.
[0099] Under sudden risk conditions, the system automatically suspends the control output of the slowly varying reinforcement learning model to prevent resource allocation imbalances caused by policy mismatch. This triggers one of two intervention measures: a remote manual alert prompting operations and maintenance personnel to take over immediately; and the system executes a pre-set safety control strategy, prioritizing key survival actions such as oxygen supply, flow control, and maintaining basic feeding, while limiting energy-intensive and precision-critical operations to maintain the system's minimum stable operating state. This mechanism effectively prevents uncontrolled decision-making by the reinforcement learning strategy under sudden risk conditions, improving the resilience and safety of the system's operations.
[0100] The embodiment of the present invention provides a marine fishery resource big data dynamic allocation system based on reinforcement learning, such as Figure 2 The figure shows a structural diagram of a marine fishery resource big data dynamic allocation system based on reinforcement learning of the present invention. The marine fishery resource big data dynamic allocation system based on reinforcement learning of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned embodiment of the marine fishery resource big data dynamic allocation method based on reinforcement learning are implemented.
[0101] The system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to run in the following units of the system:
[0102] The ocean scene initialization unit is used to initialize the marine fishery resource data scene and collect fishery data;
[0103] A reinforcement learning model building unit, used to build a slowly varying reinforcement learning model for multi-objective collaborative scheduling using fishery data;
[0104] An adaptation representation calculation unit, used to generate an adaptation representation value in real time during the running process of the slowly varying reinforcement learning model;
[0105] The imminent transition identification unit is used to construct an imminent transition model according to the adaptation characterization value and obtain an imminent transition order value.
[0106] The model execution decision unit is used to make adaptive decisions through the critical transition order value.
[0107] The reinforcement learning-based marine fishery resource big data dynamic allocation system can be run on computing devices such as desktop computers, laptops, PDAs, and cloud servers. The system can include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the example is merely an example of a reinforcement learning-based marine fishery resource big data dynamic allocation system and does not constitute a limitation on a reinforcement learning-based marine fishery resource big data dynamic allocation system. The system can include more or fewer components than the example, or a combination of certain components, or different components. For example, the reinforcement learning-based marine fishery resource big data dynamic allocation system can also include input and output devices, network access devices, buses, and the like.
[0108] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the operation system of the marine fishery resources big data dynamic allocation system based on reinforcement learning, and utilizes various interfaces and lines to connect the various parts of the entire marine fishery resources big data dynamic allocation system based on reinforcement learning.
[0109] The memory can be used to store the computer programs and / or modules. The processor implements the various functions of the reinforcement learning-based marine fishery resource big data dynamic allocation system by running or executing the computer programs and / or modules stored in the memory and accessing the data stored in the memory. The memory may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as sound playback or image playback); the data storage area may store data generated based on the use of the mobile phone (such as audio data and a phone book). Furthermore, the memory may include high-speed random access memory (RAM) and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0110] Although the present invention has been described in considerable detail and with particularity with respect to several embodiments, it is not intended to limit the present invention to any of these details or embodiments or any particular embodiment, so as to effectively encompass the intended scope of the present invention. In addition, the present invention has been described above with respect to embodiments foreseen by the inventors for the purpose of providing a useful description, and those insubstantial modifications of the present invention that are not currently foreseen may still represent equivalent modifications of the present invention.
Claims
1. A method for dynamic allocation of marine fishery resources big data based on reinforcement learning, characterized in that: The method comprises the following steps: S100, initializing the marine fishery resource data scene and collecting fishery data; S200, a slowly varying reinforcement learning model for multi-objective collaborative scheduling based on fishery data; S201, obtaining control strategies for fishery resource allocation from a slowly varying reinforcement learning model; S300, generating an adaptive representation value in real time during the running process of the slowly varying reinforcement learning model; S400, constructing an imminent transition model according to the adaptation characterization value and obtaining an imminent transition order value; S500, making an adaptive decision by using the transient transition order value; In step S200, the method for constructing a multi-objective collaborative scheduling slowly varying reinforcement learning model using fishery data is as follows: constructing a reinforcement learning model using fishery data, wherein the model uses a state compression mechanism to perform dimensionality reduction processing on the original marine environment data; The reinforcement learning model's optimization objectives were set as power consumption, cage spacing, water exchange efficiency, and feeding path length. A reinforcement learning loss function was constructed using a multi-objective reward structure. The multi-objective reward structure included feeding accuracy per unit resource usage, the degree to which cage spacing improves water exchange capacity, and the coverage of the feeding system per unit resource. The resulting reinforcement learning model was denoted as a slowly varying reinforcement learning model. Wherein, in step S200, the method for constructing a slowly-varying reinforcement learning model for multi-objective collaborative scheduling using fishery data also includes: introducing a target conflict identification function into the slowly-varying reinforcement learning model to detect the real-time directional deviation between each optimization target, the directional deviation being obtained by calculating the gradient vector angle of each target sub-reward item; when it is detected that the directional deviation between any two target sub-reward items exceeds a preset conflict threshold, the reward weight coefficient corresponding to each target in the slowly-varying reinforcement learning model is dynamically adjusted according to the state of the marine environment.
2. A method for dynamic allocation of marine fishery resources big data based on reinforcement learning according to claim 1, characterized in that: In step S100, the method for initializing the marine fishery resource data scene and collecting fishery data is as follows: a sensor system and a control device for marine environment monitoring and resource usage perception are deployed in the offshore aquaculture area, and the data collected by the sensor system and the control device are used as fishery data; The sensing system includes at least one water quality sensor, dissolved oxygen sensor, salinity meter, thermometer, wind and wave measuring instrument, energy consumption sensor and position positioning device; the control device includes a cage position control unit, an automatic feeding unit, an energy distribution execution module and a data communication module.
3. A method for dynamic allocation of marine fishery resources big data based on reinforcement learning according to claim 1, characterized in that: In step S201, the method for obtaining a control strategy from a slowly varying reinforcement learning model for fishery resource allocation is: using the control strategy output by the slowly varying reinforcement learning model to dynamically adjust key control units in the marine fishery scenario, wherein the key control units include a feeding system, a power distribution device, and a cage spacing adjustment mechanism.
4. A method for dynamic allocation of marine fishery resources big data based on reinforcement learning according to claim 1, characterized in that: In step S300, the method for generating an adaptive representation value in real time for the running process of the slowly varying reinforcement learning model is: obtaining the reward value, strategy output result and strategy entropy state of the current reinforcement learning model, and performing difference measurement with the corresponding values at the previous moment: obtaining the change amplitude of the reward value, the degree of difference in the strategy output probability distribution, and the fluctuation of the strategy entropy, and using them as three dimensions for measuring adaptability; performing a weighted combination of the above three dimensions to form an adaptive representation value that can be continuously updated over time, which is used to represent the current strategy's ability to respond to changes in environmental states.
5. The method for dynamic allocation of marine fishery resources big data based on reinforcement learning according to claim 1, characterized in that: In step S400, a transient transition model is constructed based on the adaptive characterization value. The method for obtaining the transient transition order value is as follows: any moment of obtaining the control strategy is used as a modeling point; an integer value is preset as a backtest value ST; the exponential average values of the adaptive characterization values corresponding to any modeling point to the previous ST and 2×ST modeling points are recorded as the short-term reference value STEma and the long-term reference value LTEma, respectively; the absolute value of the difference between the short-term reference value and the long-term reference value is recorded as the absolute deviation; The fluctuation response D is calculated by calculating the ratio of the absolute deviation to the long-term reference value. t ; The critical transition order value is obtained by constructing a critical transition model based on the Sigmoid function through nonlinear mapping of the fluctuation response.
6. A method for dynamic allocation of marine fishery resources big data based on reinforcement learning according to claim 1, characterized in that: In step S400, a transient transition model is constructed based on the adaptive characterization value. The method for obtaining the transient transition order value is as follows: any moment of obtaining the control strategy is used as a modeling point, any modeling point and its preceding Num modeling points are used as its dynamic modeling window, and the dynamic modeling window is constructed into a window eigenvalue sequence corresponding to the adaptive characterization value; Calculate the sea disturbance factor α through the dynamic modeling window t , environmental sudden change factor β t and the risk backlog factor γ t And bring the above factors into the critical transition model to calculate the critical transition order value Lt: L t =max(|α t |,|β t |)×exp(γ t ); The sea disturbance factor is used to determine the degree of instability and drastic changes in the parameters obtained by the sensor system in the current environment. Its mathematical expression is the standard deviation of the window eigenvalue sequence; The environmental sudden change factor is the risk gradient for the model to identify trend instability in advance: the absolute value of the difference between the adaptation characterization value step difference between the current modeling point and its previous modeling point is taken as the differential change, and the ratio of the differential change to the modeling time distance is the environmental sudden change factor; The cumulative amount of adaptive representation values exceeding the threshold within the dynamic modeling window is recorded as the risk backlog factor to quantify the degree of long-term existence of local abnormal states.
7. The method for dynamic allocation of marine fishery resources big data based on reinforcement learning according to claim 1, characterized in that: In step S500, the method for making adaptive decisions through the critical transition order value is: setting 5-15 natural days as a reference time period, obtaining all critical transition order values within the reference time period as an order value set; if the current critical transition order value is less than the lower quartile of the order value set, it is defined that a transition mark appears; when the proportion of transition marks appearing in a natural day exceeds half, it is defined that the current environment is in a mutation risk state, suspending the model control output and triggering an early warning, or executing a preset safety control strategy.
8. A marine fishery resource big data dynamic allocation system based on reinforcement learning, characterized by: The marine fishery resource big data dynamic allocation system based on reinforcement learning includes: a processor, a memory, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the steps of the marine fishery resource big data dynamic allocation method based on reinforcement learning described in any one of claims 1 to 7 are implemented. The marine fishery resource big data dynamic allocation system based on reinforcement learning runs on desktop computers, laptops, PDAs, and computing devices in cloud data centers.
Citation Information
Patent Citations
Big data dynamic allocation and optimal scheduling method based on reinforcement learning
CN119311407A
USV formation path-following method based on deep reinforcement learning
US20220004191A1