Generation strategy optimization method and device based on dynamic environment, equipment and medium

By constructing a dynamic environment state vector, generating a compliant action vector and constructing a multi-dimensional reward vector, and combining it with an adaptive strategy optimization module to update the model, the problems of insufficient multi-source heterogeneous data fusion and constraint control in existing reinforcement learning methods are solved, and the real-time, stability and compliance of strategy generation are improved.

CN120746331APending Publication Date: 2025-10-03PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510862157.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing reinforcement learning methods in the financial and medical fields lack the dynamic fusion and constraint control of multi-source heterogeneous data, resulting in the inability to generate strategies that take into account real-time performance, stability, and domain compliance.

Method used

A dynamic environment state vector is constructed based on multi-source heterogeneous data streams. An action vector is generated through a pre-trained generative model, which is then corrected in combination with domain constraint strategies to generate a compliant action vector. A multi-dimensional reward vector is constructed, and the model is updated using an adaptive strategy optimization module.

Benefits of technology

It enhances the ability to perceive changes in complex scenarios, improves the compliance and practicality of strategy execution, avoids single-target guidance errors, and achieves the stability and long-term effectiveness of strategy output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120746331A_ABST
    Figure CN120746331A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a generation strategy optimization method, device and equipment based on a dynamic environment and a medium. Correcting the generated action vector by combining a domain constraint strategy to obtain a compliant action vector, constructing a multi-dimensional reward vector according to feedback after execution, scaling the reward vector into a reward signal, and finally updating the pre-training generative model by adopting a self-adaptive strategy optimization module based on the reward signal and an interaction track to obtain a reward result. And collaborative evolution of strategy generation and environmental response is realized. According to the method, by introducing dynamic environment information and domain constraints, compliance correction and optimization updating of the generative strategy are realized, and the stability and practicability of the model in a complex environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and storage medium for optimizing a generation strategy based on a dynamic environment. Background Art

[0002] In the fintech sector, reinforcement learning methods are being applied to tasks such as quantitative trading, risk control, and asset allocation. However, existing technologies often rely on static feature sets, such as historical price series and technical indicators, for environmental modeling, lacking the ability to integrate dynamically changing factors such as macroeconomic variables, policy events, and market sentiment. This static modeling approach makes it difficult for strategies to adapt to real-time environmental fluctuations, especially during sudden policy changes or extreme market events, which can easily lead to model failure. Furthermore, in terms of reward mechanism design, traditional methods often focus solely on maximizing returns, neglecting the integration of risk control and stability indicators. This lack of a multi-dimensional reward structure makes it difficult to ensure the sustainability and stability of strategies in actual implementation. Furthermore, compliance rules, regulatory requirements, and industry experience in the financial sector are typically not incorporated into the reinforcement learning strategy generation process, resulting in the risk that the strategies generated by the models may not comply with regulatory constraints, increasing the complexity of implementation.

[0003] In the healthcare sector, reinforcement learning is being used for applications such as generating intelligent diagnosis and treatment recommendations and recommending patient intervention strategies. However, current systems mostly use structured electronic medical record data or sensor data as core state inputs, and are unable to fully integrate heterogeneous data sources such as patient behavior, policy updates, and changes in medical resources, limiting the model's environmental perception capabilities. At the same time, because the medical process itself is long-term goal-oriented and requires stability, existing reward mechanisms struggle to simultaneously capture the balance between short-term intervention effects and long-term treatment expectations. During the action generation phase, the lack of constraints on knowledge in areas such as medical norms, diagnosis and treatment pathways, or ethical restrictions leads to potential business irrationality and ethical risks in the generated actions.

[0004] Overall, existing reinforcement learning methods suffer from systemic deficiencies in dynamic environment modeling, multi-dimensional reward design, and domain constraint integration, hindering their generalization and controllability in demanding industries like finance and healthcare. Implementing dynamic state modeling driven by multi-source heterogeneous data and building a compliant, robust, and adjustable strategy generation and optimization process are key technical challenges that must be addressed in the implementation of reinforcement learning methods. Summary of the Invention

[0005] The main purpose of the present invention is to provide a generation strategy optimization method, device, equipment and storage medium based on a dynamic environment, aiming to solve the technical problem that the existing technology lacks dynamic fusion and constraint control of multi-source heterogeneous data in reinforcement learning environment modeling, resulting in the inability of strategy generation to take into account real-time, stability and domain compliance.

[0006] To achieve the above object, the present invention provides a generation strategy optimization method based on a dynamic environment, comprising:

[0007] Construct dynamic environment state vector based on multi-source heterogeneous data streams;

[0008] Based on the dynamic environment state vector, generating an action vector through a pre-trained generative model;

[0009] According to the knowledge graph containing the domain constraint strategy, the action vector is modified to generate a compliant action vector;

[0010] generating a multidimensional reward vector including a performance metric, a risk metric, and a stability metric based on feedback after the execution of the compliant action vector, and scalarizing the multidimensional reward vector into a reward signal;

[0011] Based on the reward signal and the interaction trajectory including the compliant action vector, an adaptive policy optimization module is used to update the pre-trained generative model.

[0012] Furthermore, to achieve the above-mentioned purpose, the present invention provides a generation strategy optimization device based on a dynamic environment, comprising:

[0013] The environment perception module is used to construct a dynamic environment state vector based on multi-source heterogeneous data streams;

[0014] An action generation module, configured to generate an action vector based on the dynamic environment state vector using a pre-trained generative model;

[0015] A policy constraint module, configured to modify the action vector according to a knowledge graph containing domain constraint policies to generate a compliant action vector;

[0016] a reward evaluation module, configured to generate a multidimensional reward vector including a performance metric, a risk metric, and a stability metric based on feedback after the execution of the compliant action vector, and scalarize the multidimensional reward vector into a reward signal;

[0017] A model optimization module is configured to update the pre-trained generative model using an adaptive strategy optimization module based on the reward signal and the interaction trajectory including the compliant action vector.

[0018] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a dynamic environment-based generation strategy optimization program stored in the memory and runnable on the processor. When the dynamic environment-based generation strategy optimization program is executed by the processor, the steps of the dynamic environment-based generation strategy optimization method as described above are implemented.

[0019] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a generation strategy optimization program based on a dynamic environment is stored. When the generation strategy optimization program based on a dynamic environment is executed by a processor, the steps of the generation strategy optimization method based on a dynamic environment as described above are implemented.

[0020] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. It discloses a generation strategy optimization method, device, equipment and medium based on a dynamic environment, including: constructing a dynamic environment state vector based on multi-source heterogeneous data streams, using the vector to generate an action vector, combining the domain constraint strategy to correct the generated action vector, obtaining a compliant action vector, and constructing a multi-dimensional reward vector based on the feedback after its execution, scalarizing the reward vector into a reward signal, and finally updating the pre-trained generative model based on the reward signal and the interaction trajectory using an adaptive strategy optimization module to achieve the coordinated evolution of strategy generation and environmental response. The present invention dynamically constructs the environment state by fusing multi-source heterogeneous data to enhance the perception of complex scene changes; introduces a domain constraint strategy to perform compliance correction on the generated action, thereby improving the compliance and practicality of the strategy execution; combines performance, risk and stability to construct a multi-dimensional reward vector to avoid single-target guidance errors, and realizes dynamic update and robust optimization of the generative model through adaptive strategy optimization, effectively improving the stability and long-term effectiveness of the strategy output. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0022] Figure 1 A schematic diagram of an application environment of a generation strategy optimization method based on a dynamic environment in an embodiment of the present invention;

[0023] Figure 2 This is a flow chart of an embodiment of a method for optimizing a generation strategy based on a dynamic environment according to the present invention;

[0024] Figure 3 Schematic diagram of functional modules of a preferred embodiment of a generation strategy optimization device based on a dynamic environment of the present invention;

[0025] Figure 4A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0026] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0028] The generation strategy optimization method based on dynamic environment provided by the embodiment of the present invention can be applied in Figure 1 In an application environment, the user end communicates with the server end through a network. The server end can construct a dynamic environment state vector based on multi-source heterogeneous data streams through the user end, use this vector to generate an action vector, and modify the generated action vector in combination with a domain constraint strategy to obtain a compliant action vector. Based on the feedback after its execution, a multi-dimensional reward vector is constructed, and the reward vector is scalarized into a reward signal. Finally, based on the reward signal and the interaction trajectory, an adaptive strategy optimization module is used to update the pre-trained generative model to achieve the co-evolution of strategy generation and environmental response. The present invention dynamically constructs the environment state by fusing multi-source heterogeneous data to enhance the ability to perceive complex scene changes; introduces a domain constraint strategy to modify the compliance of the generated action, thereby improving the compliance and practicality of strategy execution; combines performance, risk and stability to construct a multi-dimensional reward vector to avoid single-target guidance errors, and realizes dynamic update and robust optimization of the generative model through adaptive strategy optimization, effectively improving the stability and long-term effectiveness of the strategy output. The user end can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server end can be implemented as an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0029] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of a method for optimizing a generation strategy based on a dynamic environment provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0030] like Figure 2 As shown, the generation strategy optimization method based on dynamic environment proposed by the present invention includes the following steps:

[0031] S10, constructs a dynamic environment state vector based on multi-source heterogeneous data streams;

[0032] In this embodiment, multi-source heterogeneous data streams refer to data sets with different sources, structures, and update cycles, including structured, semi-structured, and unstructured forms, with information content that is independent of each other but potentially related. For example, real-time time series data streams originate from sensor acquisition or system logs and appear as continuously changing numerical sequences; policy text data streams typically originate from regulatory notifications, government documents, or market announcements and exist in the form of natural language text; and unstructured data streams may include text, images, or voice data such as news comments, social media content, and voice call transcriptions.

[0033] When constructing a dynamic environmental state vector, the real-time time series data stream is first connected to the environmental modeling module. This input is received and cached via a distributed data channel. A sliding window mechanism is then used to segment the data into time segments. For each time segment, statistical indicators reflecting volatility, such as standard deviation, skewness, and maximum drawdown ratio, are extracted. This extraction process typically relies on statistical analysis methods such as sliding window weighted averaging and quantile analysis to obtain features reflecting the speed and magnitude of state evolution.

[0034] After receiving the policy text data stream, it must first undergo language parsing and text cleaning, including preprocessing such as sentence segmentation, noise symbol removal, and synonym normalization. Then, a policy intent recognition model or key factor identifier is used to extract influencing factors, such as "easing monetary policy" in the financial sector or "raising the prevention and control level" in the medical context. The extraction model can use a Transformer structure, a word-graph network, or a domain-named entity recognition system. The output is structured factor labels and their impact direction values.

[0035] Domain-aware sentiment analysis models can be used to capture sentiment polarity values ​​in unstructured data streams. Emotional recognition encompasses not only positive and negative sentiment, but also multiple categories such as expected positive, expected negative, and uncertain. These models are often fine-tuned based on pre-trained language models. Contextual relevance and conceptual specificity must be considered during processing to ensure accurate and domain-relevant sentiment mapping.

[0036] Unifying and fusing these three types of feature information requires normalization and time alignment. The temporal granularity and structural representation of different data types vary significantly, necessitating a multi-source alignment mechanism for synchronization. Volatility indicators, policy factors, and sentiment polarity values ​​are encoded into a vector representation with a unified dimension. A weighted fusion mechanism is then used to calculate a vector representing the overall state of the current environment. This vector forms the foundational representation for subsequent models to perceive environmental changes, ensuring timeliness, contextual relevance, and structural diversity.

[0037] In practice, different data collection and processing methods can be used to adapt to different business scenarios. In some scenarios, simply accessing real-time time-series data streams is sufficient for preliminary modeling. For applications requiring higher generalization capabilities, policy text data streams or unstructured text data streams can be further integrated. The data update frequency can also be adjusted based on changing scenarios, for example, sampling financial markets at the minute level while synchronizing policy updates on a daily basis.

[0038] Volatility indicators can be extracted through GARCH modeling to characterize the volatility structure, or a more lightweight calculation method based on sliding standard deviation, suitable for edge computing devices. Policy factors can be extracted using keyword matching and rule trees. This allows for rapid deployment in scenarios with limited model deployment costs. When resources are sufficient, deep semantic models such as BERT can be introduced to achieve higher recognition accuracy.

[0039] To obtain sentiment polarity values, models of varying sizes can be selected based on the length and complexity of the text. For example, for short text on social platforms, a lightweight LSTM architecture can provide sufficiently accurate results. For policy commentary in the form of complete documents, a larger model can be used to extract deeply nested representations. Vector fusion strategies can be selected, including dynamic fusion based on attention weights or fixed-ratio linear weighting. The specific selection can be adjusted dynamically based on performance on historical validation sets.

[0040] Example: In the healthcare sector, real-time time-series data is collected for hospital bed utilization and ICU occupancy rates. This data is combined with regional epidemic prevention and control policy texts released by the National Health Commission, as well as patient online comments and sentiment about the epidemic on social media platforms. These three types of data are integrated to construct a state vector that drives a dynamic bed resource allocation model, improving awareness of sudden epidemic spreads or increased local medical pressures.

[0041] In the fintech sector, when constructing an environmental state vector, historical stock or bond volatility data can be used as time series input. This data is then integrated with key words from the latest central bank monetary policy or fiscal measures, along with sentiment fluctuations in investor confidence and panic indices from market commentary. This generates a dynamic vector reflecting the overall market dynamics, guiding portfolio adjustments and risk exposure management. This approach enables more timely and robust responses to drastic changes in macroeconomic variables.

[0042] This embodiment integrates structured volatility features, semantic policy factors, and unstructured sentiment information into a unified model and generates an environment vector, significantly improving the timeliness of environmental representation and external responsiveness. During model training and inference, this state vector more accurately reflects the current state of the external environment, avoiding overfitting or reduced robustness of strategies caused by a single feature.

[0043] S20, generating an action vector through a pre-trained generative model based on the dynamic environment state vector;

[0044] In this embodiment, the dynamic environment state vector is a structured representation of the environment, used to characterize the comprehensive state of the external environment at a specific point in time, generated by the fusion of multi-source heterogeneous data streams. This vector is typically a multidimensional real-valued vector that encodes various factors, including macro-trends, local anomalies, and expected changes. It has the ability to fuse features across different types and time levels. Its design goal is to provide a comprehensive, real-time state foundation for subsequent strategy generation, enabling the model to accurately respond to complex environments.

[0045] Pre-trained generative models are policy models whose parameters are initialized through large-scale offline data training. They possess transfer learning capabilities and rapid adaptability. During deployment, the model's trained parameters, including feature encoding parameters, policy network structure, activation function configuration, and normalization module, are first loaded, enabling it to directly perform action prediction tasks in new environments. The model architecture can adopt Transformer, GRU, variational autoencoder, and other formats, depending on the specific scenario and computing resources.

[0046] Before inputting the dynamic environment state vector into the model, it must be normalized and dimensionally adjusted to ensure that it meets the model's acceptance requirements in terms of numerical scale, distribution structure, and missing value imputation. The first stage after input is the feature encoding layer, which is responsible for mapping the raw state vector into an internal latent space representation. Encoding methods can include linear mapping, convolution, or a multi-head attention mechanism, aiming to enhance the representation of environmental information in high-dimensional space and capture nonlinear relationships between variables.

[0047] After the feature representation is generated, it is fed into the policy network for action generation. The policy network typically consists of several layers of feedforward networks, and its output is an action probability distribution, representing the expected quality of each executable action under the current environment state. This probability distribution is obtained using a softmax activation function, and the actual actions are selected through sampling to form an action vector. The action vector is the decision output of the model and can be a discrete action encoding, a set of continuous variables, or a policy instruction with a multi-layered nested structure, depending on the type of downstream execution mechanism.

[0048] The action vector not only serves as the output behavior instruction at the current moment, but will also be used for constraint correction, feedback evaluation, and strategy update in subsequent steps. Therefore, its semantic consistency, structural stability, and cross-step traceability must be guaranteed.

[0049] Different model architectures and processing methods can be used in different business scenarios. A simple two-layer feedforward neural network can be used as a generative model, while a transformation module with a self-attention mechanism can be introduced to enhance the ability to model long-term dependencies. When the state vector is high-dimensional or the input is redundant, dimensionality reduction can be performed using principal component analysis or autoencoders to improve the model's computational efficiency and stability.

[0050] For high-frequency scenarios, parallel batch processing techniques can be introduced to simultaneously input multiple sets of state vectors into the model for forward inference. For low-latency scenarios, lightweight model structures can be used in conjunction with quantized deployment to meet real-time requirements. The sampling mechanism for the action probability distribution can also be configured based on strategic objectives. For example, temperature-adjusted softmax sampling can be used during the exploration phase, while argmax operations can be used during the convergence phase to obtain the optimal action output.

[0051] Furthermore, a gating mechanism can be introduced to determine whether the current state meets the prerequisites for generating an action, preventing invalid actions from being generated in incomplete or abnormal environments. Generated action vectors can also be accompanied by credibility labels or coefficients of variation to assist subsequent modules in determining whether constraint checking or regeneration is necessary.

[0052] Example: In the healthcare sector, a dynamic environment state vector might include information such as current patient flow trends, hospital bed occupancy rates, and regional infection risk levels. By inputting this vector into a pre-trained generative model, a resource allocation action vector can be generated, indicating strategic actions such as opening additional wards and activating backup medical staff, thereby enabling real-time response to sudden outbreaks.

[0053] In the fintech sector, state vectors can incorporate indicators such as market volatility, policy intensity, and investor sentiment. By generating action vectors, asset reallocation strategies can be output, such as increasing the proportion of safe-haven assets or reducing positions in high-volatility assets. This helps robo-advisory systems make robust adjustments during market fluctuations, improving the stability of overall strategy returns.

[0054] This embodiment loads a pre-trained generative model and inputs a dynamic environment state vector for action generation, enabling rapid adaptive policy responses to complex environments while reducing online training resource consumption. Because the pre-training phase encompasses a variety of historical scenarios, the model possesses transferable and generalizable capabilities, enabling it to make stable decisions even in the face of sudden environmental changes, improving the system's robustness and responsiveness in complex environments.

[0055] S30, modifying the action vector according to the knowledge graph containing the domain constraint strategy to generate a compliant action vector;

[0056] In this embodiment, domain constraint policies are abstracted constraints derived from rules, practice boundaries, laws, regulations, or empirical knowledge within a specific business domain. They can include structured rules, logical inferences, upper and lower limits, combined conditions, and penalty factors. They are widely found in contexts such as financial compliance clauses, medical operating procedures, and resource scheduling specifications. These policies are typically expressed with conditional triggers and target constraints, ensuring that policy execution results adhere to regulatory boundaries or ethical standards.

[0057] In this process, the knowledge graph plays a core role in organizing domain knowledge and mapping constraints. It encodes the dependencies between domain objects using entity-relationship-attribute triples, capable of storing a variety of structured or semi-structured constraint information, including policy rules, resource constraints, and behavioral norms. Knowledge graphs can be constructed using entity recognition and relationship extraction techniques, or they can be manually maintained using a combination of rule bases and expert annotation. Their topological structure offers excellent scalability and upstream and downstream traceability.

[0058] An action vector represents the execution intent generated by the policy model based on the current environment state. This may be a parameterized sequence of operations, a resource allocation plan, a policy switching instruction, etc. If not subject to compliance verification and correction, the action may violate regulatory restrictions, incur risks, or lead to execution failure. Therefore, it is necessary to filter and correct the action in conjunction with the domain constraint policy.

[0059] The first step in correcting an action vector is to retrieve constraints relevant to the current context. This is accomplished by matching semantic features in the knowledge graph, such as task labels, object types, and time attributes, to obtain a set of applicable constraints. A logical reasoning module or rule-checking engine then determines whether the action vector violates the current constraints. Violations typically include exceeding thresholds, logical inconsistencies, conflicts in associated entities, or unmet dependencies.

[0060] Once a violation is detected, the system further calculates the degree of deviation—the distance between the action vector and the compliance boundary in vector space—and combines the constraint weights and penalty parameters in the knowledge graph to generate one or more correction direction and correction magnitude recommendations. This process can be implemented using gradient projection, symbolic rule backoff, or semantic adjustment networks.

[0061] The correction parameter value is a vectorized operation factor that can take the form of incremental offset, coefficient scaling, discrete action replacement, or dimension activation modulation. Ultimately, this parameter is applied to the original action vector to generate a compliant action vector that has undergone compliance verification and ensures that the output result is executable within the constraint space described by the knowledge graph.

[0062] In businesses driven by structured rules, a knowledge graph execution system based on a rule engine can be used to build constraint rule sets and perform matching and correction based on if-then logic. In multi-dimensional continuous space strategy scenarios, a neural network embedding approach can be used to construct a knowledge graph representation. Graph convolutional neural networks can be used to embed entity features and relationships into a unified constraint space, and then compared and corrected with action vectors in a unified semantic space.

[0063] For the graph retrieval methods relied upon during compliance verification, a subgraph traversal mechanism based on query path optimization can be used to improve the efficiency of context constraint extraction. Alternatively, the current policy environment state can be used as a query anchor, and a graph attention mechanism can be introduced to enhance the perception of relevant policy nodes. During the correction phase, for discrete action types, the maximum compliance similarity method can be used for alternative selection; for continuous action types, a penalty function-based directional adjustment method can be used to complete compliance migration.

[0064] During the actual deployment process, a dynamic adjustment mechanism for constraint strength can also be introduced to enable the system to prioritize stricter correction paths during periods of higher risk or regulatory fluctuations, while reducing constraint strength during the exploration or pre-research stage to release space for strategic innovation.

[0065] Example: In the healthcare business domain, assume the generated action vector is a resource reallocation policy for intensive care units. The knowledge graph contains entities and constraints such as current bed saturation, postoperative recovery period requirements, and patient infection isolation levels. By mapping policy intent with resource constraints, the action that consolidates multiple postoperative patients into the same area is modified to a decentralized configuration, ensuring that isolation requirements are met.

[0066] In the fintech business field, when the action vector output by the model recommends increasing the position of high-risk assets, and the regulatory rules bound to the knowledge graph stipulate that high leverage configuration should be restricted under the current volatility level, the system will verify the action based on the knowledge graph and automatically correct the suggestion to dynamic rebalancing within the low-risk asset pool, thereby avoiding potential policy violations and risk accumulation problems.

[0067] This embodiment uses the domain constraint strategy contained in the knowledge graph to perform structural corrections on the action vectors generated by the policy model, so that the final action output can meet the business execution boundaries and institutional compliance requirements based on the effectiveness of the strategy, avoiding the situation where the policy model deviates from the feasible scope when pursuing local optimality, and effectively improving the availability, controllability and compliance of the system output results.

[0068] S40, generating a multidimensional reward vector including a performance metric, a risk metric, and a stability metric based on feedback after the execution of the compliant action vector, and scalarizing the multidimensional reward vector into a reward signal;

[0069] In this embodiment, after a compliant action vector is executed in a specific environment, it generates feedback information. This feedback information carries the actual impact of the action execution results on the environment's evolution. Feedback information can include data features reflecting the execution consequences, such as environmental state transitions, payoff changes, policy execution costs, volatility parameters, and duration. Feedback information serves as an external basis for evaluating the actual effectiveness of the policy model and provides irreplaceable real-world observability.

[0070] In order to accurately characterize the value results after the execution of the strategy, it is necessary to construct a multi-dimensional reward vector based on the feedback information. Each dimension of the reward vector represents the performance of the strategy execution results on a certain type of target. The three key dimensions are performance measurement, risk measurement and stability measurement. Performance measurement represents the direct contribution of the strategy action to the target return or state improvement, which can be measured using state transfer gain, cumulative return or target indicator change rate. Risk measurement is used to describe the uncertainty fluctuations caused by the execution behavior, which can be calculated through indicators such as volatility, deviation from the historical mean, and tail loss probability. Stability measurement focuses on the persistence of the strategy effect on continuous time slices, which is achieved through means such as consistency score within the sliding window, periodic response variance or continuous effectiveness window length.

[0071] After the multidimensional reward vector is formed, it needs to be compressed into a single scalar form, namely the reward signal, for ease of use in subsequent optimization modules. The reward signal is used to guide the optimization direction of the generative model and is the basis of the objective function during the training process. The scalarization process needs to fully preserve the weight relationship of multidimensional information and cannot be simply summed or averaged. To achieve this goal, a three-layer metric weight distribution scheme can be constructed. The first layer sets the indicator priority based on environmental sensitivity or strategy importance. The second layer introduces objective function coupling terms to achieve relative weight fine-tuning. The third layer dynamically adjusts the current reward weighting strategy based on historical training feedback. Finally, through a weighted fusion strategy, the performance metric, risk metric, and stability metric are mapped into a reward signal that expresses the combined effect.

[0072] A state transition extraction module can be implemented based on a structured feedback data flow, calculating a performance improvement score by comparing it to the original state. Based on this, a sliding window risk calculator is used to assess the intensity of outcome fluctuations caused by policy execution, and a risk score is constructed based on empirical distributions. For stability metric assessment, a periodic indicator stability test can be employed, and a window-based variance threshold control mechanism can be used to ensure the steady-state consistency of the policy's continuous output. For areas where policy impact has a delayed effect, a hysteresis integration module can be introduced to account for this delayed effect in the reward function.

[0073] During reward scalarization, a weighted controller based on logic gating can be constructed to activate different weight paths as needed, for example, amplifying the weight of risk items during high-risk periods and strengthening the weight of stability items during periods of poor execution sustainability. Alternatively, a fuzzy control mechanism can be used to generate fuzzy membership functions based on input three-dimensional indicators, and output reward signals after fuzzy inference using a rule table. Furthermore, the scalarization module can be connected to a dynamic policy evaluation recorder, leveraging historical policy trajectories to perform exponential smoothing on the current weighting coefficients, thereby achieving cross-cycle adaptive optimization.

[0074] Example: In the healthcare business domain, the generated compliant action vectors can correspond to surgical scheduling optimization or ICU resource allocation recommendations. Execution feedback includes changes in patient recovery rate, postoperative complication probability, and resource load stability. By constructing a multidimensional vector consisting of performance metrics (such as a reduction in postoperative recovery days), risk metrics (such as an increase in the probability of abnormal recovery), and stability metrics (such as a change in resource utilization over three consecutive days), and combining it with scheduling priorities to generate reward signals, high-quality surgical allocation strategy training can be achieved.

[0075] In the fintech business sector, the execution of a compliance action vector might manifest as an asset reallocation plan, with feedback including portfolio returns, volatility, and monthly allocation frequency changes. These feedback metrics are converted into a multidimensional reward vector, and a weighting function is applied based on current market risk levels and investment objective preferences. This ultimately generates a dynamically adjustable reward signal to guide generative model iteration, enabling the system to seize market opportunities while mitigating short-term volatility and institutional boundary conflicts.

[0076] This embodiment constructs a multidimensional reward vector by introducing three types of measurement indicators: performance, risk, and stability in execution feedback, and uses a weighting strategy to compress this multidimensional structure into a single reward signal. This can take into account short-term benefits, long-term stability, and potential risk control requirements in model training, avoid overfitting problems under a single goal orientation, and thus improve the actual adaptability of the strategy model in complex environments.

[0077] S50: Based on the reward signal and the interaction trajectory including the compliant action vector, an adaptive strategy optimization module is used to update the pre-trained generative model.

[0078] In this embodiment, the reward signal is a compressed representation of the effectiveness of a compliant action vector. Its value lies in guiding the continuous optimization of the generative model through quantified feedback. During reinforcement training, the reward signal merely represents the target direction; the actual optimization effect depends on its synergy with the interaction trajectory. The interaction trajectory is a time series structure that records the execution history of compliant action vectors and the evolution of their corresponding environment states. It includes not only the actions themselves but also data points such as the model's response under different states, state changes, and feedback impact, forming a dynamic foundation for evaluating strategy effectiveness.

[0079] Based on this data structure, by constructing a complete interaction trajectory dataset, we can correlate reward signals with behavioral context, thereby extracting valuable optimization cues. To more effectively implement the training iteration process, we need to introduce an adaptive policy optimization module. This module dynamically receives reward signals and behavioral trajectory data from the environment, automatically generates an update policy for adjusting model parameters, and determines the update magnitude and direction based on the current environmental state and policy stability.

[0080] The optimization process involves analyzing the gradient trends, policy deviations, and goal achievement levels in the reward signal. It then extracts the coupled features of policy expressiveness and environmental response from the trajectory, constructing a joint state-action-feedback tensor as training input. Before generating the optimization gradient, adaptive parameter configuration is performed based on the current environmental state, such as adjusting the learning rate, update step size, or stability suppression coefficient. Finally, by calculating the parameter update gradient, the weights of the pre-trained generative model are partially or fully updated.

[0081] The model training process can be implemented using an interactive trajectory replay mechanism. A set of four tuples containing historically compliant action vectors, corresponding state vectors, reward signals, and next states is pre-built. During each training round, a batch of trajectories is sampled from this set and fed into the optimization module. Temporal difference learning is introduced into trajectory data processing. A gradient approximation is constructed based on the reward difference between the current and next states, and then combined with the reward signal to determine the dominant direction.

[0082] For the construction of the policy optimization module, the policy gradient method from reinforcement learning can be used to perform gradient descent updates on the probability distribution of the target policy. Alternatively, the proximal policy optimization mechanism can be combined to adjust the policy within a defined change boundary, ensuring that changes in model parameters do not cause policy mutations. If there are large environmental fluctuations or non-stationary tasks, state normalization and trajectory stability filtering mechanisms can be further introduced to compress or mark abnormal trajectory data to prevent excessive interference with the model.

[0083] The adaptive mechanism can also be achieved through a dynamic learning rate adjustment network, scaling the current gradient impact factor according to the effect of the previous round of updates, or by using an environmental change trend prediction model to construct a parameter scheduling table in advance to achieve forward-looking parameter control.

[0084] Example: In the healthcare sector, interaction trajectories can be generated by a patient diagnosis and treatment process control system, including patient status transition records, doctor's order response data, and treatment path deviations. Reward signals reflect treatment costs and the degree of prognosis improvement. Building a strategy optimization path based on this data structure can encourage generative models to continuously generate decision actions that are resource-efficient and highly stable.

[0085] In the fintech sector, interaction traces can include data such as historical trading behavior, market state responses, and compliance feedback assessments. Reward signals represent risk-return ratios, volatility control capabilities, and compliance scores. By incorporating these signals into the strategy optimization module, we can continuously adjust the action distribution of the generative model, improving the model's asset allocation decision-making effectiveness in dynamic markets and preventing strategy failures due to regulatory non-compliance or excessive trading volatility.

[0086] This embodiment dynamically adjusts the pre-trained generative model by coordinating the interaction trajectory of the reward signal and the compliant action vector into the adaptive policy optimization module. This not only improves the strategy's efficiency in achieving its goals, but also enhances its adaptability in complex dynamic environments, avoiding falling into local optimality or policy collapse problems, thereby continuously optimizing the behavior generation quality and compliance effectiveness.

[0087] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. It discloses a generation strategy optimization method, device, equipment and medium based on a dynamic environment, including: constructing a dynamic environment state vector based on multi-source heterogeneous data streams, using the vector to generate an action vector, correcting the generated action vector in combination with a domain constraint strategy, obtaining a compliant action vector, and constructing a multi-dimensional reward vector based on the feedback after its execution, scalarizing the reward vector into a reward signal, and finally updating the pre-trained generative model based on the reward signal and the interaction trajectory using an adaptive strategy optimization module to achieve the coordinated evolution of strategy generation and environmental response. The present invention dynamically constructs the environment state by fusing multi-source heterogeneous data to enhance the perception of complex scene changes; introduces a domain constraint strategy to perform compliance correction on the generated action, thereby improving the compliance and practicality of the strategy execution; constructs a multi-dimensional reward vector in combination with performance, risk and stability to avoid single-target guidance errors, and realizes dynamic update and robust optimization of the generative model through adaptive strategy optimization, effectively improving the stability and long-term effectiveness of the strategy output.

[0088] In one embodiment, the above step S10 includes:

[0089] S101, access to real-time time series data streams, policy text data streams, and unstructured data streams;

[0090] S102, extracting a volatility index from the real-time time series data stream;

[0091] S103, extracting key expected factors from the policy text data stream;

[0092] S104, determining a text sentiment polarity value based on the unstructured data stream;

[0093] S105 , fusing the volatility index, key expectation factors, and text sentiment polarity values ​​to generate a dynamic environment state vector.

[0094] In this embodiment, access to real-time time series data streams, policy text data streams, and unstructured data streams is a basic link in constructing a dynamic environment state vector, which aims to achieve the fusion of multimodal input structures. Real-time time series data streams generally refer to structured quantitative indicators that are continuously generated in a time series, covering price changes, production statistics, equipment readings, etc., and can be accessed through API interfaces to cloud databases, sensor nodes, or data middle platforms for real-time pulling. Policy text data streams refer to a collection of natural language files with structural boundaries, such as draft regulations, regulatory notices, interest rate decision statements, etc., which usually come from government public data channels, industry regulatory systems, or authoritative platforms, and are standardized through text parsing channels and uniformly input. Unstructured data streams refer to text, images, voice, or multimodal combinations that are not in a predefined format. The part used for environmental state modeling is mainly text data, such as news headlines, user comments, online public opinion summaries, etc., which are continuously updated through asynchronous crawling or subscription mechanisms.

[0095] Extracting volatility metrics from real-time time series data streams involves performing short-term and medium-term statistical analysis on the time series to derive statistical features that describe trend stability and anomaly magnitude. This process can be achieved through standard deviation calculations under a sliding window, root mean square fluctuation calculations, or exponentially weighted moving variance methods. Volatility metrics are typically used to measure the sensitivity of state variables to external perturbations. In financial markets, this can manifest as yield fluctuations, or in medical scenarios as instability in vital sign curves. The extracted results serve as quantitative signals for subsequent model state construction.

[0096] Extracting key expected factors from policy text data streams involves using natural language processing models to perform keyword extraction, syntactic structure analysis, and semantic enhancement on the text content, identifying policy content variables closely related to future behavioral directions or state transitions. This process can be accomplished using a bidirectional encoder representation network (such as BERT) or a domain fine-tuning model. By identifying policy themes, impact directions, and intervention scopes within the regulatory corpus, latent factors such as "interest rate adjustment expectations" and "signals of strengthened regulation" are extracted and used as important inputs to drive model state changes. This factor is highly interpretable and serves as a bridge between external prior information and the model's internal representation.

[0097] Determining the sentiment polarity value of text based on unstructured data streams involves performing a sentiment bias judgment on the input text, thereby helping to characterize the group attitudes or expected states of the external environment. This process can be accomplished using a sentiment lexicon or a deep sentiment analysis model (such as an LSTM-CRF structure or a Transformer sentiment discriminator), outputting quantitative values ​​of positive, negative, or neutral tendencies. Sentiment polarity values, as an important supplementary dimension of human behavioral cognition, are particularly valuable in scenarios of high public opinion sensitivity and irrational market fluctuations. Sentiment labels can reflect patients' emotional reactions in healthcare and reflect the behavioral tendencies of market entities in finance.

[0098] The fusion of volatility indicators, key expected factors, and text sentiment polarity values ​​generates a dynamic environment state vector. This involves aligning structured and unstructured features from multiple heterogeneous sources and mapping them into a shared semantic space to form a unified representation vector. This fusion process can employ methods such as early concatenation, weighted aggregation, and attention mechanism integration. Alternatively, the encoder network can be trained to encode the three types of features separately and then jointly map them in the latent space. As the input condition for the subsequent generation of action vectors, the dynamic environment state vector must be timely, consistent, and responsive. Therefore, it must be recalculated before each round of action generation and bound to the current environment observation timestamp to ensure that the input state reflects the latest external environment.

[0099] This embodiment accesses real-time time series data streams, policy text data streams, and unstructured data streams, and extracts volatility indicators, key expectation factors, and text sentiment polarity values ​​for feature fusion. This can significantly improve the dynamic environment state vector's ability to characterize complex external changes, thereby providing the generative model with a more timely and explanatory input expression, and effectively avoiding the decision-making bias and adaptability failure problems of traditional static environment representation in non-stationary task scenarios.

[0100] In one embodiment, the above step S20 includes:

[0101] S201, loading model parameters of the pre-trained generative model;

[0102] S202, inputting the dynamic environment state vector into the feature encoding layer of the pre-trained generative model to generate a feature representation;

[0103] S203, processing the feature representation through the policy network of the pre-trained generative model to generate an action probability distribution;

[0104] S204: Generate an action vector based on the action probability distribution sampling.

[0105] In this embodiment, loading the model parameters of the pre-trained generative model is a preparatory process for generating action vectors. This operation relies on the call of the pre-trained parameter file, which usually includes the network weight matrix, bias vector, regularization hyperparameters, etc. In practice, the model parameters can be stored in a persistent device in binary format and loaded into the computational graph structure through efficient deserialization tools (such as TensorFlow Checkpoint or PyTorch StateDict). The sources of pre-trained models include offline reinforcement training processes based on behavioral simulation data, historical task trajectories or expert labels, and their parameter structure must be compatible with the current runtime architecture. After loading is complete, the model can directly execute the feature processing and action generation process without retraining.

[0106] Inputting the dynamic environment state vector into the feature encoding layer to generate feature representation is the basis for achieving state understanding. The feature encoding layer generally includes components such as linear transformation, activation function, and normalization processing, which can map the original state vector to a high-dimensional latent space to form a stable and distinguishable representation. In application, the encoding structure can take the form of a fully connected neural network, a convolutional layer, a Transformer embedding module, etc., depending on the dimension and semantic distribution of the state vector. The encoding output represents the semantic structure of the current environment, which is used by the policy network to further determine the optimal behavior direction. Feature representation has two key properties: compressibility and discriminability. The former is used to reduce the model inference overhead, and the latter is used to improve the clarity of policy selection.

[0107] Processing feature representations through a pre-trained generative model's policy network to generate action probability distributions is the core process for mapping static inputs into decision-prone distributions. The policy network typically employs a multilayer perceptron or a modular structure with an attention mechanism. Its output is a probability distribution vector defined over the action space, reflecting the adaptability of each possible action in the current state. The policy network can be trained and iterated based on algorithms such as the maximum entropy principle, preferential sampling mechanisms, and Markov policy approximation, ensuring that the output balances generalization and exploration capabilities. The construction of the action probability distribution relies on the gradient adjustment of the model weights, which reflects the accumulated policy experience from previous training tasks.

[0108] Generating action vectors by sampling from the action probability distribution is the key step in converting model outputs into explicit execution instructions. This sampling process can be implemented using a variety of mechanisms, such as greedy sampling, temperature-adjusted softmax sampling, top-k restricted sampling, and Gumbel-Softmax reparameterization. The goal of sampling is to introduce a degree of randomness while ensuring policy stability, thereby enhancing the policy's exploration capabilities and avoiding local optima. The generated action vectors can contain decision information across multiple dimensions, such as asset allocation ratio vectors in financial investment advisory or resource allocation instructions in medical scheduling. The action vectors serve as direct input for downstream task execution and are subsequently used to obtain environmental feedback and update policies.

[0109] This embodiment loads the model parameters of a pre-trained generative model, receives the dynamic environment state vector, and generates action vectors through feature encoding, policy inference, and probability sampling. This enables rapid response to external dynamic environments and high-quality action generation, significantly improving the intelligent system's decision-making flexibility and policy migration capabilities in changing situations. The action vector generation process does not rely on real-time training, which helps reduce the system's computational burden and enhance deployment stability.

[0110] In one embodiment, the above step S30 includes:

[0111] S301, obtaining a domain constraint strategy related to the current task context from the knowledge graph;

[0112] S302, checking whether the motion vector violates the domain constraint strategy;

[0113] S303, when the motion vector violates the domain constraint policy, determining a violation amount of the motion vector relative to the domain constraint policy, and determining an adjustment direction for the motion vector to move toward a compliance boundary;

[0114] S304, generating a correction parameter value based on the violation amount and the adjustment direction;

[0115] S305: Process the motion vector based on the correction parameter value to generate a compliant motion vector.

[0116] In this embodiment, obtaining domain constraint strategies related to the current task context from the knowledge graph is a prerequisite for realizing a constraint-aware correction mechanism. The knowledge graph is used to represent structured semantic information related to the task, including domain rules, behavior boundaries, historical cases and other constraint contents. The current task context generally refers to the operating background described by the dynamic environment state vector and the action vector. The set of constraint strategies most relevant to the current context can be located through embedded matching, semantic association or graph subgraph query. The constraint strategy may include value domain boundaries, state dependency rules, policy legitimacy judgment conditions, etc., and its structure can be expressed in the form of RDF triples, OWL ontologies, property graph structures, etc.

[0117] Verifying action vectors for violations of domain constraint policies is a critical step in ensuring the legitimacy of generated behaviors. This process is typically performed through rule matching, logical verification, or constraint projection. Verification can encompass anomaly detection tasks such as value range violations, illegal combinations, and non-compliant sequences. For example, in the financial sector, this involves detecting whether leverage exceeds regulatory limits, or in the healthcare sector, determining whether resource scheduling violates patient priority rules. This process can be implemented through a rule engine or graph reasoning module, allowing for the automatic identification of potential non-compliant behaviors and feedback to the correction module.

[0118] When an action vector violates a domain constraint policy, it is necessary to further determine the amount of violation relative to the domain constraint policy and determine the direction in which the action vector should be adjusted toward the compliance boundary. The amount of violation indicates the distance or deviation between the current action and the legal boundary, while the direction of adjustment defines the projection direction of the correction vector. The amount of violation can be calculated using metrics such as Euclidean distance, Mahalanobis distance, and KL divergence, while the direction of adjustment can be generated based on the normal of the constrained hyperplane, the gradient of the Lagrange multiplier, or a graph-derived path. This information is directly used to parameterize the correction operation, ensuring stability and directionality.

[0119] Based on the violation amount and adjustment direction, correction parameters are generated to transform the policy-violating action vector into an intermediate representation that meets compliance requirements. These correction parameters can be a set of scalar offsets, projection matrices, or nonlinear transformation function parameters. They can be generated using weighted projections, graph-guided mapping, or constraint-based optimization algorithms based on reinforcement learning. Before executing the correction, the parameter values ​​can be normalized for stability to prevent excessive jumps or information loss during the correction process.

[0120] Correction can be achieved through linear offsets, nonlinear mapping, or generative model-based correction mechanisms. The output of a compliant action vector must maintain semantic proximity to the original policy intent while satisfying all constraints defined in the graph. This output serves as input for the next stage of environmental interaction or reward feedback, directly impacting the effectiveness and legitimacy of subsequent optimization paths.

[0121] This embodiment extracts relevant domain constraint strategies from the knowledge graph and dynamically generates correction parameters for correction based on the deviation relationship between the action vector and the compliance boundary. This can ensure that the strategy output results meet the domain specification requirements while maintaining expressiveness, effectively reduce the execution risks caused by model output violations, and enhance the usability and credibility of the intelligent system under multi-domain compliance requirements.

[0122] In one embodiment, the above step S40 includes:

[0123] S401, collecting environmental feedback data after the execution of the compliant action vector;

[0124] S402, determining a state transition effect value as a performance metric based on the environmental feedback data;

[0125] S403, determining a process fluctuation characteristic value as a risk measurement value based on the environmental feedback data;

[0126] S404, determining a periodic persistence characteristic value as a stability measurement value based on the environmental feedback data;

[0127] S405, combining the performance metric, risk metric, and stability metric to construct a multi-dimensional reward vector;

[0128] S406, configuring a three-tier metric weight distribution scheme;

[0129] S407 , executing the three-layer metric weight distribution scheme to perform weighted fusion on the multi-dimensional reward vector to generate a scalarized reward signal.

[0130] In this embodiment, collecting environmental feedback data after the execution of compliant action vectors is a key input for constructing the foundation of evaluation metrics. Environmental feedback data refers to observable response information generated by changes in the environmental state after an agent performs a specific action, including but not limited to state transition paths, fluctuation records, response delays, abnormal events, and continuous output characteristics. This data is typically obtained through embedded monitoring systems, transaction execution logs, system response modules, or business process callback interfaces, and requires time series, traceability, and contextual integrity.

[0131] The state transition effectiveness value, determined based on environmental feedback data, serves as a performance metric and is the primary indicator for measuring the effectiveness of an agent's behavior. This value reflects whether the current action guides the environment in the desired direction. It is often measured through objective function achievement, state differential value, strategy coverage, or revenue growth rate. In different application scenarios, this can be expressed using parameters such as accuracy, rate of return, and task completion rate. For example, in financial trading scenarios, this can manifest as an increase in return on assets, while in medical scheduling, it can manifest as improved patient turnover efficiency.

[0132] Determining the process fluctuation characteristic value as a risk metric based on environmental feedback data is a key dimension for measuring behavioral stability and uncertainty during execution. This process fluctuation characteristic value can be obtained by measuring the variance, mutation frequency, and policy jitter within the execution path. A higher value indicates a greater degree of uncertainty or potential risk in the system's execution. In multi-period scenarios, metrics such as moving standard deviation, maximum drawdown, and behavioral inconsistency rate can also be used to enhance the sensitivity of risk identification.

[0133] Based on environmental feedback data, we determine the cyclical persistence eigenvalue as a stability metric, used to assess the cross-period persistence of the model's output strategy. This eigenvalue measures whether the current action maintains consistent output and alignment with the target trend across multiple time windows. It is extracted using methods such as sliding window average retention rate, cross-period strategy reconstruction rate, and long-term short-term correlation coefficient. In practical applications, this metric is crucial for reducing model overfitting to short-term events and enhancing long-term return stability.

[0134] Combining performance, risk, and stability metrics to construct a multidimensional reward vector is the structural foundation for achieving long-term and short-term trade-offs in reinforcement learning. This vector is typically a three-dimensional or higher-dimensional array of values, with each dimension corresponding to a policy execution quality evaluation metric. These vectors can be constructed as vectors, nested tensors, or key-value mappings, ensuring that subsequent policy optimization modules have sufficient information to perform multi-objective policy trade-offs.

[0135] A three-tiered weight allocation scheme is configured to control the relative importance of different strategic objectives. This scheme comprises static weight configuration, a dynamic adjustment mechanism, and a closed-loop optimization strategy feedback loop. Static weight configuration allows for preset weight combinations based on different scenarios, while the dynamic adjustment mechanism automatically adjusts weights based on real-time environmental fluctuations. The closed-loop strategy feedback loop guides long-term goals through historical reward feedback.

[0136] A three-layer metric weight distribution scheme performs weighted fusion on multidimensional reward vectors to generate a scalarized reward signal, which serves as the information bridge between the policy execution and model optimization modules. This process can be implemented through weighted summation, normalized projection, nonlinear activation fusion, and attention mechanisms. The scalarized reward signal must possess high sensitivity, high discrimination, and stable trends to guide the effective subsequent parameter adjustment and policy evolution of the pre-trained generative model.

[0137] This embodiment decomposes the environmental feedback data corresponding to the compliant action vector into three evaluation dimensions: performance, risk, and stability, and combines them with a three-layer measurement weight mechanism for fusion conversion. This enables the reward signal to more comprehensively reflect the short-term effects and long-term trends of the agent's behavior, significantly enhancing the sensitivity and trade-off ability to multi-objective constraints during the strategy optimization process, thereby improving the overall decision-making robustness and cross-cycle stability of the model.

[0138] In one embodiment, the above step S50 includes:

[0139] S501, constructing an interactive trajectory dataset containing compliant action vector sequences and environmental state changes;

[0140] S502, analyzing the strategy optimization direction characteristics in the reward signal;

[0141] S503, configuring adaptive optimization parameters based on environmental conditions;

[0142] S504, determining a model parameter update gradient based on the strategy optimization direction feature, the interaction trajectory dataset, and the adaptive optimization parameters;

[0143] S505: Adjust the model weights of the pre-trained generative model based on the model parameter update gradient.

[0144] In this example, after executing multiple compliant actions, a series of time series data corresponding to changes in the environmental state is generated, forming a set of behavioral trajectories that can reflect the impact of action selection on the environment. To construct a complete interaction trajectory dataset, the corresponding environmental response state, historical reward value, behavioral selection distribution, and time label after each action execution need to be jointly encapsulated into a data structure while maintaining sequential consistency and traceability. This dataset can be organized using nested structures or sliding window tensors to meet the needs of batch processing and policy tracking during training.

[0145] The strategy optimization directional features contained in the reward signal are an important guide for model parameter updates. These directional features are typically extracted through methods such as gradient analysis, advantage function estimation, and strategy differential attribution. Their values ​​not only indicate the direction of improvement for the current strategy but also reflect the functional relationship between reward trends and behavioral deviations. The process of analyzing these features requires back-calculation based on historical strategy execution records, and can be combined with mechanisms such as moving averages and strategy volatility control to improve stability.

[0146] The dynamic nature of environmental states directly impacts the timeliness and generalization of optimization strategies. To accommodate the varying sensitivity of different state regions to strategies, adaptive optimization parameters should be set based on environmental state fluctuations, state categories, the frequency of abnormal events, or periodic trend changes. These parameters can include learning rates, regularization factors, gradient clipping thresholds, and more. Meta-learning mechanisms can also be introduced to dynamically fine-tune the strategy's response structure.

[0147] After direction extraction and parameter setting, the updated gradients of the model parameters can be derived based on this multi-dimensional information. These gradients can be calculated using techniques such as policy gradient methods, distributed value function backpropagation, and batch attribution analysis. Weighted accumulation, normalization, or self-attention mechanisms can be used to enhance the directional consistency and amplitude control of the gradients. To improve sample efficiency and training stability, multi-path trajectory fusion or confidence weighting can be used to suppress the influence of high-variance paths.

[0148] Once the updated gradient calculation is complete, it is ultimately used to adjust the model weights of the pre-trained generative model. This weight adjustment process involves not only numerical corrections to some parameters in the forward inference process, but also possible fine-tuning of the policy network structure or reconstruction of hidden layer connection weights. In actual deployments, mechanisms such as momentum optimization, adaptive learning rates, or low-rank updates can be used to enhance training convergence and long-term performance transfer. The entire optimization process must be tightly coupled with data sampling and trajectory rolling window mechanisms to ensure synchronized iteration of the training trajectory and policy weights to avoid policy drift or gradient failure.

[0149] Example: In the fintech business field, an asset management platform hopes to build an intelligent investment decision-making system with dynamic adaptability to automatically identify market changes, dynamically generate operational recommendations, and continuously optimize strategy generation models to improve robustness and compliance in highly volatile market conditions.

[0150] The system first accesses multi-source heterogeneous data streams, including real-time market data (such as stock prices, bond interest rates, and exchange rate fluctuations), regulatory policy release data (such as central bank announcements and China Securities Regulatory Commission notices), and unstructured information (such as financial news, market commentary, and social media sentiment text). By extracting price volatility indicators from real-time time-series data streams and key expectation factors (such as monetary policy direction and industry regulatory trends) from policy text data, and analyzing unstructured text using natural language processing algorithms to generate sentiment polarity values, the platform integrates these sources to generate a dynamic environmental state vector that reflects the current state of the financial market.

[0151] The system then feeds this dynamic environment state vector into a generative model pre-trained with historical market data. After loading its historically trained parameters, the model first performs feature encoding on the input state vector, converting it into an implicit feature representation. The strategy network then calculates the probability distribution of multiple possible investment actions (such as buy, hold, and sell). Based on these probabilistic sampling, the model generates the initial action vector for the current cycle, suggesting, for example, adjusting some assets or reducing exposure to certain securities.

[0152] Generated action vectors are not immediately executed but are subject to compliance review. The system integrates a set of domain constraint strategies built on the financial knowledge graph, including rules for matching various products and investors, regulatory compliance restrictions (such as leverage limits and concentration management), and institutional risk appetite templates. The model matches the currently generated action vector with contextual constraints retrieved from the knowledge graph. If a violation occurs, correction parameters are generated based on the severity and direction of the violation, adjusting the original action vector to within the compliance boundary, resulting in a final, executable, compliant action vector.

[0153] After executing a compliant action in the actual market, the system collects feedback information corresponding to that action, including data on changes in account net value after strategy adjustments, variations in risk indicators, return fluctuation ranges, and changes in fund liquidity. The system extracts state transition effects as a performance metric, period fluctuation amplitude as a risk metric, and the consistency of strategy returns across multiple time windows as a stability metric. These three metrics are combined to form a multidimensional reward vector, which is then weighted and fused according to a predetermined weighting strategy to output a scalar reward signal as an overall evaluation of the performance of the current round of operations.

[0154] Based on this reward signal, along with historical sequences of compliant action vectors and corresponding environmental state changes, the system constructs a dataset of interaction trajectories. To improve subsequent decision-making, the platform employs an adaptive optimization strategy module to analyze the directional information in the reward signal and automatically adjust hyperparameters such as the learning rate and regularization coefficient based on the current environmental state. The platform then derives the updated gradients of the model parameters using both trajectory data and reward direction. Finally, it inversely adjusts the weight configuration of the pre-trained generative model to continuously optimize its behavior generation capabilities.

[0155] To further enhance the system's responsiveness to sudden policy events and its ability to stably perceive evolving market trends, the platform incorporates two update mechanisms. Upon detecting a major policy change (such as an adjustment to the securities transaction tax rate or a shift in macroeconomic regulation targets), the system triggers an immediate environmental reconstruction mechanism, immediately refreshing factors such as volatility indicators, expectation factors, and sentiment polarity, and reconstructing the environmental state vector. A timed refresh strategy is also in place, periodically calling upon the latest data to complete environmental updates when no events occur. The system integrates these two types of update results to generate a more comprehensive and robust environmental perception vector, providing a credible foundation for the next round of action generation.

[0156] In the healthcare business, a smart hospital platform has deployed an intelligent strategy generation system for clinical intervention assistance. It is used to analyze patient status evolution in real time, predict potential risks, and provide personalized treatment suggestions, while continuously optimizing its strategy output capabilities within the constraints of medical processes.

[0157] The system first accesses multi-source heterogeneous data streams, including real-time vital signs monitoring data of patients (such as heart rate, blood pressure, and blood oxygen concentration), policy data (such as newly released diagnosis and treatment guidelines, drug use specifications, and public health emergency notifications), and unstructured text data (such as doctor's ward rounds, nursing records, and medical public opinion). The system extracts short-term volatility indicators from real-time vital signs data, such as the coefficient of variation of heart rate and the short-term change rate of blood pressure; extracts key expected factors from medical policy texts, such as the upper limit of the frequency of antibiotic use and bed classification scheduling requirements; combines sentiment recognition algorithms to analyze unstructured data such as doctor's description texts and patient complaints, and extracts text sentiment polarity values ​​in the current diagnosis and treatment process to characterize the patient's subjective discomfort level or the semantic tendency of medical intervention. After the three types of data are fused, a dynamic environmental state vector is generated to comprehensively depict the health situation in the current medical context.

[0158] The generated dynamic environment state vector is fed into a generative model pre-trained using historical hospitalization records. After loading the historical model parameters, the model is converted into a high-dimensional health state representation through a feature encoding layer. The policy network then outputs a probability distribution of multiple intervention recommendations, such as adjusting oxygen flow, modifying medication dosage, or delaying a specific examination. The system generates a recommended action vector based on sampling from the probability distribution.

[0159] To ensure that the implementation of recommendations does not violate medical standards, the platform introduces knowledge graph constraints after the action vector is generated. This knowledge graph embeds multiple domain constraint strategies, including the compatibility relationship between indications, drugs, and dosages, risk assessment rules between age and treatment methods, and the scope of use limited by the medical insurance catalog. The system invokes relevant constraint strategies based on the current patient's diagnosis and treatment context to verify whether the action vector violates the regulations. For example, when recommending the use of a certain type of drug beyond the dosage or performing a high-risk procedure on patients in a specific age group, the system assesses the degree of violation and the required adjustments, and generates correction parameters based on these, ultimately adjusting the initial recommendation to a compliant action vector that meets medical standards.

[0160] When a compliant action vector is executed in the medical process, the system collects feedback data in real time, including changes in the patient's vital signs, disease progression indicators, and post-treatment complications. Based on this feedback, the system calculates the state transition effect to measure the efficacy of the medical intervention. It uses the volatility of key indicators in the process as a risk metric and extracts stable measurement values ​​based on the stable performance of treatment continuity over a certain period of time. These values ​​are combined to form a multidimensional reward vector. The system configures a three-layer measurement weight distribution structure based on the hospital management strategy, fuses the three-dimensional reward vectors, and forms a final scalar reward signal to measure the overall performance of the current intervention strategy.

[0161] The platform constructs an interactive trajectory dataset centered around sequences of compliant action vectors and corresponding patient state change sequences, and analyzes and optimizes reward signals. For example, if the current model favors conservative intervention but the patient responds well, the strategy should be adjusted toward proactive intervention. The system adaptively sets hyperparameters such as the learning rate and regularization term based on environmental conditions. Based on trajectory data and reward feedback, it generates parameter update gradients and inversely updates model weights, continuously improving its generative capabilities.

[0162] The system also introduces a dual-update mechanism to enhance responsiveness and stability in medical scenarios. Upon detecting the release of a major medical policy or new guidelines, the system instantly reconstructs the environment based on the latest multi-source heterogeneous data, updating the patient's perceived status. In the absence of event triggers, a periodic refresh mechanism synchronizes monitoring data and policy information to generate a second update of the environment data. The data from these two sources is then integrated to produce a more robust environment update, which is then iteratively generated to generate the latest dynamic environment state vector.

[0163] Through integrated modeling of multi-source medical data, active correction of action compliance, multi-dimensional evaluation of intervention effects, and adaptive optimization of models, the system can assist doctors in achieving intelligent responses and strategy generation for complex conditions while ensuring medical safety and compliance, thereby improving the efficiency and accuracy of medical intervention and enhancing the hospital's overall risk control capabilities and service levels.

[0164] This embodiment uses a reward-signal-guided policy direction analysis and a dynamic environment response-driven adaptive optimization mechanism to achieve fine-tuning and stable convergence of policies in time-varying environments, thereby improving the behavioral generalization ability and policy adjustment sensitivity of the pre-trained generative model for complex scenarios, thereby enhancing its continuous optimization capability and long-term profit performance in actual tasks.

[0165] In one embodiment, after the above step S10, the method further includes:

[0166] S106, monitoring policy release events in multi-source heterogeneous data streams and generating event trigger signals;

[0167] S107, in response to the event trigger signal, performing real-time environment reconstruction based on the multi-source heterogeneous data stream to generate first updated environment data;

[0168] S108, starting a timed update mechanism, refreshing the environment based on the multi-source heterogeneous data stream at a preset time interval, and generating second updated environment data;

[0169] S109, fusing the first update environment data and the second update environment data to generate fused update data;

[0170] S110: Update the dynamic environment state vector based on the fused update data.

[0171] In this embodiment, when the data source of the environmental state is a multi-source heterogeneous data stream, it inherently possesses strong timeliness, event-driven, and cross-structural change characteristics. In order to maintain the dynamic accuracy of the environmental state vector, it is necessary to set up a trigger mechanism for data mutation events. This mechanism performs keyword extraction, entity recognition, and text summary comparison on policy-related content in the multi-source heterogeneous data stream to determine whether a policy release event that requires adjustment has occurred. If the event is determined to be established, an event trigger signal is generated, which is used as a control condition to drive immediate environmental updates.

[0172] Event triggers guide the system to perform immediate environmental reconstruction based on the multi-source, heterogeneous data captured at the moment. This reconstruction process, executed with high priority, includes retrieving policy text, refreshing the semantic features of unstructured data, and updating the response status of volatility indicators in time series. The resulting first-level updated environmental data is used to address the impact of temporary, sudden, or periodic policy signals on the system's decision-making logic.

[0173] In contrast to the event-triggered update mechanism, the scheduled update mechanism is used to periodically adjust environmental perception results. This mechanism schedules tasks based on preset time intervals. When no unexpected events occur, the system extracts the latest multi-source heterogeneous data at a set frequency, ensuring stable and continuous updates of the environmental state vector and preventing model cognitive staleness caused by long periods of inactivity. Secondary environmental data updates can include fluctuation trends within a time window, shifts in semantic sentiment averages, or frequency updates of external expectation factors.

[0174] To achieve a seamless transition between decision continuity and updated data, the updated data from these two sources must be fused. This fusion process must maintain both the emergency responsiveness of the first updated data and the trend evolution value of the second updated data. Dynamic fusion can be implemented using a time-priority strategy, a data credibility-weighted strategy, or a recent-moment-dominant strategy. The resulting fused updated data is both event-sensitive and trend-stable.

[0175] Finally, the environment state vector is reconstructed based on the fused updated data. This vector remains structurally consistent with the original state vector, but is numerically and semantically updated based on the fused data. The updated environment state vector can serve as direct input for subsequent model generation or behavior evaluation, achieving a state representation that balances real-time performance with stability.

[0176] This embodiment uses event-triggered updates and timed-period update mechanisms, and builds fused update data to drive environmental state vector reconstruction. This can significantly enhance the model's response speed to sudden external disturbances and its strategy adaptability, while ensuring the continuity and trend consistency of environmental state perception on a time scale, effectively supporting behavior generation and optimization under dynamic constraints and real-time feedback scenarios.

[0177] In one embodiment, a generation strategy optimization device based on a dynamic environment is provided, and the generation strategy optimization device based on a dynamic environment corresponds one-to-one to the generation strategy optimization method based on a dynamic environment in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of a dynamic environment-based strategy optimization device of the present invention. It includes an environment perception module 10, an action generation module 20, a strategy constraint module 30, a reward evaluation module 40, and a model optimization module 50. Each functional module is described in detail below:

[0178] An environment perception module 10 is used to construct a dynamic environment state vector based on multi-source heterogeneous data streams;

[0179] An action generation module 20 is configured to generate an action vector based on the dynamic environment state vector using a pre-trained generative model;

[0180] A policy constraint module 30 is configured to modify the action vector according to a knowledge graph containing domain constraint policies to generate a compliant action vector;

[0181] a reward evaluation module 40 for generating a multidimensional reward vector including a performance metric, a risk metric, and a stability metric based on feedback after the execution of the compliant action vector, and scalarizing the multidimensional reward vector into a reward signal;

[0182] The model optimization module 50 is configured to update the pre-trained generative model using an adaptive strategy optimization module based on the reward signal and the interaction trajectory including the compliant action vector.

[0183] In one embodiment, the environment perception module 10 is specifically configured to:

[0184] Access real-time time series data streams, policy text data streams, and unstructured data streams;

[0185] extracting a volatility indicator from the real-time time series data stream;

[0186] extracting key expected factors from the policy text data stream;

[0187] Determining a text sentiment polarity value based on the unstructured data stream;

[0188] The volatility index, key expectation factors and text sentiment polarity values ​​are integrated to generate a dynamic environment state vector.

[0189] In one embodiment, the action generation module 20 is specifically configured to:

[0190] Load the model parameters of the pre-trained generative model;

[0191] Inputting the dynamic environment state vector into the feature encoding layer of the pre-trained generative model to generate a feature representation;

[0192] Processing the feature representation through the policy network of the pre-trained generative model to generate an action probability distribution;

[0193] An action vector is generated based on the action probability distribution sampling.

[0194] In one embodiment, the policy constraint module 30 is specifically configured to:

[0195] Obtaining domain constraint strategies related to the current task context from the knowledge graph;

[0196] Checking whether the motion vector violates the domain constraint strategy;

[0197] When the motion vector violates the domain constraint policy, determining a violation amount of the motion vector relative to the domain constraint policy, and determining an adjustment direction for the motion vector to move toward a compliance boundary;

[0198] generating a correction parameter value based on the violation amount and the adjustment direction;

[0199] The motion vector is processed based on the correction parameter value to generate a compliant motion vector.

[0200] In one embodiment, the reward evaluation module 40 is specifically configured to:

[0201] Collect environmental feedback data after the execution of compliant action vectors;

[0202] Determining a state transition effect value as a performance metric based on the environmental feedback data;

[0203] determining a process fluctuation characteristic value as a risk metric based on the environmental feedback data;

[0204] Determining a periodic persistence characteristic value as a stability metric value based on the environmental feedback data;

[0205] Combining the performance metric, the risk metric, and the stability metric to construct a multidimensional reward vector;

[0206] Configure a three-tier metric weight distribution scheme;

[0207] The three-layer metric weight distribution scheme is executed to perform weighted fusion on the multi-dimensional reward vector to generate a scalarized reward signal.

[0208] In one embodiment, the model optimization module 50 is specifically configured to:

[0209] Construct an interactive trajectory dataset containing compliant action vector sequences and environmental state changes;

[0210] Analyzing the strategy optimization direction characteristics in the reward signal;

[0211] Configure adaptive optimization parameters based on environmental status;

[0212] Determining the model parameter update gradient based on the strategy optimization direction features, interaction trajectory dataset and adaptive optimization parameters;

[0213] The model weights of the pre-trained generative model are adjusted based on the model parameter update gradients.

[0214] In one embodiment, the environment perception module 10 is specifically configured to:

[0215] Monitor policy release events in multi-source heterogeneous data streams and generate event trigger signals;

[0216] In response to the event trigger signal, performing real-time environment reconstruction based on multi-source heterogeneous data streams to generate first updated environment data;

[0217] Start the timed update mechanism to refresh the environment based on the multi-source heterogeneous data stream at a preset time interval to generate second updated environment data;

[0218] fusing the first update environment data and the second update environment data to generate fused update data;

[0219] The dynamic environment state vector is updated based on the fused update data.

[0220] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a generation strategy optimization method based on a dynamic environment.

[0221] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a generation strategy optimization method based on a dynamic environment.

[0222] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0223] Construct dynamic environment state vector based on multi-source heterogeneous data streams;

[0224] Based on the dynamic environment state vector, generating an action vector through a pre-trained generative model;

[0225] According to the knowledge graph containing the domain constraint strategy, the action vector is modified to generate a compliant action vector;

[0226] generating a multidimensional reward vector including a performance metric, a risk metric, and a stability metric based on feedback after the execution of the compliant action vector, and scalarizing the multidimensional reward vector into a reward signal;

[0227] Based on the reward signal and the interaction trajectory including the compliant action vector, an adaptive policy optimization module is used to update the pre-trained generative model.

[0228] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0229] Construct dynamic environment state vector based on multi-source heterogeneous data streams;

[0230] Based on the dynamic environment state vector, generating an action vector through a pre-trained generative model;

[0231] According to the knowledge graph containing the domain constraint strategy, the action vector is modified to generate a compliant action vector;

[0232] generating a multidimensional reward vector including a performance metric, a risk metric, and a stability metric based on feedback after the execution of the compliant action vector, and scalarizing the multidimensional reward vector into a reward signal;

[0233] Based on the reward signal and the interaction trajectory including the compliant action vector, an adaptive policy optimization module is used to update the pre-trained generative model.

[0234] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0235] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0236] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0237] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A generation strategy optimization method based on a dynamic environment, characterized in that: The following steps are involved: Construct dynamic environment state vector based on multi-source heterogeneous data streams; Based on the dynamic environment state vector, generating an action vector through a pre-trained generative model; According to the knowledge graph containing the domain constraint strategy, the action vector is modified to generate a compliant action vector; generating a multidimensional reward vector including a performance metric, a risk metric, and a stability metric based on feedback after the execution of the compliant action vector, and scalarizing the multidimensional reward vector into a reward signal; Based on the reward signal and the interaction trajectory including the compliant action vector, an adaptive policy optimization module is used to update the pre-trained generative model.

2. The method for optimizing a generation strategy based on a dynamic environment according to claim 1, wherein: Construct a dynamic environment state vector based on multi-source heterogeneous data streams, including: Access real-time time series data streams, policy text data streams, and unstructured data streams; extracting a volatility indicator from the real-time time series data stream; extracting key expected factors from the policy text data stream; Determining a text sentiment polarity value based on the unstructured data stream; The volatility index, key expectation factors and text sentiment polarity values ​​are integrated to generate a dynamic environment state vector.

3. The method for optimizing a generation strategy based on a dynamic environment according to claim 1, wherein: Based on the dynamic environment state vector, an action vector is generated by a pre-trained generative model, including: Load the model parameters of the pre-trained generative model; Inputting the dynamic environment state vector into the feature encoding layer of the pre-trained generative model to generate a feature representation; Processing the feature representation through the policy network of the pre-trained generative model to generate an action probability distribution; An action vector is generated based on the action probability distribution sampling.

4. The method for optimizing a generation strategy based on a dynamic environment according to claim 1, wherein: According to the knowledge graph containing the domain constraint strategy, the action vector is modified to generate a compliant action vector, including: Obtaining domain constraint strategies related to the current task context from the knowledge graph; Checking whether the motion vector violates the domain constraint strategy; When the motion vector violates the domain constraint policy, determining a violation amount of the motion vector relative to the domain constraint policy, and determining an adjustment direction for the motion vector to move toward a compliance boundary; generating a correction parameter value based on the violation amount and the adjustment direction; The motion vector is processed based on the correction parameter value to generate a compliant motion vector.

5. The method for optimizing generation strategy based on dynamic environment according to claim 1, wherein: Generating a multidimensional reward vector including a performance metric, a risk metric, and a stability metric based on feedback after the execution of the compliant action vector, and scalarizing the multidimensional reward vector into a reward signal, including: Collect environmental feedback data after the execution of compliant action vectors; Determining a state transition effect value as a performance metric based on the environmental feedback data; determining a process fluctuation characteristic value as a risk metric based on the environmental feedback data; Determining a periodic persistence characteristic value as a stability metric value based on the environmental feedback data; Combining the performance metric, the risk metric, and the stability metric to construct a multidimensional reward vector; Configure a three-tier metric weight distribution scheme; The three-layer metric weight distribution scheme is executed to perform weighted fusion on the multi-dimensional reward vector to generate a scalarized reward signal.

6. The method for optimizing generation strategy based on dynamic environment according to claim 1, wherein: Based on the reward signal and the interaction trajectory including the compliant action vector, an adaptive policy optimization module is used to update the pre-trained generative model, including: Construct an interactive trajectory dataset containing compliant action vector sequences and environmental state changes; Analyzing the strategy optimization direction characteristics in the reward signal; Configure adaptive optimization parameters based on environmental status; Determining the model parameter update gradient based on the strategy optimization direction features, interaction trajectory dataset and adaptive optimization parameters; The model weights of the pre-trained generative model are adjusted based on the model parameter update gradients.

7. The method for optimizing generation strategy based on dynamic environment according to claim 1, wherein: After constructing the dynamic environment state vector based on multi-source heterogeneous data streams, it also includes: Monitor policy release events in multi-source heterogeneous data streams and generate event trigger signals; In response to the event trigger signal, performing real-time environment reconstruction based on multi-source heterogeneous data streams to generate first updated environment data; Start the timed update mechanism to refresh the environment based on the multi-source heterogeneous data stream at a preset time interval to generate second updated environment data; fusing the first update environment data and the second update environment data to generate fused update data; The dynamic environment state vector is updated based on the fused update data.

8. A generation strategy optimization device based on a dynamic environment, characterized in that: The generation strategy optimization device based on the dynamic environment includes: The environment perception module is used to construct a dynamic environment state vector based on multi-source heterogeneous data streams; An action generation module, configured to generate an action vector based on the dynamic environment state vector using a pre-trained generative model; A policy constraint module, configured to modify the action vector according to a knowledge graph containing domain constraint policies to generate a compliant action vector; a reward evaluation module, configured to generate a multidimensional reward vector including a performance metric, a risk metric, and a stability metric based on feedback after the execution of the compliant action vector, and scalarize the multidimensional reward vector into a reward signal; A model optimization module is configured to update the pre-trained generative model using an adaptive strategy optimization module based on the reward signal and the interaction trajectory including the compliant action vector.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a dynamic environment-based generation strategy optimization program stored in the memory and capable of running on the processor. When the dynamic environment-based generation strategy optimization program is executed by the processor, the steps of the dynamic environment-based generation strategy optimization method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a dynamic environment-based generation strategy optimization program, which, when executed by a processor, implements the steps of the dynamic environment-based generation strategy optimization method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent control method and system for marine ultraviolet disinfection device

    CN120939263A

  • Consistency enhancement method and device based on multi-source reward fusion, equipment and medium

    CN121052330A

  • A consistency reinforcement method, device and equipment based on multi-source reward fusion and a medium

    CN121052330B

  • SPARK cluster-oriented adaptive frequency adjustment method and system, terminal and storage medium

    CN121092334A

  • Qi liquid formula optimization method and system based on big data analysis

    CN121354733A