Method and device for generating mission decision suitable for atmospheric environment, medium and terminal

By using a decision model pre-trained on a corpus in the atmospheric environment domain and combining the total loss function of policy gradient loss and KL divergence constraint loss to fine-tune the model parameters, the subjectivity and high cost caused by manual annotation in existing technologies are solved, and stable and accurate task decision generation is achieved.

CN122453192APending Publication Date: 2026-07-24BEIJING INSIGHTS VALUE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INSIGHTS VALUE TECHNOLOGY CO LTD
Filing Date
2026-03-13
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing technologies, task decision generation methods in the atmospheric environment field rely heavily on human experience, resulting in highly subjective annotation results, inconsistent expert standards, failure to meet domain requirements, and high costs associated with manual annotation.

Method used

A decision model pre-trained on an atmospheric environment corpus is used, and the model parameters are fine-tuned using the total loss function of policy gradient loss and KL divergence constraint loss. This generates a direction adjustment with a higher probability of the optimal policy, and the model parameters are updated through multiple rounds of iteration to avoid manual annotation and ensure the stability and robustness of the model parameters.

Benefits of technology

The generated task decisions can meet the needs of the atmospheric environment field, reduce the cost of manual labeling, improve the robustness and stability of the model, and ensure the objectivity and accuracy of the decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122453192A_ABST
    Figure CN122453192A_ABST
Patent Text Reader

Abstract

The application discloses a task decision generation method and device suitable for an atmospheric environment, a medium and a terminal, relates to the technical field of artificial intelligence, and mainly aims to improve the problems that the existing method relies on artificial experience as a reward signal, the subjectivity of the labeling result is strong, different expert standards are not the same, the task decision for the atmospheric environment field cannot truly meet the field requirements, and the labeling cost is high due to a large amount of manual labeling work. The method comprises the following steps: receiving target task input by a target user about the atmospheric environment field; performing task decision analysis operation according to the target task based on an atmospheric environment field decision model that has completed model parameter fine tuning, to generate a task decision corresponding to the target task; and identifying a pollution responsible party, outputting a pollutant over-standard alarm and screening an optimal control measure based on the task decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, medium, and terminal for generating task decisions in atmospheric environments. Background Technology

[0002] As atmospheric environmental governance undergoes a profound transformation towards precision and scientific rigor, higher demands are being placed on intelligent decision-making for complex tasks such as pollution source tracing, accurate prediction of pollutant concentrations, and simulation and evaluation of the effectiveness of control measures. Large language models, with their powerful natural language understanding and generation capabilities, offer a new technological approach for processing massive amounts of environmental data and generating analytical reports. However, when general-purpose large models are directly applied to highly specialized fields like atmospheric environment, a significant gap exists between their general logical framework and the specific needs of the domain. Therefore, a fine-tuning method that can deeply integrate domain knowledge and ensure objective and reliable output is urgently needed to ensure that task decisions in the atmospheric environment domain meet the domain's requirements.

[0003] Currently, existing fine-tuning methods mainly involve first collecting various outputs of the model in the atmospheric environment field, then having domain experts score the model based on their personal experience and subjective judgment, then using this human feedback data to train a reward model, and finally using the reward model to guide the large model to generate outputs that are more in line with human preferences.

[0004] However, existing fine-tuning methods rely heavily on human experience as a reward signal, which leads to problems such as strong subjectivity in the annotation results and inconsistent standards among different experts. Consequently, task decisions in the atmospheric environment field cannot truly meet the needs of the field. Furthermore, the large amount of manual labeling work also results in high annotation costs. Summary of the Invention

[0005] In view of this, this application provides a task decision generation method, device, medium, and terminal applicable to the atmospheric environment. The main purpose is to improve the existing methods, which rely heavily on human experience as a reward signal, resulting in highly subjective annotation results and inconsistent standards among different experts. Consequently, task decisions in the atmospheric environment field cannot truly meet the needs of the field, and the annotation cost is high due to a large amount of manual annotation work.

[0006] According to one aspect of this application, a task decision generation method suitable for atmospheric environments is provided, comprising: Receive target tasks related to the atmospheric environment from the target user; Based on the atmospheric environment domain decision model with completed model parameter fine-tuning, task decision analysis is performed according to the target task to generate the task decision corresponding to the target task. The atmospheric environment domain decision model is pre-trained using atmospheric environment domain corpus and obtained by fine-tuning the model parameters using a total loss function that includes policy gradient loss and KL divergence constraint loss. The policy gradient loss is used to guide the model parameters of the atmospheric environment domain decision model to adjust in a direction with a higher probability of generating the optimal policy. The KL divergence constraint loss is used to limit the adjustment range of the model parameters. During the model parameter fine-tuning process, the optimal policy used in each iteration is generated using the atmospheric environment domain decision model obtained after the previous iteration. Based on the aforementioned task, the responsible party for pollution is identified, warnings of pollutant exceedances are issued, and the optimal control measures are selected.

[0007] Preferably, before generating the task decision corresponding to the target task by performing task decision analysis based on the fine-tuned atmospheric environment domain decision model, the method further includes: The pre-trained atmospheric environment domain decision model is used as the initial atmospheric environment domain decision model. Based on the initial atmospheric environment domain decision model, multiple first-round iteration task decision analysis operations are performed according to the preset prompt text to generate multiple first-round iteration candidate strategies. Based on the preset scoring rules, the score of each candidate strategy in the first round of iteration is calculated, and the optimal strategy in the first round of iteration is selected based on the score of each candidate strategy in the first round of iteration. The actual action probability distribution corresponding to the optimal strategy in the first round of iteration is used as the original action probability distribution in the first round of iteration; Using the initial atmospheric environment domain decision model, based on the preset prompt text, and based on the first round of iteration optimal strategy, the first round of iteration simulation task decision analysis operation is performed to obtain the simulation action probability distribution corresponding to the first round of iteration optimal strategy, and the simulation action probability distribution is used as the first round of iteration new action probability distribution; The gradient term loss of the strategy in the first iteration is calculated based on the probability distribution of the new action in the first iteration, and the KL divergence constraint term loss of the first iteration is calculated based on the probability distribution of the new action in the first iteration and the probability distribution of the original action in the first iteration. The total loss value of the first iteration is obtained by combining the gradient term loss of the strategy in the first iteration and the KL divergence constraint term loss of the first iteration. The gradient of the total loss value in the first iteration with respect to the initial model parameters of the initial atmospheric environment domain decision model is calculated, and backpropagation is performed using the gradient descent method to update the initial model parameters in the first iteration, thereby obtaining the first iteration atmospheric environment domain decision model. Based on the first-round iterative atmospheric environment domain decision model, multiple second-round iterative task decision analysis operations are performed according to the preset prompt text, and the optimal strategy for the second iteration round is selected. Using the first-round iterative atmospheric environment domain decision model, the second-round simulated task decision analysis operation is performed based on the optimal strategy for the second iteration round. The total loss value for the second iteration round is calculated based on the probability distribution of new actions and the probability distribution of original actions corresponding to the optimal strategy for the second iteration round. The parameters of the first-round iterative model of the first-round atmospheric environment domain decision model are updated in the second iteration round based on the total loss value of the second iteration round, resulting in the second-round atmospheric environment domain decision model. The model parameters are iteratively updated multiple times until a preset iteration stop condition is reached, resulting in an atmospheric environment domain decision model with fine-tuned parameters. Task decision analysis operations are then performed based on this fine-tuned atmospheric environment domain decision model.

[0008] Preferably, the step of calculating the score of each of the first-round iteration candidate strategies based on a preset scoring rule, and selecting the optimal strategy for the first-round iteration based on the score of each of the first-round iteration candidate strategies, includes: For each candidate strategy in the first round of iterations, scores are given for factual consistency, format standardization, strategy consensus, and text. Based on preset score weights that match the preset prompt text, a weighted sum of the scores for each dimension is calculated, and this weighted sum is determined as the score of the candidate strategy in the first round of iterations. Specifically, the factual consistency score assesses the degree of conformity between the candidate strategy and knowledge in the atmospheric environment domain; the format standardization score assesses the degree of conformity between the candidate strategy and the preset professional report structure; the strategy consensus score assesses the similarity between the candidate strategy and other candidate strategies in the first round of iterations; and the text score assesses the semantic fluency of the candidate strategy. Based on the scores of each candidate strategy in the first round of iteration, the optimal strategy in the first round of iteration is selected.

[0009] Preferably, the step of selecting the optimal strategy for the first round of iterations based on the score of each candidate strategy in the first round of iterations includes: All candidate strategies for the first round of iterations are sorted in descending order of their scores to generate a sequence of candidate strategies for the first round of iterations. If the score of the first-order candidate strategy in the first round of iterations is different from the score of the second-order candidate strategy in the first round of iterations, then the first-order candidate strategy in the first round of iterations shall be taken as the optimal strategy in the first round of iterations. If the score of the first-order first-round iteration candidate strategy is the same as the score of at least one subsequent-order first-round iteration candidate strategy, then the first-round iteration candidate strategy with the highest score in the factual consistency dimension is selected from multiple first-round iteration candidate strategies with the same score and is taken as the optimal strategy for the first-round iteration.

[0010] Preferably, before using the pre-trained atmospheric environment domain decision model as the initial atmospheric environment domain decision model, the method further includes: Acquire historical documents in the field of atmospheric environment within a preset historical time period, and use the historical documents as corpus in the field of atmospheric environment. Based on the atmospheric environment domain corpus, a general large model is pre-trained to obtain a pre-trained atmospheric environment domain decision model, and then the model parameters are fine-tuned based on the pre-trained atmospheric environment domain decision model.

[0011] Preferably, the preset prompt text includes atmospheric environmental data, the source of the atmospheric environmental data, the spatiotemporal coordinates of the atmospheric environmental data, and the expected analysis target.

[0012] Preferably, the preset iteration stopping condition is any one of the following: reaching a preset iteration round threshold, satisfying a preset task decision scoring threshold, or the loss of the KL divergence constraint term being continuously lower than a preset KL divergence threshold.

[0013] According to another aspect of this application, a task decision generation device suitable for atmospheric environments is provided, comprising: The task receiving module is used to receive target tasks related to the atmospheric environment from the target user. The task decision generation module is used to perform task decision analysis based on the target task according to the atmospheric environment domain decision model that has been fine-tuned, and generate the task decision corresponding to the target task. The atmospheric environment domain decision model is pre-trained using atmospheric environment domain corpus and fine-tuned using a total loss function that includes policy gradient loss and KL divergence constraint loss. The policy gradient loss is used to guide the model parameters of the atmospheric environment domain decision model to adjust in a direction with a higher probability of generating the optimal policy. The KL divergence constraint loss is used to limit the adjustment range of the model parameters. During the model parameter fine-tuning process, the optimal policy used in each iteration is generated using the atmospheric environment domain decision model obtained after the previous iteration. The task decision application module is used to identify the responsible party for pollution based on the task decision, output alarms for pollutant exceedances, and screen the optimal control measures.

[0014] Preferably, before the task decision generation module, the device further includes an atmospheric environment domain decision model training module, comprising: The first-round iteration candidate strategy generation unit is used to take the pre-trained atmospheric environment domain decision model as the initial atmospheric environment domain decision model, and based on the initial atmospheric environment domain decision model, perform multiple first-round iteration task decision analysis operations according to the preset prompt text to generate multiple first-round iteration candidate strategies. The first-round iteration optimal strategy selection unit is used to calculate the score of each of the first-round iteration candidate strategies based on the preset scoring rules, and select the first-round iteration optimal strategy based on the score of each of the first-round iteration candidate strategies. The unit for obtaining the original action probability distribution of the first iteration is used to take the actual action probability distribution corresponding to the generated optimal strategy of the first iteration as the original action probability distribution of the first iteration. The first iteration new action probability distribution acquisition unit is used to utilize the initial atmospheric environment domain decision model, according to the preset prompt text, and based on the first iteration optimal strategy to perform first iteration simulation task decision analysis operation, to obtain the simulation action probability distribution corresponding to the first iteration optimal strategy, and to use the simulation action probability distribution as the first iteration new action probability distribution; The first iteration total loss calculation unit is used to calculate the first iteration strategy gradient term loss based on the probability distribution of the new action in the first iteration, calculate the first iteration KL divergence constraint term loss based on the probability distribution of the new action in the first iteration and the probability distribution of the original action in the first iteration, and combine the first iteration strategy gradient term loss and the first iteration KL divergence constraint term loss to obtain the first iteration total loss value. The model parameter first-round iteration update unit is used to calculate the gradient of the total loss value of the first round iteration with respect to the initial model parameters of the initial atmospheric environment domain decision model, and backpropagate through the gradient descent method to perform the first round iteration update of the initial model parameters to obtain the first round iteration atmospheric environment domain decision model; The model parameter iterative update unit is used to perform multiple second-round iterative task decision analysis operations based on the first-round iterative atmospheric environment domain decision model and the preset prompt text, and to select the optimal strategy for the second-round iterative round. It also uses the first-round iterative atmospheric environment domain decision model to perform second-round simulated task decision analysis operations based on the optimal strategy for the second-round iterative round. Furthermore, it calculates the total loss value for the second-round iterative round based on the probability distribution of new actions and the probability distribution of original actions corresponding to the optimal strategy for the second-round iterative round, and updates the first-round iterative model parameters of the first-round atmospheric environment domain decision model based on the total loss value for the second-round iterative round, thus obtaining the second-round atmospheric environment domain decision model. This process is repeated multiple times until a preset iteration stop condition is reached, resulting in an atmospheric environment domain decision model with fine-tuned parameters. Task decision analysis operations are then performed based on this fine-tuned atmospheric environment domain decision model.

[0015] Preferably, the first-round iteration optimal strategy selection unit includes: A multi-dimensional scoring subunit is used to score each first-round iteration candidate strategy on the dimensions of factual consistency, format standardization, strategy consensus, and text. Based on preset scoring weights matching the preset prompt text, a weighted sum of the scores for each dimension is calculated, and this weighted sum is determined as the score of the first-round iteration candidate strategy. Specifically, the factual consistency score assesses the degree of conformity between the first-round iteration candidate strategy and knowledge in the atmospheric environment domain; the format standardization score assesses the degree of conformity between the first-round iteration candidate strategy and a preset professional report structure; the strategy consensus score assesses the similarity between the first-round iteration candidate strategy and other first-round iteration candidate strategies; and the text dimension score assesses the semantic fluency of the first-round iteration candidate strategy. The first-round iteration optimal strategy selection subunit is used to select the optimal strategy for the first round iteration based on the score of each candidate strategy in the first round iteration.

[0016] Preferably, the first-round iteration optimal strategy selection subunit is used for: All candidate strategies for the first round of iterations are sorted in descending order of their scores to generate a sequence of candidate strategies for the first round of iterations. If the score of the first-order candidate strategy in the first round of iterations is different from the score of the second-order candidate strategy in the first round of iterations, then the first-order candidate strategy in the first round of iterations shall be taken as the optimal strategy in the first round of iterations. If the score of the first-order first-round iteration candidate strategy is the same as the score of at least one subsequent-order first-round iteration candidate strategy, then the first-round iteration candidate strategy with the highest score in the factual consistency dimension is selected from multiple first-round iteration candidate strategies with the same score and is taken as the optimal strategy for the first-round iteration.

[0017] Preferably, before the first round of iterative candidate strategy generation unit, the atmospheric environment domain decision model training module further includes an atmospheric environment domain decision model pre-training unit, used for: Acquire historical documents in the field of atmospheric environment within a preset historical time period, and use the historical documents as corpus in the field of atmospheric environment. Based on the atmospheric environment domain corpus, a general large model is pre-trained to obtain a pre-trained atmospheric environment domain decision model, and then the model parameters are fine-tuned based on the pre-trained atmospheric environment domain decision model.

[0018] Preferably, the preset prompt text includes atmospheric environmental data, the source of the atmospheric environmental data, the spatiotemporal coordinates of the atmospheric environmental data, and the expected analysis target.

[0019] Preferably, the preset iteration stopping condition is any one of the following: reaching a preset iteration round threshold, satisfying a preset task decision scoring threshold, or the loss of the KL divergence constraint term being continuously lower than a preset KL divergence threshold.

[0020] According to another aspect of this application, a storage medium is provided that stores at least one executable instruction, which causes a processor to perform operations corresponding to the above-described task decision generation method applicable to atmospheric environments.

[0021] According to another aspect of this application, a terminal is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the above-described task decision generation method applicable to atmospheric environments.

[0022] By employing the above technical solutions, the technical solutions provided in the embodiments of this application have at least the following advantages: This application provides a method, apparatus, medium, and terminal for generating task decisions in the atmospheric environment field. First, it receives a target task related to the atmospheric environment from a target user. Second, based on an atmospheric environment decision model with fine-tuned parameters, it performs task decision analysis based on the target task to generate the corresponding task decision. The atmospheric environment decision model is pre-trained using atmospheric environment corpus and its parameters are fine-tuned using a total loss function that includes a policy gradient loss and a KL divergence constraint loss. The policy gradient loss guides the model parameters towards a direction with a higher probability of generating the optimal strategy, while the KL divergence constraint loss limits the adjustment range of the model parameters. During the model parameter fine-tuning process, the optimal strategy used in each iteration is generated using the atmospheric environment decision model obtained after the previous iteration. Finally, based on the task decision, it identifies the responsible party for pollution, outputs a pollutant exceedance alarm, and selects the optimal control measures. Compared with existing technologies, the embodiments of this application first pre-train a general large model using atmospheric environment domain corpus, and then fine-tune the model parameters of the pre-trained model using a total loss function that includes policy gradient loss and KL divergence constraint loss to obtain an atmospheric environment domain decision model. This model is then used to generate task decisions corresponding to the target task. Since the optimal policy used in each iteration of the model parameter fine-tuning process is generated using the atmospheric environment domain decision model obtained after the previous iteration, no manual annotation is required. This avoids the problem of low robustness of the fine-tuned model caused by the strong subjectivity and inconsistent standards of manual annotation results. As a result, the generated task decisions for the atmospheric environment domain can truly meet the domain requirements. At the same time, it saves a lot of the high annotation cost consumed by manual labeling work and reduces the model fine-tuning cost. Furthermore, by introducing KL divergence constraint loss into the total loss function to constrain the update amplitude of model parameters, the training oscillation is alleviated, and the stability of model parameter fine-tuning is ensured.

[0023] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0024] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1This application provides a flowchart of a task decision generation method applicable to atmospheric environments, as illustrated in an embodiment of this application. Figure 2 This document illustrates a flowchart of the model parameter fine-tuning process for the atmospheric environment decision-making model provided in an embodiment of this application. Figure 3 This illustration shows a block diagram of a task decision generation device suitable for atmospheric environments, provided in an embodiment of this application. Figure 4 A schematic diagram of the structure of a terminal provided in an embodiment of this application is shown. Detailed Implementation

[0025] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0026] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.

[0027] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the scope of this application and its application or use.

[0028] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0029] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0030] The embodiments of this application can be applied to computer systems / servers that can operate with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with computer systems / servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems, etc.

[0031] Computer systems / servers can be described in the general context of computer system executable instructions (such as program modules) executed by the computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are performed by remote processing devices linked through a communication network. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.

[0032] This application provides a task decision generation method suitable for atmospheric environments, such as... Figure 1 As shown, the method includes: 101. Receive target tasks related to the atmospheric environment from the target user.

[0033] The target tasks include, but are not limited to, strategy generation tasks related to the atmospheric environment, such as pollution source tracing, pollutant concentration prediction, and evaluation of the effectiveness of control measures.

[0034] In this embodiment, the current execution end can be the decision generation module of the atmospheric environment governance platform.

[0035] 102. Based on the atmospheric environment decision model with completed model parameter fine-tuning, perform task decision analysis according to the target task to generate the task decision corresponding to the target task.

[0036] The atmospheric environment domain decision model is pre-trained using atmospheric environment domain corpus and fine-tuned using a total loss function that includes policy gradient loss and KL divergence constraint loss. The policy gradient loss guides the model parameters of the atmospheric environment domain decision model towards a direction with a higher probability of generating the optimal policy, while the KL divergence constraint loss limits the adjustment range of the model parameters. During the fine-tuning of the model parameters, the optimal policy used in each iteration is generated using the atmospheric environment domain decision model obtained after the previous iteration. The target task carries at least atmospheric environment data and the analysis target. The task decision analysis operation is used to generate a policy about the analysis target based on the atmospheric environment data, i.e., the task decision.

[0037] 103. Based on task decision-making, identify the responsible party for pollution, issue warnings for pollutant exceedances, and select the optimal control measures.

[0038] In this embodiment, when the target task is pollution source tracing, the generated task decision will include the identified pollution responsible party, which will be used by law enforcement agencies for accurate verification. When the target task is pollutant concentration prediction, the prediction results can be further compared with a preset warning threshold, and when the preset warning threshold is exceeded, a pollutant exceedance alarm will be output to prevent and rectify in advance and avoid actual pollution. When the target task is control measure effectiveness evaluation, the control measure with the best effect can be selected from the evaluation results as the optimal control measure to maximize the control effect.

[0039] Compared with existing technologies, the embodiments of this application first pre-train a general large model using atmospheric environment domain corpus, and then fine-tune the model parameters of the pre-trained model using a total loss function that includes policy gradient loss and KL divergence constraint loss to obtain an atmospheric environment domain decision model. This model is then used to generate task decisions corresponding to the target task. Since the optimal policy used in each iteration of the model parameter fine-tuning process is generated using the atmospheric environment domain decision model obtained after the previous iteration, no manual annotation is required. This avoids the problem of low robustness of the fine-tuned model caused by the strong subjectivity and inconsistent standards of manual annotation results. As a result, the generated task decisions for the atmospheric environment domain can truly meet the domain requirements. At the same time, it saves a lot of the high annotation cost consumed by manual labeling work and reduces the model fine-tuning cost. Furthermore, by introducing KL divergence constraint loss into the total loss function to constrain the update amplitude of model parameters, the training oscillation is alleviated, and the stability of model parameter fine-tuning is ensured.

[0040] In one embodiment of this application, for further definition and explanation, such as Figure 2 As shown, in embodiment 102, before performing task decision analysis based on the atmospheric environment domain decision model whose model parameters have been fine-tuned and generating the task decision corresponding to the target task, the method further includes: 201. Use atmospheric environment domain corpus to pre-train a general large model to obtain a pre-trained atmospheric environment domain decision model.

[0041] Accordingly, step 201 of the embodiment specifically includes: acquiring historical documents in the atmospheric environment field within a preset historical time period, and using the historical documents as corpus in the atmospheric environment field; performing pre-training processing on a general large model based on the atmospheric environment field corpus to obtain a pre-trained atmospheric environment field decision model, so as to perform model parameter fine-tuning operation on the basis of the pre-trained atmospheric environment field decision model.

[0042] The preset historical duration can be set to nearly 10 years, etc.; historical documents include, but are not limited to, historical environmental monitoring reports, scientific papers in the field of atmospheric environment, etc.

[0043] It should be noted that the pre-trained atmospheric environment decision-making model must ensure that the coverage of basic domain knowledge exceeds 80%.

[0044] 202. The pre-trained atmospheric environment domain decision model is used as the initial atmospheric environment domain decision model. Based on the initial atmospheric environment domain decision model, multiple first-round iteration task decision analysis operations are performed according to the preset prompt text to generate multiple first-round iteration candidate strategies.

[0045] The preset prompt text includes atmospheric environmental data, the source of the atmospheric environmental data, the spatiotemporal coordinates of the atmospheric environmental data, and the expected analysis target. For example, based on the data released by a local environmental monitoring station (i.e., the source of atmospheric environmental data), 15 monitoring stations in a certain area from June to August of a certain year (i.e., the spatiotemporal coordinates of atmospheric environmental data), ozone O3 monitoring data with a daily average concentration range of 80-120 μg / m³ (i.e., atmospheric environmental data), and meteorological data of the same station during the same period (i.e., the spatiotemporal coordinates of atmospheric environmental data) with a daily average wind speed of 1.5-3 m / s and sunshine duration of 6-8 h (i.e., atmospheric environmental data), the text analyzes the temporal pattern of ozone O3 pollution peaks and the main driving factors (i.e., the expected analysis target).

[0046] In this embodiment, the multiple first-round iteration task decision analysis operations can be performed separately multiple times, i.e., only one first-round iteration candidate strategy is generated each time; or it can be performed only once, sorting multiple analysis results and selecting multiple analysis results from the sequence at once as multiple first-round iteration candidate strategies. This embodiment does not impose any specific limitations.

[0047] It should be noted that more than three candidate strategies can be selected in the first round of iterations. The specific number can be adjusted according to hardware resources, but five are recommended.

[0048] 203. Based on the preset scoring rules, calculate the score of each candidate strategy in the first round of iteration, and select the optimal strategy in the first round of iteration based on the score of each candidate strategy in the first round of iteration.

[0049] Accordingly, step 203 of the embodiment specifically includes: for each candidate strategy in the first round of iteration, scoring is performed on the dimensions of factual consistency, format standardization, strategy consensus, and text. Based on the preset scoring weights that match the preset prompt text, the weighted sum of the scores of each dimension is calculated, and the weighted sum is determined as the score of the candidate strategy in the first round of iteration. Based on the score of each candidate strategy in the first round of iteration, the optimal strategy in the first round of iteration is selected.

[0050] The factual consistency dimension score assesses the degree of conformity between the first-round iteration candidate strategies and knowledge in the atmospheric environment field. Based on an atmospheric knowledge base, which may include national standards, regional monitoring databases, and atmospheric chemistry principles, the score evaluates the conformity between the first-round iteration candidate strategies and objective facts through semantic matching or logical verification. The score range can be set from 0 to 10; for example, correctly explaining the secondary formation mechanism of ozone (O3) scores 9 points, while confusing the sources of PM2.5 and ozone (O3) scores 3 points. The format standardization dimension score assesses the degree of conformity between the first-round iteration candidate strategies and the pre-set professional report structure. The first step involves verifying whether the candidate strategies for the first round of iterations conform to the preset professional report structure in the field of atmospheric environment. This preset structure can be: data source – analysis logic – core conclusions – supporting evidence. The scoring range can be set to 0-5; for example, a complete and standardized structure earns 5 points, while missing supporting evidence earns 2 points. The strategy consensus dimension score is used to assess the similarity between the current candidate strategy for the first round of iterations and other candidate strategies for the first round of iterations. Specifically, it calculates the semantic similarity between the current candidate strategy and other candidate strategies for the first round of iterations. Similarity can be measured using BERTS score or sentence vector cosine similarity, with the mean similarity score used as the overall score. The score range can be set to 0-5; for example, a mean of 0.8 or higher earns 5 points, and a mean of 0.3 or lower earns 1 point. The text dimension score evaluates the semantic fluency of the first-round candidate strategies, determining whether the text contains redundant content such as repeated data descriptions, contradictory logic such as simultaneously claiming that high wind speed promotes pollution diffusion and that high wind speed inhibits pollution diffusion, or unsafe suggestions such as excessive emissions being harmless. Points can be deducted according to the degree of violation. The score range can be set to... The preset scoring weights can be set to -5-0, for example, 5 points are deducted for serious contradictions and 0 points are awarded for no violations; the preset scoring weights can be matched with the expected analysis objectives in the preset prompt text. For example, when the expected analysis objective is a pollution source tracing task, the preset scoring weights are 0.5 for factual consistency, 0.2 for format standardization, 0.1 for strategy consensus, and 0.2 for text. When the expected analysis objective is a concentration prediction task, the preset scoring weights are 0.6 for factual consistency, 0.1 for format standardization, 0.1 for strategy consensus, and 0.2 for text, etc.

[0051] Furthermore, based on the scores of each candidate strategy in the first round of iterations, the optimal strategy for the first round of iterations is selected, including: sorting all candidate strategies in the first round of iterations according to their scores from largest to smallest to generate a sequence of candidate strategies for the first round of iterations; if the score of the first-order candidate strategy in the first round of iterations is different from the score of the second-order candidate strategy in the first round of iterations, then the first-order candidate strategy in the first round of iterations is selected as the optimal strategy for the first round of iterations; if the score of the first-order candidate strategy in the first round of iterations is the same as the score of at least one subsequent candidate strategy in the first round of iterations, then the candidate strategy with the highest score in the factual consistency dimension is selected from multiple candidate strategies with the same score in the first round of iterations and selected as the optimal strategy for the first round of iterations.

[0052] In this embodiment, the selection rule for the optimal strategy in the first round of iteration is the highest score selection method, that is, selecting the candidate strategy with the highest score in the first round of iteration as the optimal strategy in the first round of iteration. As a possible case, when there are two or more candidates with the same score, the candidate strategy with the highest score in the factual consistency dimension is preferentially selected as the optimal strategy in the first round of iteration. That is, the factual consistency dimension scores of multiple candidate strategies with the same score in the first round of iteration are compared, and the candidate strategy with the highest score in the factual consistency dimension is selected as the optimal strategy in the first round of iteration.

[0053] It should be noted that when outputting the optimal strategy for the first iteration, a detailed scoring report should also be generated and output, which must include the scoring situation of the above multiple dimensions and the reasons for the deductions, in order to improve the traceability of the scoring scores.

[0054] 204. The actual action probability distribution corresponding to the optimal strategy in the first iteration is used as the original action probability distribution in the first iteration.

[0055] The original action probability distribution in the first iteration is the actual action probability distribution calculated and stored in step 202 of the embodiment when the optimal strategy for the first iteration is generated, and it is fixed historical data.

[0056] 205. Using the initial atmospheric environment domain decision model, based on the preset prompt text, perform the first round of iteration simulation task decision analysis operation based on the first round of iteration optimal strategy, obtain the simulation action probability distribution corresponding to the first round of iteration optimal strategy, and use the simulation action probability distribution as the new action probability distribution for the first round of iteration.

[0057] In this embodiment, the preset prompt text and the optimal strategy of the first iteration are combined as training samples and input into the initial atmospheric environment domain decision model. The initial atmospheric environment domain decision model will re-execute the forward propagation to simulate the process of generating the known sequence of the optimal strategy of the first iteration from beginning to end, and output the probability distribution of each step, that is, the probability distribution of the new action of the first iteration, and the dynamically differentiable computation graph nodes.

[0058] 206. Calculate the gradient term loss of the strategy in the first iteration based on the probability distribution of the new action in the first iteration, calculate the KL divergence constraint term loss in the first iteration based on the probability distribution of the new action in the first iteration and the probability distribution of the original action in the first iteration, and combine the gradient term loss of the strategy in the first iteration and the KL divergence constraint term loss in the first iteration to obtain the total loss value of the first iteration.

[0059] In this embodiment, the differentiable total probability for the policy gradient can be calculated based on the probability distribution of the new actions in the first iteration. Then, the loss of the policy gradient term in the first iteration is calculated based on the differentiable total probability to maximize the probability of generating the optimal policy in the first iteration, thereby guiding the model parameters of the atmospheric environment decision model to adjust towards a direction with a higher probability of generating the optimal policy. At the same time, the KL divergence is calculated based on the probability distribution of the new actions in the first iteration and the probability distribution of the original actions in the first iteration, and then multiplied by a coefficient to obtain the KL divergence constraint term loss, so as to limit the adjustment range of the model parameters. The coefficient can take the value of 0.01-0.05, preferably 0.03. Finally, the loss of the policy gradient term in the first iteration and the loss of the KL divergence constraint term in the first iteration are combined to obtain the total loss value of the first iteration.

[0060] 207. Calculate the gradient of the total loss value in the first iteration with respect to the initial model parameters of the initial atmospheric environment domain decision model, and backpropagate using the gradient descent method to update the initial model parameters in the first iteration, thereby obtaining the first iteration atmospheric environment domain decision model.

[0061] The gradient is a vector that indicates the direction of model parameter adjustment and can be calculated using automatic differentiation techniques.

[0062] 208. Iterate and update the model parameters multiple times until the preset iteration stop condition is met to obtain the atmospheric environment domain decision model with the model parameters fine-tuned.

[0063] Accordingly, step 208 of the embodiment specifically includes: based on the first-round iterative atmospheric environment domain decision model, performing multiple second-round iterative task decision analysis operations according to preset prompt text, and selecting the optimal strategy for the second-round iterative round; using the first-round iterative atmospheric environment domain decision model, performing second-round simulated task decision analysis operations based on the optimal strategy for the second-round iterative round; calculating the total loss value for the second-round iterative round based on the probability distribution of new actions and the probability distribution of original actions corresponding to the optimal strategy for the second-round iterative round; and updating the parameters of the first-round iterative atmospheric environment domain decision model based on the total loss value for the second-round iterative round to obtain the second-round atmospheric environment domain decision model, and iteratively updating the model parameters multiple times until a preset iteration stop condition is reached to obtain an atmospheric environment domain decision model with fine-tuned model parameters, and performing task decision analysis operations based on the fine-tuned atmospheric environment domain decision model.

[0064] The preset iteration stopping conditions are any one of the following: reaching a preset iteration round threshold, satisfying a preset task decision scoring threshold, or the loss of the KL divergence constraint term being continuously lower than a preset KL divergence threshold. The preset iteration round threshold can be set to 10-50 rounds and can be adjusted according to the task accuracy requirements, with 30 rounds recommended. The preset task decision scoring threshold can be set to a fact consistency dimension score ≥8 points, a format standardization dimension score ≥4 points, and a score value ≥5.5.

[0065] In this embodiment, the specific process of multi-round iterative updates is the same as steps 202-207 in the embodiment, and will not be repeated here.

[0066] In specific application scenarios, training logs can be output every 5 iterations to monitor the fine-tuning process of model parameters. These logs should include KL divergence values, the mean score of the fact consistency dimension, and training loss values. The fluctuation range of the training loss value should be controlled to be less than 10% to ensure the stability of the training process.

[0067] This application provides a task decision generation method applicable to the atmospheric environment. First, it receives a target task related to the atmospheric environment from a target user. Second, based on an atmospheric environment decision model with fine-tuned parameters, it performs task decision analysis based on the target task to generate the corresponding task decision. The atmospheric environment decision model is pre-trained using atmospheric environment corpus and its parameters are fine-tuned using a total loss function that includes a policy gradient loss and a KL divergence constraint loss. The policy gradient loss guides the model parameters towards a direction with a higher probability of generating the optimal strategy, while the KL divergence constraint loss limits the adjustment range of the model parameters. During the model parameter fine-tuning process, the optimal strategy used in each iteration is generated using the atmospheric environment decision model obtained after the previous iteration. Finally, based on the task decision, it identifies the responsible party for pollution, outputs pollutant exceedance warnings, and selects the optimal control measures. Compared with existing technologies, the embodiments of this application first pre-train a general large model using atmospheric environment domain corpus, and then fine-tune the model parameters of the pre-trained model using a total loss function that includes policy gradient loss and KL divergence constraint loss to obtain an atmospheric environment domain decision model. This model is then used to generate task decisions corresponding to the target task. Since the optimal policy used in each iteration of the model parameter fine-tuning process is generated using the atmospheric environment domain decision model obtained after the previous iteration, no manual annotation is required. This avoids the problem of low robustness of the fine-tuned model caused by the strong subjectivity and inconsistent standards of manual annotation results. As a result, the generated task decisions for the atmospheric environment domain can truly meet the domain requirements. At the same time, it saves a lot of the high annotation cost consumed by manual labeling work and reduces the model fine-tuning cost. Furthermore, by introducing KL divergence constraint loss into the total loss function to constrain the update amplitude of model parameters, the training oscillation is alleviated, and the stability of model parameter fine-tuning is ensured.

[0068] Furthermore, as a response to the above Figure 1 The implementation of the method shown in this application provides a task decision generation device suitable for atmospheric environments, such as... Figure 3 As shown, the device includes: Task receiving module 31, task decision generation module 32, task decision application module 33; Task receiving module 31 is used to receive target tasks related to the atmospheric environment input by the target user; The task decision generation module 32 is used to perform task decision analysis based on the target task according to the atmospheric environment domain decision model that has been fine-tuned, and generate the task decision corresponding to the target task. The atmospheric environment domain decision model is pre-trained using atmospheric environment domain corpus and fine-tuned using a total loss function that includes policy gradient loss and KL divergence constraint loss. The policy gradient loss is used to guide the model parameters of the atmospheric environment domain decision model to adjust in a direction with a higher probability of generating the optimal policy. The KL divergence constraint loss is used to limit the adjustment range of the model parameters. During the model parameter fine-tuning process, the optimal policy used in each iteration is generated using the atmospheric environment domain decision model obtained after the previous iteration. The task decision application module 33 is used to identify the responsible party for pollution based on the task decision, output pollutant exceedance alarms, and screen the optimal control measures.

[0069] In specific application scenarios, prior to the task decision generation module, the device further includes an atmospheric environment domain decision model training module, comprising: The first-round iteration candidate strategy generation unit is used to take the pre-trained atmospheric environment domain decision model as the initial atmospheric environment domain decision model, and based on the initial atmospheric environment domain decision model, perform multiple first-round iteration task decision analysis operations according to the preset prompt text to generate multiple first-round iteration candidate strategies. The first-round iteration optimal strategy selection unit is used to calculate the score of each of the first-round iteration candidate strategies based on the preset scoring rules, and select the first-round iteration optimal strategy based on the score of each of the first-round iteration candidate strategies. The unit for obtaining the original action probability distribution of the first iteration is used to take the actual action probability distribution corresponding to the generated optimal strategy of the first iteration as the original action probability distribution of the first iteration. The first iteration new action probability distribution acquisition unit is used to utilize the initial atmospheric environment domain decision model, according to the preset prompt text, and based on the first iteration optimal strategy to perform first iteration simulation task decision analysis operation, to obtain the simulation action probability distribution corresponding to the first iteration optimal strategy, and to use the simulation action probability distribution as the first iteration new action probability distribution; The first iteration total loss calculation unit is used to calculate the first iteration strategy gradient term loss based on the probability distribution of the new action in the first iteration, calculate the first iteration KL divergence constraint term loss based on the probability distribution of the new action in the first iteration and the probability distribution of the original action in the first iteration, and combine the first iteration strategy gradient term loss and the first iteration KL divergence constraint term loss to obtain the first iteration total loss value. The model parameter first-round iteration update unit is used to calculate the gradient of the total loss value of the first round iteration with respect to the initial model parameters of the initial atmospheric environment domain decision model, and backpropagate through the gradient descent method to perform the first round iteration update of the initial model parameters to obtain the first round iteration atmospheric environment domain decision model; The model parameter iterative update unit is used to perform multiple second-round iterative task decision analysis operations based on the first-round iterative atmospheric environment domain decision model and the preset prompt text, and to select the optimal strategy for the second-round iterative round. It also uses the first-round iterative atmospheric environment domain decision model to perform second-round simulated task decision analysis operations based on the optimal strategy for the second-round iterative round. Furthermore, it calculates the total loss value for the second-round iterative round based on the probability distribution of new actions and the probability distribution of original actions corresponding to the optimal strategy for the second-round iterative round, and updates the first-round iterative model parameters of the first-round atmospheric environment domain decision model based on the total loss value for the second-round iterative round, thus obtaining the second-round atmospheric environment domain decision model. This process is repeated multiple times until a preset iteration stop condition is reached, resulting in an atmospheric environment domain decision model with fine-tuned parameters. Task decision analysis operations are then performed based on this fine-tuned atmospheric environment domain decision model.

[0070] In specific application scenarios, the first-round iteration optimal strategy selection unit includes: A multi-dimensional scoring subunit is used to score each first-round iteration candidate strategy on the dimensions of factual consistency, format standardization, strategy consensus, and text. Based on preset scoring weights matching the preset prompt text, a weighted sum of the scores for each dimension is calculated, and this weighted sum is determined as the score of the first-round iteration candidate strategy. Specifically, the factual consistency score assesses the degree of conformity between the first-round iteration candidate strategy and knowledge in the atmospheric environment domain; the format standardization score assesses the degree of conformity between the first-round iteration candidate strategy and a preset professional report structure; the strategy consensus score assesses the similarity between the first-round iteration candidate strategy and other first-round iteration candidate strategies; and the text dimension score assesses the semantic fluency of the first-round iteration candidate strategy. The first-round iteration optimal strategy selection subunit is used to select the optimal strategy for the first round iteration based on the score of each candidate strategy in the first round iteration.

[0071] In specific application scenarios, the first-round iteration optimal strategy selection subunit is used for: All candidate strategies for the first round of iterations are sorted in descending order of their scores to generate a sequence of candidate strategies for the first round of iterations. If the score of the first-order candidate strategy in the first round of iterations is different from the score of the second-order candidate strategy in the first round of iterations, then the first-order candidate strategy in the first round of iterations shall be taken as the optimal strategy in the first round of iterations. If the score of the first-order first-round iteration candidate strategy is the same as the score of at least one subsequent-order first-round iteration candidate strategy, then the first-round iteration candidate strategy with the highest score in the factual consistency dimension is selected from multiple first-round iteration candidate strategies with the same score and is taken as the optimal strategy for the first-round iteration.

[0072] In specific application scenarios, before the first round of iterative candidate strategy generation unit, the atmospheric environment domain decision model training module further includes an atmospheric environment domain decision model pre-training unit, used for: Acquire historical documents in the field of atmospheric environment within a preset historical time period, and use the historical documents as corpus in the field of atmospheric environment. Based on the atmospheric environment domain corpus, a general large model is pre-trained to obtain a pre-trained atmospheric environment domain decision model, and then the model parameters are fine-tuned based on the pre-trained atmospheric environment domain decision model.

[0073] In specific application scenarios, the preset prompt text includes atmospheric environmental data, the source of the atmospheric environmental data, the spatiotemporal coordinates of the atmospheric environmental data, and the expected analysis target.

[0074] In specific application scenarios, the preset iteration stopping condition is any one of the following: reaching a preset iteration round threshold, satisfying a preset task decision scoring threshold, or the loss of the KL divergence constraint term being continuously lower than a preset KL divergence threshold.

[0075] This application provides a task decision generation device applicable to the atmospheric environment. First, it receives a target task related to the atmospheric environment from a target user. Second, based on an atmospheric environment decision model with fine-tuned parameters, it performs task decision analysis based on the target task to generate the corresponding task decision. The atmospheric environment decision model is pre-trained using atmospheric environment corpus and its parameters are fine-tuned using a total loss function containing policy gradient loss and KL divergence constraint loss. The policy gradient loss guides the model parameters towards a direction with a higher probability of generating the optimal strategy, while the KL divergence constraint loss limits the adjustment range of the model parameters. During the model parameter fine-tuning process, the optimal strategy used in each iteration is generated using the atmospheric environment decision model obtained after the previous iteration. Finally, based on the task decision, it identifies the responsible party for pollution, outputs pollutant exceedance warnings, and selects the optimal control measures. Compared with existing technologies, the embodiments of this application first pre-train a general large model using atmospheric environment domain corpus, and then fine-tune the model parameters of the pre-trained model using a total loss function that includes policy gradient loss and KL divergence constraint loss to obtain an atmospheric environment domain decision model. This model is then used to generate task decisions corresponding to the target task. Since the optimal policy used in each iteration of the model parameter fine-tuning process is generated using the atmospheric environment domain decision model obtained after the previous iteration, no manual annotation is required. This avoids the problem of low robustness of the fine-tuned model caused by the strong subjectivity and inconsistent standards of manual annotation results. As a result, the generated task decisions for the atmospheric environment domain can truly meet the domain requirements. At the same time, it saves a lot of the high annotation cost consumed by manual labeling work and reduces the model fine-tuning cost. Furthermore, by introducing KL divergence constraint loss into the total loss function to constrain the update amplitude of model parameters, the training oscillation is alleviated, and the stability of model parameter fine-tuning is ensured.

[0076] According to one embodiment of this application, a storage medium is provided, the storage medium storing at least one executable instruction, the computer-executable instruction being able to execute the task decision generation method applicable to atmospheric environment in any of the above method embodiments.

[0077] Based on this understanding, the technical solution of this application can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or portable hard drive), and includes several instructions to cause a computer device (such as a personal computer, server, or network device) to execute the methods described in the various implementation scenarios of this application.

[0078] Figure 4The diagram shows a structural schematic of a terminal according to one embodiment of the present application. The specific embodiments of the present application do not limit the specific implementation of the terminal.

[0079] like Figure 4 As shown, the terminal may include: a processor 402, a communications interface 404, a memory 406, and a communications bus 408.

[0080] The processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408.

[0081] Communication interface 404 is used to communicate with other network elements such as clients or other servers.

[0082] The processor 402 is used to execute program 410, specifically to perform the relevant steps in the above-described embodiment of the task decision generation method applicable to atmospheric environment.

[0083] Specifically, program 410 may include program code that includes computer operation instructions.

[0084] Processor 402 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The computer device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0085] Memory 406 is used to store program 410. Memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0086] Specifically, program 410 can be used to cause processor 402 to perform the following operations: Receive target tasks related to the atmospheric environment from the target user; Based on the atmospheric environment domain decision model with completed model parameter fine-tuning, task decision analysis is performed according to the target task to generate the task decision corresponding to the target task. The atmospheric environment domain decision model is pre-trained using atmospheric environment domain corpus and obtained by fine-tuning the model parameters using a total loss function that includes policy gradient loss and KL divergence constraint loss. The policy gradient loss is used to guide the model parameters of the atmospheric environment domain decision model to adjust in a direction with a higher probability of generating the optimal policy. The KL divergence constraint loss is used to limit the adjustment range of the model parameters. During the model parameter fine-tuning process, the optimal policy used in each iteration is generated using the atmospheric environment domain decision model obtained after the previous iteration. Based on the aforementioned task, the responsible party for pollution is identified, warnings of pollutant exceedances are issued, and the optimal control measures are selected.

[0087] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the physical device for the aforementioned task decision generation method applicable to atmospheric environments, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0088] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0089] The methods and systems of this application may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this application are not limited to the order specifically described above, unless otherwise specifically stated. Furthermore, in some embodiments, this application may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this application. Thus, this application also covers recording media storing programs for performing the methods according to this application.

[0090] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.

[0091] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A task decision generation method suitable for atmospheric environments, characterized in that, include: Receive target tasks related to the atmospheric environment from the target user; Based on the atmospheric environment domain decision model with completed model parameter fine-tuning, task decision analysis is performed according to the target task to generate the task decision corresponding to the target task. The atmospheric environment domain decision model is pre-trained using atmospheric environment domain corpus and obtained by fine-tuning the model parameters using a total loss function that includes policy gradient loss and KL divergence constraint loss. The policy gradient loss is used to guide the model parameters of the atmospheric environment domain decision model to adjust in a direction with a higher probability of generating the optimal policy. The KL divergence constraint loss is used to limit the adjustment range of the model parameters. During the model parameter fine-tuning process, the optimal policy used in each iteration is generated using the atmospheric environment domain decision model obtained after the previous iteration. Based on the aforementioned task, the responsible party for pollution is identified, warnings of pollutant exceedances are issued, and the optimal control measures are selected.

2. The method according to claim 1, characterized in that, Before generating the task decision corresponding to the target task by performing task decision analysis based on the fine-tuned atmospheric environment domain decision model, the method further includes: The pre-trained atmospheric environment domain decision model is used as the initial atmospheric environment domain decision model. Based on the initial atmospheric environment domain decision model, multiple first-round iteration task decision analysis operations are performed according to the preset prompt text to generate multiple first-round iteration candidate strategies. Based on the preset scoring rules, the score of each candidate strategy in the first round of iteration is calculated, and the optimal strategy in the first round of iteration is selected based on the score of each candidate strategy in the first round of iteration. The actual action probability distribution corresponding to the optimal strategy in the first round of iteration is used as the original action probability distribution in the first round of iteration; Using the initial atmospheric environment domain decision model, based on the preset prompt text, and based on the first round of iteration optimal strategy, the first round of iteration simulation task decision analysis operation is performed to obtain the simulation action probability distribution corresponding to the first round of iteration optimal strategy, and the simulation action probability distribution is used as the first round of iteration new action probability distribution; The gradient term loss of the strategy in the first iteration is calculated based on the probability distribution of the new action in the first iteration, and the KL divergence constraint term loss of the first iteration is calculated based on the probability distribution of the new action in the first iteration and the probability distribution of the original action in the first iteration. The total loss value of the first iteration is obtained by combining the gradient term loss of the strategy in the first iteration and the KL divergence constraint term loss of the first iteration. The gradient of the total loss value in the first iteration with respect to the initial model parameters of the initial atmospheric environment domain decision model is calculated, and backpropagation is performed using the gradient descent method to update the initial model parameters in the first iteration, thereby obtaining the first iteration atmospheric environment domain decision model. Based on the first-round iterative atmospheric environment domain decision model, multiple second-round iterative task decision analysis operations are performed according to the preset prompt text, and the optimal strategy for the second iteration round is selected. Using the first-round iterative atmospheric environment domain decision model, the second-round simulated task decision analysis operation is performed based on the optimal strategy for the second iteration round. The total loss value for the second iteration round is calculated based on the probability distribution of new actions and the probability distribution of original actions corresponding to the optimal strategy for the second iteration round. The parameters of the first-round iterative model of the first-round atmospheric environment domain decision model are updated in the second iteration round based on the total loss value of the second iteration round, resulting in the second-round atmospheric environment domain decision model. The model parameters are iteratively updated multiple times until a preset iteration stop condition is reached, resulting in an atmospheric environment domain decision model with fine-tuned parameters. Task decision analysis operations are then performed based on this fine-tuned atmospheric environment domain decision model.

3. The method according to claim 2, characterized in that, The step of calculating the score of each candidate strategy in the first round of iterations based on a preset scoring rule, and selecting the optimal strategy in the first round of iterations based on the score of each candidate strategy in the first round of iterations, includes: For each candidate strategy in the first round of iterations, scores are given for factual consistency, format standardization, strategy consensus, and text. Based on preset score weights that match the preset prompt text, a weighted sum of the scores for each dimension is calculated, and this weighted sum is determined as the score of the candidate strategy in the first round of iterations. Specifically, the factual consistency score assesses the degree of conformity between the candidate strategy and knowledge in the atmospheric environment domain; the format standardization score assesses the degree of conformity between the candidate strategy and the preset professional report structure; the strategy consensus score assesses the similarity between the candidate strategy and other candidate strategies in the first round of iterations; and the text score assesses the semantic fluency of the candidate strategy. Based on the scores of each candidate strategy in the first round of iteration, the optimal strategy in the first round of iteration is selected.

4. The method according to claim 3, characterized in that, The process of selecting the optimal strategy for the first round of iterations based on the score of each candidate strategy in the first round of iterations includes: All candidate strategies for the first round of iterations are sorted in descending order of their scores to generate a sequence of candidate strategies for the first round of iterations. If the score of the first-order candidate strategy in the first round of iterations is different from the score of the second-order candidate strategy in the first round of iterations, then the first-order candidate strategy in the first round of iterations shall be taken as the optimal strategy in the first round of iterations. If the score of the first-order first-round iteration candidate strategy is the same as the score of at least one subsequent-order first-round iteration candidate strategy, then the first-round iteration candidate strategy with the highest score in the factual consistency dimension is selected from multiple first-round iteration candidate strategies with the same score and is taken as the optimal strategy for the first-round iteration.

5. The method according to claim 2, characterized in that, Before using the pre-trained atmospheric environment domain decision model as the initial atmospheric environment domain decision model, the method further includes: Acquire historical documents in the field of atmospheric environment within a preset historical time period, and use the historical documents as corpus in the field of atmospheric environment. Based on the atmospheric environment domain corpus, a general large model is pre-trained to obtain a pre-trained atmospheric environment domain decision model, and then the model parameters are fine-tuned based on the pre-trained atmospheric environment domain decision model.

6. The method according to claim 2, characterized in that, The preset prompt text includes atmospheric environmental data, the source of the atmospheric environmental data, the spatiotemporal coordinates of the atmospheric environmental data, and the expected analysis target.

7. The method according to claim 2, characterized in that, The preset iteration stopping condition is any one of the following: reaching a preset iteration round threshold, satisfying a preset task decision scoring threshold, or the loss of the KL divergence constraint term continuously being lower than a preset KL divergence threshold.

8. A task decision generation device suitable for atmospheric environments, characterized in that, include: The task receiving module is used to receive target tasks related to the atmospheric environment from the target user. The task decision generation module is used to perform task decision analysis based on the target task according to the atmospheric environment domain decision model that has been fine-tuned, and generate the task decision corresponding to the target task. The atmospheric environment domain decision model is pre-trained using atmospheric environment domain corpus and fine-tuned using a total loss function that includes policy gradient loss and KL divergence constraint loss. The policy gradient loss is used to guide the model parameters of the atmospheric environment domain decision model to adjust in a direction with a higher probability of generating the optimal policy. The KL divergence constraint loss is used to limit the adjustment range of the model parameters. During the model parameter fine-tuning process, the optimal policy used in each iteration is generated using the atmospheric environment domain decision model obtained after the previous iteration. The task decision application module is used to identify the responsible party for pollution based on the task decision, output alarms for pollutant exceedances, and screen the optimal control measures.

9. A storage medium storing at least one executable instruction, characterized in that, The executable instructions cause the processor to perform the operations corresponding to the task decision generation method applicable to atmospheric environments as described in any one of claims 1-7.

10. A terminal, comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction, characterized in that the executable instruction causes the processor to perform the operation corresponding to the task decision generation method applicable to atmospheric environment as described in any one of claims 1-7.