Decision-making method and device based on forward diffusion and reverse denoising, equipment and medium

By generating and diffusing multimodal feature vectors and performing reverse denoising, the problems of single multimodal decision results and insufficient adaptability in existing technologies are solved, and efficient decision support in multimodal environments is achieved.

CN120952170APending Publication Date: 2025-11-14PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511060292.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-30
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing technologies lack the ability to integrate multi-dimensional information and generate goal-oriented, diverse, and adaptive dynamic decision-making results under multimodal input conditions, especially in complex environments where they struggle to provide accurate and robust decision support.

Method used

By acquiring visual data, language data, historical action data, and environmental state data, visual features, language features, action features, and environmental features are generated. These features are fused to generate a comprehensive multimodal feature vector, and Gaussian distributed noise is added to it. After performing multiple steps of forward diffusion, iterative reverse denoising starts from the pure noise endpoint state and is finally converted into action instructions through the decision mapping module.

Benefits of technology

It enables the generation of diverse and adaptive dynamic decision results under multimodal input conditions, and can dynamically adjust and output better decision results in complex environments, thereby improving the model's decision adaptability and robustness in multimodal environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952170A_ABST
    Figure CN120952170A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes such as financial science and technology and medical health, and discloses a decision-making method, device and equipment based on forward diffusion and reverse denoising and a medium. The method comprises the following steps: fusing multi-modal features to generate a comprehensive feature vector, adding Gaussian noise to the comprehensive feature vector to generate an initial noisy feature vector, executing forward diffusion to obtain a noisy state sequence, starting reverse denoising from a pure noise end point state to obtain a final decision state, and converting the final decision state into an action instruction through a decision mapping module. According to the method, through the forward diffusion and reverse denoising process guided by reinforcement learning, the feature information of the multi-modal input data is effectively combined, and the optimal decision space is gradually approached, so that the generated action instruction has higher diversity and adaptability, and a better decision result can be dynamically adjusted and output in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a decision-making method, apparatus, device, and storage medium based on forward diffusion and reverse denoising. Background Technology

[0002] In Vision-Language-Action (VLA) model-based decision generation technology, existing methods generally suffer from problems such as singular decision results, lack of flexibility and diversity, especially evident in complex robot operations or multi-scenario interaction tasks of intelligent agents. Traditional rule-based or simple neural network-based decision-making methods struggle to adapt effectively to dynamically changing environments, resulting in low task success rates and insufficient adaptability in complex tasks, making it difficult to meet the ever-increasing demands for intelligence.

[0003] In the field of fintech, existing technologies often rely on fixed rules or static models for decision-making in tasks such as intelligent risk control and automated review. These decision-making methods lack a comprehensive understanding and flexible response to dynamic multimodal inputs (such as text, images, and historical operational data), resulting in insufficient adaptability to changes in complex financial business scenarios and difficulty in providing accurate and robust decision support in scenarios with volatile risks or diverse data.

[0004] In the field of healthcare, existing technologies in applications such as intelligent diagnosis and treatment assistance and surgical robot control mainly rely on rule-driven or single neural network models to generate decisions. This makes it difficult to comprehensively process multimodal information such as visual images, verbal descriptions, and historical operational data. Consequently, when faced with complex medical scenarios such as dynamic changes in the surgical environment, the decision-making results lack diversity and flexible adjustment capabilities, making it difficult to meet the requirements of accurate and safe operation.

[0005] Furthermore, while existing technologies have gradually introduced generative models such as diffusion models to enhance data diversity, traditional diffusion models lack a clear goal orientation, making it difficult to effectively generate practical decision results in specific task scenarios. When faced with high-dimensional and complex multimodal decision spaces, the limitations of existing diffusion models and reinforcement learning are further amplified. For example, diffusion models struggle to achieve targeted decisions through pure data distribution fitting, while reinforcement learning suffers from low sample efficiency and poor convergence, lacking effective integration and utilization of multimodal data. This results in poor performance in practical multimodal decision tasks, failing to meet the real-time decision-making needs of agents in complex multi-domain environments. Summary of the Invention

[0006] The main objective of this invention is to provide a decision-making method, apparatus, device, and storage medium based on forward diffusion and reverse denoising, aiming to solve the technical problem that the prior art lacks the ability to fuse multi-dimensional information and generate goal-oriented, diverse, and adaptive dynamic decision-making results under multi-modal input conditions.

[0007] To achieve the above objectives, the present invention provides a decision-making method based on forward diffusion and reverse denoising, comprising:

[0008] Acquire visual data, language data, historical action data, and environmental state data, and generate visual features, language features, action features, and environmental features. Then, fuse the visual features, language features, action features, and environmental features to generate a comprehensive multimodal feature vector.

[0009] Gaussian noise is added to the integrated multimodal feature vector to generate an initial noisy feature vector;

[0010] Perform multi-step forward diffusion on the initial noisy feature vector to obtain a noisy state sequence;

[0011] Starting from the pure noise endpoint state of the noisy state sequence, iterative reverse denoising is performed to obtain the final decision state;

[0012] The final decision state is converted into action instructions through the decision mapping module.

[0013] Furthermore, to achieve the above objectives, the present invention provides a decision-making device based on forward diffusion and reverse denoising, comprising:

[0014] The multimodal feature encoding module is used to acquire visual data, language data, historical action data, and environmental state data, and generate visual features, language features, action features, and environmental features, and fuse the visual features, language features, action features, and environmental features to generate a comprehensive multimodal feature vector;

[0015] The noise injection module is used to add Gaussian distributed noise to the comprehensive multimodal feature vector to generate an initial noisy feature vector;

[0016] The forward diffusion module is used to perform multi-step forward diffusion on the initial noisy feature vector to obtain a noisy state sequence;

[0017] The reverse denoising module is used to iteratively perform reverse denoising starting from the pure noise endpoint state of the noisy state sequence to obtain the final decision state;

[0018] The decision mapping module is used to convert the final decision state into action instructions.

[0019] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a decision program based on forward diffusion and reverse denoising stored in the memory and executable on the processor, wherein when the decision program based on forward diffusion and reverse denoising is executed by the processor, it implements the steps of the decision method based on forward diffusion and reverse denoising as described above.

[0020] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a decision program based on forward diffusion and reverse denoising, wherein when the decision program based on forward diffusion and reverse denoising is executed by a processor, it implements the steps of the decision method based on forward diffusion and reverse denoising as described above.

[0021] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a decision-making method, apparatus, device, and medium based on forward diffusion and reverse denoising, comprising: acquiring visual data, language data, historical action data, and environmental state data and generating visual features, language features, action features, and environmental features; fusing visual features, language features, action features, and environmental features to generate a comprehensive multimodal feature vector; adding Gaussian distributed noise to the comprehensive multimodal feature vector to generate an initial noisy feature vector; performing multi-step forward diffusion to obtain a noisy state sequence; iteratively performing reverse denoising from the pure noise endpoint state of the noisy state sequence to obtain the final decision state; and converting the final decision state into action commands through a decision mapping module. This invention, through a reinforcement learning-guided forward diffusion and reverse denoising process, effectively combines the feature information of multimodal input data, gradually approximating the optimal decision space, making the generated action commands more diverse and adaptable, and able to dynamically adjust and output better decision results in complex environments. Attached Figure Description

[0022] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:

[0023] Figure 1 This is a schematic diagram of an application environment for a decision-making method based on forward diffusion and reverse denoising in one embodiment of the present invention;

[0024] Figure 2 This is a flowchart illustrating an embodiment of the decision-making method based on forward diffusion and reverse denoising of the present invention.

[0025] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the decision-making device based on forward diffusion and reverse denoising of the present invention;

[0026] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0027] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0028] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0029] The decision-making method based on forward diffusion and reverse denoising provided in this invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can acquire visual data, language data, historical action data, and environmental state data from the user terminal and generate visual features, language features, action features, and environmental features. It then fuses these features to generate a comprehensive multimodal feature vector. Gaussian noise is added to this comprehensive multimodal feature vector to generate an initial noisy feature vector. Multi-step forward diffusion is performed to obtain a noisy state sequence. Starting from the pure noise endpoint of the noisy state sequence, iterative reverse denoising is performed to obtain the final decision state. The final decision state is then converted into an action command through a decision mapping module. This invention, through reinforcement learning-guided forward diffusion and reverse denoising processes, effectively combines the feature information of multimodal input data, gradually approximating the optimal decision space. This results in more diverse and adaptable action commands, enabling dynamic adjustment and output of better decision results in complex environments. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0030] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the decision-making method based on forward diffusion and reverse denoising provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0031] like Figure 2 As shown, the decision-making method based on forward diffusion and reverse denoising proposed in this invention includes the following steps:

[0032] S10: Acquire visual data, language data, historical action data, and environmental state data, and generate visual features, language features, action features, and environmental features; fuse the visual features, language features, action features, and environmental features to generate a comprehensive multimodal feature vector.

[0033] In this embodiment, acquiring visual data, language data, historical action data, and environmental status data requires a multi-source data acquisition module. The raw data streams captured by external sensors, microphone arrays, input logs, and environmental status monitors are input into the data receiving interface. Visual data includes, but is not limited to, image frame sequences and video streams, which can be captured by a camera and transcoded into a standard image format. Language data can be acquired by an audio capture module, which can then perform speech recognition and transcribe it into text. Historical action data can be stored in a behavior log or collected by action sequence sensors. Environmental status data can be collected by a multi-dimensional environmental monitor, such as a temperature sensor, a position sensor, or other external environmental acquisition modules. All of these data undergo a standardized input channel for unified format conversion, laying the foundation for input consistency for subsequent processing.

[0034] Generating visual features involves inputting standardized visual data into a deep convolutional neural network (CNN) to extract multi-level image representations. The CNN can employ multi-layer convolution and pooling structures to capture local spatial texture patterns and gradually form high-dimensional semantic representations, suitable for recognizing semantic features such as objects, scenes, and spatial relationships in visual content. Generating linguistic features involves processing standardized language text using natural language processing models. Employing models such as word embeddings or encoders, the language text is transformed into vectorized representations, capturing contextual dependencies in word sequences and generating high-dimensional semantic embeddings, suitable for inputs such as multilingual instructions, semantic commands, and interactive task descriptions. Generating action features involves inputting historical action sequences into time-series modeling structures such as temporal convolutional networks (TCNs) or recurrent neural networks (RNNs) for sequence modeling, extracting dynamic change patterns in action trajectories. Generating environmental features involves compressing and mapping environmental state data to a standard numerical range according to statistical distribution or predefined intervals using data normalization methods, ensuring that different environmental inputs are adapted to subsequent processing modules within the same dimension and scale.

[0035] The integration of visual, linguistic, behavioral, and environmental features is used to construct a comprehensive multimodal feature vector. This involves unifying these multidimensional features into a shared feature space through vector concatenation, weighted superposition, or cross-modal transformation mapping. Vector concatenation is suitable for simple combinations, directly merging features from different modalities to form a joint vector. Weighted superposition uses learned modal weights to weight and combine features from different modalities, suitable for balancing the information contributions between modalities. Cross-modal mapping, through cross-modal attention mechanisms or alignment mapping networks, projects features from different modalities into a unified high-dimensional vector space and enhances their interrelationships. The comprehensive multimodal feature vector provides a unified, high-dimensional, and multi-information fusion input representation for subsequent reasoning and decision-making, facilitating the modeling of multimodal interaction relationships and supporting subsequent data generation and decision-making tasks.

[0036] Visual and environmental data can be acquired through a combined configuration of depth cameras and LiDAR, while language data can be collected using a distributed microphone array combined with an acoustic model. Historical motion data can be recorded using robot joint encoders and accelerometers. Different types of convolutional neural networks, such as lightweight networks, can be used to extract visual features on computationally limited devices. Language data can be processed using multilingual BERT models or Transformer structures. Motion feature extraction can be achieved using hybrid sequence networks combining one-dimensional convolutions and gated recurrent units to adapt to long-term input sequences. Environmental data normalization can be dynamically adjusted based on statistical distribution estimation to adapt to different task scenarios. Fusion can employ multilayer perceptrons or attention mechanisms to perform projection mapping and weighted superposition on each feature vector, flexibly adjusting the contribution ratio of different input modalities to the overall representation. Furthermore, multimodal feature fusion can be achieved through adaptive weighting to adapt to the multimodal input features of medical imaging and electronic health record data, or financial market charts and text reports.

[0037] Example: In the healthcare business, features are extracted and fused from patients' imaging examination results, chief complaints, past treatment procedures and current vital sign monitoring data to form a comprehensive multimodal feature vector, which is used to assist in clinical pathway recommendation and personalized treatment decision-making.

[0038] In the fintech business, feature extraction and fusion are performed on market visual chart data, market analysis report text, historical investment behavior trajectories and market environment status data to form a comprehensive multimodal feature vector, which is used to support intelligent portfolio adjustment and dynamic risk management decisions.

[0039] This embodiment transforms multi-source heterogeneous data into high-dimensional features in a unified numerical space after preprocessing, and then fuses them in the feature space to form a comprehensive multimodal feature representation. This fully utilizes the complementarity and synergy between different modal inputs, improves the integrity and expressive power of data representation, provides global information support for decision-making in complex tasks, and effectively enhances the model's decision-making adaptability and robustness in multimodal environments.

[0040] S20, Gaussian distributed noise is added to the integrated multimodal feature vector to generate an initial noisy feature vector;

[0041] In this embodiment, adding Gaussian distributed noise to the integrated multimodal feature vector to generate the initial noisy feature vector involves randomly perturbing the fused multidimensional features to introduce noise distribution characteristics and expand the range of data distribution support. The integrated multimodal feature vector is composed of high-dimensional representations of multi-source data such as vision, language, action, and environment, spliced ​​or weighted. This vector is a point in a high-dimensional continuous space, and its distribution and scale may vary depending on the modality. Gaussian distributed noise is a normally distributed random variable with zero mean, unit variance, or adjustable variance. It can be directly generated by a pseudo-random number generator using a standard normal distribution sampling function, and its value is strictly consistent with the integrated multimodal feature vector in terms of dimension. By statistically analyzing the amplitude distribution of this vector in each dimension, a noise intensity adjustment coefficient can be determined so that the added noise matches the original feature values ​​in terms of scale, thereby avoiding excessive perturbation of the feature distribution or information overload caused by noise. Based on the information entropy levels of different feature channels, a channel weight matrix can be constructed. Weighted noise is then added element-by-element to the original integrated multimodal features after being weighted by channel, enhancing the controllability and fine-grained adjustment capability of noise addition. To maintain the correlation between features in the multidimensional feature space and the stability of the overall data geometry, an orthogonal transformation can be performed on the noise addition result to optimize its representation quality in the high-dimensional space, making it suitable for the input requirements of subsequent diffusion models.

[0042] A random noise matrix with the same dimension as the comprehensive multimodal feature vector can be generated using a standard normal distribution random number generation function such as the Box-Muller method or a sampling algorithm. The noise intensity adjustment coefficient can also be dynamically adjusted by calculating the standard deviation or amplitude range of the comprehensive multimodal feature vector in real time to adapt to the amplitude distribution of different data batches. Channel weight matrices can be determined based on the information entropy or coefficient of variation of each channel, and the noise matrix can be element-wise weighted and then added to the comprehensive multimodal feature vector element-wise. Alternatively, multidimensional feature orthogonal transformation matrices such as QR decomposition or identity orthogonal matrices can be used to project the noisy vectors onto an orthogonal coordinate system to improve their distribution uniformity and reduce multidimensional redundancy and correlation. Furthermore, the noise intensity can be adjusted according to specific scenarios; for example, applying smaller weights to sensitive indicator channels in healthcare data, while applying higher noise weights to high-volatility data channels in the financial field, can enhance training robustness.

[0043] Example: In the healthcare business, combining patient multimodal feature vectors with dynamically adjusted Gaussian noise can simulate measurement errors and individual differences, improving the model's generalization and adaptability when faced with different patient data.

[0044] In the fintech business, combining market multimodal feature vectors with Gaussian noise can simulate the uncertainty of market fluctuations and improve the model's robustness and predictive ability to diverse changes in market conditions.

[0045] This embodiment adds controllable Gaussian noise to the comprehensive multimodal feature vector and performs orthogonal adjustment, which can not only effectively simulate the potential uncertainty and randomness in the data, but also enhance the distribution generalization ability of multimodal features, making the data more learnable in the subsequent diffusion process, improving the model's adaptability to diverse environments, reducing the risk of overfitting, and improving the overall generation quality.

[0046] S30, perform multi-step forward diffusion on the initial noisy feature vector to obtain a noisy state sequence;

[0047] In this embodiment, multi-step forward diffusion is performed on the initial noisy feature vector to obtain a noisy state sequence. This is achieved by iteratively adding noise to the high-dimensional feature space, gradually causing the data distribution to tend towards a random Gaussian distribution. The initial noisy feature vector serves as the starting point of the diffusion process, and forward diffusion requires sequential iterative operations according to a preset time step. In each iteration, the noise scheduling parameter for the current time step needs to be determined. This parameter can be determined by a predefined noise scheduling function or a time step mapping table to dynamically adjust the proportion of noise amplitude to be added in the current time step. During each iteration, an independent and identically distributed Gaussian random noise matrix is ​​sampled, combined with the noise scheduling parameter for the current time step, and a new feature vector is generated by weighted superposition of the current feature vector and the Gaussian noise matrix according to a weighted mixing formula. The updated feature vector needs to be stored in the state sequence to record its change trajectory. The time step counter needs to be incremented to advance the iteration until the predefined total number of forward time steps is reached, ultimately generating a complete noisy state sequence as input for the subsequent reverse denoising process.

[0048] The noise scheduling parameter sequence can be defined in linear or exponential form, allowing for slow initial noise addition followed by a gradual increase in noise intensity to accelerate diffusion. A standard Gaussian noise matrix, with dimensions identical to the current diffusion state, can be dynamically generated using a pseudo-random number generator. The current diffusion state and the noise matrix are weighted and mixed according to the noise scheduling parameters at the current time step, calculated as α_t * current diffusion state + β_t * Gaussian noise, with α_t and β_t adaptively adjusted according to the time step. Each diffusion state update can be cached to support the recording and subsequent retrieval of the complete state sequence. The noise scheduling parameter curve can be customized to meet the needs of different scenarios. For example, in the healthcare field, prioritizing the stability of key physiological indicators can reduce disturbances to key features by limiting the noise amplitude in specific dimensions; in the fintech field, the noise scheduling curve can be adjusted to more quickly simulate the high uncertainty of market conditions.

[0049] This embodiment performs multi-step forward diffusion on the initial noisy feature vector, which can gradually transform the data from the original distribution to the form of a standard Gaussian distribution, so that the subsequent reverse denoising process has good initial conditions and training stability. At the same time, through dynamic control of noise scheduling parameters, the diffusion rate can be flexibly adjusted to enhance the model's adaptability and robustness to diverse environments.

[0050] S40, starting from the pure noise endpoint state of the noisy state sequence, iteratively perform reverse denoising to obtain the final decision state;

[0051] In this embodiment, iterative reverse denoising is performed starting from the pure noise endpoint state of the noisy state sequence to obtain the final decision state. This is a progressive denoising iterative computation process used to gradually convert Gaussian noise into a usable multimodal decision representation. The pure noise endpoint state is the last state at the end of forward diffusion, and its feature vector has essentially degenerated into an approximately pure Gaussian distribution. Each iteration of reverse denoising requires the current denoised state as input to determine the current time step and obtain the noise scheduling parameters corresponding to the current time step. Based on the current time step, the current denoised state is input into the diffusion network to generate prediction results for the noise components. Simultaneously, the current denoised state is input into the policy network to output the action probability distribution. The action probability distribution is transformed to generate a policy guidance signal. The policy guidance signal is used to adjust the predicted noise. The adjusted noise, combined with the noise scheduling parameters of the current time step, is used to correct the current denoised state, generating an updated denoised state. The updated denoised state is directly used as the input for the next iteration. The time step counter decreases by a predefined step size, and the iteration process continues until the time step counter decreases to zero. The final current denoised state is the final decision state.

[0052] The noise scheduling parameters for the current time step can be dynamically calculated in each iteration. These parameters can be obtained linearly or non-linearly through a time step mapping function. A deep neural network can be used as a diffusion network, taking the current denoising state and time step embedding as input and predicting the noise components corresponding to the current time step as output. A deep policy network can be used to predict the probability of the current denoising state, outputting an action probability distribution, which can then be converted into a policy guidance signal through Softmax or other mapping functions. A weighted adjustment mechanism can be used to combine the policy guidance signal with the predicted noise, and the adjusted noise is used to correct the current denoising state. The updated denoising state can be tagged with a time step during storage to support traceability. The noise scheduling parameters or policy guidance weight coefficients can be adjusted in different application scenarios. For example, in the healthcare field, the noise adjustment magnitude for key vital signs can be reduced to decrease prediction bias; in the fintech field, the weight of the policy guidance signal can be increased to improve responsiveness to the external environment when market conditions change drastically.

[0053] This embodiment gradually guides the state of the Gaussian noise space back to a space closer to the data distribution by starting from the pure noise endpoint state and proceeding in reverse. After introducing the guiding role of the policy network, the directionality and consistency of the denoising can be effectively improved, so that the final decision state can better reflect the comprehensive decision requirements in the current multimodal environment, and improve diversity and accuracy.

[0054] S50, the final decision state is converted into action instructions through the decision mapping module.

[0055] In this embodiment, the decision mapping module converts the final decision state into action instructions, a process that transforms the denoised high-dimensional decision representation into executable low-dimensional action space instructions. The final decision state is a comprehensive representation of multimodal information output from the reverse denoising process, possessing a complex feature structure and high dimensionality. The decision mapping module typically consists of multiple sub-modules connected in series. First, the final decision state is input into a fully connected layer for linear mapping and feature compression, generating a decision feature vector to reduce dimensionality and initially extract task-related global features. The decision feature vector is then processed by an activation function module, where activation functions such as ReLU and GELU are introduced to enhance the model's expressive power, outputting activated decision features. The activated decision features are then input into a task adaptation analysis module to analyze the current task type or context, adjusting the feature dimensions and distribution to adapt to different task requirements. The task-adapted features serve as the next input to the action encoder, which maps the input features to the target action space through parameter matrix encoding operations, generating action encoding vectors. Finally, the action encoding vector enters the instruction conversion module, which performs format conversion and instruction semantic filling to convert the action encoding vector into executable action instructions for the agent to execute.

[0056] By defining the weight matrix and bias vector of the fully connected layer, the mapping from the final decision state to the decision feature vector can be accurately realized, supporting dynamic configuration of input and output dimensions. Different types of activation functions can be selected and their parameters adjusted to adapt to different nonlinear requirements. The task adaptation analysis module can perform task context analysis, performing affine transformations or normalization on feature dimensions or distributions through task identifier input or external task description parsing. The action encoder can adopt a multi-head linear transformation or single-head matrix transformation structure to flexibly adapt to different action spaces. The instruction conversion module can encapsulate the action encoding vector into action instructions according to protocol encoding, standard instruction formats, or communication protocols, based on the requirements of the target platform or execution environment. For example, in the healthcare business field, feature compression, medical task adaptation, encoding conversion, and instruction encapsulation can be performed on the final decision state to generate medical equipment control instructions; in the fintech business field, decision features can be adapted according to the task type of the financial transaction context, and then the action encoding vector can be converted into risk control adjustment instructions or transaction execution instructions.

[0057] This embodiment maps and transforms the final decision state step by step through the decision mapping module. It can compress and encode the high-dimensional complex decision representation into the executable instruction space, so that the final action instruction can be adapted to the requirements of the execution system in terms of dimension, format and semantics, thereby improving execution efficiency and accuracy and reducing information loss in the instruction conversion process.

[0058] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a decision-making method, apparatus, device, and medium based on forward diffusion and reverse denoising, comprising: acquiring visual data, language data, historical action data, and environmental state data and generating visual features, language features, action features, and environmental features; fusing visual features, language features, action features, and environmental features to generate a comprehensive multimodal feature vector; adding Gaussian distributed noise to the comprehensive multimodal feature vector to generate an initial noisy feature vector; performing multi-step forward diffusion to obtain a noisy state sequence; iteratively performing reverse denoising from the pure noise endpoint state of the noisy state sequence to obtain the final decision state; and converting the final decision state into action instructions through a decision mapping module. This invention, through a reinforcement learning-guided forward diffusion and reverse denoising process, effectively combines the feature information of multimodal input data, gradually approximating the optimal decision space, enabling the generated action instructions to have greater diversity and adaptability, and dynamically adjusting and outputting better decision results in complex environments.

[0059] In one embodiment, step S10 includes:

[0060] S101, acquire visual data, language data, historical action data, and environmental status data;

[0061] S102, The visual data is processed using a hierarchical network to generate visual features;

[0062] S103, The language data is processed using a pre-trained language model to generate language features;

[0063] S104, The historical action data is processed using a temporal convolutional network and a gated recurrent unit to generate action features;

[0064] S105, Normalize the environmental state data to generate environmental features;

[0065] S106, input the visual features, language features, action features and environmental features into the multilayer perceptron to generate a comprehensive multimodal feature vector.

[0066] In this embodiment, acquiring visual data, language data, historical action data, and environmental state data is the starting point of this operation sequence. Here, visual data refers to multidimensional tensor forms including images, video frames, or raw data from visual sensors, which can originate from image acquisition devices or sensor networks. Language data refers to data containing natural language expressions, which can be in the form of character sequences, word sequences, or encoded language embeddings. Historical action data refers to action trajectory sequences recorded by previous decision-making processes, which can be represented as time-series data and used to describe the dynamic characteristics of past actions. Environmental state data is used to describe static or dynamic state information related to the interaction with the agent in the environment, and is usually represented by environmental variable vectors or state observation tensors. The above four types of input data need to be processed independently to extract discriminative features.

[0067] When processing visual data using hierarchical networks, a multi-layered convolutional neural network (CNN) consisting of multiple convolutional, pooling, and normalization layers can be used to extract local to global spatial features from the visual data layer by layer. Convolutional operations are used to capture local patterns, pooling operations are used to reduce spatial resolution and enhance translation invariance, and batch normalization is used to alleviate gradient vanishing and accelerate the training process. Through the combination of these network units, the output visual features can express the multi-scale spatial distribution and texture details of the input visual data, while improving the discriminative power and robustness of the features.

[0068] When processing language data using pre-trained language models, a deep language encoder based on a Transformer structure can be introduced to transform the language data sequence into a context-sensitive language feature vector representation. The input language data is first processed by word segmentation or encoder embedding mapping, and then converted into an embedding vector sequence. A self-attention mechanism is then used to calculate the global dependencies between elements in the sequence, capturing long-distance semantic associations. The output language features maintain consistency in the context and reflect semantic richness.

[0069] Using temporal convolutional networks (CCNNs) and gated recurrent units (ROUs) to process historical action data is a combined sequence feature extraction process. The CCNN extracts local temporal patterns through a sliding window on the time axis using one-dimensional convolution. The gated recurrent unit further models long-term dependencies and sequence state updates on top of the CCNN's output. The gating structure controls the update ratio between the current input and the historical state, thereby effectively capturing short-term local patterns and long-term dynamic trends in historical action sequences. This ensures that the output action features accurately describe the sequence patterns and contextual dependencies of historical behavior.

[0070] Normalization of environmental status data typically involves statistically analyzing historical observation samples, calculating the mean and standard deviation of each environmental variable, and then using standard score normalization to adjust each variable to a zero-mean and unit-variance distribution. This ensures that environmental characteristics are comparable across different dimensions and reduces the interference of scale inconsistencies in subsequent processing. Normalized environmental characteristics can improve the numerical stability and feature consistency during subsequent multimodal fusion.

[0071] When visual, linguistic, action, and environmental features are input into a multilayer perceptron (MLP) for fusion processing, the MLP can be composed of multiple stacked linear fully connected layers and activation function layers. Each layer's linear transformation adjusts the linear combination relationship of the input features through matrix multiplication and addition operations, while the activation function introduces non-linearity to enhance the network's expressive power. The layer-by-layer computation of the MLP continuously compresses and transforms multimodal features, enabling the final output composite multimodal feature vector to express complex multimodal interaction patterns from vision, language, action, and environment in a unified high-dimensional space. This fusion process not only solves the problem of dimensionality inconsistency between different modalities but also, through network parameter learning, allows the fused features to capture the potential correlations and complementarities between multimodal data.

[0072] In the entire multimodal feature fusion processing flow, each module is sequentially connected and uses standardized data interfaces for efficient data transfer. Each step ensures complete adaptation in terms of dimension and data distribution to the next step. The entire process is executed in a pipelined manner, supporting batch data input and parallel computing, thereby improving overall computational efficiency and throughput performance.

[0073] This embodiment extracts and integrates features from visual, linguistic, historical action, and environmental state data step by step, enabling the establishment of cross-modal associations and information integration among multimodal inputs. This effectively solves the heterogeneity problem of the original input data in terms of dimension, scale, distribution, and context, allowing the comprehensive multimodal feature vector to express the joint patterns and semantic relationships of each modality in a single feature space. This improves the ability to perceive complex environments and multimodal contexts and the accuracy of decision-making in subsequent decision-making processes.

[0074] In one embodiment, step S20 above includes:

[0075] S201, Construct a standard normally distributed Gaussian noise generator;

[0076] S202, Generate a random noise matrix that matches the dimension of the integrated multimodal feature vector;

[0077] S203, determine the feature amplitude distribution of the integrated multimodal feature vector;

[0078] S204, determine the noise intensity adjustment coefficient based on the characteristic amplitude distribution;

[0079] S205, adjust the random noise matrix using the noise intensity adjustment coefficient to generate the adjusted random noise matrix;

[0080] S206, determine the information entropy value of each channel of the integrated multimodal feature vector;

[0081] S207, Generate a channel weight matrix based on the information entropy values ​​of each channel, and multiply the channel weight matrix element-wise with the adjusted random noise matrix, and add the element-wise multiplication result to the comprehensive multimodal feature vector to generate a fusion result;

[0082] S208, Perform a feature space orthogonal transformation on the fusion result to generate an initial noisy feature vector.

[0083] In this embodiment, adding Gaussian distributed noise to the integrated multimodal feature vector to generate an initial noisy feature vector is to introduce a controllable perturbation, allowing the feature vector to enter the diffusion process, thereby enhancing the model's ability to express diversity and generalization in high-dimensional complex feature spaces. First, a standard normal Gaussian noise generator needs to be constructed. This is achieved through a random number generation algorithm, based on a standard normal distribution with a mean of zero and a variance of one, ensuring that the generated noise sequence has the statistical characteristics of an expected value of zero and a standard deviation of one. This generator can be based on a pseudo-random number generator or a hardware-level random source, utilizing efficient sampling algorithms such as the Box-Muller transform and the Ziggurat method to meet the batch sampling requirements of high-dimensional tensor data.

[0084] Generating a random noise matrix that matches the dimensions of the integrated multimodal feature vector requires precise detection of the dimensions of the target feature tensor to ensure that the generated noise matrix is ​​strictly consistent in dimensions, avoiding computational errors caused by dimension mismatch during tensor addition. This operation directly reads the shape attribute of the input feature tensor and uses it as the input parameter of the noise matrix sampling function for dimension adaptation. Furthermore, the memory layout can adopt a data layout format consistent with the original feature tensor (such as row-major or column-major order) to improve access efficiency.

[0085] Determining the amplitude distribution of the integrated multimodal feature vector aims to obtain its amplitude statistical characteristics in the numerical space. The amplitude distribution can be described by statistics such as the global mean, variance, maximum, and minimum values, or it can be expressed by constructing a histogram or fitting a probability density function. The goal is to quantify the overall amplitude range and distribution pattern of the input features. This operation can be achieved by calculating the absolute value of each element and then obtaining the statistics, allowing the amplitude distribution information to reflect the contribution of each dimension in the data to the amplitude, providing a basis for subsequent noise intensity adjustment.

[0086] The noise intensity adjustment coefficient is determined based on the feature amplitude distribution, aiming to match the numerical scale of the added noise with that of the original features. This avoids excessive noise overwhelming low-amplitude features or excessively weak noise perturbing high-amplitude features. The adjustment coefficient can be designed as a global scalar or calculated independently for each dimension. For example, by using the standard deviation as a scaling factor to scale the noise matrix globally or dimensionally, the noise intensity can be made proportional to the local statistical characteristics of the original features. Numerical stability needs to be considered here; a very small positive number can be introduced to prevent numerical anomalies when the standard deviation is zero.

[0087] The process of adjusting the random noise matrix using noise intensity adjustment coefficients is an element-wise multiplication operation, ensuring that the noise value of each element is adjusted proportionally. The resulting adjusted random noise matrix possesses numerical distribution characteristics consistent with the amplitude of the original feature data. This significantly improves the controllability of noise addition and ensures the statistical consistency of subsequent noisy features.

[0088] Determining the information entropy value of each channel in the comprehensive multimodal feature vector is to measure the distribution uncertainty and information content of the data in each channel. The information entropy can be calculated based on the Shannon entropy formula, by discretizing the numerical distribution of each channel into a probability distribution and then summing the entropy values. This information entropy value reflects the importance and diversity of different channels in the data space, providing a data-driven quantitative basis for subsequent channel weighting.

[0089] A channel weight matrix is ​​generated based on the information entropy values ​​of each channel. The entropy values ​​are standardized or normalized and then mapped to weight values. Methods such as softmax normalization or extremum normalization can be used to ensure that each value in the weight matrix varies between 0 and 1, reflecting the relative importance of different channels. Subsequently, the channel weight matrix is ​​multiplied element-wise with the adjusted random noise matrix, adjusting the noise intensity of different channels according to their information entropy. This introduces noise perturbations that match the information complexity in different subspaces of multimodal features. This weighted noise introduction method enhances the model's ability to distinguish key information and avoids excessive contamination of low-information channels by noise.

[0090] The result of multiplying the above elements is added to the comprehensive multimodal feature vector to complete the fusion operation of noise injection and feature data. This results in an output fusion result that not only retains the global expression of the original multimodal features, but also introduces Gaussian noise interference related to the data amplitude and channel information complexity, laying the foundation for subsequent diffusion modeling.

[0091] Performing an orthogonal transformation of the feature space on the fusion result aims to further enhance the decoupling and diversity of feature representation. This orthogonal transformation can be achieved by constructing an orthogonal matrix (such as a transformation matrix generated using Gram-Schmidt orthogonalization, or by employing discrete Fourier transform, discrete cosine transform, etc.) to linearly transform the fusion result, minimizing the correlation between different feature dimensions. This orthogonalization operation not only improves the expressive power of the feature space but also provides initial noisy feature vectors with good numerical stability and statistical independence for the subsequent diffusion process.

[0092] This embodiment introduces Gaussian noise into the comprehensive multimodal feature vector, combined with feature amplitude distribution, channel information entropy weighting, and orthogonal transformation operations. This significantly improves the controllability and diversity of data perturbation, ensures the numerical consistency of different feature dimensions when noise is introduced, and enables the noisy features to not only reflect the statistical characteristics of the multimodal data itself, but also have the characteristics of dimensional decoupling and numerical stability in the subsequent diffusion process. This improves the adaptability and robustness of multimodal inputs for diverse decision generation in complex environments.

[0093] In one embodiment, step S30 above includes:

[0094] S301, Obtain the preset noise scheduling parameter sequence;

[0095] S302, the initial noisy feature vector is taken as the current diffusion state;

[0096] S303, determine the total number of forward time steps for forward diffusion, set the initial value of the time step counter, and generate the initial time step;

[0097] S304, Starting from the initial time step, obtain the noise scheduling parameters corresponding to the current time step, and generate the current noise scheduling parameters;

[0098] S305 generates independent and identically distributed noise that conforms to a Gaussian distribution;

[0099] S306, Based on the current noise scheduling parameters, mix the current diffusion state with the independent and identically distributed noise to generate a mixing result;

[0100] S307, Update the diffusion state to the mixing result and store the updated diffusion state;

[0101] S308 increments the time step counter by one step to generate a new time step;

[0102] S309, determine whether the new time step is less than or equal to the total number of forward time steps;

[0103] S310, when the new time step is less than or equal to the total number of forward time steps, repeat the steps of taking the new time step as the current time step, obtaining the noise scheduling parameters corresponding to the current time step, generating independent and identically distributed noise that conforms to a Gaussian distribution, mixing the current diffusion state with the independent and identically distributed noise based on the noise scheduling parameters, updating the diffusion state to the mixed result and storing the updated diffusion state, and incrementing the time step counter by one step.

[0104] S311, when the new time step is greater than the total number of forward time steps, a noisy state sequence is obtained based on all stored diffusion states.

[0105] In this embodiment, a multi-step forward diffusion is performed on the initial noisy feature vector to obtain a noisy state sequence. The aim is to gradually introduce noise perturbations, causing the feature vector to gradually approach the noise distribution in a high-dimensional space, forming a complete state evolution trajectory. A preset noise scheduling parameter sequence is obtained to provide noise intensity control for each time step during the diffusion process. This sequence can be generated using predefined functions, such as linear decay functions, exponential decay functions, or custom policy functions. Each parameter value corresponds to a diffusion time step, controlling the weight ratio of noise injection at each step. This sequence serves as a set of coefficients for the mixing of noise and state during the diffusion process, ensuring that the noise component gradually increases and maintains a monotonically related relationship with the time steps.

[0106] Using the initial noisy feature vector as the current diffusion state is an assignment operation on the input features during the diffusion initialization phase. This ensures that the diffusion process starts from the current input state and provides the data basis for the state update at each subsequent time step. Determining the total number of forward diffusion time steps is used to limit the step boundary of the diffusion process. This value can be set according to application requirements and is usually related to task complexity, input feature dimension, or target perturbation intensity. Setting the initial value of the time step counter and generating the initial time step are used to control the iterative logic of the entire loop process, ensuring that the diffusion state update and noise scheduling are synchronized.

[0107] Starting from the initial time step, the noise scheduling parameters corresponding to the current time step are obtained, and the current noise scheduling parameters are generated. This is an index access operation on the preset sequence, used to obtain the accurate noise intensity ratio for the current diffusion operation. Generating independent and identically distributed noise that conforms to a Gaussian distribution requires that a new noise matrix be independently sampled at each time step to maintain the independence of the noise sequence and ensure that random perturbations do not carry historical dependencies. The noise matrix is ​​sampled according to a standard normal distribution, and its dimensions must be completely consistent with the current diffusion state.

[0108] The process of generating a mixed result by mixing the current diffusion state with independent and identically distributed noise based on the current noise scheduling parameters involves using the noise scheduling parameters as linear weighting coefficients to weight and superimpose the current diffusion state with the noise at the current time step. This operation can be expressed by the formula: xt = sqrt(1-βt)*xt-1 + sqrt(βt)*εt, where βt is the noise scheduling parameter at the current time step, xt-1 is the diffusion state at the previous step, and εt is the current Gaussian noise. The diffusion state is then updated to the mixed result and stored to ensure that the diffusion trajectory retains complete time-series information in memory. This information can then be used to generate a complete noisy state sequence and as input for inverse denoising.

[0109] Incrementing the time step counter by one step to generate a new time step is an accumulation operation on the counter, typically with a step size of 1, ensuring that the time steps monotonically increase and guaranteeing that the diffusion operation proceeds sequentially. The new time step is checked to see if it is less than or equal to the total number of forward time steps, which determines whether diffusion continues and serves as the termination condition for the iteration. If the condition is met, the noise scheduling, noise sampling, state update, and storage operations starting from the current time step are repeated to ensure consistency in the diffusion logic at each time step. If the condition is not met, i.e., when the number of time steps exceeds the total number of forward time steps, a noisy state sequence is obtained based on all stored diffusion states, forming a complete diffusion process trajectory for subsequent reverse denoising.

[0110] This embodiment performs multi-step forward diffusion on the initial noisy feature vector to generate a complete noisy state sequence. This enables the original multimodal feature vector to gradually approach the noise distribution in a high-dimensional space and records the state change trajectory step by step. This ensures that the generated noisy state sequence not only has statistical diversity and sufficient perturbation, but also provides a continuous, complete, and controllable input data sequence for subsequent reverse denoising. This effectively improves the ability of decision generation to model and generalize data distribution in complex environments.

[0111] In one embodiment, step S40 above includes:

[0112] S401, Obtain the pure noise endpoint state of the noisy state sequence;

[0113] S402, take the pure noise endpoint state as the current denoising state;

[0114] S403, determine the total number of reverse time steps for reverse denoising, and set the initial value of the time step counter to the total number of reverse time steps;

[0115] S404, determines whether the value of the time step counter is greater than zero;

[0116] S405, when the value of the time step counter is greater than zero, repeatedly execute the following steps: using the value of the time step counter as the current time step, obtaining the noise scheduling parameters corresponding to the current time step, inputting the current denoising state into the diffusion network to generate predicted noise, inputting the current denoising state into the policy network to generate an action probability distribution, converting the action probability distribution into a policy guidance signal, adjusting the predicted noise based on the policy guidance signal to generate adjusted noise, updating the current denoising state based on the noise scheduling parameters and the adjusted noise to generate an updated denoising state, using the updated denoising state as the new current denoising state, and decrementing the time step counter by one step.

[0117] S406, when the value of the time step counter is equal to zero, the current denoising state is taken as the final decision state.

[0118] In this embodiment, the processing flow iteratively executes reverse denoising from the purely noisy endpoint state of the noisy state sequence to obtain the final decision state. This is mainly used to gradually restore a highly perturbed endpoint state to an effective state usable for practical decision-making through a combination of stepwise denoising and policy guidance. The purely noisy endpoint state of the noisy state sequence is obtained by retrieving the state corresponding to the last time step in the diffusion state sequence recorded during the forward diffusion process. This state serves as the denoising starting point, as it has the maximum noise coverage, ensuring that the denoising process unfolds from the high-noise space. Using the purely noisy endpoint state as the current denoising state is intended to initialize the data input for the denoising process, ensuring that each iteration uses the updated current denoising state as input, forming a strict state evolution chain.

[0119] The total number of reverse time steps for reverse denoising is determined, and the initial value of the time step counter is set to the total number of reverse time steps. This serves as an upper bound for defining the number of iterations in the denoising process. The total number of time steps is usually consistent with the total number of forward diffusion time steps to ensure time symmetry between denoising and diffusion. The initial value of the time step counter is set to the total number of time steps, and the iteration progresses by gradually decreasing it.

[0120] The termination condition for the denoising loop is whether the time step counter value is greater than zero, ensuring that the denoising operation automatically stops after reaching a predetermined number of iterations, avoiding infinite loops. When the time step counter is greater than zero, the first step is to use the current value of the counter as the current time step, ensuring that the current operation corresponds one-to-one with the noise scheduling parameters in the time dimension. Obtaining the noise scheduling parameters corresponding to the current time step requires searching by index from a preset noise scheduling parameter sequence, which is used to determine the noise cancellation intensity in the current denoising operation.

[0121] The current denoising state is input into a diffusion network to generate predicted noise. The diffusion network simulates the inverse process of forward noise generation, outputting predicted noise components under the current denoising state through network inference, which serves as the basis for subsequent noise cancellation. The current denoising state is then input into a policy network to generate an action probability distribution. The policy network analyzes the current denoising state and outputs the probability distribution of each action in a multi-dimensional action space, reflecting the rationality and weight of each decision direction under the current state. The action probability distribution is converted into a policy guidance signal. This conversion can be achieved through methods such as weighted averaging, maximum value selection, or temperature scaling, outputting a signal used to adjust the predicted noise and enhance the target orientation of the denoising process.

[0122] The adjusted noise is generated by adjusting the predicted noise based on the policy guidance signal. The adjustment process involves weighted fusion of the policy guidance signal and the original predicted noise, ensuring that the adjusted noise reduces noise components while reflecting the guidance information provided by the policy network. The updated denoising state is generated by updating the current denoising state based on the noise scheduling parameters and the adjusted noise. This operation uses the noise scheduling parameters as a scale adjustment factor to cancel out the adjusted noise from the current denoising state, updating the noise level of the current state and advancing the denoising process. The updated denoising state is used as the new current denoising state, ensuring that the next iteration starts from the latest state, guaranteeing the continuity and accumulation of the denoising trajectory. The time step counter is reduced by one step, typically 1 step, to ensure that the time steps are traversed in reverse order, forming a strictly reverse time series.

[0123] When the time step counter value equals zero, the current denoising state is taken as the final decision state, marking the completion of denoising and outputting the final decision result. The final decision state is the state after step-by-step denoising and policy-guided optimization, suitable for subsequent action instruction generation. Throughout the process, each iteration maintains strict time step control, noise scheduling, and policy guidance, ensuring that each step is built upon the current denoising state output by the previous step, forming a strict dynamic closed loop.

[0124] This embodiment achieves dynamic synergy between noise cancellation and policy optimization by iteratively reversing the noise from the end state of pure noise and introducing the target guidance signal of the policy network at each time step. This enables the final decision state to not only have the characteristics of progressive noise reduction, but also to be optimized in each step in combination with the current environment and task objectives. This ensures that the output state enhances the adaptability and pertinence of the decision while reducing the noise level, thereby improving the accuracy, robustness and efficiency of the overall decision generation and task completion.

[0125] In one embodiment, step S50 above includes:

[0126] S501, the final decision state is processed through the fully connected layer of the decision mapping module to generate a decision feature vector;

[0127] S502, the decision feature vector is processed by the activation function module of the decision mapping module to generate activated decision features;

[0128] S503, the activated decision features are processed by the task adaptation analysis module of the decision mapping module to generate task adaptation features;

[0129] S504, The action encoder of the decision mapping module processes the task adaptation features to generate an action encoding vector;

[0130] S505, the action encoding vector is converted into action instructions through the instruction conversion module of the decision mapping module.

[0131] In this embodiment, the process of converting the final decision state into action instructions through the decision mapping module aims to systematically and deeply decode and adapt multi-dimensional and multi-modal decision states, outputting directly executable action instructions. First, the final decision state is used as input data and fed into the fully connected layer of the decision mapping module to complete feature transformation. The fully connected layer is a linear transformation structure composed of a weight matrix and bias terms, which can be used to map the final decision state from the original feature space to the high-dimensional representation space inside the decision mapping module, forming a decision feature vector. The decision feature vector has the ability to compress but retain complete information, capturing the salient features and global distribution information in the original state, laying the foundation for subsequent nonlinear transformations.

[0132] Next, the decision feature vectors are processed by the activation function module of the decision mapping module to generate activated decision features. The activation function module can select nonlinear mapping units such as ReLU, LeakyReLU, Tanh, or Softmax. By utilizing the nonlinear properties of the activation function, the distribution range and expressive power of the data in the feature space are adjusted. The activated decision features can better express the nonlinear correlation structure of multidimensional states, thus improving the input quality of subsequent processing modules.

[0133] Then, the activated decision features enter the task adaptation analysis module. This module is a feature transformation unit based on task type classification, and its main function is to dynamically adjust the dimensions and distribution of features according to task context or environmental context information. The task adaptation analysis module can utilize preset task identifiers or real-time task context markers to adjust the input activated decision features in dimensions such as dimensionality compression, feature reweighting, and feature dimensionality selection, outputting task-adaptive features. These task-adaptive features are highly consistent with the specific task context and serve as a direct input source for action coding.

[0134] Task adaptation features are fed into an action encoder, which further encodes these features into action encoding vectors suitable for subsequent instruction generation. Action encoders typically employ stacked multilayer perceptron structures or feature transformation structures combining orthogonal projection and normalization. This allows them to map multidimensional task adaptation features into semantic representations within the action space, maintaining the continuity and differentiability of the action encoding vectors, making them suitable for further instruction decoding.

[0135] Finally, the action encoding vector enters the instruction conversion module. This module maps the action encoding vector into the final action instruction using a series of predefined action category mapping relationships, threshold mappings, or encoding parsing rules. The action instruction can take the form of multi-dimensional continuous motion control parameters (e.g., robotic arm control angles and speeds) or discrete instruction sequences (e.g., interactive actions or semantic instructions). Throughout the process, the modules of the decision mapping module are strictly sequentially connected, ensuring that the output of each module serves as the input to the next, forming a continuous feature decoding chain while maintaining compatibility in input and output data types and dimensions. Each step transforms and refines the information of the final decision state, ultimately outputting an executable action instruction that meets the task requirements.

[0136] This embodiment sequentially inputs the final decision state into a fully connected layer, an activation function module, a task adaptation analysis module, an action encoder, and an instruction conversion module. This allows for the gradual compression, adaptation, and encoding of complex, multi-dimensional decision states into explicit, executable action instructions, while ensuring the input state information is fully expressed. This enables accurate decoding of decision states in multimodal and multi-scenario environments. The above process improves the accuracy and adaptability of instruction generation, ensuring that the generated action instructions reflect global decision information and can be personalized for different task types, significantly enhancing the accuracy, flexibility, and environmental adaptability of task completion.

[0137] In one embodiment, after step S50 above, the method further includes:

[0138] S601, execute the action command, and monitor the task progress information, action execution cost information, and environmental safety information after the action command is executed;

[0139] S602, construct a reward function based on the task progress information, action execution cost information, and environmental safety information;

[0140] S603, update the parameters of the policy network used in the reverse denoising process using the reward function;

[0141] S604, extract the difference between the predicted noise generated by the diffusion network and the actual noise during the forward diffusion process;

[0142] S605, determine the mean square error loss based on the difference value;

[0143] S606, the gradient of the parameters of the diffusion network is determined by the mean square error loss, and the parameters of the diffusion network are updated by the gradient descent module according to the gradient.

[0144] In this embodiment, after the action instruction is generated, it is first executed. During execution, task progress information is collected, which may include multi-dimensional metrics such as task completion percentage, target achievement status, and progress rate. In parallel, action execution cost information is monitored, which may include data such as energy consumption, execution latency, and computing resource consumption. Simultaneously, environmental safety information is collected, representing the degree of environmental risk generated during execution, potential anomalies, and statistics on safety incidents.

[0145] Based on task progress information, action execution cost information, and environmental safety information, a reward function is constructed according to a preset weighting rule. This reward function uses a weighted summation or other normalized function combination to transform the indicators of the three dimensions—progress, cost, and safety—into numerical reward signals for feedback on decision-making quality. Formally, the reward function can be expressed as a weighted sum of positive progress incentives, negative cost penalties, and safety risk constraints. The weight parameters can be dynamically adjusted to adapt to different task types.

[0146] The reward function is used to update the parameters of the policy network used in the inverse denoising process. The parameter update process of the policy network is implemented through reinforcement learning optimization algorithms, such as policy proximal optimization (PPO) or other policy gradient methods. The current reward function is used as the optimization objective to calculate the gradient of the policy network parameters with respect to the reward, and the parameters are updated to improve the long-term cumulative reward.

[0147] Simultaneously, the difference between the predicted noise generated by the diffusion network and the actual noise during the forward diffusion process is extracted. This difference, calculated step-by-step, is the difference between the predicted and actual values ​​and serves as the input to the loss function. The mean squared error loss is calculated based on this difference, obtained by summing and averaging the squares of all the differences, thus forming a global quantitative index of the prediction noise accuracy.

[0148] The gradient of the diffusion network parameters is calculated using an automatic differentiation mechanism through mean squared error loss. The gradient represents the direction and magnitude of the current parameter adjustment, used to optimize the accuracy of prediction noise. Subsequently, the gradient descent algorithm is applied to update the diffusion network parameters based on this gradient. The update process adjusts the parameter values ​​with a preset step size, allowing the prediction performance of the diffusion network to continuously improve after each training iteration, gradually stabilizing and becoming more accurate. Throughout the process, a closed-loop dependency relationship is maintained between parameter updates, reward function feedback, and difference value calculation. This ensures that the reward signal is used to strengthen the policy network, and the prediction error is used to optimize the diffusion network, enabling the policy and diffusion modules to improve synergistically and enhancing the overall system performance.

[0149] Example Description: In the healthcare field, an intelligent surgical robot performs autonomous tasks in complex minimally invasive surgical scenarios. First, the system acquires visual data from the surgical scene, including real-time endoscopic image sequences, linguistic data from the patient's electronic medical record (e.g., surgical instructions and pathology reports), historical action data (recorded from past similar surgeries), and environmental status data (e.g., temperature, humidity, lighting intensity, and patient physiological parameters such as heart rate and blood pressure). A hierarchical neural network is used to extract spatial and detail-level features from the visual data, generating visual features. The linguistic data is input into a pre-trained medical language model, such as a BERT model fine-tuned based on medical knowledge, to extract linguistic features. Historical action data is processed using a combination of temporal convolutional networks and gated recurrent units to extract action sequence pattern features. Environmental status data is normalized to unify dimensions and adapt to the neural network input format, generating environmental features. These four types of features are then fused into a multilayer perceptron, outputting a comprehensive multimodal feature vector.

[0150] After fusion, Gaussian noise is added to the integrated multimodal feature vector to simulate uncertainty and perform data augmentation. Specifically, a Gaussian noise generator with a standard normal distribution is first constructed to generate a random noise matrix according to the dimensions of the integrated multimodal feature vector. The system determines the amplitude distribution of the integrated multimodal features and determines the noise intensity adjustment coefficient based on the amplitude distribution, so that the noise amplitude is adaptively adjusted to be consistent with the magnitude of the features themselves. Subsequently, the information entropy value of each channel of the integrated multimodal features is determined, and a channel weight matrix is ​​generated based on the information entropy value. The channel weight matrix is ​​multiplied element-wise with the adjusted random noise matrix, so that different modal data have differentiated weights when adding noise, and then added to the original integrated multimodal feature vector to form the fusion result. The fusion result undergoes an orthogonal transformation of the feature space to ensure the orthogonality and decorrelation between features, and finally generates the initial noisy feature vector.

[0151] The system performs a multi-step forward diffusion process on the initial noisy feature vector, gradually adding noise to form a series of noisy state sequences to simulate the uncertainty space of surgical operation decisions. A preset noise scheduling parameter sequence is obtained, the total number of time steps for forward diffusion is defined, and a time step counter is initialized. Starting from the first time step, the noise scheduling parameter corresponding to the current time step is calculated at each step, generating independent and identically distributed Gaussian noise. Based on the current noise scheduling parameter, the current diffusion state is mixed with the independent Gaussian noise of the current step to generate a new diffusion state, which is then stored. This process is repeated time steps until the total number of steps is reached. All diffusion states are stored to form a noisy state sequence.

[0152] After forward diffusion is complete, the system starts from the pure noise endpoint state in the noisy state sequence and iteratively executes the reverse denoising process to gradually restore a feasible surgical operation decision state. The reverse denoising process is performed in reverse time step order. In each step, the noise scheduling parameters for the current time step are obtained, and the current denoised state is input into the diffusion network to generate predicted noise. At the same time, it is input into the policy network to generate an action probability distribution. The action probability distribution is transformed through mapping to generate a policy guidance signal, which is used to adjust the predicted noise. The adjusted noise is combined with the noise scheduling parameters for the current time step to update the current denoised state. After each time step is completed, the time step counter is decremented until it reaches zero, and the current denoised state is taken as the final decision state.

[0153] The final decision state is converted into action commands by the decision mapping module. In this process, the final decision state is first input into the fully connected layer of the decision mapping module for linear mapping, generating a decision feature vector. The decision feature vector is then processed by the activation function module (e.g., ReLU or Softmax) to obtain activated decision features. These activated decision features are then input into the task adaptation analysis module, which adjusts the feature space according to the surgical task type (e.g., suturing, cutting, hemostasis), outputting task adaptation features. The task adaptation features are further input into the action encoder, converting them into action encoding vectors. These action encoding vectors are then parsed into low-level control signals (such as angle and speed commands for each robot joint) by the instruction conversion module, ultimately generating action commands that can directly drive the surgical robot.

[0154] After the action command is executed, the system monitors task progress information (e.g., the progress of the current surgical step), action execution cost information (e.g., energy consumption, time consumption), and environmental safety information (e.g., surgical scene stability, abnormal events) in real time. Based on this information, a reward function is constructed. This reward function comprehensively reflects the task execution effect, efficiency, and safety, and is used to optimize the decision-making process. The reward function is used to update the policy network parameters used in the reverse denoising process. A reinforcement learning optimization algorithm (e.g., PPO) is used to calculate and update the gradient of the reward with respect to the policy network parameters.

[0155] Simultaneously, during the forward diffusion process, the system extracts the difference between predicted noise and actual noise, calculates the mean squared error loss based on the difference, and automatically differentiates the gradient of the diffusion network parameters, then adjusts the diffusion network parameters using the gradient descent algorithm. Throughout the execution and monitoring cycle, the system continuously accumulates historical decision data and environmental feedback data, and periodically retrains the policy network and diffusion network, thereby continuously improving the autonomous decision-making quality and environmental adaptability of the surgical robot, achieving precise, robust, and diversified surgical operation decision support in complex scenarios in the medical and health field.

[0156] In the fintech field, intelligent investment advisory systems provide real-time multi-asset portfolio rebalancing suggestions to high-net-worth clients. First, the system acquires visual data from multiple sources, such as screenshots of investor behavior interfaces and visual charts; linguistic data from financial news and analysis reports; historical action data from client trading records; and environmental status data from current market information (such as volatility, trading volume, and liquidity indicators). Visual data is processed through hierarchical networks to extract key visual patterns, forming visual features. Linguistic data is processed by inputting it into a pre-trained language model in the financial field (e.g., FinBERT) to extract linguistic features containing market sentiment and policy signals. Historical action data undergoes pattern extraction using a temporal convolutional network combined with gated recurrent units to identify the temporal characteristics of users' historical trading behavior. Environmental status data is normalized to eliminate scale differences between different indicators, forming environmental features. The system then inputs these four types of features into a multilayer perceptron for high-dimensional spatial fusion, generating a comprehensive multimodal feature vector, which serves as the input representation for intelligent portfolio rebalancing decisions.

[0157] Subsequently, Gaussian noise is added to the integrated multimodal feature vector to simulate market uncertainty and investor behavior fluctuations. A standard normal Gaussian noise generator is constructed to generate a random noise matrix matching the dimension of the integrated multimodal feature vector. The amplitude distribution of the integrated multimodal features is determined, and a noise intensity adjustment coefficient is calculated to match the noise amplitude with the amplitude of each feature principal component. A channel weight matrix is ​​constructed based on the information entropy value of each feature channel to adapt the noise addition process to the volatility and importance of different dimensions. The channel weight matrix is ​​multiplied element-wise with the adjusted noise matrix and then added to the integrated multimodal feature vector to obtain the fusion result. Orthogonal transformation of the feature space is used to remove the correlation between different features, generating an initial noisy feature vector.

[0158] The initial noisy feature vector enters a multi-step forward diffusion process. The system iterates step by step according to a preset noise scheduling parameter sequence, generating independent and identically distributed Gaussian noise at each step, mixing the current diffusion state with the noise, updating the diffusion state and storing it, and finally forming a noisy state sequence to model the evolution path of portfolio decision state under different market conditions and uncertainty assumptions.

[0159] Starting from the pure noise endpoint of the noisy state sequence, the system iteratively performs reverse denoising. At each reverse time step, the current denoised state is input into the diffusion network to generate predictive noise. Simultaneously, the current denoised state is input into the policy network to obtain the action probability distribution, which is then converted into a policy guidance signal, serving as a market-driven offset. After adjusting the predictive noise, the denoised state is updated in conjunction with noise scheduling parameters. Iteration continues until the time step counter reaches zero; the current denoised state is the final decision state, representing the rebalancing recommendation generated under multi-factor constraints and financial objective guidance.

[0160] The final decision state is converted into specific action instructions through the decision mapping module. First, the final decision state is processed by a fully connected layer to generate a decision feature vector. This feature vector is then non-linearly mapped through an activation function module to generate activated decision features. The task adaptation analysis module analyzes the activated decision features, adjusting the feature space based on factors such as financial asset type, investor risk preference, and product limitations to generate task-adaptive features. The action encoder encodes the task-adaptive features into action encoding vectors, representing intentions such as buying, selling, and adjusting position ratios. The instruction conversion module parses the action encoding vectors into specific position adjustment instruction sets, which can be directly sent to the trading execution platform.

[0161] After the action command is executed, the system monitors task progress information in real time (such as the proportion of rebalancing executed and remaining available funds), action execution cost information (such as transaction fees and slippage losses), and environmental safety information (such as whether the transaction triggers abnormal risk control rules). A reward function is constructed to comprehensively evaluate the effectiveness, safety, and cost efficiency of the rebalancing strategy, which is used to optimize the subsequent intelligent rebalancing decision-making process. The reward function is used to update the strategy network parameters used in the reverse denoising process, and the strategy network is incrementally trained using a reinforcement learning algorithm. The system also extracts the difference between predicted noise and actual noise during the forward diffusion process, calculates the mean squared error loss, calculates the gradient of the diffusion network parameters through automatic differentiation, and applies the gradient descent algorithm for updating. Historical decision data and environmental feedback data are continuously accumulated and used for periodic retraining of the strategy network and diffusion network to ensure that the intelligent rebalancing model maintains good market adaptability and decision diversity in the long term.

[0162] This embodiment utilizes the actual execution of action commands and multi-dimensional information monitoring to construct a reward function based on task progress, action cost, and environmental safety information. The reward function is then used to update the policy network parameters, reinforcing the learning objective-oriented optimization of the reverse denoising process. Simultaneously, the mean squared error loss is calculated using the difference between predicted and actual noise. This mean squared error loss is then used to calculate the gradient of the diffusion network parameters, and the gradient descent algorithm is applied to update the parameters, forming a complete optimization loop from action execution to multi-dimensional feedback and parameter updating. This approach continuously improves the adaptability and prediction accuracy of the policy network and diffusion network, achieving more accurate, flexible, and robust multimodal decision generation, and enhancing task completion quality in complex environments.

[0163] In one embodiment, a decision-making device based on forward diffusion and reverse denoising is provided, which corresponds one-to-one with the decision-making methods based on forward diffusion and reverse denoising described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the decision-making device based on forward diffusion and reverse denoising of the present invention. The modules include a multimodal feature encoding module 10, a noise injection module 20, a forward diffusion module 30, a reverse denoising module 40, and a decision mapping module 50. Detailed descriptions of each functional module are as follows:

[0164] The multimodal feature encoding module 10 is used to acquire visual data, language data, historical action data and environmental state data and generate visual features, language features, action features and environmental features, and fuse the visual features, language features, action features and environmental features to generate a comprehensive multimodal feature vector;

[0165] Noise injection module 20 is used to add Gaussian distributed noise to the integrated multimodal feature vector to generate an initial noisy feature vector;

[0166] Forward diffusion module 30 is used to perform multi-step forward diffusion on the initial noisy feature vector to obtain a noisy state sequence;

[0167] The reverse denoising module 40 is used to iteratively perform reverse denoising starting from the pure noise endpoint state of the noisy state sequence to obtain the final decision state;

[0168] The decision mapping module 50 is used to convert the final decision state into action instructions.

[0169] In one embodiment, the multimodal feature encoding module 10 is specifically used for:

[0170] Acquire visual data, language data, historical action data, and environmental status data;

[0171] The visual data is processed using a hierarchical network to generate visual features;

[0172] The language data is processed using a pre-trained language model to generate language features;

[0173] The historical action data is processed using a temporal convolutional network and a gated recurrent unit to generate action features;

[0174] The environmental state data is normalized to generate environmental features;

[0175] The visual features, language features, action features, and environmental features are input into a multilayer perceptron to generate a comprehensive multimodal feature vector.

[0176] In one embodiment, the noise injection module 20 is specifically used for:

[0177] Construct a standard normally distributed Gaussian noise generator;

[0178] Generate a random noise matrix that matches the dimension of the integrated multimodal feature vector;

[0179] Determine the feature magnitude distribution of the integrated multimodal feature vector;

[0180] The noise intensity adjustment coefficient is determined based on the characteristic amplitude distribution;

[0181] The random noise matrix is ​​adjusted by the noise intensity adjustment coefficient to generate an adjusted random noise matrix;

[0182] Determine the entropy value of each channel of the integrated multimodal feature vector;

[0183] Based on the entropy values ​​of each channel, a channel weight matrix is ​​generated, and the channel weight matrix is ​​multiplied element-wise with the adjusted random noise matrix. The result of the element-wise multiplication is added to the comprehensive multimodal feature vector to generate a fusion result.

[0184] Perform a feature space orthogonal transformation on the fusion result to generate an initial noisy feature vector.

[0185] In one embodiment, the forward diffusion module 30 is specifically used for:

[0186] Obtain the preset noise scheduling parameter sequence;

[0187] The initial noisy feature vector is taken as the current diffusion state;

[0188] Determine the total number of forward time steps for forward diffusion, set the initial value of the time step counter, and generate the initial time step;

[0189] Starting from the initial time step, obtain the noise scheduling parameters corresponding to the current time step, and generate the current noise scheduling parameters;

[0190] Generate independent and identically distributed noise that conforms to a Gaussian distribution;

[0191] Based on the current noise scheduling parameters, the current diffusion state and the independent and identically distributed noise are mixed to generate a mixing result;

[0192] Update the diffusion state to the mixing result and store the updated diffusion state;

[0193] Increment the time step counter by one step to generate a new time step;

[0194] Determine whether the new time step is less than or equal to the total number of forward time steps;

[0195] When the new time step is less than or equal to the total number of forward time steps, repeat the following steps: take the new time step as the current time step, obtain the noise scheduling parameters corresponding to the current time step, generate independent and identically distributed noise that conforms to a Gaussian distribution, mix the current diffusion state with the independent and identically distributed noise based on the noise scheduling parameters, update the diffusion state to the mixed result and store the updated diffusion state, and increment the time step counter by one step.

[0196] When the new time step is greater than the total number of forward time steps, a noisy state sequence is obtained based on all stored diffusion states.

[0197] In one embodiment, the reverse noise reduction module 40 is specifically used for:

[0198] Obtain the pure noise endpoint state of the noisy state sequence;

[0199] The pure noise endpoint state is taken as the current denoising state;

[0200] Determine the total number of reverse time steps for reverse denoising, and set the initial value of the time step counter to the total number of reverse time steps;

[0201] Check if the value of the time step counter is greater than zero;

[0202] When the value of the time step counter is greater than zero, the following steps are repeated: using the value of the time step counter as the current time step, obtaining the noise scheduling parameters corresponding to the current time step, inputting the current denoising state into the diffusion network to generate predicted noise, inputting the current denoising state into the policy network to generate an action probability distribution, converting the action probability distribution into a policy guidance signal, adjusting the predicted noise based on the policy guidance signal to generate adjusted noise, updating the current denoising state based on the noise scheduling parameters and the adjusted noise to generate an updated denoising state, using the updated denoising state as the new current denoising state, and decrementing the time step counter by one step.

[0203] When the time step counter value is equal to zero, the current denoising state is taken as the final decision state.

[0204] In one embodiment, the decision mapping module 50 is specifically used for:

[0205] The final decision state is processed by the fully connected layer of the decision mapping module to generate a decision feature vector;

[0206] The decision feature vector is processed by the activation function module of the decision mapping module to generate activated decision features;

[0207] The activated decision features are processed by the task adaptation analysis module of the decision mapping module to generate task adaptation features.

[0208] The action encoder of the decision mapping module processes the task adaptation features to generate an action encoding vector.

[0209] The action encoding vector is converted into action instructions by the instruction conversion module of the decision mapping module.

[0210] In one embodiment, the decision mapping module 50 is specifically used for:

[0211] Execute the action command and monitor the task progress information, action execution cost information and environmental safety information after the action command is executed;

[0212] A reward function is constructed based on the task progress information, action execution cost information, and environmental safety information.

[0213] The parameters of the policy network used in the reverse denoising process are updated using the reward function;

[0214] Extract the difference between the predicted noise generated by the diffusion network and the actual noise during the forward diffusion process;

[0215] The mean squared error loss is determined based on the difference value;

[0216] The gradient of the parameters of the diffusion network is determined by the mean squared error loss, and the parameters of the diffusion network are updated by the gradient descent module based on the gradient.

[0217] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements server-side functions or steps based on a decision-making method based on forward diffusion and reverse denoising.

[0218] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides decision-making and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements user-side functions or steps based on a decision-making method based on forward diffusion and reverse denoising.

[0219] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0220] Acquire visual data, language data, historical action data, and environmental state data, and generate visual features, language features, action features, and environmental features. Then, fuse the visual features, language features, action features, and environmental features to generate a comprehensive multimodal feature vector.

[0221] Gaussian noise is added to the integrated multimodal feature vector to generate an initial noisy feature vector;

[0222] Perform multi-step forward diffusion on the initial noisy feature vector to obtain a noisy state sequence;

[0223] Starting from the pure noise endpoint state of the noisy state sequence, iterative reverse denoising is performed to obtain the final decision state;

[0224] The final decision state is converted into action instructions through the decision mapping module.

[0225] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0226] Acquire visual data, language data, historical action data, and environmental state data, and generate visual features, language features, action features, and environmental features. Then, fuse the visual features, language features, action features, and environmental features to generate a comprehensive multimodal feature vector.

[0227] Gaussian noise is added to the integrated multimodal feature vector to generate an initial noisy feature vector;

[0228] Perform multi-step forward diffusion on the initial noisy feature vector to obtain a noisy state sequence;

[0229] Starting from the pure noise endpoint state of the noisy state sequence, iterative reverse denoising is performed to obtain the final decision state;

[0230] The final decision state is converted into action instructions through the decision mapping module.

[0231] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0232] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0233] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0234] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A decision-making method based on forward diffusion and reverse denoising, characterized in that, Includes the following steps: Acquire visual data, language data, historical action data, and environmental state data, and generate visual features, language features, action features, and environmental features. Then, fuse the visual features, language features, action features, and environmental features to generate a comprehensive multimodal feature vector. Gaussian noise is added to the integrated multimodal feature vector to generate an initial noisy feature vector; Perform multi-step forward diffusion on the initial noisy feature vector to obtain a noisy state sequence; Starting from the pure noise endpoint state of the noisy state sequence, iterative reverse denoising is performed to obtain the final decision state; The final decision state is converted into action instructions through the decision mapping module.

2. The decision-making method based on forward diffusion and reverse denoising as described in claim 1, characterized in that, Acquire visual data, language data, historical action data, and environmental state data, and generate visual features, language features, action features, and environmental features. Fuse these visual features, language features, action features, and environmental features to generate a comprehensive multimodal feature vector, including: Acquire visual data, language data, historical action data, and environmental status data; The visual data is processed using a hierarchical network to generate visual features; The language data is processed using a pre-trained language model to generate language features; The historical action data is processed using a temporal convolutional network and a gated recurrent unit to generate action features; The environmental state data is normalized to generate environmental features; The visual features, language features, action features, and environmental features are input into a multilayer perceptron to generate a comprehensive multimodal feature vector.

3. The decision-making method based on forward diffusion and reverse denoising as described in claim 1, characterized in that, Adding Gaussian distributed noise to the comprehensive multimodal feature vector to generate an initial noisy feature vector includes: Construct a standard normally distributed Gaussian noise generator; Generate a random noise matrix that matches the dimension of the integrated multimodal feature vector; Determine the feature magnitude distribution of the integrated multimodal feature vector; The noise intensity adjustment coefficient is determined based on the characteristic amplitude distribution; The random noise matrix is ​​adjusted by the noise intensity adjustment coefficient to generate an adjusted random noise matrix; Determine the entropy value of each channel of the integrated multimodal feature vector; Based on the entropy values ​​of each channel, a channel weight matrix is ​​generated, and the channel weight matrix is ​​multiplied element-wise with the adjusted random noise matrix. The result of the element-wise multiplication is added to the comprehensive multimodal feature vector to generate a fusion result. Perform a feature space orthogonal transformation on the fusion result to generate an initial noisy feature vector.

4. The decision-making method based on forward diffusion and reverse denoising as described in claim 1, characterized in that, Perform multi-step forward diffusion on the initial noisy feature vector to obtain a noisy state sequence, including: Obtain the preset noise scheduling parameter sequence; The initial noisy feature vector is taken as the current diffusion state; Determine the total number of forward time steps for forward diffusion, set the initial value of the time step counter, and generate the initial time step; Starting from the initial time step, obtain the noise scheduling parameters corresponding to the current time step, and generate the current noise scheduling parameters; Generate independent and identically distributed noise that conforms to a Gaussian distribution; Based on the current noise scheduling parameters, the current diffusion state and the independent and identically distributed noise are mixed to generate a mixing result; Update the diffusion state to the mixing result and store the updated diffusion state; Increment the time step counter by one step to generate a new time step; Determine whether the new time step is less than or equal to the total number of forward time steps; When the new time step is less than or equal to the total number of forward time steps, repeat the following steps: take the new time step as the current time step, obtain the noise scheduling parameters corresponding to the current time step, generate independent and identically distributed noise that conforms to a Gaussian distribution, mix the current diffusion state with the independent and identically distributed noise based on the noise scheduling parameters, update the diffusion state to the mixed result and store the updated diffusion state, and increment the time step counter by one step. When the new time step is greater than the total number of forward time steps, a noisy state sequence is obtained based on all stored diffusion states.

5. The decision-making method based on forward diffusion and reverse denoising as described in claim 1, characterized in that, Starting from the purely noisy endpoint state of the noisy state sequence, iteratively perform reverse denoising to obtain the final decision state, including: Obtain the pure noise endpoint state of the noisy state sequence; The pure noise endpoint state is taken as the current denoising state; Determine the total number of reverse time steps for reverse denoising, and set the initial value of the time step counter to the total number of reverse time steps; Check if the value of the time step counter is greater than zero; When the value of the time step counter is greater than zero, the following steps are repeated: using the value of the time step counter as the current time step, obtaining the noise scheduling parameters corresponding to the current time step, inputting the current denoising state into the diffusion network to generate predicted noise, inputting the current denoising state into the policy network to generate an action probability distribution, converting the action probability distribution into a policy guidance signal, adjusting the predicted noise based on the policy guidance signal to generate adjusted noise, updating the current denoising state based on the noise scheduling parameters and the adjusted noise to generate an updated denoising state, using the updated denoising state as the new current denoising state, and decrementing the time step counter by one step. When the time step counter value is equal to zero, the current denoising state is taken as the final decision state.

6. The decision-making method based on forward diffusion and reverse denoising as described in claim 1, characterized in that, The final decision state is converted into action instructions through the decision mapping module, including: The final decision state is processed by the fully connected layer of the decision mapping module to generate a decision feature vector; The decision feature vector is processed by the activation function module of the decision mapping module to generate activated decision features; The activated decision features are processed by the task adaptation analysis module of the decision mapping module to generate task adaptation features. The action encoder of the decision mapping module processes the task adaptation features to generate an action encoding vector. The action encoding vector is converted into action instructions by the instruction conversion module of the decision mapping module.

7. The decision-making method based on forward diffusion and reverse denoising as described in claim 1, characterized in that, After the final decision state is converted into action instructions through the decision mapping module, the following is also included: Execute the action command and monitor the task progress information, action execution cost information and environmental safety information after the action command is executed; A reward function is constructed based on the task progress information, action execution cost information, and environmental safety information. The parameters of the policy network used in the reverse denoising process are updated using the reward function; Extract the difference between the predicted noise generated by the diffusion network and the actual noise during the forward diffusion process; The mean squared error loss is determined based on the difference value; The gradient of the parameters of the diffusion network is determined by the mean squared error loss, and the parameters of the diffusion network are updated by the gradient descent module based on the gradient.

8. A decision-making device based on forward diffusion and reverse denoising, characterized in that, The decision-making device based on forward diffusion and reverse denoising includes: The multimodal feature encoding module is used to acquire visual data, language data, historical action data, and environmental state data, and generate visual features, language features, action features, and environmental features, and fuse the visual features, language features, action features, and environmental features to generate a comprehensive multimodal feature vector; The noise injection module is used to add Gaussian distributed noise to the comprehensive multimodal feature vector to generate an initial noisy feature vector; The forward diffusion module is used to perform multi-step forward diffusion on the initial noisy feature vector to obtain a noisy state sequence; The reverse denoising module is used to iteratively perform reverse denoising starting from the pure noise endpoint state of the noisy state sequence to obtain the final decision state; The decision mapping module is used to convert the final decision state into action instructions.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a decision-making program based on forward diffusion and reverse denoising stored in the memory and executable on the processor. When executed by the processor, the decision-making program based on forward diffusion and reverse denoising implements the steps of the decision-making method based on forward diffusion and reverse denoising as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a decision program based on forward diffusion and reverse denoising, which, when executed by a processor, implements the steps of the decision method based on forward diffusion and reverse denoising as described in any one of claims 1-7.