Action representation enhancement based decision method and device, equipment and medium

By using multimodal feature fusion and dynamic clustering, the problem of being unable to identify key action features in existing technologies has been solved, enabling accurate decision-making and efficient execution in complex task environments.

CN120892776BActive Publication Date: 2026-01-16PING AN TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511403588.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-01-16
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Existing technologies cannot dynamically identify and focus on key action features, leading to the neglect of important action information in complex task environments, which affects the accuracy and efficiency of decision-making.

Method used

By acquiring visual data and language instructions, visual and language features are generated and fused with a pre-defined set of discrete action units to form an initial action representation set. Clustering is performed based on a correlation threshold to determine an importance metric, and feature enhancement is applied to the target action clusters to finally generate an optimized action representation set for input into the decision unit.

Benefits of technology

It increases the weight of key action information in the decision-making process, avoids equal processing of all actions, ensures more accurate decision-making in complex task environments, and improves the accuracy and efficiency of task completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892776B_ABST
    Figure CN120892776B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, can be applied to business scenarios such as autonomous decision of an intelligent agent, financial technology and medical health, and discloses a decision method and device based on action representation enhancement, equipment and a medium, which comprises the following steps: acquiring visual data, language instructions and a preset discrete action unit set, generating visual features and language features, fusing the initial action representation, forming an initial representation set. According to correlation threshold analysis, high-correlation-degree representations are aggregated into different action clusters, and a target action cluster is enhanced through importance measurement to generate an optimized action representation set, and finally, the decision unit is input to generate a final action decision. The application optimizes the action representation through dynamic focusing and importance measurement, improves the weight of key action information, ensures accurate decision in a complex task environment, and improves the accuracy and efficiency of task completion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a decision-making method and device based on action representation enhancement, equipment and a storage medium. BACKGROUND

[0002] In the field of financial technology business, decision support systems often need to handle complex multi-modal information, such as customer behavior data, market dynamics, financial product characteristics, etc. In traditional financial risk assessment systems, the representation of action decisions often relies on pre-set rules or fixed binning methods, which give equal weight to all possible behavior actions. However, different operations and decisions in the financial market have unequal impacts on the final results. For example, in investment decisions, specific market fluctuations or action information related to specific trading strategies play a crucial role in risk assessment results, while traditional methods cannot automatically identify and focus on these key action features. Therefore, existing systems often overlook some important factors, resulting in insufficient accuracy and efficiency of the assessment results, especially in complex market environments, which cannot effectively focus on the key feature information of decision-making.

[0003] In the field of medical health business, the development of patient treatment plans involves a large amount of multi-modal data, such as medical history, imaging data, laboratory test results, etc. Traditional medical decision support systems usually treat all possible clinical decisions as equivalent, relying on fixed binning methods to handle different treatment actions or interventions. However, in medical tasks, certain specific treatment actions, such as precise drug dosages for a particular disease or the timing of specific operations, are often critical to the success of treatment, while traditional methods cannot dynamically adjust and focus on these key action features. Inaccuracies in medical decision-making can lead to delayed diagnosis or treatment errors, which in turn affect the health of patients. Therefore, existing medical decision systems have significant shortcomings in handling complex and refined clinical tasks.

[0004] In the field of fine operation tasks of robots, traditional VLA models often use fixed and uniform discrete action spaces when processing action representation. This approach treats all action bins as equally important. However, in the process of robot operation, certain action bins (such as specific angle joint rotation, precise grip strength, etc.) are crucial to the success of the task. These key action features are often overlooked, resulting in the model failing to effectively focus on key features in actual operation, affecting the accurate execution and efficiency of the task. Existing VLA models cannot dynamically identify and focus on these key action information, resulting in less-than-expected task execution in complex environments. SUMMARY

[0005] The main purpose of the present application is to provide a decision-making method, device, equipment and storage medium based on action representation enhancement, aiming to solve the technical problem that the prior art cannot dynamically identify and focus on key action features, resulting in ignoring important action information in a complex task environment, thereby affecting the accuracy and efficiency of decision-making.

[0006] To achieve the above-mentioned purpose, the present application provides a decision-making method based on action representation enhancement, comprising:

[0007] acquiring visual data, language instructions and a preset discrete action unit set, and generating visual features and language features based on the visual data and the language instructions, respectively;

[0008] performing multi-modal feature fusion of the visual features, the language features and the initial feature vectors of each discrete action unit in the discrete action unit set, respectively, to generate initial action representations corresponding to each discrete action unit, forming an initial action representation set;

[0009] According to a preset correlation threshold, correlation analysis is performed between each initial action representation in the initial action representation set, and initial action representations with high correlation are aggregated into different action clusters;

[0010] For the different action clusters, the importance measure of each action cluster is determined;

[0011] According to the importance measure, the initial action representations in the target action cluster are processed for feature enhancement to generate an optimized action representation set;

[0012] The optimized action representation set is input into a decision unit to generate a final action decision.

[0013] Further, to achieve the above-mentioned purpose, the present application provides a decision-making device based on action representation enhancement, comprising:

[0014] A data acquisition and feature generation module is configured to acquire visual data, language instructions and a preset discrete action unit set, and generate visual features and language features based on the visual data and the language instructions, respectively;

[0015] A multi-modal feature fusion module is configured to perform multi-modal feature fusion of the visual features, the language features and the initial feature vectors of each discrete action unit in the discrete action unit set, respectively, to generate initial action representations corresponding to each discrete action unit, forming an initial action representation set;

[0016] The correlation degree analysis and clustering module is configured to perform correlation degree analysis on each initial action representation in the initial action representation set according to a preset correlation degree threshold, and aggregate initial action representations with high correlation degrees into different action clusters.

[0017] The importance metric evaluation module is configured to determine an importance metric of each action cluster.

[0018] The feature enhancement and optimization module is configured to perform feature enhancement processing on initial action representations in a target action cluster according to the importance metric, and generate an optimized action representation set.

[0019] The decision generation module is configured to input the optimized action representation set into a decision unit, and generate a final action decision.

[0020] Further, to achieve the above object, the present application also provides a computer device, which comprises a memory, a processor and a decision-making program based on action representation enhancement stored in the memory and executable on the processor, and the decision-making program based on action representation enhancement implements the steps of the decision-making method based on action representation enhancement when executed by the processor.

[0021] Further, to achieve the above object, the present application also provides a computer readable storage medium, which stores a decision-making program based on action representation enhancement, and the decision-making program based on action representation enhancement implements the steps of the decision-making method based on action representation enhancement when executed by a processor.

[0022] Beneficial effects: The present application relates to the field of artificial intelligence technology, and can be applied to business scenarios such as autonomous decision-making of intelligent agents, financial technology and medical health. The present application discloses a decision-making method, device, equipment and medium based on action representation enhancement, which comprises the following steps: obtaining visual data, language instructions and a preset discrete action unit set, generating visual features and language features, and performing multi-modal feature fusion with initial feature vectors of discrete action units to form an initial action representation set; performing correlation degree analysis on the initial action representation set according to a correlation degree threshold, aggregating initial action representations with high correlation degrees into different action clusters, determining an importance metric for each action cluster, further performing feature enhancement on initial action representations in a target action cluster to generate an optimized action representation set, and finally inputting the optimized action representation set into a decision unit to generate a final action decision. The present application can effectively improve the weight of key action information in the decision-making process by dynamically focusing on the action space and optimizing action representations based on the importance metric, avoid the equal processing of all action bins in the traditional method, ensure more accurate execution of decisions in a complex task environment, and finally improve the accuracy and efficiency of task completion. BRIEF DESCRIPTION OF DRAWINGS

[0023] The application will be further described below in conjunction with the accompanying drawings and embodiments. In the drawings:

[0024] Figure 1 An application environment diagram of the decision-making method based on action representation enhancement in an embodiment of the application;

[0025] Figure 2 A flow diagram of the decision-making method based on action representation enhancement in an embodiment of the application;

[0026] Figure 3 A functional module diagram of a preferred embodiment of the decision-making device based on action representation enhancement of the application;

[0027] Figure 4 A structure diagram of a computer device in an embodiment of the application;

[0028] Figure 5 Another structure diagram of a computer device in an embodiment of the application. DETAILED DESCRIPTION

[0029] It should be understood that the specific embodiments described herein are merely intended to explain the application and are not intended to limit the application.

[0030] The decision-making method based on action representation enhancement provided by the embodiments of the application can be applied in an application environment as shown in Figure 1 , wherein a user end communicates with a service end through a network. The service end can obtain visual data, language instructions and a preset discrete action unit set through the user end, generate visual features and language features, and perform multi-modal feature fusion with initial feature vectors of the discrete action units to form an initial action representation set. The initial action representation set is analyzed according to a correlation threshold, and initial action representations with high correlation are aggregated into different action clusters. The importance measure of each action cluster is determined, and the initial action representations in the target action cluster are further enhanced in features to generate an optimized action representation set. Finally, the optimized action representation set is input into a decision unit to generate a final action decision. The application dynamically focuses on the action space, optimizes the action representation based on the importance measure, so as to effectively improve the weight of key action information in the decision-making process, avoid the equal treatment of all action bins in the traditional method, ensure more accurate execution of the decision in a complex task environment, and finally improve the accuracy and efficiency of task completion. The user end can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The service end can be implemented by an independent server or a server cluster composed of multiple servers. The application will be described in detail below through specific embodiments.

[0031] Please refer to Figure 2 , Figure 2 The flowchart of an embodiment of the decision-making method based on action representation enhancement provided by the present application is shown. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in an order different from that shown here.

[0032] As Figure 2 shown, the decision-making method based on action representation enhancement provided by the present application includes the following steps:

[0033] S10, acquiring visual data, language instructions and a preset discrete action unit set, and respectively generating visual features and language features based on the visual data and the language instructions;

[0034] In this embodiment, visual data is acquired, which is image information derived from the execution object interacting with the environment, usually collected through visual sensors such as color cameras, depth cameras or infrared cameras, and the image data needs to be represented in a pixel matrix. In actual implementation, the resolution, frame rate or field of view of the camera can be adjusted during collection to balance the requirements for data detail and real-time performance. After the visual data is collected, it is standardized through a preprocessing module to remove the effects of brightness and color differences. Common processing methods include normalization, color space conversion, edge enhancement, etc. Language instructions are natural language inputs from the operator, derived from voice collection devices or text interfaces, such as voice collected by a microphone converted into text through an automatic speech recognition module, or directly inputted instruction text. Language instructions need to be cleaned, including removing meaningless symbols, normalizing case, simplifying numerical expressions, etc. The pre-set discrete action unit set is a pre-defined division set of the continuous action space, usually discretizing the action parameters in the action space by a certain granularity, such as discretizing the rotation angle of the robot arm joint in degrees, or dividing the movement speed into intervals to form a set. For different devices, the construction standard of the pre-set discrete action unit set is different, and needs to be combined with the execution accuracy and action limit range of the physical hardware. After the visual data and language instructions are acquired, they are used to generate visual features and language features respectively. The generation of visual features extracts the spatial and semantic information of the image through a deep convolutional network or other image feature extraction network, such as ResNet, MobileNet or ViT model, etc. The extracted visual features are represented as multi-dimensional vectors, containing spatial edges, textures, contours, target recognition results, etc. In practice, the network structure depth, convolution kernel size, etc. can be adjusted to adapt to different scenarios. The generation of language features processes the language instructions through a natural language processing module, including word segmentation, embedding encoding, context modeling, etc. For example, using BERT or Transformer model to extract multi-dimensional context semantic representation from instruction text, language features are represented as high-dimensional dense vectors, containing multi-level information such as words, phrases, semantic associations, command intent, etc. After the generation of visual features and language features, they need to be fused with the initial feature vector of each action unit in the pre-set discrete action unit set in the subsequent steps. The initial feature vector is a structured encoding representation of each action unit, usually in the form of a fixed-length vector, encoding the parameter information of the action unit such as direction, amplitude, rate, etc. In this way, the perception information and action information are connected, providing basic data support for subsequent action representation generation, analysis and optimization. In the entire process, the input order and synchronization of visual data, language instructions and discrete action unit set need to be strictly managed to ensure consistent timestamps and context associations between data, avoiding inconsistencies caused by time alignment errors or data delays.

[0035] The visual data input in a complex environment can be realized by multi-camera synchronous acquisition, and the scene coverage and perspective diversity are enhanced. Low-power camera networking mode can also be used to reduce energy consumption and support long-time acquisition. For language instruction input, a combination of local speech recognition model and remote cloud service can be used to improve recognition accuracy and processing speed. The construction of the discrete action unit set can be dynamically adjusted according to the action execution range and accuracy requirements of different hardware platforms to adapt to the action granularity requirements in different application scenarios. The visual feature generation module can use a lightweight convolutional neural network to adapt to devices with limited computing resources, or use a multi-scale convolutional module to enhance the feature's ability to perceive different target sizes. The language feature generation module can introduce a domain-specific vocabulary or a self-defined embedding model to improve the understanding of instructions in specific fields, especially in the medical health or financial technology scenarios, to adapt to the diversity and complexity of professional terms.

[0036] By obtaining visual data, language instructions and a preset discrete action unit set, and combining the generation and synchronous input of visual features and language features, the embodiment can ensure the integrity and semantic consistency of data in the multi-source heterogeneous data input stage, provide higher quality basic data for subsequent action representation enhancement, avoid information fragmentation and semantic mismatch problems, and thus improve the action decision accuracy and execution efficiency of the model in complex task scenarios.

[0037] S20, respectively, the visual features, the language features and the initial feature vectors of each discrete action unit in the discrete action unit set are subjected to multi-modal feature fusion to generate an initial action representation corresponding to each discrete action unit, thereby forming an initial action representation set;

[0038] In this embodiment, after obtaining the visual features, which are spatial and semantic representations extracted from the original visual data, usually generated by models such as deep convolutional neural networks, including quantitative descriptions of object, scene, location, etc. Language features come from the analysis of language instructions, and the context vector sequence formed by word segmentation, part-of-speech tagging, and semantic association capture is used to depict task requirements or user intent. On this basis, the visual features and language features are respectively fused with the initial feature vectors of each discrete action unit in the discrete action unit set. Each unit in the discrete action unit set represents a discrete representation of an executable action, such as each predefined angle or force combination in robot operation. The initial feature vector is used to represent the original attributes or default behavior tendency of each action unit. The specific fusion operation is realized by splicing or tensor combination, which combines the visual features, language features, and the initial feature vector of each action unit into a high-dimensional feature vector. Further, through the fully connected layer of the multilayer perceptron, linear transformation and nonlinear mapping are performed to improve the expression ability of the feature space, and the fused vector is obtained as the initial action representation corresponding to the current action unit. After performing this operation on all discrete action units, these initial action representations are combined into a whole, called the initial action representation set. The whole process ensures that visual information, language instruction information, and action unit original attributes are simultaneously encoded into a unified high-dimensional space for subsequent action relevance analysis and importance measurement.

[0039] Different convolutional network architectures can be used in the acquisition of visual features to adapt to the visual perception needs of high-resolution visual input scenes or low-light environments. Different pre-trained language models can also be selected according to the diversity of language instructions for more complex multilingual or dialect analysis. The implementation of multi-modal feature fusion can be adapted to different business scenarios through different fusion strategies, such as using element-wise weighted summation to reduce feature dimensions, or using cross-modal attention mechanisms to enhance multi-modal information interaction. The configuration of the fully connected layer can adjust the hidden layer dimension and activation function type to adapt to the computing resource and real-time requirement. The result of feature fusion can be stored in the form of a tensor and used for subsequent correlation analysis and action aggregation.

[0040] In the field of autonomous decision-making of agents, agents collect scene visual images and policy task instructions, fuse these inputs with the feature vectors of each predefined motion or interaction action unit, and form a dynamic action representation set, so that the agent can prepare a candidate action representation set that better supports context awareness based on the current environment and task context, improving the depth and flexibility of action understanding before task execution.

[0041] The embodiment can improve the expression richness and context sensitivity of the action unit representation by simultaneously fusing the initial feature vectors of visual features, language features and discrete action units, solving the problem that the action unit representation in the prior art cannot reflect the environment and task semantics based on the uniform distribution of the action space, laying a foundation for subsequent efficient action aggregation and decision making, and improving the pertinence of action selection and the accuracy of decision making.

[0042] In the embodiment, based on the preset correlation threshold, the preset correlation threshold refers to a numerical parameter set in advance, which is used to determine whether any two initial action representations satisfy the sufficient correlation requirement. The initial action representation set is generated by the previous step and contains multiple vectors of fused visual, language and action unit features. The correlation analysis between each initial action representation in the set is performed. First, the similarity or correlation between any two initial action representations in the set is calculated by pairwise combination. The correlation is measured by self-attention mechanism, and the dot product similarity is usually used as the basis for calculation. In specific implementation, a query vector and a key vector are constructed for each initial action representation. The dot product of the two vectors is calculated and normalized to obtain a standardized correlation score. Then, the correlation score is compared with the preset correlation threshold to determine whether the two initial action representations are correlated. If the correlation is higher than the threshold, a connection is established between the two. All connections that meet the conditions form an implicit correlation relationship network or graph structure. On the correlation graph, the initial action representations associated with each other are extracted as an action cluster by a connected component recognition algorithm. After performing this process on all initial action representations in the set, multiple independent action clusters are formed, and each cluster represents a group of highly similar or semantically related action representations in the multi-modal context.

[0043] In the embodiment, based on the preset correlation threshold, the preset correlation threshold refers to a numerical parameter set in advance, which is used to determine whether any two initial action representations satisfy the sufficient correlation requirement. The initial action representation set is generated by the previous step and contains multiple vectors of fused visual, language and action unit features. The correlation analysis between each initial action representation in the set is performed. First, the similarity or correlation between any two initial action representations in the set is calculated by pairwise combination. The correlation is measured by self-attention mechanism, and the dot product similarity is usually used as the basis for calculation. In specific implementation, a query vector and a key vector are constructed for each initial action representation. The dot product of the two vectors is calculated and normalized to obtain a standardized correlation score. Then, the correlation score is compared with the preset correlation threshold to determine whether the two initial action representations are correlated. If the correlation is higher than the threshold, a connection is established between the two. All connections that meet the conditions form an implicit correlation relationship network or graph structure. On the correlation graph, the initial action representations associated with each other are extracted as an action cluster by a connected component recognition algorithm. After performing this process on all initial action representations in the set, multiple independent action clusters are formed, and each cluster represents a group of highly similar or semantically related action representations in the multi-modal context.

[0044] The correlation threshold can be adjusted in different application environments to adapt to different decision requirements, such as using a higher threshold to filter low-correlation action representations in a scene with high visual noise, or using an adaptive threshold adjustment method to dynamically determine the correlation determination standard in a multi-task mixed environment. Different measurement standards can be used to optimize the calculation efficiency and distinguish degree, such as Euclidean distance, cosine similarity or learned similarity function. The construction of the correlation relationship graph can support sparse matrix optimization storage and parallel computing to improve processing performance, and the extraction of connected components can use a depth-first search algorithm or a graph convolution network.

[0045] In the field of autonomous decision-making of intelligent agents, it can be used for robots or virtual agents to identify highly relevant candidate actions in the current task phase in the task environment, for example, representing and aggregating the most relevant navigation, control or interaction actions to the current task goal into an action cluster, so that the intelligent agent can focus on a set of actions most closely related to the current scene, thereby improving the efficiency and adaptability of autonomous decision-making.

[0046] By performing correlation degree analysis on the initial action representation set based on the preset correlation degree threshold, the embodiment dynamically aggregates the initial action representations with high relevance into action clusters, which can avoid the uniform division and static processing limitations of the action space, realize automatic grouping of key action information, ensure that subsequent action evaluation can be performed on a structured relevant action set, improve the action space organization efficiency and decision flexibility, and solve the problem of dynamic organization of action information in a multi-modal context in the prior art.

[0047] S40, determining an importance measure of each action cluster for the different action clusters;

[0048] In the embodiment, determining the importance measure of each action cluster for the different action clusters involves aggregation of the representations within the cluster and relevance analysis with the task goal. The action cluster is formed by correlation degree analysis in the previous step, representing a grouping of initial action representations with high relevance in the multi-modal context. The task goal is usually a target expression extracted from the language features, and a unified feature space is established by further converting the language features into a task goal feature vector. For each action cluster, first, the set of all initial action representations in the cluster needs to be extracted, and then all initial action representations in the set are aggregated. Aggregation can be performed by mean pooling, maximum pooling or weighted fusion, etc. to generate the overall representation of the action cluster, i.e. the action cluster feature vector. Then, the action cluster feature vector and the task goal feature vector are spliced to form a cluster-task joint feature. This splicing operation ensures that the task semantics and the overall information of the action grouping are integrated in the same representation. Then, the cluster-task joint feature is input into the importance analysis module, which is usually composed of a small neural network or a lightweight model. Its function is to output the importance measure value of the action cluster based on the matching degree of the task goal and the cluster action. The measure value can be represented as a probability score or a normalized weight, which is used for subsequent action selection or priority ordering of focus.

[0049] The aggregation manner of the action cluster features can be adjusted according to different application scenarios, for example, adaptive weighted aggregation is used in a multi-task environment to reflect the difference in contribution of each action representation. The generation of the task target feature vector can be implemented by different language encoders, for example, a bidirectional recurrent neural network, a Transformer encoder or a pre-trained language model. The design of the importance analysis module can select different network structures, for example, a two-layer fully connected network, a gate unit or a graph neural network, to ensure adaptation to action cluster feature data of different sizes. The output of the importance measure value can be a scalar score or a multi-dimensional vector to support multi-index decision-making.

[0050] By determining the importance measure for different action clusters respectively, the embodiment can realize dynamic matching and highlighting of different action aggregation units in the multi-modal context with the current task target, automatically strengthen the priority of key action clusters, avoid the decision generalization problem caused by equal treatment of all clusters in the subsequent processing process, and solve the problem that the prior art cannot dynamically distinguish the contribution degree of each action cluster in combination with the task target characteristics.

[0051] S50, performing feature enhancement processing on the initial action representations in the target action cluster according to the importance measure, to generate an optimized action representation set;

[0052] In the embodiment, the process of performing feature enhancement processing on the initial action representations in the target action cluster according to the importance measure includes selection of the target action cluster, mapping of the importance measure value and the initial action representation, feature adjustment and formation of the optimized set. The target action cluster is selected from all action clusters by the preceding importance analysis link, and the importance measure value is a measure of the importance of the target action cluster relative to the current task target. In specific implementation, a weight adjustment coefficient can be determined according to the importance measure value as a quantitative basis for the enhancement strength, and the feature vector of each initial action representation in the target action cluster is adjusted by weighting, for example, by multiplying a weight factor greater than 1 to amplify the value of the action representation in the target cluster. The initial action representations in the non-target action cluster can be selected not to be processed, or optionally a reduction factor less than 1 can be applied. The action representation after weighting adjustment needs to be re-normalized or standardized to ensure the scale consistency between different action representations and the numerical stability of subsequent processing. Finally, all adjusted action representations are integrated to form an optimized action representation set, which has higher distinguishability than the initial action representation set, so that the action representation highly related to the task target has a higher weight in the set.

[0053] Static or dynamic weight adjustment strategies can be adopted. The static adjustment strategy can set a predefined fixed coefficient, for example, the weight of the target action cluster is fixedly enhanced by 1.2 times. The dynamic adjustment strategy can adaptively determine the weight coefficient according to the real-time calculated target cluster importance metric value, for example, by a linear function, a sigmoid function or an exponential function, to map the metric value into a preset weight adjustment range. The specific calculation process of feature enhancement can be realized by matrix operation, which is suitable for efficient parallel processing of GPU or special AI acceleration hardware. Feature normalization processing can adopt L2 regularization or batch normalization, and optimization of action representation set integration can adopt sequential splicing, reordering according to cluster labels or sparse matrix storage to adapt to the calculation requirements of different scales and scenes.

[0054] The embodiment effectively improves the expression weight of key action information in the action representation set by performing importance metric-based feature enhancement on the initial action representations in the target action cluster, so that the model can more accurately focus on the action distribution most relevant to the task target in the subsequent decision-making link, improve the discriminability of the action representation and the discriminability of the model, and solve the problem of insufficient target focus caused by equalization of all action representation weights in the prior art.

[0055] S60, inputting the optimized action representation set into a decision unit to generate a final action decision.

[0056] In the embodiment, the optimized action representation set is a set of action representations integrated after importance enhancement processing, and the process of inputting the set into the decision unit needs to pass the set to the decision module through a data interface or a model calling interface. Each action representation in the optimized action representation set can be represented by a multi-dimensional feature vector, including visual features, language features and action features after multi-modal feature fusion, etc. The decision unit receives the set as input and uses internal decision logic to make action selection. The decision unit can use a pre-trained deep neural network model, a rule engine or a reinforcement learning strategy network as a decision subject, use the enhanced feature distribution provided by the optimized action representation set, and select the desired action output from multiple candidate actions through pattern recognition, probability inference or value function calculation. The generation of action decision can be represented as a scalar value (such as action ID), action parameters (such as angle or force), or multi-dimensional control instructions (such as robot end pose instructions). In order to ensure the stability of numerical processing, the input optimized action representation set can be subjected to feature normalization, dimension correction or data type conversion to fully meet the expected format of the decision unit input.

[0057] The optimized action representation set can be packaged into a tensor format and input into a deep reinforcement learning policy network. The network calculates the action selection probability distribution through algorithms such as policy gradient or Q-learning, and then outputs the final action decision through the maximum value selection or sampling method. A traditional rule engine can also be used to convert the optimized action representation set into an action index sequence, which is parsed by the pre-set decision logic in the rule base to directly generate an action execution instruction. When adapting to different scenarios, a lightweight deep neural network model can be used as a decision unit for tasks with high real-time requirements, such as real-time control of robot operations, to reduce latency. For tasks with high complexity requirements, such as multi-step interactive decision-making, a multi-level decision tree or a stacked neural network structure can be used to process the optimized action representation set layer by layer.

[0058] Example: In the medical health business field, it can be applied to intelligent surgical assistance systems. The optimized action representation set is input to enable the system to preferentially select the most precise action matching the current surgical site and operation step, achieving high-precision and low-latency surgical assistance. In the financial technology business field, it can be applied to intelligent risk control robots. The optimized action representation set is input to select the response or verification action of customer interaction with higher confidence, reducing the risk of misjudgment. In the field of autonomous decision-making of agents, it can be applied to dynamic task scheduling of service robots. The optimized action representation set is input into the task scheduling decision network to enable the robot to select the most suitable behavior instruction for the current environment and task requirements in real time, achieving autonomous and efficient task completion.

[0059] This embodiment inputs the optimized action representation set into the decision unit, uses the optimized feature weight distribution to enhance the expression of key action information, enables the decision unit to accurately perceive the action features most relevant to the task target, avoids the problem of insensitive selection of key actions caused by equalization of weights, and improves the correctness of action selection and the overall efficiency of task execution.

[0060] The application relates to the technical field of artificial intelligence, can be applied to business scenarios such as autonomous decision-making of intelligent agents, financial technology and medical health, and discloses a decision-making method and device based on action representation enhancement, equipment and a medium, which comprises the following steps: acquiring visual data, language instructions and a preset discrete action unit set, generating visual features and language features, and performing multimodal feature fusion on the initial feature vectors of the discrete action units to form an initial action representation set; performing correlation degree analysis on the initial action representation set according to a correlation degree threshold, aggregating initial action representations with high correlation degrees into different action clusters, determining the importance measure of each action cluster, further performing feature enhancement on the initial action representations in the target action cluster, generating an optimized action representation set, and finally inputting the optimized action representation set into a decision unit to generate a final action decision. The application can effectively improve the weight of key action information in the decision-making process by dynamically focusing on the action space and optimizing the action representation based on the importance measure, can avoid the equal processing of all action bins in the traditional method, can ensure that the decision-making is more accurate in a complex task environment, and finally can improve the accuracy and efficiency of task completion.

[0061] In one embodiment, the step S10 comprises:

[0062] S101, acquiring visual data, language instructions and a preset discrete action unit set;

[0063] S102, extracting basic semantic features of the visual data;

[0064] S103, performing key object region feature enhancement processing on the basic semantic features to generate visual features;

[0065] S104, performing word segmentation processing on the language instructions to generate a word segmentation result;

[0066] S105, performing part-of-speech tagging on the word segmentation result to generate part-of-speech tagging information;

[0067] S106, generating a word vector sequence based on the part-of-speech tagging information;

[0068] S107, performing semantic correlation capturing processing on the word vector sequence to generate language features.

[0069] In this embodiment, the acquisition of visual data, language instructions and a preset discrete action unit set is the input basis for the whole action representation enhancement. The visual data is usually provided in the form of images or video frames, which can be collected by optical sensors such as RGB cameras, has time and space distribution characteristics, and needs to be sampled by a pixel array to form a standardized input matrix. The language instruction is provided in the form of natural language text, which can be obtained through user voice transcription or text input interface, and contains semantic clues and context prompts. The preset discrete action unit set represents a set of candidate actions. The action unit can be defined as a standardized operation template, such as a mechanical arm joint angle adjustment unit, a specific gripping force unit, etc. It is uniformly represented by vector coding to ensure compatibility with the dimensions and formats of visual and language data.

[0070] Extracting the basic semantic features of visual data requires preprocessing the original visual data input, including noise filtering, size normalization and color standardization. Then, a deep convolutional neural network (CNN) is used to extract features from the input image. The convolutional layer processes and extracts low-level features such as edges, textures and color distributions layer by layer. Then, high-level convolutional units are used to extract high-level features such as object contours and spatial relationships. A multi-dimensional feature tensor is output as a specific representation of the basic semantic features of the visual data.

[0071] Performing key object region feature enhancement processing on the basic semantic features, mainly through object detection networks such as Faster R-CNN or YOLO framework, to locate and mask the key object regions in the visual data. Within the key object mask range, the feature vector in the region is amplified in amplitude or selectively enhanced in information channel through weighted convolution operation or attention weight promotion mechanism, so that the feature response of the key object region has a higher weight in subsequent multi-modal fusion. This processing can be flexibly adapted to different scenarios, such as using an adaptive ROI (Region of Interest) selection method in dynamic scenarios to dynamically determine the boundaries of the key region.

[0072] Performing word segmentation on the language instruction, the input natural language text is segmented by a word segmentation algorithm to obtain the basic granularity of semantic units. The word segmentation algorithm can combine dictionary matching and statistical modeling methods, and can flexibly adjust the window length and segmentation strategy for long texts and short instructions to ensure the accuracy and coverage of the word segmentation results. After generating the word segmentation results, part-of-speech tagging operation needs to be performed on the word segmentation results. Use a part-of-speech tagging model based on statistical learning or deep learning such as HMM (Hidden Markov Model) or BiLSTM-CRF model to assign each word segmentation result to its corresponding part-of-speech label, providing structural information support for subsequent semantic extraction.

[0073] The generation of the word vector sequence based on the word tagging information is a process of vectorizing and encoding the results of word segmentation and part-of-speech tagging. An embedding vector model such as Word2Vec or BERT is used to map each word to a fixed-length vector, preserving semantic proximity and context information, and a word vector sequence matrix is formed by stacking. The word vector sequence not only contains the independent semantics of a single word, but also embodies the word order relationship and context semantic dependence.

[0074] The semantic association capturing process performed on the word vector sequence can model the dependency relationship between each word vector in the sequence through a self-attention mechanism (Self-Attention) to capture long-distance semantic relationships. This process specifically includes generating query (Query), key (Key), and value (Value) vectors, calculating an attention weight matrix through the dot product between the query and the key, and then weighting and summing the value vectors to obtain global semantic association features. This process ensures that important context clues contained in the language instruction are effectively integrated, generating language features as key inputs for downstream multi-modal fusion.

[0075] During the entire process, low-to-high layer semantic extraction is first performed on the visual data, and then key object region enhancement is performed on the basic visual features to enhance the importance of task-related regions. The language instruction processing starts from the input text, goes through semantic unit segmentation, part-of-speech structure tagging, and semantic embedding encoding, and finally extracts context information through global semantic capturing. All sub-processing processes are organized in a modular manner and can be optimized independently on different hardware platforms and system architectures, such as using lightweight deep models in edge computing environments and using large-scale Transformer architecture processing on high-performance server sides.

[0076] The embodiment extracts visual features and language features through joint processing of visual data, language instructions, and discrete action unit sets, and performs multi-level structured enhancement. In the visual data, the key object region is enlarged to strengthen the representation strength of the task-related visual target, and in the language data, deep semantic associations are extracted through part-of-speech tagging and context dependence modeling. This processing process improves the expression accuracy and semantic discrimination of visual and language inputs, making the subsequent fusion process with discrete action units more discriminative, reducing the coupling of irrelevant features, improving the overall input quality and multi-modal interaction consistency, thereby providing more accurate and context-aware basic inputs for the final action decision, and improving the performance of the multi-modal action decision system in complex task environments.

[0077] In one embodiment, the above step S20 includes:

[0078] S201, for each discrete action unit in the discrete action unit set, concatenating the visual feature, the language feature and the initial feature vector of the discrete action unit to generate a first concatenated feature vector;

[0079] S202, inputting the first concatenated feature vector into a fully connected layer of a multi-layer perception to generate a fusion feature vector;

[0080] S203, for each discrete action unit, concatenating the fusion feature vector and the initial feature vector of the discrete action unit to generate a second concatenated feature vector;

[0081] S204, performing a linear transformation operation on the second concatenated feature vector to generate an initial action representation of the discrete action unit;

[0082] S205, integrating the initial action representations of all discrete action units to form an initial action representation set.

[0083] In the embodiment, in the multi-modal action decision, in order to make full use of the complementarity of the visual feature, the language feature and the initial feature vector of the discrete action unit, first, the visual feature, the language feature and the initial feature vector of each discrete action unit in the discrete action unit set need to be jointly processed. The specific operation is to perform dimension alignment and merging on the visual feature vector, the language feature vector and the initial feature vector of the discrete action unit in a predetermined order by feature vector concatenation, to form a first concatenated feature vector, so as to ensure that different modal information is expressed in a shared feature space and its independent semantics is retained. The concatenation operation can be realized by tensor-level connection, for example, using the concatenate operation in the Numpy or PyTorch framework to concatenate three input tensors in the feature dimension, and the feature dimension after concatenation is equal to the sum of the feature dimensions of the three input features.

[0084] The first concatenated feature vector is processed by a fully connected layer of a constructed multi-layer perception (MLP) to generate a fusion feature vector. The fully connected layer is used to learn the nonlinear relationship between the visual, language and action initial features, and contains a weight matrix and a bias term, which realizes cross-modal feature fusion through matrix multiplication and activation function combination operation. ReLU or GELU can be selected as the activation function to ensure the expression ability of the nonlinear transformation, and batch normalization and Dropout and other regularization methods are supported to improve the training stability and generalization performance.

[0085] For each discrete action unit, after the fusion feature vector is generated, it also needs to be spliced with the initial feature vector of the discrete action unit again to form a second spliced feature vector. The purpose of this operation is to maintain the integrity of the original action unit features and combine the multi-modal fusion result with its original action representation to strengthen the context association. The splicing order is consistent with the first spliced feature, and linear transformation needs to be used for unified mapping after dimension expansion. Linear transformation operation is completed through matrix multiplication and addition, which maps the second spliced feature vector to a predefined dimension space and outputs as the initial action representation of the discrete action unit.

[0086] The initial action representations of all discrete action units are integrated to form an initial action representation set for subsequent correlation analysis and clustering processing. The integration operation is realized by batch-level combination in the tensor dimension to ensure that each initial action representation in the set is consistent in data format and dimension, facilitating subsequent calculations. The continuity and rigor of this processing flow require that the input and output of each feature path be strictly dimensionally matched during multi-modal fusion to avoid calculation errors caused by inconsistent tensors.

[0087] This embodiment can significantly improve the collaborative expression ability between cross-modal features and enhance the interaction between visual, language and action representations by multi-stage splicing and fusion of visual features, language features and initial feature vectors of discrete action units and adding a multi-layer perceptron nonlinear mapping. It solves the problem that traditional methods simply splicing cannot depict high-order interactions of multi-modal. By re-splicing the original initial features after fusion, it ensures that the original properties of the action unit are fully preserved in the fusion features, reducing feature deviation or information dilution caused by multi-modal processing.

[0088] In one embodiment, the above step S30 comprises:

[0089] S301, combining each initial action representation in the initial action representation set in pairs to generate an initial action representation pair set;

[0090] S302, for each initial action representation pair in the initial action representation pair set, generating a query vector based on the first initial action representation in the initial action representation pair;

[0091] S303, generating a key vector based on the second initial action representation in the initial action representation pair;

[0092] S304, determining the dot product of the query vector and the key vector, performing scaling processing on the dot product, and applying a normalization function to the scaling processing result to generate an attention weight between the first initial action representation and the second initial action representation;

[0093] S305, comparing the attention weight between the first initial action representation and the second initial action representation with a preset correlation threshold;

[0094] S306, when the attention weight between the first initial action representation and the second initial action representation is greater than the correlation threshold, establishing a connection between the first initial action representation and the second initial action representation;

[0095] S307, constructing an initial action representation correlation graph based on the established connection of all initial action representations satisfying that the attention weight is greater than the correlation threshold;

[0096] S308, identifying a connected component in the initial action representation correlation graph;

[0097] S309, aggregating all initial action representations in each connected component into an action cluster to form different action clusters.

[0098] In this embodiment, the specific process of correlation analysis first needs to combine each initial action representation in the initial action representation set in pairs to form a set of initial action representation pairs. This combination is completed by enumerating all possible two-by-two combinations in the set, ensuring that each pair of initial action representations is considered. The combination can be achieved by matrix broadcasting or double loop, and specifically, the index pair list can be generated by nested index operation to ensure the calculation order and coverage integrity.

[0099] For each initial action representation pair in the set of initial action representation pairs, a query vector needs to be generated based on the first initial action representation in the pair. The query vector generation can be achieved by linear mapping, which usually includes matrix multiplication of the first initial action representation and the query weight matrix and adding a bias term. The weight matrix and bias term can be obtained by training and learning, and the mapping output dimension is usually consistent with the input dimension or adjusted according to the model design requirements. This operation ensures that the query vector can capture the correlation information in the action features represented by the first initial action representation.

[0100] At the same time, a key vector needs to be generated based on the second initial action representation in the initial action representation pair. The key vector generation process is the same as the query vector generation, using a separate weight matrix and bias parameter to ensure the independence and correlation of the query and the key in the feature space. The key vector obtained by linear mapping is used for matching calculation with the corresponding query vector.

[0101] The dot product between the query vector and the key vector is calculated to measure the similarity of the two in the vector space, and the dot product is calculated by summing the element-wise products. Then, the dot product result is scaled according to a scaling factor, which is usually the inverse of the square root of the dimension of the query vector or the key vector, in order to prevent the gradient from vanishing or exploding due to the large dot product of high-dimensional vectors. The scaled value is input into a normalization function, usually a Softmax function or a Sigmoid function, to limit the output to the [0, 1] interval as the attention weight, reflecting the relative correlation strength between the two initial action representations.

[0102] The correlation strength is compared with a preset correlation threshold. The correlation threshold can be pre-set according to the statistical distribution of the training data or dynamically adjusted by the model to control the sensitivity of aggregation. If the attention weight between a pair of initial action representations is greater than the threshold, a connection relationship is established between them, which can be recorded by an adjacency matrix or an edge set.

[0103] All eligible connection relationships are used to construct an initial action representation association graph, in which each node represents an initial action representation and the edge represents the high correlation between them. After the graph is constructed, graph theory algorithms are used to identify the connected components, which can effectively identify all subsets of initial action representations that are directly or indirectly associated with each other. All initial action representations in each connected component are considered as action units with strong correlation and are directly aggregated into an action cluster. The set of action clusters corresponding to all connected components is the final output of different action clusters.

[0104] The embodiment can measure the similarity of semantics and feature space between any two action representations in detail by comprehensively analyzing all pairs of initial action representations. By combining the query and key vector mechanism with the scaled dot product attention weight, the correlation perception ability for different action pairs is dynamically adjusted, effectively overcoming the limitations of traditional fixed threshold or static feature comparison. By constructing an association graph based on the connection relationship of the correlation threshold and extracting connected components, the initial action representation groups with inherent connections can be adaptively discovered, and high-quality action clustering can be achieved. The whole association analysis and clustering process improves the modeling accuracy of the context relationship between key action units, enabling the subsequent decision-making process to focus on strongly correlated action groups and reduce the interference of irrelevant or weakly related action representations, thereby improving the accuracy and efficiency of subsequent action representation enhancement and decision-making stage.

[0105] In one embodiment, the above step S40 comprises:

[0106] S401, generating a task target feature vector based on the language features;

[0107] S402, for each action cluster, extracting all initial action representations within the action cluster;

[0108] S403, aggregate all initial action representations in the action cluster to generate an action cluster feature vector;

[0109] S404, concatenate the action cluster feature vector and the task target feature vector to generate a cluster task joint feature;

[0110] S405, input the cluster task joint feature into an importance analysis module to generate an importance measure value of the action cluster through the importance analysis module.

[0111] In this embodiment, first, a task target feature vector is generated based on language features. Specifically, language features extracted from original language instructions are input into a specific target representation network, for example, a sequence of linear layers and nonlinear activation functions or optional Transformer encoder blocks, to extract components closely related to the current task intent from the language features and generate a fixed-length vector as the task target feature vector. This vector is used to measure the matching relationship between the action cluster and the task semantics. The generation of the task target feature vector can combine context information, keyword extraction, instruction subject-predicate-object structure analysis, and other technologies to ensure that the feature vector accurately expresses the core intent of the task. This operation not only captures explicit instruction semantics in the language instructions, but also maps implicit target requirements through context relationships.

[0112] Subsequently, for each action cluster, all initial action representations in the cluster need to be extracted. The initial action representation is formed after the multi-modal feature fusion and contains the comprehensive representation of visual features, language features, and action features. The extraction operation is to traverse all members in the current action cluster, collect them one by one from the storage structure (such as a tensor list or a matrix set) in order, form the initial action representation set corresponding to the cluster, and use it for further feature aggregation.

[0113] All initial action representations in the action cluster are aggregated to generate an action cluster feature vector. Aggregation can be achieved in various ways, such as element-wise averaging (average pooling) by dimension, or taking the maximum value of each dimension (maximum pooling), or through weighted averaging combined with weight coefficients to represent the importance of different action representations. The action cluster feature vector formed by this aggregation result serves as a compressed representation of the cluster as a whole.

[0114] After aggregation is complete, the action cluster feature vector and the task target feature vector need to be concatenated. The concatenation operation cascades by vector dimension and can be realized by directly expanding the dimension in memory or through a tensor concatenation function to generate a cluster task joint feature. This joint feature contains not only the multi-modal comprehensive features of the current action cluster, but also the semantic intent representation of the task target, providing rich context for subsequent analysis.

[0115] Finally, the clustering task joint feature is input into the importance analysis module, which can be a lightweight neural network containing several fully connected layers, activation functions, normalization layers, etc., and is specifically used to learn the importance measure value of the action cluster under the current task semantics. The module can be trained using supervised learning, and the output is a scalar value representing the importance of the action cluster relative to the task goal, which can be normalized to [0, 1] by the Sigmoid function, facilitating the connection with the subsequent weight adjustment mechanism.

[0116] By introducing the task target feature vector in the action cluster analysis, the embodiment can directly associate the representation of the current action cluster with the task semantic target, enhancing the context awareness capability of the importance analysis. Extracting and aggregating the initial action representations within the action cluster makes the action cluster features more representative and compact, avoiding the bias that a single representation may bring. Splicing the action cluster features with the task target features can achieve cross-modal deep joint modeling, enabling the importance analysis module to perceive the correlation between the action cluster and the task target. The measure value generated by the importance analysis module not only provides a basis for distinguishing the importance of the cluster, but also supports subsequent feature enhancement operations for high-importance clusters.

[0117] In one embodiment, the above step S50 comprises:

[0118] S501, determining a target action cluster and a weight adjustment coefficient based on the importance measure value;

[0119] S502, applying the weight adjustment coefficient to increase the feature weight of the feature vector of each initial action representation within the target action cluster, and applying the weight adjustment coefficient to reduce the feature weight of the feature vector of each initial action representation within the non-target action cluster, to generate updated action representations;

[0120] S503, integrating all updated action representations to generate an optimized action representation set.

[0121] In this embodiment, the goal is to adjust the feature weights of the initial action representations so that the optimized action representation set can more effectively serve key action recognition and high-precision decision-making in multi-modal decision-making scenarios.

[0122] First, the importance measure value is the result of the previous clustering analysis and task target semantic association analysis. Its source is the multi-dimensional feature coupling degree between the aggregated feature vector of the initial action representation within each action cluster and the task target feature vector, which is usually output by an embedded importance analysis module and normalized to a standard interval. This value not only quantifies the semantic matching strength of the cluster and the task target, but also reflects the explainability of the action cluster in multi-modal tasks.

[0123] In the specific processing flow, firstly, all current action clusters and their corresponding importance measure values need to be retrieved, and one or more are selected as "target action clusters" using ranking or threshold screening mechanism. This screening action needs to dynamically adapt to the task context, which can use absolute threshold screening (such as selecting clusters with importance measure values greater than 0.8) or relative indicators (such as selecting the top-k high-value clusters), to ensure that the selected target action clusters are truly dominant and necessary in the current task decision.

[0124] Subsequently, a weight adjustment coefficient needs to be determined for each target action cluster. The weight adjustment coefficient is derived from a monotonic mapping of the importance measure value, and the mapping relationship allows flexible adjustment to adapt to the fine degree of different tasks. For example, linear mapping (such as directly using the measure value itself as the coefficient) or non-linear mapping (such as using sigmoid, exponential, or log functions to compress or amplify it) can be used. The weight adjustment coefficient should be strictly greater than 1 to ensure that the feature dimensions represented in the target action cluster are amplified as a whole, enhancing their distinguishability in the feature space.

[0125] The processing of the initial action representation in the target action cluster is implemented using matrix operations. The feature vectors of all initial action representations in the cluster are scaled and adjusted one by one, that is, each dimension is traversed, and the original value is multiplied by the weight adjustment coefficient to form a new enhanced feature vector. The entire operation can be completed in batch in a high-performance tensor computing framework, ensuring computational efficiency and parallelism.

[0126] The initial action representation of the non-target action cluster is "reversed adjusted", that is, a weight adjustment coefficient less than 1 is applied. Such a coefficient can be the difference between the adjustment coefficient of the target cluster and 1, ensuring that the adjusted weight proportion is still in the ordered interval, and the degree of weakening is in a controllable symmetric relationship with the degree of enhancement of the target cluster. Such a design not only ensures the prominence of the target action cluster, but also avoids the complete loss of information contribution of the non-target cluster, maintaining overall diversity and feature integrity.

[0127] All adjusted action representations are integrated into an optimized action representation set through a unified integration operation. The integration mechanism uses ordered organization at the data structure level, commonly using tensor format, to ensure that the optimized action representation set maintains the same dimensional structure as the original set, compatible with the batch input requirements of subsequent decision modules.

[0128] It is important to note that this process must ensure consistency between the feature enhancement operation and the interfaces of upstream and downstream modules. In other words, the adjusted feature vectors do not change the vector length or arrangement order, but only reflect enhancement or reduction on the numerical scale. This allows subsequent decision-making units to directly read and utilize the optimized action representation set without adaptation, ensuring the overall closed-loop and efficient operation of the end-to-end system.

[0129] In implementation, the aforementioned adjustment strategy supports fine-grained parameter tuning. Parameterized design includes importance thresholds, amplification / compression factors for weight adjustments, and selection of nonlinear mapping functions, enabling it to adapt to the heterogeneity of target action importance distributions in different scenarios. This parameterized flexibility gives the processing highly adaptable and transferable capabilities, such as fine-grained focusing on complex physiological actions in the healthcare field, and enhancing the dominant features of risk-driven decision-making action sequences in fintech businesses.

[0130] The final result is an optimized set of action representations. The feature weight distribution in the set has been reconstructed according to the importance of the actions. The geometric distribution of the set in the multimodal feature space is more in line with the task objective, enabling the subsequent decision-making stage to focus on the key action instructions required by the task more quickly and accurately.

[0131] This embodiment dynamically determines target action clusters using importance metrics and employs refined weight adjustment coefficients to weight and enhance or weaken the initial action representations of different clusters, achieving adaptive reconstruction of action features in a multimodal space. This process avoids the drawback of treating all action representations equally in traditional methods, enabling the optimized action representation set to centrally express the key action information most relevant to the task objective. This improves the decision-making unit's ability to identify and select key actions in complex task environments, strengthens the system's ability to differentiate non-homogeneous features in the scene, and significantly improves the accuracy and efficiency of task decision-making without increasing the overall data processing burden.

[0132] In one embodiment, after step S60 above, the method further includes:

[0133] S701, execute the action decision and collect action execution result information;

[0134] S702, Based on the action execution result information and the visual data, language instructions and discrete action unit set, generate training samples;

[0135] S703, based on the training samples, update the feature fusion network parameters and self-attention mechanism parameters;

[0136] S704, adjusts the multimodal feature fusion process based on the updated feature fusion network parameters;

[0137] S705, based on the updated self-attention mechanism parameters, adjusting the correlation degree analysis process.

[0138] In this embodiment, after the decision unit completes the final action decision, an adaptive optimization process characterized by a decision-making closed loop is constructed. First, the action decision needs to be executed, and the execution link of the action decision corresponds to mapping the action instruction output by the decision unit into the operation behavior in the actual environment, such as the control signal of the mechanical arm, the operation command of the virtual agent, or the action response of the service system. The action execution result information is the feedback data reflecting the execution effect collected from the operation behavior, including but not limited to the action execution success rate, the completion state, the response delay, the deviation degree from the target, etc. Such feedback information is an important basis for measuring the decision accuracy and a key element for generating training samples.

[0139] Subsequently, the action execution result information needs to be combined with the previously input visual data, language instructions, and discrete action unit set to generate training samples. The composition of the training samples provides supervised learning data for subsequent model parameter optimization. The visual data and language instructions come from the task input, and the pre-set discrete action unit set provides a discrete expression framework for the action space. The action execution result information serves as a label or constraint signal. Through data preprocessing operations such as data alignment, time sequence association, and label mapping, the combination of the three forms a multi-dimensional labeled data set for model training.

[0140] After the training samples are generated, the gradient backpropagation algorithm based on the training samples is used to update the feature fusion network parameters and the self-attention mechanism parameters. The feature fusion network parameters refer to the weight matrix and bias term used for multi-modal information integration in the multi-layer perceptron and its related layers, and the self-attention mechanism parameters refer to the mapping matrix of the query, key, and value used to construct the attention weight distribution. The update process is realized by minimizing the loss function on the training samples, which can specifically include mean square error loss, cross-entropy loss, or weighted combination loss. Parameter updating is based on batch gradient descent or adaptive optimizers such as Adam and RMSprop, ensuring training convergence and global stability.

[0141] Based on the updated feature fusion network parameters, the multi-modal feature fusion process needs to be adjusted. The adjustment mainly reflects changes in the weight matrix of the fusion operator, so that the subsequent multi-modal feature fusion results can better align with the target task requirements and reduce redundant information. Similarly, based on the updated self-attention mechanism parameters, the correlation degree analysis process is adjusted. The adjustment makes the subsequent attention distribution more accurately express the real relevance between action representations, improving the accuracy of clustering analysis. The above adjustments do not require modification of the overall structure and only rely on parameter updates to ensure algorithm stability and reusability.

[0142] Example: In the medical health field, a system for assisting surgical robot autonomous decision-making can be built, aiming to automatically determine the next best action decision according to the real-time scene in the operating room, the doctor's voice instructions, and the preset standard operation action set, to improve the accuracy and safety of robot autonomous operation in complex minimally invasive surgery.

[0143] First, the visual data, language instructions, and preset discrete action unit set are obtained. The visual data comes from the real-time image sequence captured by the high-precision 3D camera installed in the operating room, which contains the visual information of the patient's anatomical site; the language instruction is the natural language command issued by the lead surgeon through the voice system, which may include target site description, action prompt, etc.; the discrete action unit set is a set of standardized basic action units defined according to surgical operation experience, such as "0.5 millimeter left cutting" "clamp tightening 2 Newton" etc.

[0144] Subsequently, visual features and language features are generated. Visual features are generated by performing basic semantic feature extraction on visual data, including segmentation and labeling of patient skin, blood vessels, and organ surfaces, using region feature enhancement algorithms to enhance the image region feature expression of key anatomical structures (such as specific arteries, nerves), making the visual features more sensitive to key parts. Language feature generation includes word segmentation of the doctor's instructions, part-of-speech tagging using a medical terminology dictionary, extracting action targets, surgical tools, and anatomical structure nouns involved in the instructions, constructing word vector sequences, and further using a semantic association capture model to convert the implicit operation intent in the instruction context into a compact language feature vector.

[0145] Next, the visual features and language features are respectively fused with the initial feature vector of each action unit in the discrete action unit set. For each action unit, the visual features, language features, and the initial feature vector of the action unit are concatenated to generate a first concatenated feature vector, which is input into a multi-layer perceptron fully connected layer to output a fused feature vector, which is then concatenated with the initial feature vector of the action unit to form a second concatenated feature vector, and finally the linear transformation is performed to obtain the initial action representation corresponding to the action unit. After integrating the initial action representations of all action units, an initial action representation set is formed.

[0146] The correlation degree analysis is performed in the initial action representation set. All initial action representations in the set are combined in pairs to construct a set of initial action representation pairs. For each pair of initial action representations, a query vector is generated based on the first representation, a key vector is generated based on the second representation, a scaled dot product of the two is calculated and normalized to obtain an attention weight. If the weight is greater than a preset correlation degree threshold, it is considered that the correlation degree is high, and a connection between them is established. Based on all connection relationships, an initial action representation correlation graph is constructed, and different action clusters are aggregated through graph connected component recognition to ensure that strongly related action units are effectively aggregated, providing a basis for subsequent action filtering.

[0147] For different action clusters, the importance measure of each action cluster is determined. First, a task target feature vector is generated based on language features, such as the feature vector of the "locate artery suture" task target. For each action cluster, all initial action representations are extracted and aggregated to generate an action cluster feature vector. The action cluster feature vector is concatenated with the task target feature vector to obtain a cluster-task joint feature, which is input into the importance analysis module. The importance measure value is calculated by a neural network and is used to represent the criticality of the current action cluster to the surgical task.

[0148] According to the importance measure result, the initial action representations in the target action cluster are enhanced. Based on the importance measure value, the target action cluster and the weight adjustment coefficient are determined. For each initial action representation in the target action cluster, the weight adjustment coefficient is applied to increase its feature weight, thereby improving the model's attention to key actions. For initial action representations in non-target action clusters, the weight adjustment coefficient is applied to reduce their feature weights, thereby reducing the influence of redundant actions. Updated action representations are generated, and an optimized action representation set is formed after integration.

[0149] After the optimized action representation set is input into the decision unit, a final action decision is generated, such as "0.5 millimeters left cut" as the next operation instruction. The action decision is directly executed, and the robot collects action execution result information in real time after execution, such as cutting accuracy, bleeding control, etc. Based on the action execution result information, visual data, language instructions, and discrete action unit set, training samples are generated. These training samples are used to continuously update the feature fusion network parameters and self-attention mechanism parameters, and the model is optimized through gradient descent method. The updated parameters are used to adjust the multi-modal feature fusion process and correlation degree analysis process, so that the system is more accurate and meets the actual needs of the operation environment in the next round of decision-making, and has self-adaptive optimization capability.

[0150] In the field of financial technology, an autonomous decision-making module for an intelligent financial advisor system can be constructed. The goal is to automatically generate personalized asset allocation decisions based on user risk preferences, market data, investor language consultation content, and a pre-defined set of standardized financial operations, in order to improve the intelligence level and dynamic response capability of portfolio adjustment.

[0151] First, collect visual data, language instructions, and a pre-defined set of discrete action units. Visual data comes from dynamic visual financial market data streams, including real-time K-line charts, price trend charts, etc. Language instructions are investor consultation content input through natural language, such as "how to balance the risk and return in the current investment portfolio". The set of discrete action units is a pre-defined set of standard investment operation units for financial advisors, such as "allocate 20% of funds to low-risk bonds" and "reduce holdings of high-volatility technology stocks".

[0152] Subsequently, generate visual features and language features. Visual features are extracted from the basic semantic features of visual data, analyzing trend direction, volatility intensity, key support and pressure areas in market trend images, and enhancing regional features to make visual features more sensitive to specific market risk indicators. The generation of language features includes word segmentation of investor consultation sentences, part-of-speech tagging combined with a financial terminology dictionary, extraction of keywords related to risk, return, and asset categories to form word vector sequences, and semantic association capture of investment goals and intentions through context modeling to convert into language feature representations for subsequent decision analysis.

[0153] Next, perform multi-modal feature fusion of visual features and language features with the initial feature vector of each action unit in the set of discrete action units. For each action unit, concatenate the visual features, language features, and the initial feature vector of the action unit to generate a first concatenated feature vector, input it into a multi-layer perceptron fully connected layer to output a fusion feature vector, then concatenate it with the initial feature vector of the action unit to generate a second concatenated feature vector, and finally obtain the initial action representation of the action unit through linear transformation. The set of initial action representations of all action units is formed at this stage.

[0154] Perform correlation analysis within the set of initial action representations. Combine all initial action representations in pairs to form a set of initial action representation pairs. For each pair, generate a query vector based on the first representation and a key vector based on the second representation. Calculate the scaled dot product and normalize it to generate an attention weight. If the weight is greater than a correlation threshold, establish a connection between the two. All connections are built into an initial action representation correlation graph. Through connected component identification, high-correlation action representations are aggregated into different action clusters, such as "reduce holdings of high-volatility technology stocks" and "transfer to low-volatility utility stocks" are aggregated into a risk mitigation cluster.

[0155] For different action clusters, the importance measure of each action cluster is determined. First, the current user investment intention feature vector is generated based on the language features, such as the "hope to reduce the volatility of the portfolio" target vector. For each action cluster, all initial action representations therein are extracted and aggregated to generate an action cluster feature vector, which is spliced with the user investment intention feature vector into a joint feature of the clustering task, input into the importance analysis module, and output the importance measure value of the action cluster to represent the key degree of the action cluster to the current investment target.

[0156] According to the importance measure value, the initial action representations in the target action cluster are enhanced in features to improve the attention of the system to the key investment operations. By determining the weight adjustment coefficient, for each initial action representation in the target action cluster, the weight adjustment coefficient is applied to amplify its weight, and for the action representations in the non-target action cluster, the weight adjustment coefficient is applied to reduce its weight, to generate an updated action representation set, which is integrated to form an optimized action representation set.

[0157] The optimized action representation set is input into the decision unit to generate the final personalized investment action decision, such as "reduce 30% high-volatility assets, increase 20% bond funds". After the automatic execution of the decision instruction, the system collects the execution feedback, including the risk and return ratio after the change of the investment portfolio, the user investment experience score, the market response, etc. Based on these execution feedbacks, combined with the original visual data, language instructions and discrete action unit set, new training samples are generated to update the feature fusion network parameters and self-attention mechanism parameters, so that the model is more sensitive to subsequent user investment behavior and market environment. The updated parameters are used to optimize the multi-modal feature fusion process and correlation analysis process, supporting the system to gradually adapt to different users, different risk preferences and market dynamics, and realizing more refined, personalized and interpretable investment suggestion generation.

[0158] In the field of autonomous decision-making of intelligent agents, considering the scenario of autonomous mobile robots completing multi-task path planning and dynamic obstacle avoidance in complex environments, the goal is to dynamically decide key actions through autonomous perception and interpretation of environmental data, combined with task instructions, to improve environmental adaptability and task execution efficiency.

[0159] First, the robot acquires visual data, language instructions and a preset discrete action unit set. The visual data includes the surrounding environment image stream collected by the robot camera, such as indoor space layout, obstacle distribution, etc. The language instructions are issued by the operator in natural language form, such as "find the nearest exit and avoid high-risk areas". The discrete action unit set is predefined as standard motion action units that the robot can execute, such as "turn left 15 degrees", "move forward 0.5 meters", "slow down through narrow lane", etc.

[0160] The robot generates visual features and language features based on visual data and language instructions. The visual feature generation includes basic semantic feature extraction on environment images, such as target region identification of walls, doorways, obstacles, etc., and then feature enhancement on key object regions related to the task, such as narrow passages or dangerous areas, so that they have higher feature weights in subsequent processing. The language feature generation includes word segmentation processing on operator natural language instructions, part-of-speech tagging combined with a task dictionary to generate word tagging information, further construction of word vector sequences, and capture of intent, spatial target description and constraint conditions through context dependency modeling, and finally form semantic task instruction representation.

[0161] The visual features and language features are respectively fused with the initial feature vectors of each action unit in the discrete action unit set. For each action unit, the robot concatenates the visual features, language features and its initial feature vector to generate a first concatenated feature vector, inputs it into a multi-layer perceptron fully connected layer to obtain a fusion feature vector, and then concatenates it with the initial feature vector to generate a second concatenated feature vector. Finally, after linear transformation, the initial action representation corresponding to the action unit is obtained, and finally an initial action representation set is formed to describe the semantic and environmental adaptability of each selectable action in the robot's current action space under multi-modal information.

[0162] The robot performs relevance analysis on the initial action representation set. By combining each pair of initial action representations, an action representation pair set is generated. For each pair of action representations, a query vector is generated based on the first representation, and a key vector is generated based on the second representation. The scaled dot product is calculated and normalized to generate an attention weight. When the attention weight is greater than a relevance threshold, a connection relationship between the two is established. The robot constructs an initial action representation association graph based on all connection relationships, identifies connected components in the association graph, and aggregates high-relevance initial action representations into action clusters, such as "slow down and turn to avoid obstacles" and "pass through narrow passages" which can be aggregated into a safe passage action cluster.

[0163] For different action clusters, the robot determines the importance measure of each action cluster. First, a task target feature vector is generated based on the language features, such as "preferentially selecting a safe path and saving energy consumption". For each action cluster, the robot extracts all initial action representations and aggregates them to generate an action cluster feature vector, which is concatenated with the task target feature vector to generate a cluster task joint feature, which is input into an importance analysis module. The importance measure value of the action cluster is output by the module, which measures the criticality of each cluster to the current task target.

[0164] The robot performs feature enhancement on the target action cluster according to the importance measure value. The target action cluster and the weight adjustment coefficient are determined based on the importance measure value. For each initial action representation feature vector in the target action cluster, the weight adjustment coefficient is applied to increase the weight thereof, while the weight adjustment coefficient is applied to the representation feature vector of the non-target action cluster to decrease the weight thereof, to obtain an updated action representation. After integration, an optimized action representation set is formed, so that the robot pays more attention to key path planning and obstacle avoidance actions in the subsequent decision-making process.

[0165] The optimized action representation set is input into a decision unit to generate a final action decision. The robot selects an optimal set of action sequences according to the optimized representation set, such as "turn 30 degrees, then decelerate 0.5 meters and pass through the narrow channel". After executing the action decision, the robot collects execution feedback, such as whether the obstacle avoidance is successful, whether the path is smooth, whether the energy consumption meets the standard, etc. Based on these execution feedbacks, combined with the original visual data, language instructions and discrete action unit set, new training samples are generated to update the feature fusion network parameters and the self-attention mechanism parameters, so that the robot has stronger action selection adaptability in similar complex environments in the future. The updated parameters are used to optimize the multi-modal feature fusion process and the correlation analysis process, respectively, to improve the self-adaptation and refinement level of the robot's perception-decision closed loop.

[0166] In this embodiment, the action execution result information is combined with the visual data, language instructions and discrete action unit set to generate training samples, which drive the continuous optimization of the feature fusion network parameters and the self-attention mechanism parameters, thereby realizing the adaptive adjustment of multi-modal feature fusion and correlation analysis. Unlike traditional static models, this process can dynamically correct the weight parameters in the model according to the execution feedback, so that the model gradually improves its ability to capture key action features and the accuracy of relatedness modeling in different task iterations. Ultimately, the overall adaptive optimization of scene input data and historical execution results in the action decision-making link is realized, effectively improving the accuracy and robustness of the decision-making link and enhancing the adaptability and continuous optimization capability of the system in complex task environments.

[0167] In an embodiment, a decision-making device based on action representation enhancement is provided, which corresponds to the decision-making method based on action representation enhancement in the above embodiments. Referring to Figure 3 , Figure 3 A functional module schematic diagram of a preferred embodiment of the decision-making device based on action representation enhancement of the present application is shown in FIG. 1. The data acquisition and feature generation module 10, the multi-modal feature fusion module 20, the correlation analysis and clustering module 30, the importance measure evaluation module 40, the feature enhancement and optimization module 50, and the decision generation module 60. The detailed description of each functional module is as follows:

[0168] The data acquisition and feature generation module 10 is configured to acquire visual data, language instructions and a preset discrete action unit set, and generate visual features and language features based on the visual data and the language instructions, respectively.

[0169] The multi-modal feature fusion module 20 is configured to perform multi-modal feature fusion between the visual features, the language features and initial feature vectors of each discrete action unit in the discrete action unit set, to generate initial action representations corresponding to each discrete action unit, and form an initial action representation set.

[0170] The correlation analysis and clustering module 30 is configured to perform correlation analysis between initial action representations in the initial action representation set according to a preset correlation threshold, and aggregate initial action representations with high correlation into different action clusters.

[0171] The importance measure evaluation module 40 is configured to determine an importance measure of each action cluster for the different action clusters.

[0172] The feature enhancement and optimization module 50 is configured to perform feature enhancement processing on initial action representations in a target action cluster according to the importance measure, to generate an optimized action representation set.

[0173] The decision generation module 60 is configured to input the optimized action representation set into a decision unit, to generate a final action decision.

[0174] In an embodiment, the data acquisition and feature generation module 10 is specifically configured to:

[0175] acquire visual data, language instructions and a preset discrete action unit set;

[0176] extract basic semantic features of the visual data;

[0177] perform key object region feature enhancement processing on the basic semantic features, to generate visual features;

[0178] perform word segmentation processing on the language instructions, to generate a word segmentation result;

[0179] perform part-of-speech tagging on the word segmentation result, to generate part-of-speech tagging information;

[0180] generate a word vector sequence based on the part-of-speech tagging information;

[0181] perform semantic correlation capturing processing on the word vector sequence, to generate language features.

[0182] In an embodiment, the multi-modal feature fusion module 20 is specifically configured to:

[0183] concatenate the visual feature, the language feature and the initial feature vector of the discrete action unit, to generate a first concatenated feature vector;

[0184] input the first concatenated feature vector into a fully connected layer of a multi-layer perception machine, to generate a fusion feature vector;

[0185] for each discrete action unit, concatenate the fusion feature vector and the initial feature vector of the discrete action unit, to generate a second concatenated feature vector;

[0186] perform a linear transformation operation on the second concatenated feature vector, to generate an initial action representation of the discrete action unit;

[0187] integrate the initial action representations of all discrete action units, to form an initial action representation set.

[0188] In an embodiment, the correlation analysis and clustering module 30 is specifically configured to:

[0189] combine each initial action representation in the initial action representation set in pairs, to generate an initial action representation pair set;

[0190] for each initial action representation pair in the initial action representation pair set, generate a query vector based on a first initial action representation in the initial action representation pair;

[0191] generate a key vector based on a second initial action representation in the initial action representation pair;

[0192] determine a dot product of the query vector and the key vector, perform scaling processing on the dot product, and apply a normalization function to the scaling processing result, to generate an attention weight between the first initial action representation and the second initial action representation;

[0193] compare the attention weight between the first initial action representation and the second initial action representation with a preset correlation threshold;

[0194] when the attention weight between the first initial action representation and the second initial action representation is greater than the correlation threshold, establish a connection between the first initial action representation and the second initial action representation;

[0195] based on the connections established by all initial action representation pairs satisfying that the attention weight is greater than the correlation threshold, construct an initial action representation correlation graph;

[0196] identify connected components in the initial action representation correlation graph;

[0197] Aggregate all initial action representations in each connected component into an action cluster to form different action clusters.

[0198] In an embodiment, the importance measure evaluation module 40 is specifically configured to:

[0199] generate a task target feature vector based on the language features;

[0200] For each action cluster, extract all initial action representations within the action cluster;

[0201] aggregate all initial action representations within the action cluster to generate an action cluster feature vector;

[0202] concatenate the action cluster feature vector and the task target feature vector to generate a cluster-task joint feature;

[0203] input the cluster-task joint feature into the importance analysis module to generate an importance measure value of the action cluster through the importance analysis module.

[0204] In an embodiment, the feature enhancement and optimization module 50 is specifically configured to:

[0205] determine a target action cluster and a weight adjustment coefficient based on the importance measure value;

[0206] apply the weight adjustment coefficient to increase the feature weight of the feature vector of each initial action representation within the target action cluster, and apply the weight adjustment coefficient to decrease the feature weight of the feature vector of each initial action representation within a non-target action cluster to generate updated action representations;

[0207] integrate all updated action representations to generate an optimized action representation set.

[0208] In an embodiment, the decision generation module 60 is specifically configured to:

[0209] execute the action decision and collect action execution result information;

[0210] generate training samples based on the action execution result information and the visual data, the language instructions, and the discrete action unit set;

[0211] update the feature fusion network parameters and the self-attention mechanism parameters based on the training samples;

[0212] adjust the multi-modal feature fusion process based on the updated feature fusion network parameters;

[0213] adjust the correlation degree analysis process based on the updated self-attention mechanism parameters.

[0214] In an embodiment, a computer device is provided, which can be a server, and an internal structure diagram thereof can be as shown in Figure 4 The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external client through a network connection. The computer program, when executed by the processor, implements the functions or steps of a server side of a decision-making method based on action representation enhancement.

[0215] In an embodiment, a computer device is provided, which can be a client, and an internal structure diagram thereof can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external server through a network connection. The computer program, when executed by the processor, implements the functions or steps of a client side of a decision-making method based on action representation enhancement.

[0216] In an embodiment, a computer device is provided, which includes a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the following steps:

[0217] Obtaining visual data, language instructions and a preset discrete action unit set, and generating visual features and language features based on the visual data and the language instructions, respectively;

[0218] Performing multi-modal feature fusion on the visual features, the language features and initial feature vectors of each discrete action unit in the discrete action unit set, respectively, to generate initial action representations corresponding to each discrete action unit, and forming an initial action representation set;

[0219] According to a preset correlation threshold, performing correlation analysis on each initial action representation in the initial action representation set, and aggregating initial action representations with high correlation into different action clusters;

[0220] determine an importance measure of each action cluster;

[0221] perform feature enhancement processing on the initial action representations in the target action cluster according to the importance measure, to generate an optimized action representation set;

[0222] input the optimized action representation set into a decision unit to generate a final action decision.

[0223] In one embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the following steps:

[0224] obtain visual data, language instructions and a preset discrete action unit set, and generate visual features and language features based on the visual data and the language instructions, respectively;

[0225] perform multi-modal feature fusion between the visual features, the language features and initial feature vectors of each discrete action unit in the discrete action unit set, respectively, to generate initial action representations corresponding to each discrete action unit, to form an initial action representation set;

[0226] perform correlation degree analysis between the initial action representations in the initial action representation set according to a preset correlation degree threshold, and aggregate initial action representations with high correlation degrees into different action clusters;

[0227] determine an importance measure of each action cluster;

[0228] perform feature enhancement processing on the initial action representations in the target action cluster according to the importance measure, to generate an optimized action representation set;

[0229] input the optimized action representation set into a decision unit to generate a final action decision.

[0230] It should be noted that the functions or steps that the computer readable storage medium or the computer device can implement correspond to the descriptions of the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0231] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0232] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0233] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for example introduction and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A method for decision making based on action representation enhancement, characterized in that, The method comprises the following steps: acquiring visual data, language instructions and a preset discrete action unit set, and generating visual features and language features based on the visual data and the language instructions respectively; performing multi-modal feature fusion on the visual features, the language features and initial feature vectors of each discrete action unit in the discrete action unit set to generate initial action representations corresponding to each discrete action unit, and forming an initial action representation set; performing correlation degree analysis on each initial action representation in the initial action representation set according to a preset correlation degree threshold, and aggregating initial action representations with high correlation degrees into different action clusters; determining the importance measure of each action cluster; performing feature enhancement processing on the initial action representations in the target action cluster according to the importance measure to generate an optimized action representation set; inputting the optimized action representation set into a decision unit to generate a final action decision.

2. The action representation-based reinforcement enhanced decision method of claim 1, wherein, The method comprises the following steps: acquiring visual data, language instructions and a preset discrete action unit set, and generating visual features and language features based on the visual data and the language instructions respectively; acquiring visual data, language instructions and a preset discrete action unit set; extracting basic semantic features of the visual data; performing key object region feature enhancement processing on the basic semantic features to generate visual features; performing word segmentation processing on the language instructions to generate a word segmentation result; performing part-of-speech tagging on the word segmentation result to generate part-of-speech tagging information; generating a word vector sequence based on the part-of-speech tagging information; 3. The action representation based reinforcement enhanced decision method of claim 1, wherein, performing semantic correlation capturing processing on the word vector sequence to generate language features. The method comprises the following steps: for each discrete action unit in the discrete action unit set, concatenating the visual features, the language features and the initial feature vector of the discrete action unit to generate a first concatenated feature vector; inputting the first concatenated feature vector into a fully connected layer of a multi-layer perception machine to generate a fusion feature vector; for each discrete action unit, concatenating the fusion feature vector and the initial feature vector of the discrete action unit to generate a second concatenated feature vector; performing linear transformation on the second concatenated feature vector to generate the initial action representation of the discrete action unit; 4. The action representation based reinforcement enhanced decision method of claim 1, wherein, integrating the initial action representations of all discrete action units to form an initial action representation set. The method comprises the following steps: performing pairwise combination on each initial action representation in the initial action representation set to generate an initial action representation pair set; for each initial action representation pair in the initial action representation pair set, generating a query vector based on a first initial action representation in the initial action representation pair; generate a key vector based on the second initial action representation in the pair of initial action representations; determine a dot product of the query vector and the key vector, scale the dot product, and apply a normalization function to the scaled result to generate an attention weight between the first initial action representation and the second initial action representation; compare the attention weight between the first initial action representation and the second initial action representation with a preset correlation threshold; when the attention weight between the first initial action representation and the second initial action representation is greater than the correlation threshold, establish a connection between the first initial action representation and the second initial action representation; construct an initial action representation correlation graph based on the connections established between all pairs of initial action representations whose attention weights are greater than the correlation threshold; identify connected components in the initial action representation correlation graph; aggregate all initial action representations in each connected component into an action cluster to form different action clusters.

5. The action representation based reinforcement enhanced decision method of claim 1, wherein, For the different action clusters, determine an importance measure of each action cluster, including: generate a task target feature vector based on the language features; for each action cluster, extract all initial action representations within the action cluster; aggregate all initial action representations within the action cluster to generate an action cluster feature vector; concatenate the action cluster feature vector and the task target feature vector to generate a cluster-task joint feature; input the cluster-task joint feature into an importance analysis module to generate an importance measure value of the action cluster through the importance analysis module.

6. The action representation-based reinforcement enhanced decision method of claim 5, wherein, According to the importance measure, perform feature enhancement processing on the initial action representations within the target action cluster to generate an optimized action representation set, including: determine a target action cluster and a weight adjustment coefficient based on the importance measure value; for each initial action representation within the target action cluster, apply the weight adjustment coefficient to increase the feature weight, and for each initial action representation within the non-target action cluster, apply the weight adjustment coefficient to reduce the feature weight to generate an updated action representation; integrate all updated action representations to generate an optimized action representation set.

7. The action representation based reinforcement enhanced decision method of claim 1, wherein, After inputting the optimized action representation set into a decision unit to generate a final action decision, further including: execute the action decision and collect action execution result information; based on the action execution result information and the visual data, language instructions, and discrete action unit set, generate training samples; based on the training samples, update the feature fusion network parameters and the self-attention mechanism parameters; based on the updated feature fusion network parameters, adjust the multi-modal feature fusion process; based on the updated self-attention mechanism parameters, adjust the correlation analysis process.

8. A decision-making device based on action representation enhancement, characterized in that, The decision device based on action representation enhancement includes: a data acquisition and feature generation module for acquiring visual data, language instructions, and a preset discrete action unit set, and generating visual features and language features based on the visual data and the language instructions, respectively; The multi-modal feature fusion module is configured to perform multi-modal feature fusion between the visual feature and the language feature and an initial feature vector of each discrete action unit in the discrete action unit set, to generate an initial action representation corresponding to each discrete action unit, and to form an initial action representation set; The correlation analysis and clustering module is configured to perform correlation analysis between each initial action representation in the initial action representation set according to a preset correlation threshold, and to aggregate initial action representations with high correlation into different action clusters; The importance measure evaluation module is configured to determine an importance measure of each action cluster for the different action clusters; The feature enhancement and optimization module is configured to perform feature enhancement processing on the initial action representations in a target action cluster according to the importance measure, to generate an optimized action representation set. The decision generation module is configured to input the optimized action representation set into a decision unit, to generate a final action decision.

9. A computer device, comprising: The computer device includes a memory, a processor, and an action representation enhancement-based decision program stored on the memory and executable on the processor. When the action representation enhancement-based decision program is executed by the processor, the steps of the action representation enhancement-based decision method according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that, The storage medium stores an action representation enhancement-based decision program. When the action representation enhancement-based decision program is executed by the processor, the steps of the action representation enhancement-based decision method according to any one of claims 1-7 are implemented.

Citation Information

Patent Citations

  • Motion decision-making method and device, medium and computing equipment

    CN114781646A

  • Intelligent agent key decision action detection method and device

    CN117437510A