Decision generation method and device based on knowledge distillation, equipment and medium

By extracting multimodal features and constructing a hierarchical knowledge graph based on knowledge distillation, the accuracy and efficiency issues of multimodal data decision-making in existing technologies are solved, enabling efficient processing and accurate decision-making for complex tasks.

CN120930800APending Publication Date: 2025-11-11PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511104779.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing vision-language-action (VLA) technology lacks the ability to mine deep knowledge when making decisions based on multimodal data, resulting in high error rates and low efficiency in task execution. Furthermore, it lacks a hierarchical reasoning mechanism, making it difficult to handle complex tasks.

Method used

This paper describes a method based on knowledge distillation to extract multimodal features, construct a hierarchical knowledge graph, and perform hierarchical reasoning and decision-making. The method includes extracting initial multimodal features, knowledge distillation, and constructing a hierarchical knowledge graph. The method also utilizes Transformer and graph neural networks for reasoning and decision-making.

Benefits of technology

It improves the accuracy of multimodal data decision-making and the ability to handle complex tasks, and enables effective reasoning from local features to global decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930800A_ABST
    Figure CN120930800A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, provides a decision generation method, device and equipment based on knowledge distillation and a medium, is applied to financial and medical health care service scenes, can extract multi-modal features in target input data, and realizes extraction and preliminary integration of multi-modal effective features. Knowledge distillation is carried out on the multi-modal initial features, more implicit knowledge is learned for the multi-modal initial features, and the quality and expression ability of the features are improved; the hierarchical knowledge graph is constructed according to the multi-modal knowledge enhancement features, so that the multi-modal knowledge can be subjected to structured representation from different hierarchies, and a clear and organized knowledge basis is provided for subsequent reasoning and decision making; the hierarchical reasoning decision is carried out according to the multi-modal knowledge enhancement features and the hierarchical knowledge graph, analysis and decision making can be carried out on multi-modal information from different granularities, the processing capacity and decision making accuracy of complex tasks are improved, and effective reasoning from local features to global decision making is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a decision generation method, apparatus, device, and medium based on knowledge distillation. Background Technology

[0002] Existing Vision-Language-Action (VLA) technologies, when making decisions based on multimodal data, are often limited to the extraction and fusion of surface features, exhibiting weak capabilities in mining the deep knowledge implicit in the data. Furthermore, traditional models typically employ simple feature concatenation or weighted fusion methods, failing to effectively extract knowledge such as physical laws from visual images, common-sense logic from verbal text, and operational experience from action data. For example, in assembly tasks of robots widely used in financial and medical settings, models struggle to extract spatial constraints for parts assembly from visual images and cannot combine the assembly sequence logic in verbal instructions with actual action execution, resulting in a high task execution error rate.

[0003] Furthermore, existing models lack hierarchical reasoning mechanisms, making it impossible to analyze and make decisions on multimodal information at different granularities when faced with complex tasks. Since most models employ a single reasoning path, they cannot flexibly adjust the depth and breadth of reasoning according to task requirements, thus performing poorly when handling complex tasks requiring multi-step reasoning. Simultaneously, the ability to transfer and reuse knowledge between models is poor, making it difficult to leverage existing knowledge and experience to improve the efficiency and accuracy of handling new tasks. Summary of the Invention

[0004] In view of the above, it is necessary to provide a decision generation method, apparatus, device and medium based on knowledge distillation, which aims to solve the problems of high error rate and low efficiency when making decisions based on multimodal data.

[0005] A decision generation method based on knowledge distillation, the decision generation method based on knowledge distillation includes:

[0006] In response to a decision generation instruction based on target input data, multimodal features are extracted from the target input data to obtain initial multimodal features;

[0007] Knowledge distillation is performed on the initial multimodal features to obtain multimodal knowledge-enhanced features;

[0008] A hierarchical knowledge graph is constructed based on the multimodal knowledge enhancement features;

[0009] Based on the multimodal knowledge enhancement features and the hierarchical knowledge graph, hierarchical reasoning and decision-making are performed to obtain target decision data based on the target input data.

[0010] A decision generation device based on knowledge distillation, the decision generation device based on knowledge distillation includes:

[0011] An extraction unit is configured to extract multimodal features from the target input data in response to a decision generation instruction based on the target input data, thereby obtaining initial multimodal features;

[0012] The distillation unit is used to perform knowledge distillation on the initial multimodal features to obtain multimodal knowledge-enhanced features;

[0013] The construction unit is used to construct a hierarchical knowledge graph based on the multimodal knowledge enhancement features;

[0014] The decision-making unit is used to perform hierarchical reasoning and decision-making based on the multimodal knowledge enhancement features and the hierarchical knowledge graph to obtain target decision data based on the target input data.

[0015] A computer device, the computer device comprising:

[0016] Memory, storing at least one instruction; and

[0017] The processor executes the instructions stored in the memory to implement the knowledge distillation-based decision generation method.

[0018] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the knowledge distillation-based decision generation method.

[0019] As can be seen from the above technical solutions, this invention can extract multimodal features from target input data to obtain multimodal initial features, realizing the extraction and preliminary integration of effective multimodal features; it performs knowledge distillation on the multimodal initial features to obtain multimodal knowledge-enhanced features, enabling the multimodal initial features to learn more implicit knowledge, thus improving the quality and expressive power of the features; it constructs a hierarchical knowledge graph based on the multimodal knowledge-enhanced features, which can represent multimodal knowledge in a structured way from different levels, providing a clear and organized knowledge foundation for subsequent reasoning and decision-making; and it performs hierarchical reasoning and decision-making based on the multimodal knowledge-enhanced features and the hierarchical knowledge graph to obtain target decision data based on the target input data, enabling the analysis and decision-making of multimodal information at different granularities, improving the processing ability and decision accuracy of complex tasks, and realizing effective reasoning from local features to global decisions. Attached Figure Description

[0020] Figure 1 This is a flowchart of a preferred embodiment of the decision generation method based on knowledge distillation of the present invention.

[0021] Figure 2This is a functional block diagram of a preferred embodiment of the decision generation device based on knowledge distillation of the present invention.

[0022] Figure 3 This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the knowledge distillation-based decision generation method of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0024] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the decision generation method based on knowledge distillation of the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.

[0025] The knowledge distillation-based decision generation method is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0026] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.

[0027] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0028] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0029] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0030] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0031] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).

[0032] S10, in response to the decision generation instruction based on the target input data, extract the multimodal features from the target input data to obtain the initial multimodal features.

[0033] In this embodiment, the target input data may be the collected visual images, voice text, and action execution data of the robot.

[0034] In this embodiment, the decision generation instruction can be automatically triggered when the target input data is detected to be input to the designated platform.

[0035] In this embodiment, extracting multimodal features from the target input data to obtain initial multimodal features includes:

[0036] Feature extraction is performed on the image data in the target input data using the EfficientNet-V2 (Efficient Neural Network Version 2) network to obtain an initial visual feature vector; global semantic features of the initial visual feature vector are extracted using the Vision Transformer (ViT) to obtain the target visual feature vector; wherein, when extracting the global semantic features, a spatial attention mechanism is used to highlight the key object regions of the initial visual feature vector.

[0037] The text data in the target input data is feature extracted using a BERT-Large (Bidirectional Encoder Representations from Transformers Large-scale version) pre-trained language model to obtain the semantic vector representation of each word; the syntactic structure and semantic relationship corresponding to the semantic vector representation of each word are extracted through syntactic analysis and semantic role labeling strategies to obtain the target text feature vector.

[0038] The action data in the target input data is extracted using a Temporal Convolutional Network (TCN) to obtain the target action feature vector;

[0039] The target visual feature vector, the target text feature vector, and the target action feature vector are time-stamp aligned.

[0040] The target visual feature vector, the target text feature vector, and the target action feature vector are concatenated and aligned to obtain the multimodal initial features.

[0041] The initial visual feature vector may include, but is not limited to, object appearance, texture, spatial location, etc.

[0042] Among these, highlighting key object regions through spatial attention mechanisms can enhance the expressive power of visual features.

[0043] Before using the BERT-Large pre-trained language model to extract features from the text data in the target input data, preprocessing such as word segmentation and magnetic annotation can be performed on the text data in the target input data to improve data quality.

[0044] The motion data can be collected using sensor devices such as inertial sensors and joint angle sensors.

[0045] Before using a temporal convolutional network to extract features from the action data in the target input data, the action data can be filtered and smoothed to improve data continuity and reduce data noise.

[0046] The target action feature vector can characterize action speed, acceleration changes, and action sequence patterns.

[0047] For example, in financial risk control scenarios, features can be extracted from customer facial expression images (i.e., image data), loan application text (i.e., text data), and gestures made when filling out the application (i.e., action data), and the extracted multimodal features can be fused. In the medical field, features can be extracted from patient medical images (i.e., image data), medical record text (i.e., text data), and patient limb movements (such as walking posture and other action data), and the extracted multimodal features can be fused.

[0048] Through the above embodiments, effective feature extraction and preliminary integration of visual, linguistic text, and action modal data were achieved, providing basic feature support for subsequent knowledge mining and reasoning decision-making, enhancing the expressive power of each modal feature, and making the features better reflect the essential information of the data.

[0049] S11, perform knowledge distillation on the initial multimodal features to obtain multimodal knowledge-enhanced features.

[0050] In this embodiment, in order to further improve data quality, knowledge distillation is also required for the initial multimodal features.

[0051] Specifically, the step of performing knowledge distillation on the initial multimodal features to obtain multimodal knowledge-enhanced features includes:

[0052] A knowledge distillation network is constructed based on the target visual feature vector to obtain a visual knowledge distillation network.

[0053] A knowledge distillation network is constructed based on the target text feature vector to obtain the text knowledge distillation network;

[0054] A knowledge distillation network is constructed based on the target action feature vector to obtain the action knowledge distillation network.

[0055] The multimodal initial features are input into the visual knowledge distillation network to obtain visual knowledge enhanced features;

[0056] The multimodal initial features are input into the text knowledge distillation network to obtain text knowledge enhancement features;

[0057] The multimodal initial features are input into the action knowledge distillation network to obtain action knowledge enhanced features;

[0058] The multimodal knowledge enhancement features are obtained by fusing the visual knowledge enhancement features, the text knowledge enhancement features, and the action knowledge enhancement features.

[0059] Specifically, the step of constructing a knowledge distillation network based on the target visual feature vector to obtain a visual knowledge distillation network includes:

[0060] Obtain the student network and the pre-built teacher network;

[0061] The target visual feature vector is input into the teacher network and the student network respectively to obtain the first output feature of the teacher network and the second output feature of the student network;

[0062] Construct a distillation loss function based on the difference between the first output feature and the second output feature;

[0063] The student network and the teacher network are trained based on the distillation loss function;

[0064] When the value of the distillation loss function no longer decreases, training stops, and the currently obtained student network is determined as the visual knowledge distillation network.

[0065] The teacher network can be a pre-built ResNeXt-101 network (ResNeXt-101 Convolutional Network, the 101st layer of the residual network convolutional neural network).

[0066] The student network refers to the original feature extraction network that needs to be trained based on the teacher network so that it possesses the knowledge of the teacher network.

[0067] The difference between the first output feature and the second output feature can be calculated using the mean square error function.

[0068] During training, the value of the distillation loss function is minimized so that the student network learns the knowledge from the teacher network.

[0069] The text knowledge distillation network and the action knowledge distillation network can be trained in a similar manner, which will not be elaborated here.

[0070] For example, in financial stock prediction tasks, by utilizing the knowledge of existing high-performance stock analysis models (i.e., teacher networks), student networks can learn more about the implicit patterns of stock market trends through knowledge distillation, such as the deep correlation between different economic indicators and stock price fluctuations, which can improve the accuracy of student networks in predicting stock prices. In tasks that assist in medical testing, student networks learn the knowledge of mature disease prediction models (i.e., teacher networks), which enables student networks to better understand the correlation between subtle features in medical images and diseases, as well as the hidden symptom information in medical records, thereby improving the accuracy of testing.

[0071] Through the above embodiments, student networks can learn from the rich knowledge contained in teacher networks, allowing multimodal initial features to contain more implicit knowledge, enhancing the model's ability to extract and learn deep knowledge from multimodal data, thereby improving the quality and expressive power of features.

[0072] S12, construct a hierarchical knowledge graph based on the multimodal knowledge enhancement features.

[0073] In this embodiment, constructing a hierarchical knowledge graph based on the multimodal knowledge enhancement features includes:

[0074] Entity recognition is performed on the multimodal knowledge enhancement features to obtain multiple entities; wherein, the multiple entities include object entities, concept entities, and operation object entities;

[0075] Explore the relationships between the multiple entities;

[0076] The hierarchical knowledge graph is obtained by using graph neural networks to aggregate and abstract knowledge about the multiple entities and the relationships between them.

[0077] Among them, named entity recognition and object detection technologies can be used for entity recognition.

[0078] Among these methods, syntactic analysis, visual relationship detection (such as spatial relationships between objects), and action logic analysis (such as the sequence of actions and causal relationships) can be used to uncover relationships between entities, such as "on," "cause," and "execution object."

[0079] Based on the identified low-level entities and the mined mid-level relationships, graph neural networks can be used for knowledge aggregation and abstraction to extract higher-level knowledge concepts, such as "assembly process" and "operation specifications," thereby forming a hierarchical knowledge graph.

[0080] Through the above embodiments, a hierarchical knowledge graph is formed, which can represent multimodal knowledge in a structured way at different levels, providing a clear and organized knowledge foundation for subsequent reasoning and decision-making.

[0081] S13, perform hierarchical reasoning and decision-making based on the multimodal knowledge enhancement features and the hierarchical knowledge graph to obtain target decision data based on the target input data.

[0082] In this embodiment, the step of performing hierarchical reasoning and decision-making based on the multimodal knowledge enhancement features and the hierarchical knowledge graph to obtain target decision data based on the target input data includes:

[0083] The multimodal knowledge enhancement features are input into the Transformer-based low-level inference network for inference, and the hierarchical knowledge graph is combined with the hierarchical knowledge graph for local feature inference during the inference process to obtain the low-level inference result.

[0084] Based on the underlying reasoning results and the hierarchical knowledge graph, mid-level relationship reasoning is performed using a graph attention network to obtain mid-level reasoning results; wherein, the mid-level reasoning results include the interaction relationships between objects, the logical relationships in language descriptions, and the connection relationships between actions;

[0085] The mid-level reasoning results are input into the high-level decision network for reasoning, and the hierarchical knowledge graph is combined with the hierarchical knowledge graph to make global decisions, thereby obtaining the target decision data.

[0086] Specifically, the multimodal knowledge enhancement features can be input into a Transformer-based low-level inference network, combining low-level conceptual information from a hierarchical knowledge graph for local feature inference, such as identifying the category of specific objects, the basic semantics of language words, and the meaning of individual actions. Further, based on the low-level inference results and mid-level relational information from the knowledge graph, mid-level relational inference is performed through a graph attention network to analyze the interaction relationships between objects, the logical relationships in language descriptions, and the connection relationships between actions, resulting in a deeper semantic understanding. Going further, the mid-level inference results are input into a high-level decision network, combining high-level knowledge abstraction information from the knowledge graph, and using reinforcement learning algorithms for global decision-making inference.

[0087] For example, in the financial and medical fields, robots are being used more and more widely. This embodiment, through deep and accurate reasoning and decision-making, can generate optimal action sequence decisions or verbal responses, thereby improving the robot's responsiveness. For instance, financial service robots can respond promptly to customers based on generated verbal responses, and surgical assistance robots can more accurately deliver surgical tools to medical staff through generated action sequence decisions.

[0088] Through the above embodiments, the model is able to analyze and make decisions on multimodal information at different granularities, and can flexibly adjust the inference strategy according to task requirements, thereby improving the model's processing ability and decision accuracy in complex tasks and realizing effective inference from local features to global decisions.

[0089] In this embodiment, after obtaining the target decision data based on the target input data, the method further includes:

[0090] At preset time intervals, the execution result feedback data and newly added multimodal features of the target decision data are obtained;

[0091] The hierarchical knowledge graph, the low-level inference network, the graph attention network, and the high-level decision network are optimized using the execution result feedback data and the newly added multimodal features.

[0092] The preset time interval can be 30 days, and can be configured according to the actual needs of the scenario.

[0093] For example, feedback on the decision-making results of the model used in hierarchical reasoning can be collected, such as whether the action was successfully executed and whether the verbal response was accurate. Based on the feedback information and the task objective, the model's loss function is calculated, including feature extraction error, inference result error, and decision execution error. Furthermore, using the backpropagation algorithm, the model is trained based on the loss function, updating the parameters of modules such as the multimodal knowledge distillation network and the hierarchical reasoning network, thereby optimizing model performance. Moreover, with the continuous input of new data and the accumulation of task experience, the hierarchical knowledge graph can be updated, adding new concept nodes and relational edges, and correcting erroneous knowledge information.

[0094] Through the above embodiments, the model can continuously absorb new knowledge and experience, continuously optimize its performance, enhance the model's generalization ability and practicality, and enable it to adapt to the changing needs of different scenarios and tasks.

[0095] As can be seen from the above technical solutions, this invention can extract multimodal features from target input data to obtain multimodal initial features, realizing the extraction and preliminary integration of effective multimodal features; it performs knowledge distillation on the multimodal initial features to obtain multimodal knowledge-enhanced features, enabling the multimodal initial features to learn more implicit knowledge, thus improving the quality and expressive power of the features; it constructs a hierarchical knowledge graph based on the multimodal knowledge-enhanced features, which can represent multimodal knowledge in a structured way from different levels, providing a clear and organized knowledge foundation for subsequent reasoning and decision-making; and it performs hierarchical reasoning and decision-making based on the multimodal knowledge-enhanced features and the hierarchical knowledge graph to obtain target decision data based on the target input data, enabling the analysis and decision-making of multimodal information at different granularities, improving the processing ability and decision accuracy of complex tasks, and realizing effective reasoning from local features to global decisions.

[0096] like Figure 2 The diagram shown is a functional block diagram of a preferred embodiment of the decision generation device based on knowledge distillation of the present invention. The decision generation device 11 based on knowledge distillation includes an extraction unit 110, a distillation unit 111, a construction unit 112, and a decision unit 113. The module / unit referred to in this invention refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.

[0097] The extraction unit 110 is used to extract multimodal features from the target input data in response to a decision generation instruction based on the target input data, and obtain multimodal initial features.

[0098] In this embodiment, the target input data may be the collected visual images, voice text, and action execution data of the robot.

[0099] In this embodiment, the decision generation instruction can be automatically triggered when the target input data is detected to be input to the designated platform.

[0100] In this embodiment, the extraction unit 110 extracts multimodal features from the target input data to obtain initial multimodal features, including:

[0101] Feature extraction is performed on the image data in the target input data using the EfficientNet-V2 (Efficient Neural Network Version 2) network to obtain an initial visual feature vector; global semantic features of the initial visual feature vector are extracted using the Vision Transformer (ViT) to obtain the target visual feature vector; wherein, when extracting the global semantic features, a spatial attention mechanism is used to highlight the key object regions of the initial visual feature vector.

[0102] The text data in the target input data is feature extracted using a BERT-Large (Bidirectional Encoder Representations from Transformers Large-scale version) pre-trained language model to obtain the semantic vector representation of each word; the syntactic structure and semantic relationship corresponding to the semantic vector representation of each word are extracted through syntactic analysis and semantic role labeling strategies to obtain the target text feature vector.

[0103] The action data in the target input data is extracted using a Temporal Convolutional Network (TCN) to obtain the target action feature vector;

[0104] The target visual feature vector, the target text feature vector, and the target action feature vector are time-stamp aligned.

[0105] The target visual feature vector, the target text feature vector, and the target action feature vector are concatenated and aligned to obtain the multimodal initial features.

[0106] The initial visual feature vector may include, but is not limited to, object appearance, texture, spatial location, etc.

[0107] Among these, highlighting key object regions through spatial attention mechanisms can enhance the expressive power of visual features.

[0108] Before using the BERT-Large pre-trained language model to extract features from the text data in the target input data, preprocessing such as word segmentation and magnetic annotation can be performed on the text data in the target input data to improve data quality.

[0109] The motion data can be collected using sensor devices such as inertial sensors and joint angle sensors.

[0110] Before using a temporal convolutional network to extract features from the action data in the target input data, the action data can be filtered and smoothed to improve data continuity and reduce data noise.

[0111] The target action feature vector can characterize action speed, acceleration changes, and action sequence patterns.

[0112] For example, in financial risk control scenarios, features can be extracted from customer facial expression images (i.e., image data), loan application text (i.e., text data), and gestures made when filling out the application (i.e., action data), and the extracted multimodal features can be fused. In the medical field, features can be extracted from patient medical images (i.e., image data), medical record text (i.e., text data), and patient limb movements (such as walking posture and other action data), and the extracted multimodal features can be fused.

[0113] Through the above embodiments, effective feature extraction and preliminary integration of visual, linguistic text, and action modal data were achieved, providing basic feature support for subsequent knowledge mining and reasoning decision-making, enhancing the expressive power of each modal feature, and making the features better reflect the essential information of the data.

[0114] The distillation unit 111 is used to perform knowledge distillation on the initial multimodal features to obtain multimodal knowledge-enhanced features.

[0115] In this embodiment, in order to further improve data quality, knowledge distillation is also required for the initial multimodal features.

[0116] Specifically, the distillation unit 111 performs knowledge distillation on the initial multimodal features to obtain multimodal knowledge-enhanced features, including:

[0117] A knowledge distillation network is constructed based on the target visual feature vector to obtain a visual knowledge distillation network.

[0118] A knowledge distillation network is constructed based on the target text feature vector to obtain the text knowledge distillation network;

[0119] A knowledge distillation network is constructed based on the target action feature vector to obtain the action knowledge distillation network.

[0120] The multimodal initial features are input into the visual knowledge distillation network to obtain visual knowledge enhanced features;

[0121] The multimodal initial features are input into the text knowledge distillation network to obtain text knowledge enhancement features;

[0122] The multimodal initial features are input into the action knowledge distillation network to obtain action knowledge enhanced features;

[0123] The multimodal knowledge enhancement features are obtained by fusing the visual knowledge enhancement features, the text knowledge enhancement features, and the action knowledge enhancement features.

[0124] Specifically, the step of constructing a knowledge distillation network based on the target visual feature vector to obtain a visual knowledge distillation network includes:

[0125] Obtain the student network and the pre-built teacher network;

[0126] The target visual feature vector is input into the teacher network and the student network respectively to obtain the first output feature of the teacher network and the second output feature of the student network;

[0127] Construct a distillation loss function based on the difference between the first output feature and the second output feature;

[0128] The student network and the teacher network are trained based on the distillation loss function;

[0129] When the value of the distillation loss function no longer decreases, training stops, and the currently obtained student network is determined as the visual knowledge distillation network.

[0130] The teacher network can be a pre-built ResNeXt-101 network (ResNeXt-101 Convolutional Network, the 101st layer of the residual network convolutional neural network).

[0131] The student network refers to the original feature extraction network that needs to be trained based on the teacher network so that it possesses the knowledge of the teacher network.

[0132] The difference between the first output feature and the second output feature can be calculated using the mean square error function.

[0133] During training, the value of the distillation loss function is minimized so that the student network learns the knowledge from the teacher network.

[0134] The text knowledge distillation network and the action knowledge distillation network can be trained in a similar manner, which will not be elaborated here.

[0135] For example, in financial stock prediction tasks, by utilizing the knowledge of existing high-performance stock analysis models (i.e., teacher networks), student networks can learn more about the implicit patterns of stock market trends through knowledge distillation, such as the deep correlation between different economic indicators and stock price fluctuations, which can improve the accuracy of student networks in predicting stock prices. In tasks that assist in medical testing, student networks learn the knowledge of mature disease prediction models (i.e., teacher networks), which enables student networks to better understand the correlation between subtle features in medical images and diseases, as well as the hidden symptom information in medical records, thereby improving the accuracy of testing.

[0136] Through the above embodiments, student networks can learn from the rich knowledge contained in teacher networks, allowing multimodal initial features to contain more implicit knowledge, enhancing the model's ability to extract and learn deep knowledge from multimodal data, thereby improving the quality and expressive power of features.

[0137] The construction unit 112 is used to construct a hierarchical knowledge graph based on the multimodal knowledge enhancement features.

[0138] In this embodiment, the construction unit 112 constructs a hierarchical knowledge graph based on the multimodal knowledge enhancement features, including:

[0139] Entity recognition is performed on the multimodal knowledge enhancement features to obtain multiple entities; wherein, the multiple entities include object entities, concept entities, and operation object entities;

[0140] Explore the relationships between the multiple entities;

[0141] The hierarchical knowledge graph is obtained by using graph neural networks to aggregate and abstract knowledge about the multiple entities and the relationships between them.

[0142] Among them, named entity recognition and object detection technologies can be used for entity recognition.

[0143] Among these methods, syntactic analysis, visual relationship detection (such as spatial relationships between objects), and action logic analysis (such as the sequence of actions and causal relationships) can be used to uncover relationships between entities, such as "on," "cause," and "execution object."

[0144] Based on the identified low-level entities and the mined mid-level relationships, graph neural networks can be used for knowledge aggregation and abstraction to extract higher-level knowledge concepts, such as "assembly process" and "operation specifications," thereby forming a hierarchical knowledge graph.

[0145] Through the above embodiments, a hierarchical knowledge graph is formed, which can represent multimodal knowledge in a structured way at different levels, providing a clear and organized knowledge foundation for subsequent reasoning and decision-making.

[0146] The decision unit 113 is used to perform hierarchical reasoning and decision-making based on the multimodal knowledge enhancement features and the hierarchical knowledge graph to obtain target decision data based on the target input data.

[0147] In this embodiment, the decision unit 113 performs hierarchical reasoning decision-making based on the multimodal knowledge enhancement features and the hierarchical knowledge graph to obtain target decision data based on the target input data, including:

[0148] The multimodal knowledge enhancement features are input into the Transformer-based low-level inference network for inference, and the hierarchical knowledge graph is combined with the hierarchical knowledge graph for local feature inference during the inference process to obtain the low-level inference result.

[0149] Based on the underlying reasoning results and the hierarchical knowledge graph, mid-level relationship reasoning is performed using a graph attention network to obtain mid-level reasoning results; wherein, the mid-level reasoning results include the interaction relationships between objects, the logical relationships in language descriptions, and the connection relationships between actions;

[0150] The mid-level reasoning results are input into the high-level decision network for reasoning, and the hierarchical knowledge graph is combined with the hierarchical knowledge graph to make global decisions, thereby obtaining the target decision data.

[0151] Specifically, the multimodal knowledge enhancement features can be input into a Transformer-based low-level inference network, combining low-level conceptual information from a hierarchical knowledge graph for local feature inference, such as identifying the category of specific objects, the basic semantics of language words, and the meaning of individual actions. Further, based on the low-level inference results and mid-level relational information from the knowledge graph, mid-level relational inference is performed through a graph attention network to analyze the interaction relationships between objects, the logical relationships in language descriptions, and the connection relationships between actions, resulting in a deeper semantic understanding. Going further, the mid-level inference results are input into a high-level decision network, combining high-level knowledge abstraction information from the knowledge graph, and using reinforcement learning algorithms for global decision-making inference.

[0152] For example, in the financial and medical fields, robots are being used more and more widely. This embodiment, through deep and accurate reasoning and decision-making, can generate optimal action sequence decisions or verbal responses, thereby improving the robot's responsiveness. For instance, financial service robots can respond promptly to customers based on generated verbal responses, and surgical assistance robots can more accurately deliver surgical tools to medical staff through generated action sequence decisions.

[0153] Through the above embodiments, the model is able to analyze and make decisions on multimodal information at different granularities, and can flexibly adjust the inference strategy according to task requirements, thereby improving the model's processing ability and decision accuracy in complex tasks and realizing effective inference from local features to global decisions.

[0154] In this embodiment, after obtaining the target decision data based on the target input data, the execution result feedback data and newly added multimodal features of the target decision data are obtained at preset time intervals.

[0155] The hierarchical knowledge graph, the low-level inference network, the graph attention network, and the high-level decision network are optimized using the execution result feedback data and the newly added multimodal features.

[0156] The preset time interval can be 30 days, and can be configured according to the actual needs of the scenario.

[0157] For example, feedback on the decision-making results of the model used in hierarchical reasoning can be collected, such as whether the action was successfully executed and whether the verbal response was accurate. Based on the feedback information and the task objective, the model's loss function is calculated, including feature extraction error, inference result error, and decision execution error. Furthermore, using the backpropagation algorithm, the model is trained based on the loss function, updating the parameters of modules such as the multimodal knowledge distillation network and the hierarchical reasoning network, thereby optimizing model performance. Moreover, with the continuous input of new data and the accumulation of task experience, the hierarchical knowledge graph can be updated, adding new concept nodes and relational edges, and correcting erroneous knowledge information.

[0158] Through the above embodiments, the model can continuously absorb new knowledge and experience, continuously optimize its performance, enhance the model's generalization ability and practicality, and enable it to adapt to the changing needs of different scenarios and tasks.

[0159] As can be seen from the above technical solutions, this invention can extract multimodal features from target input data to obtain multimodal initial features, realizing the extraction and preliminary integration of effective multimodal features; it performs knowledge distillation on the multimodal initial features to obtain multimodal knowledge-enhanced features, enabling the multimodal initial features to learn more implicit knowledge, thus improving the quality and expressive power of the features; it constructs a hierarchical knowledge graph based on the multimodal knowledge-enhanced features, which can represent multimodal knowledge in a structured way from different levels, providing a clear and organized knowledge foundation for subsequent reasoning and decision-making; and it performs hierarchical reasoning and decision-making based on the multimodal knowledge-enhanced features and the hierarchical knowledge graph to obtain target decision data based on the target input data, enabling the analysis and decision-making of multimodal information at different granularities, improving the processing ability and decision accuracy of complex tasks, and realizing effective reasoning from local features to global decisions.

[0160] like Figure 3 The diagram shown is a schematic representation of the structure of a computer device that implements a preferred embodiment of the knowledge distillation-based decision generation method of the present invention.

[0161] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a decision generation program based on knowledge distillation.

[0162] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.

[0163] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.

[0164] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as the code of a decision generation program based on knowledge distillation, but also to temporarily store data that has been output or will be output.

[0165] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., executing decision generation programs based on knowledge distillation) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.

[0166] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the various knowledge distillation-based decision generation method embodiments described above, for example... Figure 1 The steps are shown.

[0167] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing a specific function, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into an extraction unit 110, a distillation unit 111, a construction unit 112, and a decision-making unit 113.

[0168] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the knowledge distillation-based decision generation method described in the various embodiments of this invention.

[0169] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.

[0170] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.

[0171] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.

[0172] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0173] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.

[0174] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0175] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish a communication connection between the computer device 1 and other computer devices.

[0176] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.

[0177] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.

[0178] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0179] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement a decision generation method based on knowledge distillation, and the processor 13 can execute the multiple instructions to achieve:

[0180] In response to a decision generation instruction based on target input data, multimodal features are extracted from the target input data to obtain initial multimodal features;

[0181] Knowledge distillation is performed on the initial multimodal features to obtain multimodal knowledge-enhanced features;

[0182] A hierarchical knowledge graph is constructed based on the multimodal knowledge enhancement features;

[0183] Based on the multimodal knowledge enhancement features and the hierarchical knowledge graph, hierarchical reasoning and decision-making are performed to obtain target decision data based on the target input data.

[0184] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.

[0185] It should be noted that all data involved in this case was legally obtained. Software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0186] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.

[0187] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0188] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0189] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0190] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0191] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.

[0192] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.

[0193] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A decision generation method based on knowledge distillation, characterized in that, The decision generation method based on knowledge distillation includes: In response to a decision generation instruction based on target input data, multimodal features are extracted from the target input data to obtain initial multimodal features; Knowledge distillation is performed on the initial multimodal features to obtain multimodal knowledge-enhanced features; A hierarchical knowledge graph is constructed based on the multimodal knowledge enhancement features; Based on the multimodal knowledge enhancement features and the hierarchical knowledge graph, hierarchical reasoning and decision-making are performed to obtain target decision data based on the target input data.

2. The decision generation method based on knowledge distillation as described in claim 1, characterized in that, The step of extracting multimodal features from the target input data to obtain initial multimodal features includes: The image data in the target input data is used to extract features based on the EfficientNet-V2 network to obtain an initial visual feature vector; the global semantic features of the initial visual feature vector are extracted based on the visual Transformer to obtain the target visual feature vector; wherein, when extracting the global semantic features, the key object regions of the initial visual feature vector are highlighted through a spatial attention mechanism. The BERT-Large pre-trained language model is used to extract features from the text data in the target input data to obtain the semantic vector representation of each word; the syntactic structure and semantic relationship corresponding to the semantic vector representation of each word are extracted through syntactic analysis and semantic role labeling strategies to obtain the target text feature vector. The action data in the target input data is extracted using a temporal convolutional network to obtain the target action feature vector; The target visual feature vector, the target text feature vector, and the target action feature vector are time-stamp aligned. The target visual feature vector, the target text feature vector, and the target action feature vector are concatenated and aligned to obtain the multimodal initial features.

3. The decision generation method based on knowledge distillation as described in claim 2, characterized in that, The step of performing knowledge distillation on the initial multimodal features to obtain multimodal knowledge-enhanced features includes: A knowledge distillation network is constructed based on the target visual feature vector to obtain a visual knowledge distillation network. A knowledge distillation network is constructed based on the target text feature vector to obtain the text knowledge distillation network; A knowledge distillation network is constructed based on the target action feature vector to obtain the action knowledge distillation network. The multimodal initial features are input into the visual knowledge distillation network to obtain visual knowledge enhanced features; The multimodal initial features are input into the text knowledge distillation network to obtain text knowledge enhancement features; The multimodal initial features are input into the action knowledge distillation network to obtain action knowledge enhanced features; The multimodal knowledge enhancement features are obtained by fusing the visual knowledge enhancement features, the text knowledge enhancement features, and the action knowledge enhancement features.

4. The decision generation method based on knowledge distillation as described in claim 3, characterized in that, The step of constructing a knowledge distillation network based on the target visual feature vector to obtain a visual knowledge distillation network includes: Obtain the student network and the pre-built teacher network; The target visual feature vector is input into the teacher network and the student network respectively to obtain the first output feature of the teacher network and the second output feature of the student network; Construct a distillation loss function based on the difference between the first output feature and the second output feature; The student network and the teacher network are trained based on the distillation loss function; When the value of the distillation loss function no longer decreases, training stops, and the currently obtained student network is determined as the visual knowledge distillation network.

5. The decision generation method based on knowledge distillation as described in claim 1, characterized in that, The construction of a hierarchical knowledge graph based on the multimodal knowledge enhancement features includes: Entity recognition is performed on the multimodal knowledge enhancement features to obtain multiple entities; wherein, the multiple entities include object entities, concept entities, and operation object entities; Explore the relationships between the multiple entities; The hierarchical knowledge graph is obtained by using graph neural networks to aggregate and abstract knowledge about the multiple entities and the relationships between them.

6. The decision generation method based on knowledge distillation as described in claim 1, characterized in that, The hierarchical reasoning and decision-making process based on the multimodal knowledge enhancement features and the hierarchical knowledge graph, to obtain target decision data based on the target input data, includes: The multimodal knowledge enhancement features are input into the Transformer-based low-level inference network for inference, and the hierarchical knowledge graph is combined with the hierarchical knowledge graph for local feature inference during the inference process to obtain the low-level inference result. Based on the underlying reasoning results and the hierarchical knowledge graph, mid-level relationship reasoning is performed using a graph attention network to obtain mid-level reasoning results; wherein, the mid-level reasoning results include the interaction relationships between objects, the logical relationships in language descriptions, and the connection relationships between actions; The mid-level reasoning results are input into the high-level decision network for reasoning, and the hierarchical knowledge graph is combined with the hierarchical knowledge graph to make global decisions, thereby obtaining the target decision data.

7. The decision generation method based on knowledge distillation as described in claim 6, characterized in that, After obtaining the target decision data based on the target input data, the method further includes: At preset time intervals, the execution result feedback data and newly added multimodal features of the target decision data are obtained; The hierarchical knowledge graph, the low-level inference network, the graph attention network, and the high-level decision network are optimized using the execution result feedback data and the newly added multimodal features.

8. A decision generation device based on knowledge distillation, characterized in that, The decision generation device based on knowledge distillation includes: An extraction unit is configured to extract multimodal features from the target input data in response to a decision generation instruction based on the target input data, thereby obtaining initial multimodal features; The distillation unit is used to perform knowledge distillation on the initial multimodal features to obtain multimodal knowledge-enhanced features; The construction unit is used to construct a hierarchical knowledge graph based on the multimodal knowledge enhancement features; The decision-making unit is used to perform hierarchical reasoning and decision-making based on the multimodal knowledge enhancement features and the hierarchical knowledge graph to obtain target decision data based on the target input data.

9. A computer device, characterized in that, The computer device includes: Memory, storing at least one instruction; and The processor executes instructions stored in the memory to implement the knowledge distillation-based decision generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the knowledge distillation-based decision generation method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Mechanical arm self-adaptive control method and system, readable storage medium and computer

    CN121200034A