Humanoid robot knowledge enhancement decision-making system

By combining the deep integration of industry knowledge graphs and real-time sensor data, the Q-learning algorithm based on knowledge constraints is adopted to build a dynamic knowledge constraint mechanism, which solves the problem of inaccurate decision-making of humanoid robots, improves the accuracy of independent decision-making and operation, and is suitable for scenarios such as home care and medical assistance.

CN120347750APending Publication Date: 2025-07-22HUIZHOU BEIJIABAO ROBOT CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510695069.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Due to decision-making problems, existing humanoid robots are unable to accurately complete instructions, resulting in insufficient reliability and operational accuracy of independent decision-making in complex environments.

Method used

The industry knowledge graph module is used to store structured industry knowledge, combine real-time sensor data to extract knowledge constraints through the SPARQL query engine, and dynamically adjust the Q-learning algorithm based on knowledge constraints, output optimization action instructions, and complete operations through the robotic arm execution module. The effect evaluation module generates feedback data to update the knowledge graph.

Benefits of technology

It effectively solves the problem of instruction execution errors caused by semantic understanding deviations in traditional methods, improves the reliability and operation accuracy of the robot's independent decision-making in complex dynamic environments, and realizes cross-scene migration capabilities, which are suitable for home care, medical assistance and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120347750A_ABST
    Figure CN120347750A_ABST
Patent Text Reader

Abstract

The invention provides a humanoid robot knowledge enhancement decision-making system. A dynamic knowledge constraint mechanism is constructed through fusion of an industry knowledge graph and real-time sensor data. The system comprises an industry knowledge graph module for storing structured operation knowledge, an SPARQL query engine for extracting knowledge constraints, sensor data for generating environment vectors through state feature coding, a Q-learning algorithm combined with the knowledge constraints for dynamically adjusting a Q value updating strategy, balancing exploration efficiency and a security boundary, and outputting an optimization action instruction. The action parameter adjusting module converts an instruction into a mechanical arm control signal, an execution result generates feedback data through the effect evaluation module, and the graph increment updating module is driven to achieve knowledge base iterative optimization. According to the architecture, the problem of execution errors caused by semantic deviation in a traditional method is effectively solved, and reliability and operation accuracy of autonomous decision making in a complex dynamic environment are enhanced through knowledge-guided reinforcement learning and closed-loop verification of physical operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robot intelligence, and particularly relates to a knowledge-enhanced decision-making system for humanoid robots. Background Art

[0002] With the rapid development of modern technology, the improvement of people's living standards, and the increasing requirements for the service industry, service robots have been rapidly developed and applied. Service robots can be roughly divided into household service robots, professional service robots, and entertainment service robots according to their uses. Currently, various service robots mainly provide services for specific scenarios. For example, the aging of the population in society is becoming more and more serious, and the number of disabled people in society is also high. Household robots are needed to provide daily care for the elderly and disabled people, and cleaning robots are needed to perform cleaning, washing, and other tasks of household hygiene. Due to decision-making problems, existing humanoid robots cannot accurately complete instructions. Summary of the Invention

[0003] The purpose of the present invention is to provide a knowledge-enhanced decision-making system for humanoid robots to solve the problems raised in the above background art.

[0004] To achieve the above purpose, the present invention provides the following technical solution: A knowledge-enhanced decision-making system for humanoid robots, comprising: an industry knowledge graph module for storing and updating structured industry knowledge related to robot operations; a real-time sensor data acquisition module for obtaining dynamic data of the environment and the robot body; a SPARQL query engine connected to the industry knowledge graph module for performing semantic queries to extract knowledge constraints; a state feature encoding module connected to the real-time sensor data acquisition module for generating machine-parsable state feature vectors; a Q-learning decision optimizer based on knowledge constraints for dynamically adjusting the Q-value update strategy according to the state feature vectors and combining the knowledge constraints, and outputting optimized action instructions; an action parameter adjustment module for converting the action instructions into robotic arm control parameters; a robotic arm execution module for completing physical operations based on the control parameters; an effect evaluation module for evaluating the execution results of the robotic arm and generating feedback data; and a graph incremental update module for realizing iterative update of the industry knowledge graph through version snapshot management based on the feedback data.

[0005] Preferably, the industry knowledge graph module stores multi-source heterogeneous data using a graph database.

[0006] Preferably, the SPARQL query engine generates context-related knowledge constraint conditions by dynamically binding real-time sensor data with entity relationships in the knowledge graph.

[0007] Preferably, in the knowledge-constrained Q-learning decision optimizer, the knowledge constraint is embedded in the Q-learning algorithm in the form of a reward function, specifically including: restricting the optional range of the action space through knowledge constraints; dynamically adjusting the reward weight according to the priority rules in the knowledge graph.

[0008] Preferably, the graph incremental update module uses a difference comparison algorithm to identify the knowledge change part, and realizes the rollback and traceability functions through version snapshots.

[0009] Preferably, the effect evaluation module includes a multi-modal data fusion unit for quantifying the execution effect by combining visual, haptic and environmental feedback data.

[0010] Preferably, the state feature encoding module uses an attention mechanism to dynamically allocate the fusion weights of sensor data and knowledge constraints.

[0011] Preferably, the knowledge-constrained Q-learning decision optimizer further includes: generating knowledge constraints through real-time sensor data and the knowledge graph; encoding the sensor data and knowledge constraints into state feature vectors; optimizing action decisions based on the knowledge-constrained Q-learning algorithm; and iteratively updating the industry knowledge graph according to the execution effect feedback.

[0012] Preferably, the generation of the knowledge constraint includes: using a SPARQL query engine to extract entity relationships related to the current scenario; and converting the entity relationships into boundary conditions and reward rules in the decision-making process.

[0013] Compared with the prior art, the beneficial effects of the present invention are:

[0014] Through the deep fusion of the industry knowledge graph and real-time sensor data, the present invention constructs a dynamic knowledge constraint mechanism, enabling the robot to perform decision-making reasoning based on structured industry knowledge, and effectively solving the problem of instruction execution errors caused by semantic understanding deviation in traditional methods.

[0015] The present invention adopts a knowledge-constrained Q-learning algorithm to realize the reward function, achieving the balance between exploration efficiency and safety boundary in the decision-making process.

[0016] The present invention constructs the continuous evolution ability of the knowledge graph through a closed-loop feedback mechanism, automatically updates the operation knowledge base using the execution effect evaluation data, and improves the task adaptation ability of the system in scenarios such as home care and precision operation. The system also realizes cross-scene migration through modular design, can be quickly adapted to 6 major fields such as medical assistance and smart home, and significantly improves the autonomous decision-making reliability and operation accuracy of service robots in complex dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a schematic diagram of the process of the present invention. Specific embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment 1:

[0020] As Figure 1As shown in the figure, a knowledge-enhanced decision-making system for a humanoid robot includes: an industry knowledge graph module for storing and updating structured industry knowledge related to robot operations; a real-time sensor data acquisition module for obtaining dynamic data of the environment and the robot body; a SPARQL query engine connected to the industry knowledge graph module for performing semantic queries to extract knowledge constraints; a state feature encoding module connected to the real-time sensor data acquisition module for generating machine-parsable state feature vectors; a Q-learning decision optimizer based on knowledge constraints that dynamically adjusts the Q-value update strategy according to the state feature vectors and combines the knowledge constraints to output optimized action instructions; an action parameter adjustment module for converting the action instructions into robotic arm control parameters; a robotic arm execution module for completing physical operations based on the control parameters; an effect evaluation module for evaluating the execution results of the robotic arm and generating feedback data; and a graph incremental update module for realizing iterative updates of the industry knowledge graph through version snapshot management based on the feedback data. The industry knowledge graph module stores multi-source heterogeneous data using a graph database. The SPARQL query engine generates context-related knowledge constraint conditions by dynamically binding real-time sensor data with entity relationships in the knowledge graph. In the Q-learning decision optimizer based on knowledge constraints, the knowledge constraints are embedded in the Q-learning algorithm in the form of a reward function, specifically including: restricting the optional range of the action space through knowledge constraints; dynamically adjusting the reward weights according to the priority rules in the knowledge graph. The graph incremental update module uses a difference comparison algorithm to identify the knowledge change parts and realizes the rollback and traceability functions through version snapshots. The effect evaluation module includes a multi-modal data fusion unit for quantifying the execution effect by combining visual, force sense, and environmental feedback data. The state feature encoding module uses an attention mechanism to dynamically allocate the fusion weights of sensor data and knowledge constraints. The Q-learning decision optimizer based on knowledge constraints further includes: generating knowledge constraints through real-time sensor data and the knowledge graph; encoding the sensor data and knowledge constraints into state feature vectors; optimizing action decisions based on the Q-learning algorithm with knowledge constraints; and iteratively updating the industry knowledge graph according to the execution effect feedback. The generation of the knowledge constraints includes: using the SPARQL query engine to extract entity relationships related to the current scenario; converting the entity relationships into boundary conditions and reward rules in the decision-making process.

[0021] Through the above technical solutions, the present invention deeply integrates the industry knowledge graph with real-time sensor data, constructs a dynamic knowledge constraint mechanism, enables the robot to perform decision-making reasoning based on structured industry knowledge, and effectively solves the problem of incorrect instruction execution caused by semantic understanding deviation in traditional methods.

[0022] The present invention adopts a Q-learning algorithm based on knowledge constraints and a reward function to achieve the balance between exploration efficiency and safety boundary in the decision-making process.

[0023] The present invention constructs the continuous evolution ability of the knowledge graph through a closed-loop feedback mechanism, automatically updates the operation knowledge base by using the execution effect evaluation data, and improves the task adaptation ability of the system in scenarios such as home care and precision operation. The system also realizes cross-scenario migration through modular design, can be quickly adapted to six major fields such as medical assistance and smart home, and significantly improves the reliability of autonomous decision-making and the operation accuracy of service robots in complex dynamic environments.

[0024] Embodiment 2:

[0025] As Figure 1 shown, the present invention includes an industry knowledge graph module, a real-time sensor data acquisition module, a SPARQL query engine, a state feature encoding module, a Q-learning decision optimizer based on knowledge constraints, an action parameter adjustment module, a robotic arm execution module, an effect evaluation module, and a graph incremental update module. The industry knowledge graph module stores the structured knowledge related to robot operations in a Neo4j graph database, including operation specifications, safety constraints, and best practices in different scenarios. The real-time sensor data acquisition module obtains environmental information and robot state data through various sensors such as infrared, ultrasonic, and vision, such as object position, temperature, humidity, and robot joint angles. The SPARQL query engine is connected to the industry knowledge graph module to execute semantic queries to extract knowledge constraints. The state feature encoding module converts the sensor data into a 256-dimensional state feature vector to accurately represent the current environment and robot state. The Q-learning decision optimizer based on knowledge constraints adopts a dual Q-network structure, dynamically adjusts the Q-value update strategy according to the state feature vector and knowledge constraints, and outputs optimized action instructions. The action parameter adjustment module converts the action instructions into specific robotic arm control parameters, such as movement trajectories, speeds, and forces. The robotic arm execution module receives the control parameters and drives the robotic arm to complete physical operations. The effect evaluation module evaluates the execution results through visual feedback and force feedback, and generates feedback data including indicators such as task completion rate, operation accuracy, and execution time. The graph incremental update module realizes the iterative update of the knowledge graph based on the execution feedback data by using the version snapshot management method to ensure that the knowledge base is synchronized with the actual operation environment.

[0026] In practical applications, the system can be used in the scenario of home care robots. When performing the task of "bringing tea to the elderly", the system first extracts relevant knowledge constraints from the knowledge graph, such as the tea water temperature not exceeding 60°C and the pouring height not exceeding 10 cm. Combining with real-time sensor data, the system generates an action sequence that meets the safety constraints, such as controlling the robotic arm to grab the teacup, moving it in front of the elderly, and slowly tilting to pour water. During the execution process, the system continuously monitors the environmental changes and dynamically adjusts the parameters to ensure the safety of the operation. After the task is completed, the system evaluates the execution effect and updates the knowledge graph to optimize the decision-making process for future similar tasks. Through this knowledge-driven decision-making mechanism, the system can safely and efficiently complete various service tasks in a complex and changing home environment, significantly improving the autonomous decision-making ability and operation accuracy of humanoid robots.

[0027] Embodiment III:

[0028] Such as Figure 1As shown in the figure, the industry knowledge graph module of the present invention uses a Neo4j graph database to store multi-source heterogeneous data. This graph database supports the property graph model and can flexibly represent complex entity relationship networks. During the construction of the knowledge graph, the system collects heterogeneous data from multiple sources, including structured operation manuals, semi-structured expert experience summaries, unstructured operation videos, etc. For structured data, the system directly maps it into nodes and relationships in the graph. For example, "cup" is used as a node, and related attributes such as "material" and "capacity" are used as attributes of the node. For semi-structured data, the system uses natural language processing technology to extract key information and transform it into a graph structure. For example, the knowledge "the distance between the cup mouth and the water surface should be kept at 5 - 10 cm when pouring water" is extracted from expert experience and represented as a relationship between the "pouring water" node and the "distance" node, while adding a numerical range attribute. For unstructured data, the system uses computer vision and deep learning technologies to identify key actions and objects from operation videos and transform them into nodes and edges in the graph. For example, the "grasping" action is identified and used as a node, which is associated with related object nodes. In the graph database, the system uses the RDF (Resource Description Framework) model to represent knowledge triples and uses OWL (Web Ontology Language) to define the ontology structure. This representation method allows the system to perform complex semantic reasoning, such as inferring the operation method of an object based on the type inheritance relationship. To process temporal information, the system introduces a timestamp attribute in the graph to support the temporal representation and query of the operation process. In addition, the system also establishes a multi-level index structure to optimize the query performance. Through the flexibility of the graph database, the system can effectively integrate knowledge from different sources and in different formats, providing a rich knowledge base for robot decision-making. In practical applications, when a robot needs to execute a new task, it can quickly retrieve relevant knowledge from the graph database, such as object attributes, operation constraints, best practices, etc., providing comprehensive reference information for decision-making. At the same time, the scalability of the graph database also enables the system to easily integrate new knowledge sources and continuously enrich and update the knowledge base. This multi-source heterogeneous data storage method based on the graph database provides a powerful and flexible knowledge management ability for the humanoid robot knowledge-enhanced decision-making system, effectively supporting the intelligent decision-making process in complex environments.

[0029] Embodiment 4:

[0030] As Figure 1 shown in the figure, the SPARQL query engine of the present invention generates context-related knowledge constraint conditions by dynamically binding real-time sensor data with entity relationships in the knowledge graph. The SPARQL query engine first receives the data stream from the real-time sensor data acquisition module, including environmental parameters and robot state information. After preprocessing, these data are converted into a series of key-value pairs, where the key represents the parameter type and the value represents the specific measurement result.

[0031] Meanwhile, the query engine maintains a dynamic variable mapping table that maps sensor data types to corresponding entities or relationships in the knowledge graph. When executing a query, the SPARQL query engine uses this mapping table to dynamically replace variables in the query statement. Specifically, the query engine adopts parameterized query technology, using placeholders to represent dynamic parameters in a predefined SPARQL template. When a query needs to be executed, the engine fills the real-time sensor data into these placeholders to generate a complete SPARQL query statement.

[0032] The query engine executes this dynamically generated SPARQL statement to obtain operation constraints related to the current environment from the knowledge graph. In addition to simple numerical comparisons, the query engine also supports more complex semantic associations.

[0033] For example, based on the current position of the robot, it can query the properties and operation methods of nearby objects. This involves the dynamic binding of spatial relationships, and the query engine uses a geographic coordinate system to convert real-time position data into spatial query conditions in the knowledge graph.

[0034] In addition, the query engine implements a query result caching mechanism, locally storing frequently used query results to reduce the overhead of repeated queries. The caching policy takes into account the timeliness of the data. For rapidly changing environmental parameters, a shorter caching period is adopted. For query results, the SPARQL query engine performs post-processing to convert the returned RDF triples into a structured data format that can be directly used by the robot decision-making system. This includes converting text descriptions into numerical ranges and simplifying complex relationship chains into direct constraints. Through this dynamic binding and query mechanism, the SPARQL query engine can generate knowledge constraint conditions highly relevant to the current environment and task in real time, providing accurate and timely knowledge support for the robot's decision-making process. This method significantly improves the efficiency and relevance of knowledge utilization, enabling the robot to quickly adjust its behavior strategy according to real-time situations and enhancing the system's adaptability in complex and dynamic environments.

[0035] Embodiment 5:

[0036] As Figure 1 shown, the Q-learning decision optimizer based on knowledge constraints of the present invention optimizes the decision-making process by embedding knowledge constraints in the form of a reward function into the Q-learning algorithm. The optimizer first extracts relevant domain knowledge from the knowledge graph, including action limit conditions, priority rules, etc. In the state-action value function Q(s, a) of Q-learning, a knowledge-based reward function R(s, a) is introduced, such that Q(s, a) = Q(s, a) + λR(s, a), where λ is the knowledge constraint weight.

[0037] The design of R(s, a) considers two aspects: action space constraint and reward weight adjustment. For action space constraint, R(s, a) assigns a large negative penalty to actions that violate the knowledge constraint, effectively excluding these actions from the optional range. For example, in a medical assistance scenario, the knowledge constraint may stipulate that the robot should not perform specific actions in certain body positions, and R(s, a) will assign a reward value of -∞ to these state-action pairs.

[0038] For reward weight adjustment, R(s, a) dynamically adjusts the reward weights of different actions according to the priority rules in the knowledge graph. The priority rules may come from expert knowledge or historical data analysis. For example, in a rehabilitation training task, the suitable training intensity for patients varies at different stages, and R(s, a) will adjust the reward weights of actions with different intensities according to the current rehabilitation stage.

[0039] In the iterative process of Q-learning, the optimizer continuously updates Q(s, a). In each iteration, the agent selects an action according to the ε-greedy policy, executes the action and observes the environmental feedback, and then updates the Q value: Q(s, a) ← Q(s, a) + α[r + γmax Q(s', a') - Q(s, a)] + λR(s, a). Where α is the learning rate, γ is the discount factor, and r is the environmental feedback reward.

[0040] In this way, while retaining the exploration ability of Q-learning, the optimizer uses knowledge constraints to guide the agent to avoid unreasonable actions and adjusts the decision-making preference according to the task requirements. This significantly improves the learning efficiency and decision-making quality. For example, in a home service robot scenario, the optimizer can quickly learn to avoid dangerous areas (such as stairs, fire sources) and give priority to action sequences preferred by users.

[0041] To adapt to the dynamic environment, the optimizer also includes an online update mechanism for knowledge constraints. When detecting environmental changes or receiving new expert guidance, the system updates the knowledge graph and accordingly adjusts the calculation method of R(s, a). This ensures that the decision-making process is always based on the latest domain knowledge.

[0042] In practical applications, the performance of the optimizer is affected by multiple hyperparameters, such as the knowledge constraint weight λ, the learning rate α, the exploration rate ε, etc. The system automatically tunes these parameters through methods such as cross-validation to achieve the best results in different tasks. For example, in a precision operation task, a larger λ may be required to strengthen the knowledge constraint; while in an open environment exploration task, a smaller λ may be required to retain more autonomous learning space.

[0043] The Q-learning decision optimizer based on knowledge constraints realizes intelligent decision-making guided by knowledge by integrating domain knowledge into the reinforcement learning process, significantly improving the task execution ability and adaptability of humanoid robots in complex environments.

[0044] Example Six:

[0045] As Figure 1 shown, the graph incremental update module of the present invention uses a difference comparison algorithm to identify the changed parts of knowledge and realizes the rollback and traceability functions through version snapshots. This module mainly consists of three units: a difference identification unit, a version management unit, and an update execution unit.

[0046] The difference identification unit is responsible for comparing the old and new knowledge graphs and identifying the changed parts. Specifically, this unit represents the knowledge graph as a set of triples (subject, relationship, object). For two versions of the knowledge graphs G1 and G2, the difference identification process is as follows:

[0047] 1. Construct hash tables H1 and H2, which store all the triples in G1 and G2 respectively.

[0048] 2. Traverse each triple t in H1:

[0049] - If t is not in H2, mark t as "deleted".

[0050] - If t is in H2 but the attribute values are different, mark t as "modified".

[0051] 3. Traverse each triple t in H2:

[0052] - If t is not in H1, mark t as "added".

[0053] 4. Summarize all the triples marked as "deleted", "modified", and "added" to form a difference report.

[0054] The version management unit is responsible for creating and managing version snapshots of the knowledge graph. Before each update operation, this unit creates a complete snapshot of the current knowledge graph and assigns a unique version number to it. The version information includes: version number, creation time, creation reason, summary of changed content, etc. All version snapshots are organized in a tree structure to support branch management.

[0055] The version management unit also provides a rollback function. When it is detected that the knowledge update causes a decline in system performance, the knowledge graph can be restored to any previous version through a simple instruction. The rollback operation includes:

[0056] 1. Load the snapshot of the specified version.

[0057] 2. Replace the current knowledge graph with the loaded snapshot.

[0058] 3. Update the version information and record the rollback operation.

[0059] The update execution unit is responsible for applying the identified changes to the current knowledge graph. The update process is as follows:

[0060] 1. Back up the current knowledge graph.

[0061] 2. Update the knowledge graph one by one according to the "delete", "modify", and "add" instructions in the difference report.

[0062] 3. Perform a consistency check on the updated knowledge graph to ensure that no contradictions or conflicts are introduced.

[0063] 4. If the consistency check passes, submit the update; otherwise, roll back to the backup version and report an error.

[0064] To improve the update efficiency, this module adopts the incremental indexing technology. A multi-level indexing structure, including entity index, relationship index, and attribute index, is constructed on the basis of the knowledge graph. The update operation only needs to modify the affected index items instead of reconstructing the entire index, which greatly improves the update speed of large-scale knowledge graphs.

[0065] This module also implements a concurrency control mechanism to support multiple clients to request updates simultaneously. Adopting the optimistic locking strategy, it allows concurrent reads, but checks for version conflicts when submitting updates. If a conflict is detected, the system will automatically merge the non-conflicting parts and submit the conflicting parts to manual arbitration.

[0066] To support knowledge traceability, this module maintains a change history for each knowledge point. By querying the change history, system administrators can trace the source and evolution process of any knowledge point, which helps to diagnose and optimize system behavior.

[0067] Generally speaking, the graph incremental update module enables the knowledge base of the humanoid robot to keep pace with the times through an efficient difference identification algorithm and a powerful version management function, while ensuring the stability and traceability of the system.

[0068] Example Seven:

[0069] As Figure 1 shown, the effect evaluation module of the present invention includes a multi-modal data fusion unit for quantifying the execution effect by combining visual, force sense, and environmental feedback data. This unit adopts a hierarchical fusion architecture, including a data preprocessing layer, a feature extraction layer, a data alignment layer, and a decision fusion layer.

[0070] In the data preprocessing layer, the system first performs noise reduction and standardization on the data from various modal sensors. For visual data, Gaussian filtering is used to remove image noise, followed by color space conversion and brightness equalization. Tactile data undergoes median filtering to eliminate spike noise and then zero-drift correction. Environmental feedback data (such as temperature, humidity, air pressure, etc.) is linearly calibrated according to the characteristics of each sensor.

[0071] The feature extraction layer extracts key features for different modal data. Convolutional neural network (CNN) models are used for visual feature extraction to extract high-level visual features including edges, textures, color distributions, etc. Time-frequency analysis methods are used for tactile feature extraction to calculate the time-domain statistical features (mean, variance, peak, etc.) and frequency-domain features (power spectral density, main frequency, etc.) of the force signal. The feature extraction of environmental feedback data mainly considers time series features such as trends, periodicity, and mutation points.

[0072] The data alignment layer addresses the issues of asynchrony and sampling rate differences among different modal data. The dynamic time warping (DTW) algorithm is used in this layer to map different modal data onto a unified time axis. For missing data points, interpolation algorithms are used to complete them. In addition, this layer also deals with the spatial registration problem between different sensors to ensure that all data corresponds to the correct positions in the robot coordinate system.

[0073] The decision fusion layer uses deep learning models to integrate multi-modal features and generate the final evaluation results of the execution effect. Specifically, an LSTM network with a multi-modal attention mechanism can dynamically adjust the weights of different modal data according to the task context. For example, in visually guided fine operation tasks, the model assigns higher weights to visual features; while in flexible object operation tasks dominated by force control, it relies more on tactile features.

[0074] To improve the interpretability of the model, the decision fusion layer also integrates a rule-based reasoning module. This module contains a series of evaluation rules defined by domain experts, such as "if the force feedback exceeds threshold X and the object position deviation is greater than Y, then the operation is judged as failed". The output of the deep learning model is compared with the results of rule-based reasoning. When significant differences occur, the system marks the relevant samples for key analysis.

[0075] The output of the effect evaluation includes multiple dimensions: task completion, operation accuracy, stability, efficiency, etc. Each dimension has corresponding quantitative indicators. For example, operation accuracy may include sub-indicators such as position error, attitude error, and force control error. The system dynamically adjusts the weights of each indicator according to the task type and calculates the comprehensive score.

[0076] To handle different task scenarios, this unit adopts transfer learning technology. First, the model is pre-trained on a large-scale general dataset, and then fine-tuned for specific tasks. This greatly reduces the training data requirements for new tasks and improves the adaptability of the system.

[0077] In addition, this unit also includes an adaptive learning mechanism. During continuous operation, the system will continuously accumulate evaluation samples and regularly update model parameters. This enables the evaluation model to adapt to environmental changes and dynamic changes in robot performance.

[0078] The effect evaluation results are not only used to judge the task completion situation, but also fed back to the decision optimizer and the knowledge graph update module. For example, if a certain type of action frequently leads to low scores, the decision optimizer will adjust the corresponding reward function; if there are significant differences between the evaluation results and the knowledge graph predictions, the system will trigger the knowledge update process.

[0079] The multi-modal data fusion unit provides a comprehensive and accurate execution effect evaluation for the humanoid robot by comprehensively analyzing visual, tactile, and environmental data, which is crucial for improving the robot's task adaptability and learning ability.

[0080] Embodiment VIII:

[0081] As Figure 1 shown, the state feature encoding module of the present invention adopts an attention mechanism to dynamically allocate the fusion weights of sensor data and knowledge constraints. Specifically, the state feature encoding module includes a multi-head attention layer and a feed-forward neural network layer. The multi-head attention layer consists of 8 attention heads, and the dimension of each attention head is 64. The inputs of the attention mechanism include real-time data from various sensors of the robot and relevant constraint information extracted from the knowledge graph.

[0082] During the processing, the multi-head attention layer first maps the input data into three matrices: query, key, and value. Then, by calculating the dot product between the query and the key, attention scores are obtained. These scores are normalized by the softmax function to generate attention weights. Next, these weights are multiplied by the value matrix to obtain the weighted feature representation. The outputs of the 8 attention heads are concatenated together and then passed through a linear transformation layer to form the final attention output.

[0083] The feed-forward neural network layer includes two fully-connected layers, using ReLU and linear activation functions respectively. The number of neurons in the first fully-connected layer is 4 times the input dimension, and the second fully-connected layer restores the dimension to the original input dimension. This structure allows the model to capture more complex feature interactions.

[0084] During the training process, the Adam optimizer is used for parameter updating. To prevent overfitting, dropout layers are added after both the attention layer and the feed-forward layer. In addition, layer normalization technology is applied to stabilize the training process.

[0085] The output of the state feature encoding module is a vector of a fixed dimension, which contains the fused information of sensor data and knowledge constraints. This vector will be used as the input for the subsequent decision optimizer. Through the attention mechanism, the system can dynamically adjust the attention to different input sources according to the current task and environment, so as to generate the most relevant and informative state representation.

[0086] The state feature encoding module realizes the dynamic fusion of sensor data and knowledge constraints through the attention mechanism, providing a high-quality state representation for subsequent decision optimization. This method not only improves the system's adaptability to complex environments, but also enhances the accuracy and efficiency of decision-making.

[0087] Embodiment Nine:

[0088] As Figure 1 shown, the working process of the Q-learning decision optimizer based on knowledge constraints of the present invention includes: knowledge constraint generation, state feature encoding, action decision optimization, and knowledge graph update.

[0089] In the knowledge constraint generation stage, the system comprehensively analyzes real-time sensor data and a pre-constructed industry knowledge graph. The sensor data includes multi-modal information such as vision, touch, and force feedback, which is collected in real time by various sensors of the robot. The knowledge graph stores the prior knowledge encoded by domain experts, including object attributes, operation rules, and safety restrictions. The system uses the knowledge graph reasoning engine to extract relevant entities and relationships from the knowledge graph based on the current sensor input, forming specific constraint conditions for the current scenario.

[0090] In the state feature encoding stage, the system converts the sensor data and the generated knowledge constraints into a unified vector representation. This process uses a deep neural network, including convolutional layers to process visual information, recurrent neural networks to process temporal data, and fully connected layers to fuse features of different modalities. The knowledge constraints are encoded through graph neural networks to capture the relationships between entities. Finally, all features are integrated into a state feature vector of a fixed dimension.

[0091] The third step is to optimize the action decision based on the knowledge constraint Q-learning algorithm. This algorithm is an improvement of traditional Q-learning, introducing knowledge constraints as an additional reward term.

[0092] Iteratively update the industry knowledge graph based on the feedback of the execution effect. The system records the results of each decision, including task completion, resource consumption, and safety metrics. After analysis, these data are used to update the entity attributes and relationship weights in the knowledge graph. The update process uses incremental learning methods to maintain the consistency and timeliness of the knowledge.

[0093] The Q-learning decision optimizer based on knowledge constraints realizes efficient and safe decision optimization by integrating real-time perception and prior knowledge. The dynamic knowledge update mechanism ensures that the system can continuously learn and adapt to new environments. This method greatly improves the autonomous decision-making ability of humanoid robots in complex and dynamic environments.

[0094] Example Ten:

[0095] As Figure 1 shown, the knowledge constraint generation of the present invention includes two main steps: using the SPARQL query engine to extract entity relationships related to the current scenario, and converting the extracted entity relationships into boundary conditions and reward rules in the decision-making process.

[0096] The system uses the SPARQL query engine to extract entity relationships related to the current scenario from the pre-constructed industry knowledge graph. The knowledge graph is stored in the RDF (Resource Description Framework) format and contains a large number of triples (subject-predicate-object). The SPARQL query language allows the system to construct complex query statements according to the current task and environmental state.

[0097] The system converts the extracted entity relationships into boundary conditions and reward rules in the decision-making process. This step involves semantic parsing and rule generation. The system uses predefined templates to interpret the SPARQL query results and generate corresponding constraint conditions.

[0098] These generated boundary conditions and reward rules directly affect the state transition function and reward function in the Q-learning algorithm. They ensure that the robot takes into account the characteristics of objects and the safety of operations during the decision-making process.

[0099] The dynamic generation of knowledge constraints greatly improves the adaptability of the system. When new objects or tasks are introduced into the environment, the system can automatically extract relevant information from the knowledge graph and generate appropriate constraints without manual intervention. This method enables the robot to safely and efficiently handle various complex scenarios.

[0100] The method of extracting entity relationships through the SPARQL query engine and converting them into decision constraints realizes the effective integration of the knowledge graph and real-time decision-making. This not only improves the accuracy and safety of decision-making but also enhances the interpretability and credibility of the system.

[0101] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0102] As described above, it is only used to illustrate the technical solution of the present invention rather than to limit it. Other modifications or equivalent replacements made by those of ordinary skill in the art to the technical solution of the present invention should be covered within the scope of the claims of the present invention as long as they do not depart from the spirit and scope of the technical solution of the present invention.

Claims

1. A knowledge-enhanced decision-making system for a humanoid robot, characterized in that, Including: An industry knowledge graph module for storing and updating structured industry knowledge related to robot operations; A real-time sensor data acquisition module for obtaining dynamic data of the environment and the robot body; A SPARQL query engine, connected to the industry knowledge graph module, for performing semantic queries to extract knowledge constraints; A status feature encoding module, connected to the real-time sensor data acquisition module, for generating machine-parsable status feature vectors; A knowledge-constraint-based Q-learning decision optimizer that dynamically adjusts the Q-value update strategy according to the status feature vectors and combines the knowledge constraints, and outputs optimized action instructions; An action parameter adjustment module for converting the action instructions into robotic arm control parameters; A robotic arm execution module that completes physical operations based on the control parameters; An effect evaluation module for evaluating the execution results of the robotic arm and generating feedback data; A graph incremental update module that realizes iterative update of the industry knowledge graph through version snapshot management based on the feedback data.

2. The knowledge enhanced decision-making system for a humanoid robot according to claim 1, characterized in that, The industry knowledge graph module stores multi-source heterogeneous data using a graph database.

3. A knowledge-enhanced decision-making system for a humanoid robot according to claim 1, characterized in that, The SPARQL query engine generates context-related knowledge constraint conditions by dynamically binding real-time sensor data and entity relationships in the knowledge graph.

4. The knowledge enhanced decision-making system of a humanoid robot according to claim 1, characterized in that, In the knowledge-constraint-based Q-learning decision optimizer, the knowledge constraints are embedded in the Q-learning algorithm in the form of a reward function, specifically including: Restricting the optional range of the action space through knowledge constraints; Dynamically adjusting the reward weights according to the priority rules in the knowledge graph.

5. The knowledge-enhanced decision-making system for a humanoid robot according to claim 1, wherein The graph incremental update module uses a difference comparison algorithm to identify the knowledge change part and realizes the rollback and traceability functions through version snapshots.

6. The knowledge enhanced decision-making system of a humanoid robot according to claim 1, wherein The effect evaluation module includes a multi-modal data fusion unit for quantifying the execution effect by combining visual, force sense, and environmental feedback data.

7. The knowledge enhanced decision-making system for a humanoid robot according to claim 1, wherein The status feature encoding module uses an attention mechanism to dynamically allocate the fusion weights of sensor data and knowledge constraints.

8. A knowledge-enhanced decision-making system for a humanoid robot according to claim 1, wherein The knowledge-constraint-based Q-learning decision optimizer further includes: Generating knowledge constraints through real-time sensor data and the knowledge graph; Encoding sensor data and knowledge constraints into status feature vectors; Optimizing action decisions based on the knowledge-constraint-based Q-learning algorithm; Iteratively updating the industry knowledge graph according to the execution effect feedback.

9. The knowledge enhanced decision-making system for a humanoid robot according to claim 8, wherein The generation of the knowledge constraints includes: Using the SPARQL query engine to extract entity relationships related to the current scenario; Converting the entity relationships into boundary conditions and reward rules in the decision-making process.

Citation Information

Cited By

  • Robot control method and device based on security enhancement, equipment and medium

    CN120871708A

  • Attention mechanism-based humanoid robot control method and related equipment

    CN121879126A

  • Robot electricity meter replacement operation decision-making method based on positive and negative reasoning of knowledge graph and electronic equipment thereof

    CN122242778A

  • A robot battery table operation decision-making method based on knowledge graph forward and backward reasoning and electronic equipment thereof

    CN122242778B