Robot visual identification decision control method based on deep learning

By constructing a multi-level thinking decision tree and multi-modal features, the dynamic correlation problem between natural language instructions and visual perception is solved, the robot can respond quickly and make accurate decisions in complex environments, and the execution stability of service robots is improved.

CN120244983AActive Publication Date: 2025-07-04青岛冠成软件有限公司

Patent Information

Application Number
CN202510597897.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-04
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

The existing technology has failed to effectively establish a dynamic correlation between natural language instructions and real-time visual perception, resulting in untimely responses from robots and rigid behavioral planning in complex environments.

Method used

Build a multi-level thinking decision tree structure, combine multimodal feature fusion and deep learning engine, integrate text instructions and visual image information by integrating search engines and semantic alignment mechanisms, establish a linear mapping relationship between top-level semantic nodes and middle-level perceptual nodes, and perform execution state optimization of the underlying behavior nodes.

Benefits of technology

It realizes the coordinated processing of accurate analysis of task intentions and environment perception, improves the real-time response and decision-making accuracy of service robots, enhances the dynamic adaptability of behavioral planning strategies, and improves execution stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244983A_ABST
    Figure CN120244983A_ABST
Patent Text Reader

Abstract

The invention discloses a robot visual identification decision control method based on deep learning, and relates to the technical field of control systems. Comprising the steps of constructing a thinking decision tree according to an input instruction and a real-time visual image, and obtaining a linear relationship between a middle-layer leaf node as a main body and a top-layer leaf node for realizing association between a current visual feature and a task intention. According to the method, deep fusion and unified representation of multi-source heterogeneous data are realized by constructing a multi-layer thinking decision tree structure and combining multi-modal feature fusion and a deep learning engine, and text instructions, visual images and other different modal information are effectively integrated by constructing a fusion search engine and a semantic alignment mechanism, so that the multi-source heterogeneous data fusion and unified representation are realized. Cooperative processing of accurate analysis of task intentions and environment perception is realized, and real-time response and decision accuracy of the service robot to complex tasks are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of control systems, and specifically provides a robot vision recognition decision control method based on deep learning. Background Art

[0002] With the continuous development of artificial intelligence, computer vision, and human-computer interaction technologies, service robots, industrial robots, and special operation robots are increasingly widely deployed in various application scenarios. In actual task execution, service robots need to perceive the environment according to natural language instructions, understand task objectives, and plan reasonable behavior sequences to complete complex operations. Therefore, a robot vision recognition decision control method based on deep learning is required.

[0003] After retrieval, the Chinese invention patent with the publication number "CN114298244A" discloses "A decision control method, device, and system for intelligent agent group interaction". This application constructs an initial decision control model including a top-level learning model and a bottom-level learning model, and performs top-level and bottom-level fusion training on the initial decision control model to obtain a final decision control model for decision control. This decision control method, device, and system improve the effectiveness of decision control during intelligent agent group interaction.

[0004] In addition, the Chinese invention patent with the publication number "CN108985463A" discloses "An artificial intelligence combat method and robot system based on a knowledge base and deep learning". This application automatically generates combat case samples and combines real combat case samples to solve the problem of insufficient combat case samples for effective deep learning and auxiliary decision-making, improving the effect of deep learning and the ability of combat auxiliary decision-making, and enhancing the accuracy of the deep learning model, as well as the subjective initiative and intelligence of combat robots.

[0005] However, in actual use, the above-mentioned disclosed patent methods and similar patent methods do not establish a unified and dynamic association deep reasoning mechanism between natural language instructions and real-time visual perception, resulting in problems such as untimely response and rigid behavior planning for robots when facing complex and changing environments. Summary of the Invention

[0006] The purpose of the present invention is to provide a robot vision recognition decision control method based on deep learning to solve the problems raised in the above background art.

[0007] To achieve the above purpose, the present invention provides the following technical solution: A robot vision recognition decision control method based on deep learning, including: Construct a thinking decision tree based on the input instructions and real-time visual images. The trunk of the thinking decision tree is the structural representation subject integrating interconnected big data and deep learning search engines. Each leaf node of the thinking decision tree represents a specific decision logic node and can be divided into three types of functional modules. The top-level leaf node corresponds to the task objective analysis layer, which is responsible for semantic decomposition and intention recognition. The middle-level leaf nodes form the feature vector understanding layer of the real-time visual images, and the bottom-level leaf nodes form the action generation layer to complete the planning and execution of the behavior sequence. Obtain the linear relationship between the middle-level leaf nodes as the main body and the top-level leaf nodes, which is used to realize the association between the current visual features and the task intention, and ensure that the task semantic objective can be correctly inferred according to the real-time image content. Obtain the content of the corresponding bottom-level leaf nodes according to the linear relationship. According to the attribute characteristics of the decision objective and the content of the input instructions, construct corresponding decision adjustment suggestions for each bottom-level leaf node, which are used to dynamically optimize the behavior generation strategy and improve the adaptability in complex environments.

[0008] As a further optimization of this technical solution, the construction method of the thinking decision tree includes: Construct a data acquisition library for receiving multi-source heterogeneous data. Construct a fusion search engine for retrieving and analyzing multi-source heterogeneous data. Based on the data acquisition library and the fusion search engine, form the trunk framework of the thinking decision tree. Construct the leaf node framework of the thinking decision tree. The leaf node framework of the thinking decision tree assigns content to the top-level leaf nodes according to the input instructions. The trunk framework divides the functions of the middle-level leaf nodes according to the feature vectors in the real-time visual images.

[0009] As a further optimization of this technical solution, the construction method of the fusion search engine includes: Obtain the feature vector representation content of multi-source heterogeneous data for the standardized mapping from the data space to the vector space. Establish a multi-modal feature fusion framework to complete the deep interaction and semantic alignment of heterogeneous features. Construct a co-evolution paradigm of semantic indexing and deep learning models, and design a dynamic update method based on incremental learning. Implement the optimized deployment of the joint retrieval scheduling strategy and establish a multi-engine dynamic allocation mechanism based on reinforcement learning.

[0010] As a further optimization of this technical solution, the dynamic update method allows the data acquisition library and the deep learning model parameters to be optimized synchronously. After receiving new visual or semantic input data, the existing deep learning model is adjusted based on the online learning algorithm, and at the same time, the semantic index items are updated to adapt to the new task requirements.

[0011] As a further optimization of this technical solution, the multi-engine dynamic allocation mechanism includes: Taking the text semantic feature vector as the reference template, by obtaining the area in the visual feature vector that conforms to the text semantic elements, counting the proportion of the number of matching feature vectors, and establishing a semantic matching degree calculation formula: matching degree = Σ(matching feature vectors) / Σ(total feature vectors)×100%; When the matching degree exceeds the preset threshold of 50%, the semantic search engine is preferentially called, otherwise the online deep learning engine is activated for incremental learning.

[0012] As a further optimization of this technical solution, the method for obtaining the content corresponding to the underlying leaf nodes includes: Obtaining the linear path index relationship, by selecting the path graph in the thinking decision tree, and the selection basis follows the path graph from the "top-level task node → middle-level perception node"; Determining the content of the corresponding underlying leaf node in the thinking decision tree according to the linear path index relationship, by obtaining the corresponding relationship between the top-level task node and the middle-level perception node, and determining the content of the underlying leaf node based on the corresponding relationship.

[0013] As a further optimization of this technical solution, the construction method of the decision adjustment suggestion includes: Initializing and assigning values to the execution bottleneck scores of the underlying leaf nodes, which is used to detect problems such as performance degradation, execution failure, or response delay during the task execution process in the underlying leaf nodes of the thinking decision tree, so as to identify bottlenecks or potential error sources; Sorting the underlying leaf nodes according to the execution bottleneck scores, which is used to sort each leaf node according to its execution bottleneck score. The sorting result of the execution bottleneck scores determines the priority of the optimization process, and the nodes with higher scores will be preferentially optimized; Constructing adjustment suggestions corresponding to the underlying leaf nodes based on the sorting result.

[0014] As a further optimization of this technical solution, the method for initializing and assigning values to the underlying leaf nodes includes: Defining an execution bottleneck evaluation function: ; Used to represent the bottleneck score of the nth underlying leaf node; Used to represent the task execution error rate, which reflects the failure situation of the service robot during task execution, that is, represents the ratio of the number of recognition errors to the total number of attempts. Specifically ; For task completion delay, which represents the time delay between when the service robot receives an instruction and starts to execute the task. It is obtained by the service robot monitoring the timestamps of task execution and recording the time difference between instruction input and the start of task execution, that is ; For resource occupancy rate, which represents the consumption of resources within the service robot during task execution. It is obtained by monitoring the usage of hardware resources within the service robot, that is ; They are the task weight coefficient, resource limit coefficient, and environment adaptation coefficient respectively.

[0015] As a further preference of this technical solution, , the number of emergency tasks is used to represent tasks containing time within the executed tasks, and the number of complex tasks refers to tasks with more than one execution command within the tasks; , where the available resource amount is calculated through real-time collection and analysis of the service robot's monitoring data, and the resource limit coefficient decreases, indicating sufficient resources, while when the resource usage is high, increases, indicating that resource allocation needs to be optimized when executing tasks; , the number of features within the image is obtained through the service robot's vision system.

[0016] Compared with the prior art, the beneficial effects of the present invention are: This robot vision recognition decision control method based on deep learning realizes the deep fusion and unified representation of multi-source heterogeneous data by constructing a multi-level thinking decision tree structure, combining multi-modal feature fusion and a deep learning engine, and effectively integrates different modal information such as text instructions and visual images by constructing a fusion search engine and a semantic alignment mechanism, achieving accurate parsing of task intentions, collaborative processing of environmental perception, and improving the real-time response and decision-making accuracy of the service robot for complex tasks; Furthermore, by establishing a linear mapping relationship between the top-level semantic nodes and the middle-level perception nodes, the service robot can quickly locate the key areas related to the task target based on real-time images, thereby accurately inferring semantic intentions and enhancing the dynamic adaptation ability of the behavior planning strategy; Finally, by performing bottleneck scoring and generating adjustment suggestions for the execution status of the bottom-level behavior nodes, it can optimize the action sequence in real time during task execution and improve the execution stability under resource-constrained or scenario-changing conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is the step flow chart of the disclosed method of the present invention; Figure 2 It is an auxiliary explanatory diagram for step S100 of the present invention; Figure 3 It is a flowchart of the method for constructing the thinking decision tree of the present invention; Figure 4 It is a flowchart of the method for constructing the decision adjustment suggestion of the present invention. Specific embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] Before understanding the technical solutions proposed in this application, it should be clear that the control method proposed in this application is mainly applied in the fields of service robots and industrial robots, and is particularly suitable for intelligent robot systems that need to combine natural language instruction parsing and real-time visual perception for autonomous decision-making.

[0020] Specifically, as Figure 1 shown, the present invention provides a technical solution: a robot vision recognition decision control method based on deep learning, including: step S100-step S400.

[0021] Step S100: Construct a thinking decision tree based on the input instruction and the real-time visual image.

[0022] It should be noted that, referring to Figure 2 it can be seen that in step S100 of this application, the trunk part of the thinking decision tree is the structural representation main body integrating interconnected big data and a deep learning search engine. Each leaf node of the thinking decision tree represents a specific decision logic node and can be divided into three types of functional modules. The top-level leaf node corresponds to the task objective analysis layer, which is responsible for semantic decomposition and intention recognition. The middle-level leaf nodes form the feature vector understanding layer of the real-time visual image, and the bottom-level leaf nodes form the action generation layer, which completes the planning and execution of the behavior sequence.

[0023] It should be added that in this application, the trunk part of the thinking decision tree serves as the processing center for decision control, integrating multi-source heterogeneous data, a real-time online deep learning engine, and a semantic search engine.

[0024] It is worth noting that, referring to Figure 3 it can be seen that the method for constructing the thinking decision tree in step S100 includes: step S101-step S105.

[0025] Step S101: Construct a data acquisition library.

[0026] It should be noted that the data acquisition library in step S101 is used to receive multi-source heterogeneous data. Specifically, for the technical solution proposed in this application, it is mainly used to obtain multi-source heterogeneous data of visual data and instruction data. In addition, it should be noted that the data acquisition library constructed in step S101 in this application provides data support for the construction of the thinking decision tree.

[0027] Step S102: Construct a fusion search engine.

[0028] It should be noted that the fusion search engine in step S102 includes a real-time online deep learning engine and a semantic search engine. It should be added that both the real-time online deep learning engine and the semantic search engine are common search tools in the prior art. Through their combination, the rapid retrieval and analysis of multi-source heterogeneous data are realized, providing technical support for the construction of the thinking decision tree.

[0029] It should be added that the trunk framework of the thinking decision tree is composed of the data acquisition library disclosed in step S101 and the fusion search engine disclosed in step S102.

[0030] It is worth noting that the combination methods of the real-time online deep learning engine and the semantic search engine include: steps S102A - S102D.

[0031] Step S102A: Obtain the feature vector representation content of multi-source heterogeneous data to achieve the standardized mapping from the data space to the vector space.

[0032] It should be noted that in this application, since the multi-source heterogeneous data are visual data and instruction data respectively, the feature vector representation method of the multi-source heterogeneous data in step S102A is the image feature vector, and the instruction data is the text feature vector.

[0033] Step S102B: Establish a multi-modal feature fusion framework.

[0034] It should be clear that after the feature vector representation is completed, in order to realize the deep fusion of the semantics in the text feature vector and the image feature vector, a neural network architecture suitable for multi-modal data fusion needs to be constructed. In actual construction, the graph neural network (GNN) in the prior art is used to complete the alignment of multi-modal information in the shared semantic space, and then a feature representation with enhanced context association is obtained. It should be further noted that step S102B not only improves the complementarity between different modalities but also provides guarantee for the joint retrieval of the subsequent search engine.

[0035] As a preferred embodiment, this embodiment selects step S102B for specific illustration. In this embodiment, a cross-modal attention model based on the Transformer structure in the prior art is adopted, and the visual and semantic vectors generated in step S102A are used as Key-Value and Query inputs respectively. Taking the input instruction of "pick up the red cup on the table" as an example, the visual features will locate the red object area in the image feature vector, and "pick up", "red", and "cup" in the semantic vector are used as tokens to guide the visual recognition to focus on the target area (the red object area).

[0036] Step S102C: Construct a co-evolution paradigm of semantic indexing and deep learning model, and design a dynamic update method based on incremental learning.

[0037] It should be noted that considering the dynamic changes and real-time requirements of the task environment, this step proposes a dynamic update method that allows the data acquisition library and the deep learning model parameters to be optimized synchronously. The specific implementation method is: after receiving new visual or semantic input data, adjust the existing deep learning model based on online learning algorithms (such as incremental learning or meta-learning), and at the same time update the semantic index items to adapt to the new task requirements.

[0038] As a preferred embodiment, this embodiment selects step S102C for illustration. Specifically, during the process of the service robot performing tasks, for example, when identifying a new type of tool (pliers), the service robot cannot correctly identify it. At this time, since the service robot cannot perform the corresponding actions, the action feedback after the service robot fails to execute is used to enter this sample (pliers) as a new sample into the data acquisition library in step S101, and the visual system of the service set robot is updated during the low-load period at night through the incremental training method in the prior art, and at the same time, add the matching labels of "pliers" and related task actions to the semantic index.

[0039] Step S102D: Implement the optimized deployment of the joint retrieval scheduling strategy and establish a multi-engine dynamic allocation mechanism based on reinforcement learning.

[0040] It should be pointed out that to establish an efficient decision support mechanism, a joint retrieval scheduling strategy needs to be constructed to coordinate the query call order of semantic search and deep learning engines. Specifically, this embodiment stipulates that it is necessary to dynamically adjust the engine call priority according to the quantization score results by evaluating the matching degree of the input information in the visual space.

[0041] As an optimized implementation solution, this part selects step S102D for technical analysis. When the service robot receives the voice command "Please hand me the black notebook on the desk", its execution process includes: First, the service robot extracts the core semantic elements ("notebook", "black", "desk") through the semantic search engine, and derives the task intention of "fetching and delivering" based on the verb "hand". At the same time, the visual system of the service robot, according to the multi-modal feature fusion mechanism in step S102B, matches the collected visual feature vectors with the above semantic elements. Specifically, taking the text semantic feature vector as the reference template, by locating the area in the visual feature vector that conforms to the text semantic elements, counting the proportion of the number of matching feature vectors, and establishing a semantic matching degree calculation formula: matching degree = Σ (number of matching feature vectors) / Σ (total number of feature vectors) × 100%. When the matching degree exceeds the preset threshold of 50%, the semantic search engine is preferentially called, otherwise the online deep learning engine is activated for incremental learning.

[0042] It should be noted that the online deep learning engine performs incremental learning by updating the visual system of the service set robot in the low-load period at night through the incremental training method in the existing technology in step S102C, and at the same time adding the matching tags of "notebook", "black", "desk" and related task actions (fetching and delivering) to the semantic index.

[0043] Step S103: Construct the leaf node framework of the thinking decision tree.

[0044] It should be noted that in step S103, the thinking decision tree framework uses the graph neural network (GNN) in the existing technology to model the tree structure. The leaf nodes are connected by semantic relationships, and the trunk part is based on the data collection library structure to construct linear connections with each leaf node.

[0045] As a preferred implementation solution, this implementation solution selects step S103 for explanation. Specifically, when the input command is "clean the window", the top leaf node of the thinking decision tree is "cleaning task", the middle-level leaf nodes are "identifying the window" and "locating the stain area", and the bottom leaf nodes include the action behaviors of the service robot such as "moving the arm" and "wiping action".

[0046] Step S104: The leaf node framework of the thinking decision tree assigns content to the top leaf node according to the input command.

[0047] It should be noted that in step S104, the input command is semantically parsed through the natural language processing model (T5 or GPT) in the existing technology, extracting the action verb, target object, and location phrase, and setting a linear relationship. According to the linear relationship, the action verb, target object, and location phrase form the corresponding task and are mapped into the thinking decision tree framework to endow the top leaf node with functions.

[0048] Step S105: The trunk part of the thinking decision tree divides the functions of the middle-layer leaf nodes according to the feature vectors in the real-time visual image.

[0049] It should be noted that in this step, the visual acquisition system of the service robot captures the surrounding environment image, extracts the image features through a deep learning model (multi-modal feature fusion mechanism), and these feature vectors are then input into the middle-layer leaf nodes to identify and understand the real-time visual information, such as object category, position, and shape.

[0050] As a preferred implementation, this implementation is described mainly based on Step S105. Specifically, the service robot acquires the current environment image through the visual system, that is, acquires the current environment image through the integrated camera. The image data is input into the visual feature extraction module of the cross-modal Transformer in the prior art to extract the multi-dimensional feature vectors of the image (including color, texture, edge, shape, etc.), and then input into the middle-layer leaf nodes. For example, when the input instruction is "grab the red water cup", the visual system will focus on the object area in the image with the color feature of red and the shape of "cup", and map this feature to the middle-layer nodes such as "water cup area" and "red area" nodes.

[0051] Step S200: Obtain the linear relationship between the middle-layer leaf nodes as the main body and the top-layer leaf nodes.

[0052] It should be noted that Step S200 is used to judge the executability of the input instruction by obtaining the linear relationship. The linear relationship in Step S200 corresponds to the treetop part in the thinking decision tree.

[0053] Step S300: Obtain the content of the corresponding bottom-layer leaf nodes according to the linear relationship.

[0054] It should be noted that Step S300 is used to quantify the input instruction and the visual result into control parameters, and provide a specific instruction set (grasping point and angle), so as to improve the operability of the input instruction.

[0055] Specifically, the method for obtaining the content of the corresponding bottom-layer leaf nodes includes: Step S301 - Step S302.

[0056] Step S301: Obtain the linear path index relationship.

[0057] It should be noted that by selecting the path map in the thinking decision tree, the selection basis follows the path map from the "top-level task node → middle-level perception node", where the path map represents the connection relationship and transmission direction between different layer leaf nodes, so as to determine the path from the top-level task node to the bottom-level execution node.

[0058] Step S302: Determine the corresponding content of the bottom-level leaf node in the thinking decision tree according to the linear path index relationship.

[0059] It should be noted that in step S302, when determining the position of the bottom-level leaf node, the corresponding relationship between the top-level task node and the middle-level perception node is first obtained, and the corresponding relationship includes: space and time. The artificial intelligence association model in the prior art determines the content of the bottom-level leaf node by analyzing the corresponding relationship.

[0060] As a preferred implementation, this implementation method selects step S302 for elaboration. Specifically, when the content corresponding to the top-level task node is "Keep the room tidy before 12 o'clock" and the content of the middle-level perception node is "Identify that there are toys on the floor in the room", at this time, based on the artificial intelligence association model in the prior art, the constraint condition of "before 12 o'clock" in the time corresponding relationship will be compared with the "current time", and at the same time, the overlapping range of the coordinate system of the "floor area" in the space corresponding relationship and the working range of the service robot will be analyzed, and the content of the bottom-level leaf nodes obtained are "speed adjustment", "position movement", and "cleaning" respectively.

[0061] Step S400: Construct corresponding decision adjustment suggestions for each bottom-level leaf node according to the attribute characteristics of the decision target and the content of the input instruction.

[0062] It should be noted that the decision adjustment suggestion in step S400 is used to change the content of the bottom-level leaf node.

[0063] Specifically, referring to Figure 4 It can be seen that the construction method of the decision adjustment suggestion includes: step S401-step S403.

[0064] Step S401: Initialize and assign a value to the execution bottleneck score of the bottom-level leaf node.

[0065] It should be added that the goal of step S401 is to detect problems such as performance degradation, execution failure, or response delay during the task execution process in the bottom-level leaf nodes of the thinking decision tree, so as to identify bottlenecks or potential error sources.

[0066] Specifically, for how to identify the initialization assignment of the bottom-level leaf node in step S401, an execution bottleneck evaluation function is defined.

[0067] That is , it should be added that is used to represent the bottleneck score of the nth bottom-level leaf node, is used to represent the task execution error rate, which reflects the failure situation of the service robot during task execution, that is, it represents the ratio of the number of recognition errors to the total number of attempts. Specifically , For task completion delay, specifically referring to the time delay between when the robot receives an instruction and starts to execute the task, which is obtained by monitoring the timestamps of task execution by the service robot and recording the time difference between instruction input and the start of task execution, that is , is the resource occupancy rate, indicating the consumption of resources within the service robot during task execution, which is obtained by monitoring the usage of hardware resources within the service robot. For example, existing resource monitoring tools (such as top and htop tools) in the prior art can be used to collect the system resource usage in real time, calculate the CPU usage rate and memory occupancy rate, that is , are the task weight coefficient, resource limit coefficient, and environmental adaptation coefficient respectively. It should be supplemented that , it should be noted that the number of urgent tasks refers to tasks containing time within the task, and the number of complex tasks refers to tasks with more than one execution command within the task , where it should be noted that the "available resource amount" is calculated through real-time collection and analysis of the monitoring data of the service robot. For example, when the CPU usage rate is low, the resource limit coefficient decreases, indicating sufficient resources, while when the resource usage is high increases, indicating that resource allocation needs to be optimized during task execution. Finally , the number of features in the image is obtained through the visual system of the service robot, and the target area is represented by the multi-modal feature fusion mechanism disclosed in step S102B. The total number of features in the image is captured and counted in real time by the visual recognition system of the service robot. The more features there are, the more complex the environment is, which poses higher requirements for the visual recognition and processing capabilities of the service robot. Therefore, the environmental adaptation coefficient will increase accordingly

[0068] As a preferred implementation, this implementation elaborates on the execution bottleneck evaluation function. Specifically, when performing the "move the water cup" task, the service robot fails to accurately identify the position of the water cup and has 3 grasping failures. A total of 5 tasks are performed. Therefore = 0.6, and the task delay time is 4s. Therefore = 4, and the (cpu occupancy rate of the service robot is 70%). Therefore = 0.7, and the total number of tasks is 5, and the task content is the same. Therefore, both the number of urgent tasks and the number of complex tasks are 1. Therefore, the task weight coefficient is 0.2, the resource limit coefficient is 0.3 (cpu occupancy rate is 70%), the environmental adaptation coefficient is 0.2, the total number of features in the image content is 5, and the number of target features is only one water cup. Therefore = 0.2×0.6 + 0.3×4 + 0.2×0.7 = 0.12 + 1.2 + 0.14 = 1.46. Thus, it can be known that under this implementation, the bottleneck score of the underlying leaf node is 1.46.

[0069] It should be noted that in step S401, by recording and analyzing the situations of task execution failure or inefficiency in the underlying leaf nodes, bottleneck factors affecting decision-making effectiveness are identified, such as identification failures, execution errors, and environmental interferences. In addition, it should be added that there is a linkage between step S401 and step S102C.

[0070] As a preferred implementation, this implementation elaborates on the linkage relationship between step S401 and step S102C. For example: when the service robot receives a task instruction and starts to execute, the robot monitors the status of the task in real time and calculates the bottleneck score (Bn). In addition, it is assumed that in the grasping task, the service robot finds that there are many grasping failures, resulting in an increase in the error rate (Er), thus generating a relatively high bottleneck score. The evaluation result triggers the collaborative update mechanism of semantic indexing and deep learning model in step S102C. The service robot will input the images of grasping failures and related data (such as pliers images) as samples into the data acquisition library. At the same time, the semantic search engine and the deep learning model perform incremental training and update based on the new data to improve the recognition ability of the item (such as pliers). Therefore, between step S401 and step S102C, first, the execution bottleneck is evaluated through step S401 to discover the execution bottleneck in real time, and then through the model update mechanism of step S102C, the visual recognition and semantic understanding capabilities of the robot are dynamically adjusted to solve the fundamental problems caused by the bottleneck, realizing a bottleneck feedback mechanism, that is, optimizing the behavior of the robot through online learning, so as to execute tasks more efficiently and accurately.

[0071] Step S402: Sort the underlying leaf nodes according to the execution bottleneck scores.

[0072] It should be noted that in step S402, by sorting each leaf node according to its execution bottleneck score, the sorting result of the execution bottleneck scores determines the priority of the optimization process. The nodes with higher scores will be preferentially optimized. Step S402 is used to provide a clear priority order for subsequent optimization to ensure that the execution bottlenecks are solved.

[0073] Step S403: Construct adjustment suggestions corresponding to the underlying leaf nodes based on the sorting results.

[0074] It should be noted that in step S403, based on the sorting result of step S402, a specific optimization decision plan is generated for each leaf node. Since the content represented by the execution bottleneck evaluation function includes three parameters: task execution error rate, task completion delay and resource occupancy rate, the adjustment suggestion needs to set three types of optimization strategies in a targeted manner: for high error rate nodes, activate the incremental learning mechanism of step S102C to supplement the target object feature template; for high delay nodes, send instructions to the service robot to shorten the motion trajectory; for high resource occupancy nodes, reduce them by shutting down unnecessary functions in the service robot.

[0075] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is limited by the attached embodiments and their equivalents.

Claims

1. A robot vision recognition decision control method based on deep learning, characterized in that, Including: Construct a thinking decision tree based on input instructions and real-time visual images. The trunk part of the thinking decision tree is the structural representation subject integrating interconnected big data and deep learning search engines. Each leaf node of the thinking decision tree represents a specific decision logic node and can be divided into three types of functional modules. The top-level leaf nodes correspond to the task objective analysis layer, responsible for semantic decomposition and intention recognition. The middle-level leaf nodes form the feature vector understanding layer of the real-time visual image. The bottom-level leaf nodes form the action generation layer to complete the planning and execution of the behavior sequence. Obtain the linear relationship between the middle-level leaf nodes as the main body and the top-level leaf nodes, which is used to realize the association between the current visual features and task intentions, and ensure that the task semantic objectives can be correctly inferred according to the real-time image content. Obtain the content of the corresponding bottom-level leaf nodes according to the linear relationship. According to the attribute characteristics of the decision-making objectives and the content of the input instructions, construct corresponding decision adjustment suggestions for each bottom-level leaf node, which are used to dynamically optimize the behavior generation strategy and improve the adaptability in complex environments.

2. The method for robot vision recognition decision control based on deep learning according to claim 1, characterized in that: The construction method of the thinking decision tree includes: Construct a data acquisition library for receiving multi-source heterogeneous data. Construct a fusion search engine for retrieving and analyzing multi-source heterogeneous data. Form the trunk framework of the thinking decision tree based on the data acquisition library and the fusion search engine. Construct the leaf node framework of the thinking decision tree. The leaf node framework of the thinking decision tree assigns content to the top-level leaf nodes according to the input instructions. The trunk framework divides the functions of the middle-level leaf nodes according to the feature vectors in the real-time visual image.

3. A method for robot vision recognition decision control based on deep learning according to claim 2, characterized in that: The construction method of the fusion search engine includes: Obtain the feature vector representation content of multi-source heterogeneous data for the standardized mapping from the data space to the vector space. Establish a multi-modal feature fusion framework to complete the deep interaction and semantic alignment of heterogeneous features. Construct a co-evolution paradigm of semantic indexing and deep learning models, and design a dynamic update method based on incremental learning. Implement the optimized deployment of the joint retrieval scheduling strategy and establish a multi-engine dynamic allocation mechanism based on reinforcement learning.

4. A robot vision recognition decision control method based on deep learning according to claim 3, characterized in that: The dynamic update method allows the synchronous optimization of the data acquisition library and the deep learning model parameters. After receiving new visual or semantic input data, the existing deep learning model is adjusted based on the online learning algorithm, and at the same time, the semantic index items are updated to adapt to the new task requirements.

5. A decision control method for robot vision recognition based on deep learning according to claim 3, characterized in that: The multi-engine dynamic allocation mechanism includes: Taking the text semantic feature vector as the reference template, by obtaining the regions in the visual feature vector that conform to the text semantic elements, counting the proportion of the number of matching feature vectors, and establishing a semantic matching degree calculation formula: matching degree = Σ (number of matching feature vectors) / Σ (total number of feature vectors) × 100%. When the matching degree exceeds the preset threshold of 50%, the semantic search engine is preferentially called, otherwise the online deep learning engine is activated for incremental learning.

6. A robot vision recognition decision control method based on deep learning according to claim 1, characterized in that: The method for obtaining the content of the corresponding bottom-level leaf nodes includes: Obtain the linear path index relationship by selecting a path map in the thinking decision tree. The selection basis follows the path map from "top-level task node → middle-level perception node". Determine the corresponding content of the underlying leaf nodes within the thinking decision tree according to the linear path index relationship, by obtaining the corresponding relationship between the top-level task nodes and the middle-level perception nodes, and determining the content of the underlying leaf nodes based on the corresponding relationship.

7. A robot vision recognition decision control method based on deep learning according to claim 1, characterized in that: The construction method of the decision adjustment suggestion includes: Initialize and assign the execution bottleneck score for the underlying leaf nodes, which is used to detect problems such as performance degradation, execution failure, or response latency during the task execution in the underlying leaf nodes of the thinking decision tree, so as to identify the bottlenecks or potential error sources; Sort the underlying leaf nodes according to the execution bottleneck score, which is used to sort each leaf node according to its execution bottleneck score. The sorting result of the execution bottleneck score determines the priority of the optimization process, and the nodes with higher scores will be preferentially optimized; Construct adjustment suggestions corresponding to the underlying leaf nodes based on the sorting result.

8. A method for robot vision recognition decision control based on deep learning according to claim 7, characterized in that: The method for initializing and assigning values to the underlying leaf nodes includes: Define the execution bottleneck evaluation function: ; Used to represent the bottleneck score of the nth bottom-level leaf node; Used to represent the task execution error rate, which reflects the failure situation of the service robot during task execution, that is, it represents the ratio of the number of recognition errors to the total number of attempts. Specifically ; Is the task completion delay, used to represent the time delay between the service robot receiving an instruction and starting to execute the task. It is obtained by the service robot monitoring the timestamp of task execution and recording the time difference between the instruction input and the start of task execution, that is ; Is the resource occupancy rate, which represents the consumption of resources in the service robot during task execution and is obtained by monitoring the hardware resource usage in the service robot, that is ; They are the task weight coefficient, the resource constraint coefficient, and the environment adaptation coefficient respectively.

9. The method for robot vision recognition decision control based on deep learning according to claim 8, wherein: , the number of urgent tasks is used to represent tasks containing time within the execution tasks, and the number of complex tasks refers to tasks with more than one execution command within the tasks; , where the available resource quantity is calculated by real-time collection and analysis of the monitoring data of the service robot, and the resource limit coefficient decreases, indicating sufficient resources, while when the resource usage is high, increases, indicating that resource allocation needs to be optimized when executing tasks; , the number of features in the image is obtained through the vision system of the service robot.

Citation Information

Patent Citations

  • Artificial intelligence combat method and robot system based on knowledge base and deep learning

    CN108985463A

  • Decision control method, device and system for agent group interaction

    CN114298244A

  • Personnel reidentification method based on deep learning and distance metric learning

    CN108345860A

  • Intelligent elderly accompanying system based on computer vision technology

    CN116945156A

  • Workshop personnel on-duty detection method based on multi-view gait space-time repair network

    CN119206787A

Cited By

  • Control method and device based on task understanding representation, equipment and medium

    CN120871706A

  • Hierarchical robot operation strategy generation method, device and equipment

    CN120886274A

  • Cloud side-end collaborative unmanned aerial vehicle cluster intelligent sensing and decision-making system

    CN121477980A

  • A cloud-edge-device collaborative intelligent sensing and decision-making system for drone swarms

    CN121477980B

  • Robot control method and system thereof, medium, equipment and program product

    CN121515155A