Visual identification-based intelligent decision-making control method and system for body

Through the improved YOLO model for visual recognition, the problem of insufficient accuracy and slow response of the embodied intelligent system in the determination of action completion is solved, real-time and accurate judgment of robotic arm tasks is achieved, the overall efficiency and reliability of the system are improved, and its wide application in the fields of industry, medical and service is promoted.

CN120279536AActive Publication Date: 2025-07-08SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510766271.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing embodied intelligent system has problems of insufficient accuracy and slow response in determining action completion, which affects its application in actual scenarios.

Method used

The improved YOLO model is used for visual recognition, and the occlusion perception attention module and feature enhancement network are embedded in the backbone network, and the dynamic mask and dynamic threshold loss function are combined to achieve real-time accurate judgment of robotic arm tasks.

Benefits of technology

It improves the accuracy and response speed of the action completion judgment of the embodied intelligent system, improves the overall efficiency and reliability of the system, and is suitable for complex tasks in the fields of industry, medical care and service.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279536A_ABST
    Figure CN120279536A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent decision-making control method and system based on visual identification, and belongs to the technical field of intelligent ownership, and the method comprises the steps: collecting image data in real time in the process that a mechanical arm executes a current task; the image data is input into a pre-trained improved YOLO model, reasoning information features of the target object and the pre-trained improved YOLO model are obtained, and the construction process comprises the steps that an occlusion perception attention module is added into a backbone network of the YOLO model, and a mask of the occlusion perception attention module is a dynamic mask; and comparing the reasoning information features of the target object with preset features in a feature library, and if the reasoning information features of the target object are consistent with the preset features in the feature library, stopping the current task and executing the next task. According to the method, the working state of the mechanical arm can be accurately judged, the real-time requirement is met, and the overall efficiency and reliability of the intelligent system with the mechanical arm are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of embodied intelligence technology, and particularly relates to an embodied intelligence decision-making control method and system based on visual recognition. Background Art

[0002] In the current rapid development of technology, embodied intelligence, as a highly potential and challenging frontier field of artificial intelligence, is attracting the attention of global researchers. Its aim is to endow intelligent entities with the ability to perceive, make decisions and act in the real physical environment, requiring intelligent entities to be able to understand complex and dynamic environmental information, and accordingly make appropriate decisions and accurately execute actions to achieve efficient interaction with the environment.

[0003] Among the many technical branches of embodied intelligence, imitation learning is crucial. It enables intelligent entities to learn to perform specific tasks by observing human demonstrations or high-quality examples. However, there are significant problems with imitation learning currently, and the accuracy is difficult to reach an ideal state, with frequent mistakes. Taking industrial manufacturing as an example, when a robotic arm performs complex assembly tasks through imitation learning, problems such as incorrect part assembly positions and disordered sequences often occur due to misunderstandings of operation details or environmental interference; when a service robot imitates a human to deliver items, it is also prone to misjudging the target location or dropping the item when grasping it. These mistakes seriously affect the reliability and stability of the embodied intelligence system and hinder its wide application in actual scenarios.

[0004] For an embodied intelligence system, the ability to judge whether a model has successfully completed a given action is crucial. In the embodied intelligence architecture, the ability to drive behavior execution is similar to the human cerebellum, precisely regulating the body movements of intelligent entities to achieve expected actions; the ability to drive judgment and decision-making is like the human brain, making reasonable decisions based on environmental information and task goals to guide the behavior of intelligent entities. An efficient and reliable embodied intelligence system requires the close cooperation and precise coordination of the brain and the cerebellum. However, due to the lack of an effective action completion judgment mechanism, even if the cerebellum can prompt an intelligent entity to execute an action, it is impossible to determine whether the action meets the expected goal, greatly weakening the system performance.

[0005] Taking the application of a robotic arm in industrial production as an example, the robotic arm on the production line needs to complete complex operations such as material handling and part processing. When the existing model controls the robotic arm to perform these operations, it is difficult to accurately judge whether the operation has been successfully completed. For example, after a part is installed in the specified position, the model cannot determine whether the part is correctly and firmly installed, which may lead to quality problems in subsequent production processes and even cause equipment failures. This uncertainty is continuously amplified in large-scale production, seriously affecting production efficiency and product quality. There are deficiencies such as insufficient accuracy and slow response in the action completion judgment of the embodied intelligence system. Summary of the Invention

[0006] In view of the deficiencies of the prior art, the present application proposes a method and system for embodied intelligent decision-making control based on visual recognition.

[0007] In a first aspect, the present application proposes a method for embodied intelligent decision-making control based on visual recognition, including:

[0008] Step S1: The robotic arm sequentially executes tasks according to a preset imitation learning instruction sequence;

[0009] Step S2: During the execution of the current task by the robotic arm, image data is collected in real time;

[0010] Step S3: Input the image data into a pre-trained improved YOLO model to obtain the inference information features of the target object. The construction process of the pre-trained improved YOLO model includes: adding an occlusion-aware attention module to the backbone network of the YOLO model, where the mask of the occlusion-aware attention module is a dynamic mask;

[0011] Step S4: Compare the inference information features of the target object with the preset features in the feature library. If the inference information features of the target object are consistent with the preset features in the feature library, stop the current task and return to Step S1 to execute the next task.

[0012] The dynamic mask is calculated as follows:

[0013] ;

[0014] where M is the dynamic mask, X is the image data, is the first neural network, is the sigmoid activation function.

[0015] The first neural network is an N-layer neural network, and the size of the convolution kernel of each layer of the neural network is: , where is the length of the image data, is the width of the image data, N is an integer greater than 3, and the training process of the first neural network includes: using the occluded image data as a sample and training with the target mask as a label to obtain the first neural network.

[0016] The construction process of the pre-trained improved YOLO model further includes: a feature extraction network and a feature enhancement network;

[0017] The feature extraction network is used to extract features from the output of the backbone network to obtain red-channel features, green-channel features, and blue-channel features;

[0018] The feature enhancement network is used to calculate the covariance matrix of the three-channel features based on the red-channel features, green-channel features, and blue-channel features, and use the red-channel features, green-channel features, blue-channel features, and the covariance matrix of the three-channel features as feature input data. The second neural network is used to reduce the dimension of the feature input data to obtain the first-channel features and the second-channel features; based on the first-channel features and the second-channel features, the inference information features of the target object are calculated.

[0019] The second neural network is a one-layer convolutional network, and the convolutional kernel size is .

[0020] The calculation of the inference information features of the target object based on the first-channel features and the second-channel features is as follows:

[0021] ;

[0022] where, is the inference information feature of the target object, i is the i-th channel, i is 1 or 2, is the weight parameter of the i-th channel, is the feature value of the i-th channel, and b is the bias parameter.

[0023] The construction process of the pre-trained improved YOLO model further includes:

[0024] Step S100: Calculate the light intensity of the image data according to the image data, and select the corresponding dynamic threshold of the loss function of the improved YOLO model according to the light intensity;

[0025] Step S101: When the loss function of the improved YOLO model is less than the selected dynamic threshold, stop training the improved YOLO model, output the final improved YOLO model, and use the final improved YOLO model as the pre-trained improved YOLO model;

[0026] Step S102: When the loss function of the improved YOLO model is greater than or equal to the selected dynamic threshold, increase the number of samples, reconstruct the improved YOLO model again, and return to Step S100;

[0027] where, the loss function of the improved YOLO model is the sum of the weights of the mean square error and the absolute error, and the calculation formula is as follows:

[0028] ;

[0029] where, is the loss function of the improved YOLO model, and α is a weight parameter between 0 and 1, is the mean square error between the sample and the label, is the absolute error between the sample and the label.

[0030] In a second aspect, the present application proposes an embodied intelligent decision control system based on visual recognition, including:

[0031] A task execution module, configured to execute tasks in sequence by the robotic arm according to a preset imitation learning instruction sequence;

[0032] An image acquisition module, configured to collect image data in real time during the execution of the current task by the robotic arm;

[0033] A model inference module, configured to input the image data into a pre-trained improved YOLO model to obtain the inference information features of the target object. The construction process of the pre-trained improved YOLO model includes: adding an occlusion-aware attention module to the backbone network of the YOLO model, where the mask of the occlusion-aware attention module is a dynamic mask;

[0034] A visual recognition module, configured to compare the inference information features of the target object with the preset features in the feature library. If the inference information features of the target object are consistent with the preset features in the feature library, stop the current task and return to the task execution module to execute the next task.

[0035] In a third aspect, the present application proposes an electronic device, including: one or more processors, and a memory. The memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the above-mentioned embodied intelligent decision control method based on visual recognition.

[0036] In a fourth aspect, the present application proposes a computer-readable storage medium, which stores executable instructions. When the instructions are executed, the processor executes the above-mentioned embodied intelligent decision control method based on visual recognition.

[0037] Beneficial effects:

[0038] The present application proposes an embodied intelligent decision control method and system based on visual recognition. Compared with the traditional action completion determination method, it can accurately determine the working state of the robotic arm, effectively make up for the deficiencies in accuracy and slow response in the action completion determination of the existing embodied intelligent system, and significantly improve the overall efficiency and reliability of the embodied intelligent system. Description of the Drawings

[0039] Figure 1 is a flowchart of an embodied intelligent decision control method based on visual recognition according to an embodiment of the present application;

[0040] Figure 2 It is a schematic flow diagram of an embodied intelligent decision-making control method based on visual recognition according to an embodiment of the present application;

[0041] Figure 3 It is a block diagram of the principle of an embodied intelligent decision-making control system based on visual recognition according to an embodiment of the present application. Specific Embodiments

[0042] The following further describes the specific embodiments of the present application in detail in conjunction with the accompanying drawings and embodiments.

[0043] To solve the technical problems existing in the background art, visual recognition technology, as a powerful perception means, can provide rich and accurate information for the decision-making of the embodied intelligent system. Through visual recognition, the system can obtain key information such as the position, shape, and posture of objects in the environment in real time, so as to more accurately judge whether the current action achieves the expected goal. For example, using visual recognition technology to monitor the assembly status of parts during the operation of a robotic arm can timely detect problems such as position deviation and unstable installation, providing a reliable basis for subsequent decision-making. Therefore, making decisions based on visual recognition methods has become a key way to solve the current dilemmas of the embodied intelligent system, improve the overall efficiency and reliability of the system, and is expected to promote the wide application of embodied intelligence in many fields such as industry, healthcare, and services. However, because visual recognition technology requires the application of model algorithms, for complex models with large amounts of data, usually a huge amount of computing is required, and real-time performance needs to be noted during the process of the robotic arm completing each task. How to balance the relationship between accuracy and time-consuming is also a huge challenge.

[0044] Embodiment 1:

[0045] This embodiment proposes an embodied intelligent decision-making control method based on visual recognition, as shown in Figure 1 、 Figure 2 and includes:

[0046] Step S1: The robotic arm sequentially executes tasks according to a pre-set imitation learning instruction sequence;

[0047] In this embodiment, first, it enters the start stage, starts the entire robotic arm task execution and determination system, and completes system initialization work, including loading the parameters of the pre-trained improved YOLO model, initializing visual acquisition devices such as cameras, establishing a communication connection with the robotic arm, etc., to ensure that each component is in a workable state. This step belongs to the prior art and will not be elaborated in this application.

[0048] After the start stage, this embodiment enters the task execution stage, and the robotic arm starts to execute tasks according to a pre-set imitation learning strategy or instruction. For example, in an industrial assembly scenario, it grabs and installs parts, and in a service scenario, it grabs and delivers items, etc.

[0049] Step S2: During the process of the robotic arm executing the current task, image data is collected in real time;

[0050] In this embodiment, during the process of the robotic arm executing the task, vision acquisition devices such as cameras deployed in the working scenario collect image data in real time, including: capturing the operation process of the robotic arm and the status information of the target object.

[0051] In the actual application scenario, advanced vision acquisition devices such as high-resolution cameras are deployed omnidirectionally and multi-angularly in the working area of the robotic arm. These devices have excellent image capture capabilities, can clearly record subtle movements and object features. At the same time, they can be flexibly adjusted according to factors such as the light change and spatial layout of the working environment to ensure that complete and accurate image data is collected, providing a solid data basis for subsequent target recognition.

[0052] Step S3: Input the image data into a pre-trained improved YOLO model to obtain the inference information features of the target object. The construction process of the pre-trained improved YOLO model includes: adding an occlusion-aware attention module to the backbone network of the YOLO model, where the mask of the occlusion-aware attention module is a dynamic mask;

[0053] In this embodiment, the collected image data is input into a pre-trained improved YOLO model for inference operation. The model identifies the target object in the image and analyzes key information such as its position, shape, and posture. Among them, the position, shape, and posture are the inference information features of the target object.

[0054] In the embodied intelligent decision-making control method, in order to balance the relationship between accuracy and time consumption, avoid the problems of insufficient accuracy and slow response in the existing technology, and solve the technical problem of accurate recognition and task completion determination in the occlusion environment, this embodiment improves the existing YOLO model. The YOLOv8 model (You Only Look Once version 8) is adopted in this embodiment. Other versions of the YOLO model can also be modified accordingly to achieve the effect of the pre-trained improved YOLO model mentioned in this embodiment. However, after verification, other versions of the YOLO model do not have as good an effect as the modified YOLOv8 model. The improved YOLO model can perform high-precision visual recognition of the operation process of the robotic arm and the status information of the target object.

[0055] In this embodiment, the YOLOv8 model is improved in three aspects, so that the method of this embodiment can accurately perform visual recognition even in the occlusion environment. And in the improvement of the algorithm, attention is paid to using as few calculations as possible to meet the real-time requirement.

[0056] Improvement in the first aspect:

[0057] In this embodiment, an occlusion-aware attention module (OAA) of ORCTrack (Occlusion-Aware Detection and Re-ID Calibrated Network for Multi-Object Tracking) is embedded in the CBS module (Convolutional Block with Shortcut) of the backbone network of the YOLOv8 model, fully integrating the advantages of ORCTrack in handling occlusion scenarios with the efficient object detection architecture of YOLOv8, effectively enhancing the adaptive extraction ability of spatial features in each convolutional stage. This module utilizes the high-order statistical features of the overall representation to highlight the spatial details of the feature channels, strengthen the attention to the foreground visible target regions, and simultaneously suppress the interference of the occluded background regions.

[0058] In specific implementation, since occlusion situations are often encountered, but the masks of previous object recognition methods are mostly static, while the dynamic mask (VAM, Vehicle Allocation Matrix) proposed in this embodiment can change continuously according to the real-time state of the target and context information. This enables the model to more accurately focus on foreground targets in different occlusion scenarios, enhancing the model's adaptability to complex environments. The dynamic mask is generated by a first neural network, and the calculation formula is as follows:

[0059] ;

[0060] where M is the dynamic mask, X is the image data, is the first neural network, is the sigmoid activation function.

[0061] The first neural network is an N-layer neural network, and the size of the convolutional kernel of each layer of the neural network is: , where is the length of the image data, is the width of the image data, N is an integer greater than 3, and the training process of the first neural network includes: using the occluded image data as samples and training with the target mask as labels to obtain the first neural network.

[0062] Improvement in the second aspect:

[0063] In the YOLO model, there is not only a backbone network but also a feature extraction network. In this embodiment, a feature enhancement network is added after the feature extraction network, so as to more accurately extract the inference information features of the target object on the premise of meeting real-time performance.

[0064] The construction process of the pre-trained improved YOLO model further includes: a feature extraction network and a feature enhancement network;

[0065] The feature extraction network is used to extract features from the output of the backbone network to obtain the red channel feature R, the green channel feature G, and the blue channel feature B;

[0066] The feature enhancement network is used to calculate the covariance matrix C of the three-channel features according to the red channel feature R, the green channel feature G, and the blue channel feature B, and use the red channel feature R, the green channel feature G, the blue channel feature B, and the covariance matrix C of the three-channel features as feature input data, and perform dimensionality reduction on the feature input data through the second neural network to obtain the first channel feature X1 and the second channel feature X2; according to the first channel feature X1 and the second channel feature X2, calculate the inference information feature Y of the target object.

[0067] The second neural network is a one-layer convolutional network, and the convolution kernel size is .

[0068] The calculation formula for calculating the inference information feature of the target object according to the first channel feature and the second channel feature is as follows:

[0069] ;

[0070] where, is the inference information feature of the target object, i is the i-th channel, i is 1 or 2, is the weight parameter of the i-th channel, is the feature value of the i-th channel, and b is the bias parameter. In this embodiment, the value range of b is: [0.01, 0.1], The value range of the weight parameter is: [-0.414, 0.414].

[0071] In this embodiment, in terms of feature processing, feature enhancement based on the covariance matrix is adopted. After adjusting the number of channels by using 1×1 convolution, the association of different channel features is reflected by calculating the covariance matrix. Among them, the calculation formula of the covariance matrix is as follows:

[0072] ;

[0073] where, is the covariance between the red channel feature R and the green channel feature G, is the covariance between the red channel feature R and the blue channel feature B, is the covariance between the green channel feature G and the blue channel feature B.

[0074] In this embodiment, in order to extract the inference information features of the target object more precisely, a feature enhancement network is added after the feature extraction network of the YOLO model. After extracting the features of the output of the backbone network, the features of three channels are obtained, including: the red channel feature R, the green channel feature G, and the blue channel feature B. According to these three features, the covariance matrix C of the three-channel features is calculated, and the red channel feature R, the green channel feature G, the blue channel feature B, and the covariance matrix C of the three-channel features are used as the feature input data, and the feature input data is reduced in dimension through 1×1 convolution to obtain the first channel feature X1 and the second channel feature X2. According to the first channel feature X1 and the second channel feature X2, the inference information feature Y of the target object is calculated. Through the above operations, the inference information features of the target object are obtained more precisely. At the same time, the computational complexity of the above steps is not large, which can meet the real-time requirements.

[0075] Improvement in the third aspect:

[0076] The construction process of the pre-trained improved YOLO model further includes:

[0077] Step S100: Calculate the light intensity of the image data according to the image data, and select the corresponding dynamic threshold of the loss function of the improved YOLO model according to the light intensity;

[0078] Step S101: When the loss function of the improved YOLO model is less than the selected dynamic threshold, stop training the improved YOLO model, output the final improved YOLO model, and use the final improved YOLO model as the pre-trained improved YOLO model;

[0079] Step S102: When the loss function of the improved YOLO model is greater than or equal to the selected dynamic threshold, increase the number of samples, reconstruct the improved YOLO model again, and return to Step S100;

[0080] Among them, the loss function of the improved YOLO model is the sum of the weights of the mean square error and the absolute error, and the calculation formula is as follows:

[0081] ;

[0082] Among them, is the loss function of the improved YOLO model, and α is a weight parameter between 0 and 1, is the mean squared error between the sample and the label, is the absolute error between the sample and the label.

[0083] In this embodiment, in the fields of object detection and computer vision, the quality of solving regression problems is directly related to the accuracy of the model's prediction of object attributes such as position and size. Traditional loss functions, such as mean squared error (MSE) and mean absolute error (MAE), although performing well in many scenarios, are sensitive to outliers, which may cause the model to be overly influenced by a few abnormal samples during training, thereby reducing the overall generalization ability and prediction accuracy. To overcome this problem, this embodiment proposes the DTIOU (Dynamic Thresholded Intersection over Union) loss function, which not only combines the advantages of mean squared error and absolute error, but also effectively reduces the sensitivity to outliers by introducing a dynamic threshold mechanism, making the model more robust in regression problems. The dynamic threshold mechanism dynamically sets a specified threshold according to the actual situation of the light identified in the real-time collected image data. For example, if the light is very poor (i.e., the light recognition result is less than the first set light threshold) and there are many occlusions, the threshold is set to 0.5 to 0.99. If the light is relatively good (i.e., the light recognition result is greater than or equal to the second set threshold) and there are fewer occlusions, the threshold is set to 0.01 - 0.49. Multiple light thresholds can also be set according to specific needs to achieve the effect of dynamic thresholds.

[0084] The improvements in the above three aspects enable the pre-trained improved YOLO model to not only accurately obtain the inference information features of the target object but also meet the real-time requirements. It can perform in-depth processing on the working scene images during the task execution of the robotic arm at an extremely fast speed. After the robotic arm completes each task, the visual recognition system will parse the acquired images within milliseconds.

[0085] The system can accurately identify the position information of the target object. Whether it is a millimeter-level position deviation or the relative position relationship in a complex space, it can be accurately presented; for the shape of the target object, even if it has an extremely complex and irregular contour, it can be clearly defined through advanced image recognition technology.

[0086] Step S4: Compare the inference information features of the target object with the preset features in the feature library. If the inference information features of the target object are consistent with the preset features in the feature library, stop the current task and return to step S1 to execute the next task.

[0087] In this embodiment, the visual recognition module carefully compares the key information of the target object identified with the preset features in the completion state feature library carefully constructed in the system in advance. For example, in the part assembly task, it determines whether the part is installed at the specified position and whether the installation angle is correct, etc. This feature library is constructed by deeply analyzing and learning a large number of samples that have successfully completed tasks, covering multi-dimensional features such as the standard position, shape contour, and posture angle that the target object should present when various tasks are completed.

[0088] If the comparison result meets the preset task completion standard, it is determined as "Y" (yes), indicating that the task is successfully completed. When the determination result is "Y", the system sends a task completion instruction to the robotic arm. The robotic arm stops the current task and can enter the next task process according to the subsequent arrangement. If it does not meet the preset standard, it is determined as "N" (no), indicating that the task is not successfully completed. When the determination result is "Y", the system sends a task completion instruction to the robotic arm. The robotic arm stops the current task and can enter the next task process according to the subsequent arrangement. When the determination result is "N", the system issues an alarm to indicate that the task execution has failed. At the same time, it sends an instruction to the robotic arm to return to the initial state or restart the task according to a specific strategy, and then enters the model inference stage again, repeating the above process until the task is successfully completed.

[0089] If, after strict comparison, it is determined that the current working state conforms to the preset completion state features, the system will quickly and accurately send a completion instruction to the robotic arm through a high-speed and stable communication protocol, enabling the robotic arm to receive and respond in a timely manner, and then orderly enter the next task process. On the contrary, once it is determined that the current imitation learning has not successfully completed the work, the system will immediately trigger the alarm mechanism. This alarm will not only be prominently displayed on the local operation interface but also be synchronously transmitted to the remote monitoring terminal through the network so that relevant technical personnel can promptly understand the situation.

[0090] Meanwhile, the system will automatically start the task restart process, and the robotic arm will execute the task again according to the preset initial action sequence. During this loop process, the system will record and analyze the data generated each time the task is executed, which is used to continuously optimize the pre-trained improved YOLO model and the task execution strategy, thereby gradually increasing the success rate of task completion.

[0091] Compared with the traditional method for judging the completion of actions, this embodiment proposes an embodied intelligent decision-making control method based on visual recognition, which can accurately judge the working state of the robotic arm, effectively make up for the deficiencies such as insufficient accuracy and slow response in the existing embodied intelligent systems in judging the completion of actions. And it will significantly improve the overall efficiency and reliability of the embodied intelligent system, lay a solid foundation for its wide application in many complex and critical application scenarios such as precision assembly in industrial production, surgical assistance in the medical field, and personalized services in the service industry, and strongly promote the embodied intelligent technology from theoretical research to large-scale practical applications.

[0092] Embodiment 2:

[0093] This embodiment proposes an embodied intelligent decision-making control system based on visual recognition, as Figure 3 shown, including: a task execution module, an image acquisition module, a model inference module, and a visual recognition module; the task execution module is connected to the image acquisition module, the image acquisition module is connected to the model inference module, the model inference module is connected to the visual recognition module, and the visual recognition module is connected to the task execution module;

[0094] The task execution module is used for the robotic arm to execute tasks in sequence according to a pre-set imitation learning instruction sequence;

[0095] The image acquisition module is used for real-time acquisition of image data during the execution of the current task by the robotic arm;

[0096] The model inference module is used for inputting the image data into a pre-trained improved YOLO model to obtain the inference information features of the target object. The construction process of the pre-trained improved YOLO model includes: adding an occlusion-aware attention module to the backbone network of the YOLO model, where the mask of the occlusion-aware attention module is a dynamic mask;

[0097] The visual recognition module is used for comparing the inference information features with the preset features in the feature library. If the inference information features are consistent with the preset features in the feature library, the current task is stopped and the task execution module is returned to execute the next task.

[0098] Embodiment 3:

[0099] This embodiment proposes an electronic device, including: one or more processors, and a memory, where the memory is used for storing instructions, and when the instructions are executed by the one or more processors, the one or more processors execute the above-mentioned embodied intelligent decision-making control method based on visual recognition.

[0100] The electronic device can be a mobile phone, a computer, a tablet computer, etc., including a memory and a processor. A computer program is stored on the memory, and when the computer program is executed by the processor, it implements a vision recognition-based embodied intelligent decision-making control method as described in the embodiments. It can be understood that the electronic device may further include an input / output (I / O) interface and a communication component.

[0101] Among them, the processor is used to execute all or part of the steps in a vision recognition-based embodied intelligent decision-making control method as described in the above embodiments. The memory is used to store various types of data, which may include, for example, instructions of any application program or method in the electronic device, as well as data related to the application program.

[0102] The processor can be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components, and is used to execute a vision recognition-based embodied intelligent decision-making control method as described in the above embodiments.

[0103] Embodiment 4:

[0104] This embodiment provides a computer-readable storage medium, which stores executable instructions. When the instructions are executed and implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.

[0105] The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of a vision recognition-based embodied intelligent decision-making control method described in various embodiments of the present application.

[0106] The foregoing storage medium includes: flash memory, hard disk, multimedia card, card-type memory (e.g., SD (Secure Digital Memory Card) or DX (abbreviation for Memory Data Register, MDR, memory data register) memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, server, APP (abbreviation for Application, application software) application mall, and other various media that can store program verification codes. A computer program is stored thereon, and when the computer program is executed by a processor, each step of the foregoing method for embodied intelligent decision-making control based on visual recognition can be implemented.

[0107] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.

[0108] The protection scope of this application is not limited to the above embodiments. Obviously, those skilled in the art can make various changes and deformations to the present disclosure without departing from the scope and spirit of the present disclosure. If these changes and deformations fall within the scope of the claims of the present disclosure and their equivalent technologies, the intention of the present disclosure also includes these changes and deformations.

Claims

1. An embodied intelligent decision-making control method based on visual recognition, characterized in that, Including: Step S1: The robotic arm sequentially executes tasks according to a pre-set imitation learning instruction sequence; Step S2: During the execution of the current task by the robotic arm, image data is collected in real time; Step S3: The image data is input into a pre-trained improved YOLO model to obtain the inference information features of the target object. The construction process of the pre-trained improved YOLO model includes: adding an occlusion-aware attention module to the backbone network of the YOLO model, where the mask of the occlusion-aware attention module is a dynamic mask; Step S4: Compare the inference information features of the target object with the preset features in the feature library. If the inference information features of the target object are consistent with the preset features in the feature library, stop the current task and return to Step S1 to execute the next task.

2. The embodied intelligent decision-making control method based on visual recognition according to claim 1, wherein, The dynamic mask is calculated as follows: ; where M is a dynamic mask, and X is image data, is the first neural network, is the sigmoid activation function.

3. The embodied intelligent decision-making control method based on visual recognition according to claim 2, wherein The first neural network is an N-layer neural network, and the size of the convolution kernel of each layer of the neural network is: , where is the length of the image data, is the width of the image data, N is an integer greater than 3, and the training process of the first neural network includes: using the occluded image data as a sample and training with the target mask as a label to obtain the first neural network.

4. A method for embodied intelligent decision-making control based on visual recognition according to claim 1, characterized in that, The construction process of the pre-trained improved YOLO model further includes: a feature extraction network and a feature enhancement network; The feature extraction network is used to extract features from the output of the backbone network to obtain red channel features, green channel features, and blue channel features; The feature enhancement network is used to calculate the covariance matrix of the three-channel features based on the red channel features, green channel features, and blue channel features. The red channel features, green channel features, blue channel features, and the covariance matrix of the three-channel features are used as feature input data, and the feature input data is dimensionally reduced through a second neural network to obtain the first channel feature and the second channel feature; based on the first channel feature and the second channel feature, the inference information features of the target object are calculated.

5. A method for embodied intelligent decision-making control based on visual recognition according to claim 4, characterized in that The second neural network is a one-layer convolutional network, and the size of the convolutional kernel is .

6. The embodied intelligent decision-making control method based on visual recognition according to claim 4, characterized in that The calculation of the inference information features of the target object based on the first channel feature and the second channel feature is as follows: ; Among them, is the inference information feature of the target object, i is the i-th channel, and i is 1 or 2, is the weight parameter of the i-th channel, is the feature value of the i-th channel, and b is the bias parameter.

7. A method for embodied intelligent decision-making control based on visual recognition according to claim 1, characterized in that The construction process of the pre-trained improved YOLO model further includes: Step S100: Calculate the light intensity of the image data according to the image data, and select the corresponding dynamic threshold of the loss function of the improved YOLO model according to the light intensity; Step S101: When the loss function of the improved YOLO model is less than the selected dynamic threshold, stop training the improved YOLO model, output the final improved YOLO model, and use the final improved YOLO model as the pre-trained improved YOLO model; Step S102: When the loss function of the improved YOLO model is greater than or equal to the selected dynamic threshold, increase the number of samples, reconstruct the improved YOLO model, and return to Step S100; Among them, the loss function of the improved YOLO model is the sum of the weights of the mean square error and the absolute error, and the calculation formula is as follows: ; Among them, is the loss function of the improved YOLO model, and α is a weight parameter between 0 and 1. is the mean square error between the sample and the label. is the absolute error between the sample and the label.

8. An embodied intelligent decision-making control system based on visual recognition, characterized in that, Including: A task execution module for the robotic arm to sequentially execute tasks according to a pre-set imitation learning instruction sequence; An image acquisition module for collecting image data in real time during the execution of the current task by the robotic arm; A model inference module, configured to input image data into a pre-trained improved YOLO model to obtain inference information features of target objects. The construction process of the pre-trained improved YOLO model includes: adding an occlusion-aware attention module to the backbone network of the YOLO model, where the mask of the occlusion-aware attention module is a dynamic mask; A visual recognition module, configured to compare the inference information features of the target object with the preset features in the feature library. If the inference information features of the target object are consistent with the preset features in the feature library, the current task is stopped and the task execution module is returned to execute the next task.

9. An electronic device, characterized in that, Comprising: One or more processors, and a memory for storing instructions. When the instructions are executed by the one or more processors, the one or more processors execute a method for embodied intelligent decision-making control based on visual recognition according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores executable instructions, and when the instructions are executed, the processor executes a method for embodied intelligent decision-making control based on visual recognition according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Attention neural network method based on multiple paths of dynamic masks

    CN110516065A

  • Improved YOLOv8 tower foundation target detection method and device

    CN119048736A

  • Broiler target detection method, system and device based on occlusion perception and medium

    CN119763009A

  • Target detection method for image collected by AR wearable device based on improved YOLOv8

    CN119942059A

  • Estimation model for interaction detection by a device

    US20230360425A1

Cited By

  • Big and small brain AI visual fusion control method and system of intelligent agent with body

    CN121742309A