Position Execution Control Method, Device, Equipment, Medium and Product of Agent

By using the initial visual image in the agent to determine the haptic detection range, and correct the touch detection process at the detection end by the contact detection image to generate the contact positioning image, the problems of high cost and poor versatility of the positioning operation of the agent in the prior art are solved, and the positioning execution effect with high accuracy, low cost and high versatility are achieved.

CN119831821BActive Publication Date: 2025-06-27BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510300942.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

In the prior art, the cost of performing operations of the agent positioning is high, and the data processing method is not universal enough, resulting in poor universality and complex operation.

Method used

By receiving the initial visual image, the tactile detection range of the agent detection end is determined, and the detection end is controlled to perform touch detection within the tactile detection range, generate a contact detection image, and correct the touch detection process of the detecting end by the contact detection image, and generate a contact positioning image to control the execution end of the agent to perform positioning.

Benefits of technology

It realizes high-precision agent positioning and execution operations on a lower cost basis, simplifies equipment installation calibration and operation settings, improves positioning efficiency and operating performance, and achieves rapid adaptation across different agent systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119831821B_ABST
    Figure CN119831821B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a positioning execution control method for an intelligent agent, which can be applied to the field of artificial intelligence technology. The positioning execution control method for the intelligent agent includes: determining the tactile detection range of the detection end of the intelligent agent according to the received initial visual image; controlling the detection end to perform tactile detection within the tactile detection range to generate a contact detection image; and generating a contact positioning image through the correction of the tactile detection process of the detection end, where the contact positioning image is used to control the execution end of the intelligent agent to perform positioning execution. An embodiment of the present invention also provides a positioning execution control device, equipment, storage medium and program product for the intelligent agent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of intelligent agent control technology, and more specifically to a positioning execution control method, device, equipment, medium and product for an intelligent agent. Background Art

[0002] Artificial Intelligence (AI for short) is an important driving force for the new round of scientific and technological revolution and industrial transformation. It is a new key technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. As an important part of intelligent science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine (i.e., an intelligent agent) that can react in a way similar to human intelligence.

[0003] For embodied intelligence, the positioning operation to achieve the accuracy of complex action execution is one of the important aspects reflecting the intelligent level. For example, for an embodied intelligent agent to pick up a tiny object (such as a steel nail) on the ground or imitate a human to unlock a lock by inserting a key into the keyhole, the intelligent agent must be able to accurately locate and perceive the positions of the steel nail, the keyhole, etc., and then perform action positioning and execution to achieve accurate operations. For an intelligent agent, visual and tactile perception, as the "eyes" and "skin", are important sensory functions when the intelligent agent interacts with the external environment and are one of the key technologies for intelligent robot operations. Visual information can provide the robot with information about the surrounding environment, the number and distribution of existing objects, and the contours, sizes, etc. of each object; tactile information can provide the robot with information about the surface characteristics of the object and the contact state information between the robot and the object.

[0004] In the prior art, in order to obtain more accurate positioning operation execution, expensive detection devices or sensing sensors (such as high-end cameras or radars, etc.) are usually adopted, resulting in a high cost for the perception of visual and tactile information in traditional technologies. In addition, using such expensive detection devices and sensing sensors also requires a further complete set of software systems for matching, further increasing the amount of data required, making the generality worse, and the data processing methods are often not general enough, and each intelligent agent can only be dedicated and used specifically. Summary of the Invention

[0005] In view of at least one of the above problems, embodiments of the present invention aim to provide a positioning execution control method, device, equipment, medium and product for an intelligent agent with higher precision, which can achieve more general data processing and simpler operation at a lower cost. Thus, a method is provided that can use low-cost detection devices such as tactile sensors and detection cameras to achieve high-precision measurement of six-dimensional information of a target object in space, and can make full use of visual and tactile perception feedback to improve positioning efficiency and operation performance. At the same time, it can achieve rapid adaptation across different intelligent agent systems, thereby maintaining a higher level of intelligence.

[0006] An aspect of an embodiment of the present invention provides a positioning execution control method for an intelligent agent, which includes: determining the tactile detection range of the detection end of the intelligent agent according to the received initial visual image; controlling the detection end to perform tactile detection within the tactile detection range to generate a contact detection image; and generating a contact positioning image through the correction of the tactile detection process of the detection end by the contact detection image, where the contact positioning image is used to control the execution end of the intelligent agent to perform positioning execution.

[0007] According to an embodiment of the present invention, before determining the tactile detection range of the detection end of the intelligent agent according to the received initial visual image, it further includes: performing visual detection on the target detection space where the intelligent agent is located to generate an initial visual image.

[0008] According to an embodiment of the present invention, in determining the tactile detection range of the detection end of the intelligent agent according to the received initial visual image, it includes: performing object position analysis on the initial visual image according to the target execution object; and determining the tactile detection range according to the object position analysis.

[0009] According to an embodiment of the present invention, in controlling the detection end to perform tactile detection within the tactile detection range to generate a contact detection image, it includes: obtaining the tactile detection execution position within the tactile detection range; and controlling the detection end to perform tactile detection on the tactile detection execution position to generate a contact detection image.

[0010] According to an embodiment of the present invention, in generating a contact positioning image through the correction of the tactile detection process of the detection end by the contact detection image, it includes: obtaining plane position correction data and attitude angle correction data through the contact detection image; updating the tactile detection execution position of the detection end within the tactile detection range according to the plane position correction data; and adjusting the tactile detection angle when the detection end performs tactile detection at the updated tactile detection execution position according to the attitude angle correction data; controlling the detection end to perform tactile detection on the updated tactile detection execution position according to the tactile detection angle until a contact positioning image is generated.

[0011] According to an embodiment of the present invention, before obtaining the plane position correction data and the attitude angle correction data through the contact detection image, it further includes: identifying the target imaging area of the target execution object in the contact detection image.

[0012] According to an embodiment of the present invention, in obtaining the plane position correction data through the contact detection image, it includes: comparing the target imaging area and the imaging field of view area of the contact detection image to obtain the plane position correction data.

[0013] According to an embodiment of the present invention, in obtaining the attitude angle correction data through the contact detection image, it includes: extracting the filtered contact image of the contact detection image; obtaining the corresponding relationship between each pixel point on the filtered contact image and the preset simulated contact image to obtain the attitude angle correction data.

[0014] Another aspect of the embodiments of the present invention provides a positioning execution control device for an intelligent agent, which includes a range determination module, an image generation module, and a touch measurement correction module. The range determination module is used to determine the tactile detection range of the detection end of the intelligent agent according to the received initial visual image; the image generation module is used to control the detection end to perform touch measurement within the tactile detection range to generate a contact detection image; and the touch measurement correction module is used to correct the touch measurement process of the detection end through the contact detection image to generate a contact positioning image, and the contact positioning image is used to control the execution end of the intelligent agent to perform positioning execution.

[0015] Another aspect of the embodiments of the present invention provides an electronic device, including one or more processors and a memory, and the memory is used to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned positioning execution control method for the intelligent agent.

[0016] Another aspect of the embodiments of the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned positioning execution control method for the intelligent agent.

[0017] Another aspect of the embodiments of the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-mentioned positioning execution control method for the intelligent agent is implemented.

[0018] The positioning execution control method for the intelligent agent provided by the embodiments of the present invention can at least partially solve at least one of the technical problems existing in the related art, and thus can at least achieve one of the following technical effects:

[0019] First, it is possible to achieve high-precision positioning and execution operations by only using lower-cost visual and tactile perception devices. The installation and calibration operations of these perception devices are simpler and more convenient, and there is no need to perform frequent setting operations for different objects and different scenarios, saving time and effort.

[0020] Secondly, for the processing of positioning and execution operations based on visual and tactile information, it is no longer necessary to strictly rely on the end-to-end processing method based on neural networks, enabling it to achieve higher versatility for different scenarios and different execution tasks, and no longer requiring a larger amount of data in the traditional method.

[0021] Therefore, it is possible to use low-cost detection devices such as tactile sensors and detection cameras to achieve high-precision measurement of the six-dimensional information of the target object in space, and can make full use of visual and tactile perception feedback to improve positioning efficiency and operation performance. At the same time, it can achieve rapid adaptation across different intelligent agent systems, has higher versatility, a wider application range, and can maintain a higher level of intelligence.

[0022] It should be understood that the above general description and the following specific embodiments are only exemplary and explanatory, and do not limit the scope of what the present invention intends to claim. Description of the Drawings

[0023] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0024] Figure 1 Schematically shows an application scenario diagram of a positioning execution control method, device, equipment, medium, and program product of an intelligent agent according to an embodiment of the present invention;

[0025] Figure 2 Schematically shows a flowchart of a positioning execution control method of an intelligent agent according to an embodiment of the present invention;

[0026] Figure 3 Schematically shows an application scenario flowchart of a positioning execution control method of an intelligent agent according to an embodiment of the present invention;

[0027] Figure 4A Schematically shows an application scenario diagram of plane position correction of a positioning execution control method of an intelligent agent according to an embodiment of the present invention;

[0028] Figure 4B Schematically shows an application scenario diagram of attitude angle correction of a positioning execution control method of an intelligent agent according to an embodiment of the present invention;

[0029] Figure 5A structural block diagram of a positioning execution control device of an agent according to an embodiment of the present invention is schematically shown; and

[0030] Figure 6 A block diagram of an electronic device suitable for implementing a positioning execution control method of an agent according to an embodiment of the present invention is schematically shown.

[0031] The above-mentioned drawings are a part of the specification of the embodiments of the present invention, which illustrate the exemplary embodiments of the present invention. The attached drawings and the description of the specification are used together to explain the principles of the embodiments of the present invention. It should be understood that the above general description of the drawings and the following specific embodiments are only exemplary and explanatory, and they do not limit the scope that the present invention intends to claim. Detailed Embodiments

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the spirit of the content disclosed by the present invention will be clearly described below with reference to the drawings and in detail. After any person skilled in the art understands the embodiments of the content of the present invention, they can make changes and modifications based on the technology taught by the content of the present invention, which do not depart from the spirit and scope of the content of the present invention.

[0033] The exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention. In addition, the same or similar reference numerals of elements / components used in the drawings and embodiments are used to represent the same or similar parts.

[0034] Regarding the "first", "second",... etc. used in the present invention, they do not particularly refer to the meaning of order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.

[0035] Regarding the directional terms used in the present invention, such as: up, down, left, right, front or back, etc., they are only references to the directions in the drawings. Therefore, the directional terms used are for explanation and not for limiting this creation.

[0036] Regarding the "comprising", "including", "having", "containing", etc. used in the present invention, they are all open-ended terms, that is, they mean including but not limited to.

[0037] Regarding the "and / or" used in the present invention, it includes any one or all combinations of the described things.

[0038] Regarding the "multiple" in the present invention, it includes "two" and "more than two"; regarding the "multiple groups" in the present invention, it includes "two groups" and "more than two groups".

[0039] Regarding the terms "substantially", "about", etc. used in the present invention, they are used to modify any quantity or error that can vary slightly, but these slight variations or errors do not change their essence. Generally, the range of such slight variations or errors modified by such terms can be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments, or other values. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.

[0040] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0041] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In cases where expressions similar to "at least one of A, B, or C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). Those skilled in the art should also understand that substantially any disjunctive conjunction and / or phrase representing two or more alternative items, whether in the specification, claims, or drawings, should be understood as presenting the possibility of including one of these items, either of these items, or both items. For example, the phrase "A or B" should be understood as including the possibility of "A" or "B", or "A and B".

[0042] As described in the background art, in the prior art, the solution of using visual-tactile fusion to achieve agent positioning and execution operations usually uses a camera for high-precision positioning, often requiring very expensive depth cameras or lidar. Moreover, the pre-calibration and installation steps before use are complex and cumbersome. In addition, for different objects and different sites, frequent settings and operations are required, which is time-consuming and laborious, resulting in high costs and complex operations for this method. Furthermore, most of the existing methods for positioning and operation processing using visual and tactile information are based on neural networks for end-to-end processing. For each scenario and each task, a large amount of data sets are often required. However, the models trained from the data sets of each specific scenario and specific object have poor generalization ability and can often only handle the same tasks, resulting in a larger amount of required data and poor versatility. Moreover, existing tactile sensors often need to use a specific neural network for filtering processing to obtain highly available tactile feedback information. However, this method will also further result in a large amount of required data and poor generality. For example, when changing different tactile sensors, re-training is required to be used, making the data processing method not universal.

[0043] For example, there is also a robot operation pose control method based on visual-tactile multi-scale positioning in the prior art. This method mainly includes three stages: rough positioning of the operation target object based on a visual sensor, fine positioning of the operation target object based on a tactile sensor, and calculation of the robot operation pose. First, the state of the operation environment and the target object to be operated is measured and obtained through a visual sensor. According to the collected image information and considering the safety distance, the estimated pose of the target object to be operated and the estimated target pose of the robotic arm system are initially obtained. Then, the relative pose deviation between the end of the robotic arm system and the target object to be operated is measured and obtained through a tactile sensor. Finally, according to the robot operation accuracy requirements, the estimated target pose of the robotic arm system is iteratively controlled and adjusted. The main purpose of this existing solution is to accurately control the robot operation pose through the visual-tactile multi-scale positioning method, making up for the problems existing in single visual measurement, such as being affected by light, background clutter interference, and overlap and occlusion between objects. By using tactile perception to increase local information, the pose estimation accuracy of the target object to be operated is improved. Among them, this existing technology still uses neural network modeling to complete the processing of displacement deviation and rotation deviation to achieve precise positioning of the operation target object. It is mainly used to improve the adaptability of the agent to complex environments, background interference, and object overlap and occlusion during positioning and execution, and cannot achieve the technical effects of reducing the agent positioning and execution cost and improving data versatility.

[0044] In view of at least one of the above problems, embodiments of the present invention aim to provide a positioning execution control method, device, equipment, medium and product of an intelligent agent with higher precision, which can achieve more general data processing and simpler operation at a lower cost. Thus, a method is provided that can use low-cost detection devices such as tactile sensors and detection cameras to achieve high-precision measurement of six-dimensional information of a target object in space, and can make full use of visual and tactile perception feedback to improve positioning efficiency and operation performance. At the same time, it can achieve rapid adaptation across different intelligent agent systems, thereby maintaining a higher level of intelligence.

[0045] One aspect of an embodiment of the present invention provides a positioning execution control method for an intelligent agent, which includes: determining the tactile detection range of the detection end of the intelligent agent according to the received initial visual image; controlling the detection end to perform tactile detection within the tactile detection range to generate a contact detection image; and generating a contact positioning image through the correction of the tactile detection process of the detection end, where the contact positioning image is used to control the execution end of the intelligent agent to perform positioning execution.

[0046] Figure 1 Schematically shows an application scenario diagram of a positioning execution control method, device, equipment, medium and program product of an intelligent agent according to an embodiment of the present invention.

[0047] As Figure 1 shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0048] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0049] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0050] Server 105 may be a server that provides various services. For example, it may be a background management server (only for example) that supports the websites browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0051] It should be noted that the positioning execution control method of the agent provided in the embodiments of the present invention can generally be executed by server 105. Correspondingly, the positioning execution control device of the agent provided in the embodiments of the present invention can generally be set in server 105. The positioning execution control method of the agent provided in the embodiments of the present invention can also be executed by a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105. Correspondingly, the positioning execution control device of the agent provided in the embodiments of the present invention can also be set in a server or a server cluster different from server 105 and capable of communicating with terminal devices 101, 102, 103 and / or server 105.

[0052] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0053] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 Based on the Figures 2 to 4B scenario described below, the positioning execution control method of the agent in the disclosed embodiments will be described in detail through

[0054] Figure 2 FIG. schematically shows a flowchart of the positioning execution control method of the agent according to an embodiment of the present invention.

[0055] As Figure 2 shown, one aspect of the embodiments of the present invention provides a positioning execution control method of an agent, which includes operations S201 to S203.

[0056] In operation S201, determine the tactile detection range of the detection end of the agent according to the received initial visual image;

[0057] In operation S202, control the detection end to perform tactile detection within the tactile detection range to generate a contact detection image; and

[0058] In operation S203, generate a contact positioning image through the correction of the tactile detection process of the detection end, and the contact positioning image is used to control the execution end of the agent to perform positioning execution.

[0059] The agent can be the execution subject of the above-mentioned positioning execution control method of the embodiments of the present invention, or the execution party controlled by the positioning execution control method. Specifically, it can be a humanoid intelligent robot or other artificial intelligence devices, which usually have actuators to complete specific action tasks by themselves. For example, a humanoid robot can use a mechanical dexterous hand to complete the action task of picking up an object, and even perform action tasks such as product packaging on an industrial production line. Specifically, embodied intelligence can perform operations similar to human lock-opening, accurately locate the position of the lock, and perform actions such as precisely inserting and turning the key to unlock.

[0060] In the embodiments of the present invention, the agent can respond to the generated action execution control instruction and perform the target action task. The target action task can be provided to the agent by the user according to the actual scenario requirements through the data input end of the agent (such as a microphone, a remote control, and a touch display device), or an operation task independently created by the agent according to its own working scenario requirements, based on the changes in the actual scenario or the action needs. According to the execution task, the agent can generate the above-mentioned action execution control instruction to control the actuator to complete the execution task.

[0061] The initial visual image can be a computer vision image obtained by the agent detecting the current environmental information of the current space (such as indoor or outdoor) according to the execution requirements of the target task, and its acquisition method can be through a visual sensor (such as a simple camera). The current environmental information can include the type, shape, size, position, quantity, color and other attribute information of the objects, people and other agents in the current space, and can also include information such as environmental brightness. For example, when a home humanoid robot follows the user home and arrives at the door, it can receive the voice command "open the door" issued by the user through the microphone, and respond to the voice command to create an execution task of "unlock and open the door", and detect the object information such as the position (such as the position of the door and the position of the lock on the door), type, size and dimensions of the objects in the current environment, as well as the information of people such as the number, appearance, body shape and action behavior of the people around. In addition, it can also detect the brightness and visibility of the current door.

[0062] The detection end of the agent can be the actuator of the agent with contact perception, such as a mechanical dexterous hand and an arm structure with simulated skin, etc., and there is no specific limitation.

[0063] The tactile detection range can be the range of the object area where the agent performs contact perception based on the target action task, and it is the operation area where the operation position for executing the operation is located. For example, for the target action task of "unlock and open the door", the tactile detection range can be the area occupied by the door lock in the area of the door. Through the image analysis of the initial visual image, the agent can distinguish the observed objects and people, and initially realize the recognition of the area where the execution position of the target action task is located. The area where the execution position of the target animal task is located can be understood as the above-mentioned tactile detection range. Therefore, through the acquisition and recognition of the initial visual image, the approximate spatial area and range where the positioning execution position is located can be quickly obtained.

[0064] In the embodiment of the present invention, the detection end structure can have a visual-tactile sensing structure including a visual sensor (such as a camera) and an optically designed structure. Of course, it can also include a corresponding piezoelectric sensor to meet the pressure perception, which can be specifically adjusted according to actual needs. For example, the optically designed structure can be an annular light belt composed of multiple LED lights. At least one camera is arranged inside the optically designed structure as a visual sensor to obtain the optical path change image of the optically designed structure. In addition, a silicone layer or a rubber layer can be laid outside the optically designed structure as a contact layer. The contact layer can further design some elastic recovery structures on the basis of the elastic recovery of the silicone layer or the rubber layer to achieve the rebound effect after contact separation, imitating real skin. When the contact surface of the contact layer contacts the object surface, the contact layer deforms, causing a physical change in the optically designed structure of the annular light belt, and then generating an optical path change. The camera can detect this optical path change in time and generate a contact detection image.

[0065] Among them, the detection end in contact with the object surface can have a contact surface, which is formed by the outer surface of the contact layer. During the tactile measurement process, the central position of the contact surface needs to fall within the above-mentioned tactile detection range.

[0066] Therefore, when the detection end performs tactile measurement within the tactile detection range, the detection image of the contact layer contacting a certain position within the tactile detection range can be obtained through the above-mentioned visual-tactile sensing structure. Among them, the contact layer changes due to contact force, driving the optically designed structure to generate an optical path change. This optical path change enables the visual sensor in the detection state to generate a corresponding detection image, that is, a contact detection image. In other words, the above method in the embodiment of the present invention can image the tactile information, thereby protecting the optically designed structure and the internal sensing sensors by means of the contact layer, and at the same time making it difficult for the tactile sensation of the visual-tactile sensor to be interfered by the external environment, ensuring the stability of the perception process and having better robustness.

[0067] It can be seen that the visualization of the sense of touch is realized by means of the contact detection image, ensuring the accurate positioning of the contact position between the detection end and the object surface corresponding to the tactile detection range in the case of contact perception. Therefore, if the contact detection image generated by the current touch measurement cannot be determined as the operation position for the final positioning execution, the touch measurement position can be further corrected, and the touch measurement can be performed again on the object surface within the tactile detection range to generate a new contact detection image, so as to further determine whether the contact position corresponding to the contact detection image is the operation position for the positioning execution.

[0068] Among them, although the initial visual image can provide the tactile detection range of the detection end and roughly determine the operation range for the positioning execution, it still cannot accurately locate the execution position. By further combining the touch measurement of the detection end on the tactile detection range, the position of the tactile detection range can be located and identified to ensure that the accurate operation position for the positioning execution can be obtained in the shortest time. For example, in the above-mentioned target action task of "unlocking and opening the door" by the intelligent agent, the spatial area of the door and the spatial area of the door lock on the door (i.e., the tactile detection range) can be determined through the initial visual image, and then the specific finger of the dexterous hand is used to repeatedly perform touch measurements in the spatial area of the door lock. If the final operation position for the positioning execution cannot be located during each touch measurement, the touch measurement position and the detection posture are corrected, and the touch measurement is performed at other positions within the tactile detection range. Finally, the final contact detection image can be obtained through the visual-tactile sensor as the contact positioning image, and this contact positioning image can be the imaging of the object surface where the target action task to be executed at the operation position for the positioning execution is formed. Therefore, the data continuously parsed from the contact detection image can be used as the basis for correcting the touch measurement position and the detection posture of the detection end during the touch measurement process.

[0069] In summary, according to the above positioning execution control method of the embodiment of the present invention, by generating the contact detection image and correcting the touch measurement process of the detection end according to the contact detection image, it is possible to obtain the initial visual image through a low-cost camera to roughly estimate the object pose to obtain the tactile detection range, and then cooperate with the visual-tactile sensing structure of the detection end to continuously correct and adjust the touch measurement within the tactile detection range, so that the object and the visual-tactile sensing structure are fully contacted to form a detection image, that is, the accurate calculation of the operation position for the positioning execution can be realized, and thus the execution end can be controlled to perform the task at this operation position. For example, the keyhole can be the execution position of the intelligent agent, and the intelligent agent can control the key to be inserted into the keyhole and perform a rotation operation to achieve the target action task of "opening the door".

[0070] Among them, the execution end of the intelligent agent can be another actuator relative to the detection end, or the detection end itself (i.e., the case where the detection end serves as the execution end at the same time). For example, the left hand of the intelligent agent can be used as the detection end, the right hand can be used as the execution end, or the same hand can be used as the detection end and the execution end at the same time, without specific restrictions.

[0071] In summary, compared with the traditional method of realizing positioning and execution based on neural network computing, the above positioning and execution control method of the embodiments of the present invention can adopt only lower-cost visual and tactile perception devices (visual-tactile perception structures) to achieve high-precision positioning and execution operations. Among them, operations such as installation and calibration of these perception devices are simpler and more convenient, and there is no need to perform frequent setting operations for different objects and different scenarios, saving time and effort. Moreover, the processing of positioning and execution operations based on visual information and tactile information no longer needs to strictly rely on the end-to-end processing method based on neural networks, making it more versatile for different scenarios and different execution tasks, and no longer requiring a larger amount of data in the traditional method. Therefore, it can use low-cost tactile sensors and detection cameras and other detection devices to achieve high-precision measurement of the six-dimensional information of the target object in space, and can make full use of visual and tactile perception feedback to improve the positioning efficiency and operation performance. At the same time, it can quickly adapt across different intelligent agent systems, with higher stability and versatility, a wider application range, and can maintain a higher level of intelligence.

[0072] Figure 3 Schematically shows a flowchart of an application scenario of the positioning and execution control method of the intelligent agent according to an embodiment of the present invention.

[0073] Such as Figure 2 and Figure 3 As shown, according to an embodiment of the present invention, before determining the tactile detection range of the detection end of the intelligent agent according to the received initial visual image in operation S201, it further includes:

[0074] Perform visual detection on the target detection space where the intelligent agent is located to generate an initial visual image.

[0075] Such as Figure 3 In operation 310 shown, the target detection space can be the space where the intelligent agent performs visual detection for the positioning and execution of the target action task. For example, if the target action task of the intelligent agent is "unlock and open the door", and the intelligent agent stands in the corridor outside the door at this time, the space of this corridor can be used as the target detection space. Visual detection can be understood as the "watching" process of the intelligent agent, that is, the process of the intelligent agent using its own visual sensor (such as a simple low-cost RGB camera) to perform image detection on the surrounding environment or the target object.

[0076] In response to the target action task, the agent can obtain the corresponding initial visual image I at the current moment through "observing" the target detection space. camera , for example, in the space where the agent is currently located, the detection image of the target object in front of the agent. For example, for the target action task of "unlock and open the door", the initial visual image may include the door in front of the agent, the wall around the door, and the extended environment of part of the corridor. It may also include some humans, objects, or even pets at the same time.

[0077] Therefore, even very inexpensive visual sensors can easily obtain the initial visual image, thus greatly reducing the image acquisition cost compared to the existing visual perception technologies that usually adopt high costs and complex settings. At the same time, it also reduces the complexity of parameter settings and greatly simplifies the detection method.

[0078] As Figure 2 and Figure 3 shown, according to an embodiment of the present invention, in operation S201, determining the tactile detection range of the detection end of the agent according to the received initial visual image includes:

[0079] Performing object position parsing on the initial visual image according to the target execution object;

[0080] Determining the tactile detection range according to the object position parsing.

[0081] The target execution object can be the target object that the agent operates to complete the positioning execution. For example, in the case where the target action task is "unlock and open the door" as described above, the lock can be the target execution object. The agent itself has corresponding image parsing and processing capabilities, and can realize the segmentation and recognition of the objects in the initial visual image, and recognize the target objects in the image. At the same time, the agent can also obtain the position information of each object recognized from the initial visual image in the target detection space coordinate system (and can also recognize its type, size, shape and other related information) according to the "observation" data (such as space coordinate data) of the target detection space, that is, perform object position parsing.

[0082] Among them, for the recognition of the target execution object in the initial visual image, the agent can compare the feature of the image data related to the target execution object crawled from the network in real time or stored in the historical database with the object image segmented and extracted from the initial visual image, so as to recognize the object with the closest feature to the target execution object and obtain the spatial position of the object in the target detection space, which will not be elaborated here.

[0083] Furthermore, since the target execution object can be obtained by parsing and extracting the initial visual image, and at the same time, the position information of the target execution object in the target detection space can be determined. Based on these position information, the position and size of the space area occupied by the target execution object in the target detection space can be determined, that is, the tactile detection range can be obtained.

[0084] Specifically, the position coordinates of each pixel point of the target execution object in the initial visual image mapped to the space can be obtained. Since a camera with poor performance is used, it is possible that light interference phenomena such as shadows generated by the target execution object will also be generally regarded by the agent as being within the area of the tactile detection range.

[0085] Therefore, through the above object position parsing of the initial visual image, a relatively simple object detection technology can be adopted, that is, the tactile detection range can be roughly determined directly, quickly and effectively.

[0086] Such as Figure 2 and Figure 3 As shown, according to an embodiment of the present invention, in operation S202, controlling the detection end to perform tactile measurement within the tactile detection range to generate a contact detection image includes:

[0087] Obtaining the tactile measurement execution position within the tactile detection range;

[0088] Controlling the detection end to perform tactile measurement on the tactile measurement execution position to generate a contact detection image.

[0089] Such as Figure 3 As shown in operation S320, after determining the tactile detection range in which the agent detection end can perform tactile measurement, it is also necessary to consider the tactile measurement position (i.e., the tactile measurement execution position) on the surface of the object within the area defined by the tactile detection range each time the detection end performs tactile measurement, and the tactile measurement execution position must fall within the tactile detection range. Those skilled in the art should understand that if the tactile measurement execution position is a contact surface, at least the center point of the contact surface must fall within the tactile detection range.

[0090] Generally, the agent can roughly determine a contact position that is most likely to approach the positioning execution position as the starting tactile measurement position of the tactile measurement based on the comparison result between the overall shape of the tactile detection range and the overall shape of the target execution object.

[0091] For example, in the case where the target action task is "unlock and open the door", the lock, as the target execution object, can be a spherical object. Through the recognition and extraction of the initial visual image, the agent can roughly determine an approximately elliptical or circular image area as the tactile detection range based on the general position of the lock on the door (e.g., the lock is generally located in the lower middle of the door, near the edge of the door frame, and below the doorknob). Further, the keyhole can be the operation position for the positioning execution of the target execution object. Therefore, the agent can further determine the approximate position of the keyhole in the defined area of the tactile detection range according to the central position of the keyhole in the spherical lock, and this approximate position can be used as the starting tactile detection position corresponding to the tactile execution position.

[0092] For the tactile detection process after data correction, the tactile execution position is the tactile detection position generated after data correction based on the previous tactile execution position, and this tactile detection position will also fall within the tactile detection range.

[0093] According to the above tactile detection range and corresponding tactile execution position and other tactile detection information, the controller of the detection end (such as Exploration Arm Controler) can further perform tactile detection on the tactile execution position within the tactile detection range according to the current execution state of the detection end (such as Exploration Arm). Among them, the current execution state of the detection end can include state information related to the posture of the detection end, such as the name of the action currently executed by the detection end, motion parameters (such as the rotation angles of each joint of the finger), spatial position, and the angle of the contact surface.

[0094] Therefore, according to the current execution state of the detection end, the detection end can perform accurate calculation on the required posture adjustment data to meet the tactile detection requirements for the tactile execution position, and thus generate a corresponding contact detection image (i.e., Contact Image) I according to the visual-tactile sensing structure of the tactile contact surface. tactile 。

[0095] Among them, the contact detection image can be the image obtained by the detection end performing tactile detection based on continuous posture position correction. Only the last contact detection image obtained when the detection end no longer needs to perform posture position correction can be used as the contact positioning image. Therefore, the contact detection image can be the basis for continuously performing tactile correction on the above detection end.

[0096] Therefore, by determining the above-mentioned touch measurement execution position, it is possible to intelligently determine a touch measurement closer to the positioning execution position, thereby greatly reducing the number of corrections to the touch measurement process and the correction amplitude, enabling the detection resource of the detection end to be less occupied, and accurately obtaining the contact positioning image in the shortest time. Further, performing touch measurement correction according to the current execution state of the detection end can enable each touch measurement to obtain a more accurate and stable contact detection image.

[0097] As Figure 2 and Figure 3 shown, according to an embodiment of the present invention, in the operation S203 of generating a contact positioning image by correcting the touch measurement process of the detection end through a contact detection image, it includes:

[0098] Obtaining plane position correction data and attitude angle correction data through the contact detection image;

[0099] According to the plane position correction data, updating the touch measurement execution position of the detection end within the tactile detection range; and according to the attitude angle correction data, adjusting the touch measurement angle when the detection end performs touch measurement at the updated touch measurement execution position;

[0100] Controlling the detection end to perform touch measurement on the updated touch measurement execution position according to the touch measurement angle until a contact positioning image is generated.

[0101] In Figure 3 the operations S330 - S331 and S332 shown, the touch measurement position and touch measurement attitude of the next touch measurement of the detection end can be corrected according to the analysis of the contact detection image. Among them, a visual touch image of a target execution object with a target execution position can be obtained in advance through simulation (such as Taxim Simulation) using a model such as Model Mesh, and this visual touch image can be used as a detection comparison image (Target Image) of the target execution object.

[0102] Compared with directly filtering the contact detection image through a neural network, in the embodiment of the present invention, a differential amplification process can be performed on the contact detection image and the above-mentioned detection comparison image to generate a detection process image. Since the foreground and background of this detection process image are significantly different and the sensitivity can be adjusted, it has higher versatility and strong migration ability, and also avoids the situation that the contact detection image after neural network filtering requires a large amount of data to be re - collected and retrained for each sensor.

[0103] Considering that in practical applications, it is necessary to consider that the initial visual image can only provide a rough position for the positioning and execution of the target execution object. Moreover, when applied to different detection ends of different intelligent agents, due to differences in the structural design of the visual-tactile sensing structure, differences in the hardness of the contact surface, etc., it will bring huge computational errors to the theoretical data, thus seriously affecting the accurate judgment of the touch position, increasing the number of touch measurement processes, and prolonging the touch measurement time.

[0104] Therefore, it is necessary to consider making the object surface area in the tactile detection range of the target execution object and the visual-tactile sensing structure achieve full contact during the touch measurement process to avoid the consequences brought by the above errors. And the full contact of the touch measurement needs to meet the following two situations: The first is that a partial area corresponding to the tactile detection range of the object surface exceeds the measurement range of the visual-tactile sensing structure. At this time, it is necessary to correct the touch measurement position of the visual-tactile sensing structure of the detection end in the plane direction to make it possible for the two to be in full contact; The second is that although there is contact between the object surface corresponding to the tactile detection range and the contact surface of the visual-tactile sensing structure, there is an inclined angle. At this time, it is necessary to adjust the posture of the detection end so that the inclined angle formed between the contact surfaces of the visual-tactile sensing structure is 0 to make it possible for the two to be in full contact.

[0105] By performing image analysis on the detection processing image generated after processing the above contact detection image and the above detection comparison image (Target Image) respectively, and comparing and estimating (Compare and Evaluate) the image analysis data of the two, the plane position deviation and attitude angle deviation of the detection end corresponding to the touch measurement position during the acquisition process of the current contact detection image are obtained. For example, as the comparison object of the contact detection image, the detection comparison image can be the simulation image closest to the contact positioning image. By performing image analysis on it, the standard position of each pixel point can be obtained. Therefore, through the difference between the detection position of each pixel point obtained by the image analysis of the contact detection image and this standard position, it can be determined whether there are graphic deficiencies (such as partial contact areas not within the tactile detection range) and offsets (such as the center of the contact area being offset), and whether there are deformations in the graphics in the image (such as the existence of an angle due to the uneven contact surface) in the contact detection image. Thus, the corresponding plane position correction data (i.e., Plane Correction) (Δx, Δy) and attitude angle correction data (i.e., Orientation) R(θ, v) can be obtained very accurately.

[0106] When the values of the above plane position correction data (Δx, Δy) exceed a certain threshold, it can be determined that a position deviation has occurred. At this time, such as Figure 3For the operation S332 shown above, the above planar position correction data (Δx, Δy) is sent to the probe controller (such as Exploration Arm Controler), so that the probe controller can correct the execution position during the next touch measurement according to the current execution state of the probe at the current moment. The corrected position can be the touch measurement position that the probe needs to execute during the next touch measurement execution process.

[0107] Correspondingly, when the value of the above attitude angle correction data R(θ, v) exceeds a certain threshold, it can be determined that an angle deviation has occurred. At this time, as Figure 3 For the operation S331 shown above, the above attitude angle correction data R(θ, v) is sent to the probe controller, so that the probe controller can correct the touch measurement angle of the probe during the next touch measurement according to the current execution state of the probe at the current moment. Specifically, it can be the angle adjustment of the contact surface of the visual touch sensing structure. The corrected angle can be the touch measurement angle of the contact surface of the visual touch sensor that the probe needs to execute during the next touch measurement execution process.

[0108] Among them, the above planar position correction and attitude angle correction can be carried out simultaneously. The probe will update the execution position of the next touch measurement according to the planar position correction data to obtain a new touch measurement execution position. At the same time, it will also correct the inclination angle of the contact surface of the visual touch sensing structure when performing the next touch measurement at the new touch measurement execution position, so as to ensure that the next touch measurement execution is carried out strictly according to the new touch measurement execution position and the inclination angle of the new contact surface.

[0109] Of course, the planar position correction and attitude angle correction can also be carried out separately. This situation generally occurs when one aspect does not need to be corrected. For example, if there is a planar position deviation but no angle deviation, only the planar position needs to be corrected.

[0110] Therefore, the probe only needs to perform the next touch measurement according to the corrected touch measurement angle and the corresponding touch measurement execution position, and generate a new contact detection image. Further, based on the above deviation judgment and deviation correction of the new contact detection image, perform another touch measurement execution according to the latest correction data... until a contact detection image without any deviation can be generated as the contact positioning image. During this process, the probe also needs to continuously adjust the magnitude of the required contact force according to the needs of deviation correction, so as to ensure a more sufficient contact effect with the surface of the touch measurement execution position on the basis of deviation correction.

[0111] Therefore, through the analysis and judgment of the contact detection image each time, the next touch measurement process can be corrected according to the judgment result, so as to achieve repeated correction of the touch measurement process of the detection end, making the contact surface of the visual touch sensor of the detection end continuously approach the target area on the target execution object, generating the final contact positioning image, and at this time the touch measurement correction process can be terminated.

[0112] Therefore, by means of the above repeated correction of the touch measurement process, even without using the traditional calculation method based on the large data volume neural network model, high-precision positioning detection can be obtained through a lower-cost sensing technology, effectively avoiding the traditional end-to-end processing solution, and the positioning detection process can be realized for different scenarios and different tasks based on different agents, achieving higher versatility, and at the same time avoiding more data volume and no longer requiring frequent parameter settings and data changes.

[0113] Figure 4A The application scenario diagram of the planar position correction of the positioning execution control method of the agent according to the embodiment of the present invention is schematically shown.

[0114] Such as Figure 2 and Figure 3 shown, according to an embodiment of the present invention, before obtaining the planar position correction data and the attitude angle correction data through the contact detection image, it further includes:

[0115] Identifying the target imaging area of the target execution object in the contact detection image.

[0116] Such as Figure 4A shown, before obtaining the planar position correction data, it is necessary to perform image analysis on the contact detection image to obtain the area in the contact detection image that reflects the target operation position, and this area is the area where the target execution object performs the positioning execution operation, that is, the target imaging area. By analyzing and extracting the features of the image, the position of the target imaging area in the contact detection image can be determined, such as Figure 4A shown as the concentric circle area (Circle 0.97, that is, circle 0.97) enclosed by the blue square frame shown. Identifying the target imaging area can be realized by using a YOLO model trained with a small amount of data.

[0117] It should be noted that, since the visual touch sensing structure has an optical path structure (such as a circular LED light strip), the imaging of the contact detection image within the tactile detection range of the target execution object is not consistent with the actual object surface of the target execution object, but the contact detection image is generated based on the image generated by the change of the optical path of the optical path structure. For example Figure 4A shown as the color concentric ring image.

[0118] Therefore, the target imaging area can be quickly identified in a very simple manner, so as to determine the imaging position of the target execution object in the contact detection image, enabling the subsequent determination of its deviation correction data more quickly and ensuring the accuracy of its deviation correction data.

[0119] As Figures 2 to 4A shown, according to an embodiment of the present invention, in obtaining the planar position correction data through the contact detection image, it includes:

[0120] Comparing the target imaging area and the imaging field of view area of the contact detection image to obtain the planar position correction data.

[0121] As Figure 4A shown, by means of the identification of the target imaging area, the position of the target execution object in the contact detection image formed by the visual-tactile sensing structure can be determined. Among them, the imaging range of the contact detection image can be determined according to the imaging field of view of the visual-tactile sensing structure. As Figure 4A shown, the largest black background area can be regarded as the imaging area of the imaging field of view of the visual-tactile sensing structure, that is, the imaging field of view area. Among them, as Figure 4A shown, the center of the target imaging area is B, and the center of the imaging field of view area is A. When the two do not coincide, it can indicate that there is a planar position deviation in the corresponding touch measurement execution.

[0122] Therefore, the planar position correction data can be expressed by the following formula (1):

[0123]

[0124] Among them, b refers to the reference coordinate system of the detection end, ee1 represents the end of the detection end, tac refers to the visual-tactile sensing structure, respectively represent the position coordinates of the center of the target imaging area and the center of the imaging field of view area in the image, and represents the distance corresponding to each pixel point in reality. The T in the upper left corner represents the reference coordinate system, the lower left corner represents the target coordinate system, and T represents the conversion relationship from the reference coordinate system to the target coordinate system.

[0125] Among them, the position of the center B of the target imaging area in the imaging field of view area can also represent the plane position correction data required in the next touch measurement correction process. Moreover, the calculation result of the area size of the target imaging area in the imaging field of view area can determine whether the target imaging area exceeds the imaging field of view. In other words, when the corresponding detection end touches, whether at least part of it exceeds the boundary of the tactile detection range or exceeds the boundary of the imaging field of view. Therefore, for the next touch measurement execution process after corresponding correction, from the imaging visual angle, the detection end should move obliquely from the center A of the detection imaging field of view towards the target imaging area B to complete the position correction. The essence of the above position correction is to make the center B of the target imaging area in the contact detection image generated by the next touch measurement execution coincide with the center A of the imaging field of view area.

[0126] Therefore, through a very simple technical principle, the accurate acquisition of plane position correction data can be quickly achieved, the high-precision measurement of the plane information of the target object in space can be realized, thereby ensuring the positioning efficiency and subsequent operation performance, and the rapid adaptation across different intelligent agent systems can be achieved, with higher versatility.

[0127] Figure 4B The application scenario diagram of the attitude angle correction of the positioning execution control method of the intelligent agent according to the embodiment of the present invention is schematically shown.

[0128] Such as Figures 2 to 4B As shown, according to an embodiment of the present invention, in obtaining the attitude angle correction data through the contact detection image, it includes:

[0129] Extract the filtered contact image of the contact detection image;

[0130] Obtain the corresponding relationship between each pixel point on the filtered contact image and the preset simulated contact image, and obtain the attitude angle correction data.

[0131] Such as such Figure 4B As shown, after filtering the contact detection image through a custom filter (Custom Filter), a filtered contact image can be generated. Further image analysis processing is performed on the filtered contact image to extract the pixel information of each imaging pixel point in the filtered contact image, and the pixel information may include the position information of the corresponding pixel point. Among them, the graphics in the filtered contact image may be due to the existence of an angle on the contact surface of the detection end that cannot be fully contacted, resulting in incomplete imaging and possible graphic distortion. As Figure 4B As shown, only the upper left corner pattern of a partial ring is shown.

[0132] The preset simulated contact image can be the detection comparison image (Target Image) mentioned above. Similar to the image analysis and processing process of the above-mentioned filtered contact image, the pixel information of each imaging pixel point of the preset simulated contact image can be extracted.

[0133] Furthermore, compare the pixel information of the pixel points of the filtered contact image with the pixel information of the pixel points of the preset simulated contact image and establish a corresponding relationship between the pixel points. For example, the pixel information of the preset simulated contact image can be used as a reference standard to calculate the comparison score of each pixel point of the filtered contact image. Then, the positional relationship between the filtered contact image and the preset simulated contact image can be determined, that is, the angular difference between the filtered contact image and the imaging requirements actually to be achieved can be obtained. Furthermore, the above-mentioned attitude angle correction data can be obtained, and the angle by which the detection end needs to adjust the contact surface of the visual and tactile sensing structure according to the current execution state during the next touch measurement execution process can be determined.

[0134] Among them, the attitude angle correction data satisfies the following formulas (2)-(4):

[0135]

[0136] Among them, β is the inclination angle of rotation:

[0137] As Figure 4B shown, the figure circled by the white square at the bottom right can be the imaging of the target execution object parsed from the filtered contact image. At this time, if angle correction is to be performed, it is necessary to rotate the figure to the other side according to a virtual rotation axis to complete the imaging, and the actual required rotation angle is the correction value of the attitude angle.

[0138] Therefore, the attitude angle correction data can be obtained quickly and accurately, realizing high-precision measurement of the deviation angle information of the target object in space. Furthermore, the positioning efficiency and subsequent operation performance can be guaranteed, and rapid adaptation across different intelligent agent systems can be achieved, with higher versatility. During this process, there is no longer a need for large amounts of data training and processing, the cost is lower, and the accuracy of the data can be guaranteed, without any negative impact on the intelligence level of the intelligent agent.

[0139] In summary, the positioning execution control method of the intelligent agent provided by the embodiments of the present invention can at least partially solve at least one of the technical problems existing in the related art, and thus can at least achieve one of the following technical effects:

[0140] First of all, only lower-cost visual and tactile perception devices can be used to achieve high-precision positioning execution operations. Among them, the installation and calibration operations of these perception devices are simpler and more convenient, and there is no need to perform frequent setting operations for different objects and different scenarios, saving time and effort.

[0141] Secondly, the processing of performing positioning operations based on visual information and tactile information no longer needs to strictly rely on the end-to-end processing method based on neural networks, enabling it to achieve higher generality for different scenarios and different execution tasks, and no longer requiring a larger amount of data in the traditional method.

[0142] Therefore, even with low-cost detection devices such as tactile sensors and detection cameras, it is possible to achieve high-precision measurement of the six-dimensional information of the target object in space, and the visual and tactile perception feedback can be fully utilized to improve the positioning efficiency and operation performance. At the same time, it can achieve rapid adaptation across different intelligent agent systems, has higher generality, a wider application range, and can maintain a higher level of intelligence.

[0143] Based on the above positioning execution control method of the intelligent agent, the present invention also provides a positioning execution control device for the intelligent agent. The following will be combined with Figure 5 to describe the device in detail.

[0144] Figure 5 The structural block diagram of the positioning execution control device for the intelligent agent according to an embodiment of the present invention is schematically shown.

[0145] As Figure 5 shown, the positioning execution control device 500 for the intelligent agent in this embodiment includes a range determination module 510, an image generation module 520, and a touch measurement correction module 530.

[0146] The range determination module 510 is used to determine the tactile detection range of the detection end of the intelligent agent according to the received initial visual image. In one embodiment, the range determination module 510 can be used to perform the operation S201 described above, which will not be elaborated here.

[0147] The image generation module 520 is used to control the detection end to perform touch measurement within the tactile detection range to generate a contact detection image. In one embodiment, the image generation module 520 can be used to perform the operation S202 described above, which will not be elaborated here.

[0148] The touch measurement correction module 530 is used to correct the touch measurement process of the detection end through the contact detection image to generate a contact positioning image, and the contact positioning image is used to control the execution end of the intelligent agent to perform positioning and execution. In one embodiment, the touch measurement correction module 530 can be used to perform the operation S203 described above, which will not be elaborated here.

[0149] According to an embodiment of the present invention, any of the range determination module 510, the image generation module 520, and the touch measurement correction module 530 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the range determination module 510, the image generation module 520, and the touch measurement correction module 530 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the range determination module 510, the image generation module 520, and the touch measurement correction module 530 may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.

[0150] Figure 6 A block diagram of an electronic device suitable for implementing the positioning execution control method of an agent according to an embodiment of the present invention is schematically shown.

[0151] The above-mentioned electronic device provided by the embodiment of the present invention includes one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned positioning execution control method of the agent.

[0152] As Figure 6 shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which may execute various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 601 may also include on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for executing different actions of the method flow according to an embodiment of the present invention.

[0153] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present invention by executing programs in the ROM 602 and / or the RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiments of the present invention by executing programs stored in the one or more memories.

[0154] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input part 606 including a keyboard, a mouse, etc.; an output part 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 608 including a hard disk, etc.; and a communication part 609 including a network interface card such as a LAN card, a modem, etc. The communication part 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage part 608 as needed.

[0155] The present invention also provides a computer-readable storage medium having executable instructions stored thereon, and when the instructions are executed by a processor, the processor is caused to execute the above-mentioned positioning execution control method of the agent.

[0156] Among them, the computer-readable storage medium may be included in the device / apparatus / system described in the above embodiments; or it may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0157] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the above-described ROM 602 and / or RAM 603 and / or one or more memories other than ROM 602 and RAM 603.

[0158] An embodiment of the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the above-mentioned positioning execution control method of the intelligent agent.

[0159] Wherein, the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the method provided by the embodiment of the present invention.

[0160] When the computer program is executed by the processor 601, it executes the above-mentioned functions defined in the system / device of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, modules, units, etc. can be implemented by computer program modules.

[0161] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 609, and / or be installed from the removable medium 611. The program code contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0162] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or be installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above-mentioned functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0163] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0165] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present invention can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly recited in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features recited in the various embodiments and / or claims of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0166] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A positioning execution control method for an intelligent agent, characterized in that: include: Determining the tactile detection range of the detection end of the agent according to the received initial visual image, wherein the initial visual image is a computer vision image obtained by the agent through the visual sensor to detect the current environment information of the current space according to the execution requirements of the target action task; Controlling the detection end to perform touch detection on a touch detection execution position within the tactile detection range to generate a touch detection image, wherein the touch detection image is an image obtained by the detection end performing touch detection based on continuous posture position correction; as well as Correcting the touch detection process of the detection end through the contact detection image to generate a contact positioning image, and the contact positioning image is used to control the execution end of the intelligent body to perform positioning execution, which includes: obtaining plane position correction data and posture angle correction data through the contact detection image; According to the plane position correction data, the touch detection execution position of the detection end within the tactile detection range is updated; and according to the posture angle correction data, the touch detection angle of the detection end when performing touch detection at the touch detection execution position after the update is adjusted; the detection end is controlled to perform touch detection on the touch detection execution position after the update according to the touch detection angle until the contact positioning image is generated, wherein the contact positioning image is an imaging of the surface of the object to be performed by the target action task formed by the positioning execution operation position.

2. The method according to claim 1, characterized in that Before determining the tactile detection range of the detection end of the intelligent agent according to the received initial visual image, the method further includes: Perform visual detection on the target detection space where the intelligent agent is located to generate the initial visual image.

3. The method according to claim 1, characterized in that: In determining the tactile detection range of the detection end of the intelligent agent according to the received initial visual image, the method includes: performing object position parsing on the initial visual image according to a target execution object; The tactile detection range is determined according to the object position analysis.

4. The method according to claim 1, characterized in that Before acquiring the plane position correction data and the posture angle correction data through the contact detection image, the method further includes: A target imaging region of a target execution object in the contact detection image is identified.

5. The method according to claim 4, characterized in that In the step of acquiring the plane position correction data through the contact detection image, the method comprises: The target imaging area and the imaging field area of ​​the contact detection image are compared to obtain the plane position correction data.

6. The method according to claim 1, characterized in that In the step of acquiring posture angle correction data through the contact detection image, the method includes: extracting a filtered contact image of the contact detection image; The corresponding relationship between each pixel point on the filtered contact image and the preset simulated contact image is obtained to obtain the posture angle correction data.

7. A positioning execution control device for an intelligent body, characterized in that: include: A range determination module, used to determine the tactile detection range of the detection end of the agent according to the received initial visual image, wherein the initial visual image is a computer vision image obtained by the agent through the visual sensor to detect the current environment information of the current space according to the execution requirements of the target action task; An image generation module, used for controlling the detection end to perform touch detection on the touch detection execution position within the tactile detection range to generate a touch detection image, wherein the touch detection image is an image obtained by the detection end performing touch detection based on continuous posture position correction; as well as A touch detection correction module is used to correct the touch detection process of the detection end through the contact detection image to generate a contact positioning image, and the contact positioning image is used to control the execution end of the intelligent body to perform positioning execution, including: obtaining plane position correction data and posture angle correction data through the contact detection image; According to the plane position correction data, the touch detection execution position of the detection end within the tactile detection range is updated; and according to the posture angle correction data, the touch detection angle of the detection end when performing touch detection at the touch detection execution position after the update is adjusted; the detection end is controlled to perform touch detection on the touch detection execution position after the update according to the touch detection angle until the contact positioning image is generated, wherein the contact positioning image is an imaging of the surface of the object to be performed by the target action task formed by the positioning execution operation position.

8. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Mobile vision robot and measurement and control method thereof

    CN106607907A

  • Tactile sensor, robot, elastomer, object sensing method and computing equipment

    CN112304248A