Task processing method and device for robot, electronic equipment and robot

By calling large language models to generate operation sequences and combining images and multimodal operation information to generate instruction sequences, the problem of robots being difficult to understand and perform flexible and variable tasks is solved, and higher task processing accuracy and user experience are achieved.

CN120116232AActive Publication Date: 2025-06-10JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510608124.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-10
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The difficulty of robots to understand and accurately perform flexible and changeable tasks, which affects the user experience.

Method used

By calling the large language model, an operation sequence is generated based on the task information and text operation information input by the user, and a command sequence is generated in combination with the control panel image and multimodal operation information to control the robot to accurately operate each area.

Benefits of technology

It realizes that the robot can flexibly understand and execute complex task information, improving the accuracy and user experience of task processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120116232A_ABST
    Figure CN120116232A_ABST
Patent Text Reader

Abstract

The invention provides a task processing method and device for a robot, electronic equipment and the robot, and relates to the technical fields of artificial intelligence, trusted artificial intelligence, intelligent robots, intelligent equipment and the like. The method comprises the steps that in response to received task information input by a user, a first control panel image of an operated object indicated by the task information and multi-modal operation information are obtained, and the multi-modal operation information comprises text operation information; calling a large language model, and generating an operation sequence according to the task information and the text operation information; generating an instruction sequence according to the operation sequence, the first control panel image and the multi-mode operation information; and controlling the robot to operate each operation area according to the instruction sequence so as to process a task corresponding to the task information. The technical problem that the use experience is limited due to the fact that the robot difficultly and accurately executes the task can be solved, it is guaranteed that the robot can accurately control the operation area, and the technical effect of improving the use experience is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence, trustworthy artificial intelligence, intelligent robots, intelligent equipment, etc., and relates to a task processing method, device, electronic device and robot for a robot. Background Art

[0002] With the rapid development of artificial intelligence, robots are increasingly being applied in many fields such as industrial and agricultural production, construction, logistics, and daily life. The field of embodied intelligence has developed rapidly, and robots with embodied intelligence functions can interact with the environment based on the user's task information to realize the operation of equipment by the robot. The equipment that can be operated by robots in daily life usually includes a variety of control panels, and the control panels usually have various operation logics.

[0003] In the process of implementing the inventive concept of the present invention, the inventors found that there are at least the following problems in the related art: Robots usually learn to perform tasks by learning pre-constructed end-to-end instances. However, for flexible and variable task information, it is difficult for robots to understand or accurately execute the tasks corresponding to the task information, thus affecting the user experience. Summary of the Invention

[0004] In view of this, the present invention provides a task processing method, device, electronic device and robot for a robot.

[0005] One aspect of the present invention provides a task processing method for a robot, including: in response to receiving task information input by a user, obtaining a first control panel image of an object to be operated indicated by the task information and multimodal operation information, where the multimodal operation information includes text operation information for operating the object to be operated; calling a large language model, and generating an operation sequence according to the task information and the text operation information, where the operation sequence includes at least one operation area; generating an instruction sequence according to the operation sequence, the first control panel image and the multimodal operation information, where the instruction sequence includes control instructions corresponding to each operation area; and controlling the robot to operate each operation area according to the instruction sequence to process the task corresponding to the task information.

[0006] According to an embodiment of the present invention, calling a large language model and generating an operation sequence according to the task information and the text operation information includes: obtaining prompt information, where the prompt information includes an example operation sequence, at least one first blank field corresponding to the object to be operated, and a second blank field corresponding to the task information; respectively filling the first blank field and the second blank field with the text operation information and the task information to obtain updated prompt information; and inputting the updated prompt information into the large language model to output an operation sequence in the same format as the example operation sequence.

[0007] According to an embodiment of the present invention, the first vacant field and the second vacant field are respectively filled with text operation information and task information to obtain updated prompt information, including: determining a target field value matching each first vacant field from the text operation information according to the field name of the first vacant field in the prompt information; and filling the target field value into the first vacant field and filling the task information into the second vacant field to obtain updated prompt information.

[0008] According to an embodiment of the present invention, the multimodal operation information further includes a pre-collected second control panel image, and the first control panel image is collected by the robot in response to receiving the input task information; an instruction sequence is generated according to the operation sequence, the first control panel image, and the multimodal operation information, including: determining a position transformation relationship according to the first control panel image and the second control panel image, where the position transformation relationship represents the position transformation relationship between the coordinate system of the second control panel image and the coordinate system of the first control panel image; determining the depth information of each operation area according to the first control panel image; and generating a control instruction according to the depth information, the position transformation relationship, and the text operation information.

[0009] According to an embodiment of the present invention, determining the position transformation relationship according to the first control panel image and the second control panel image includes: performing feature registration on the first control panel image and the second control panel image to obtain at least one pair of feature points; and determining the position transformation relationship according to the two-dimensional position information of the two feature points in the pair of feature points.

[0010] According to an embodiment of the present invention, the text operation information includes the two-dimensional position information of each operation area in the second control panel image; generating a control instruction according to the depth information, the position transformation relationship, and the text operation information includes: for each operation area, determining the two-dimensional position information of the operation area in the robot coordinate system according to the two-dimensional position information of the operation area in the second control panel image and the position transformation relationship; determining the three-dimensional position information of the operation area according to the depth information and the two-dimensional position information of the operation area in the robot coordinate system; and generating a control instruction according to the three-dimensional position information.

[0011] According to an embodiment of the present invention, the method further includes: calling a region detection model to perform region detection on the second control panel image to obtain first identification information of at least one operation area; calling a fine-tuned vision-language model to generate text operation information according to the first identification information and the second control panel image; and forming the text operation information and the second control panel image into multimodal operation information and storing it in a database so that the multimodal operation information is obtained from the database in response to receiving the input task information.

[0012] According to an embodiment of the present invention, the vision-language model is fine-tuned by the following method: calling a region detection model to perform region detection on a sample control panel image to obtain second identification information of at least one sample operation region; calling a multi-modal large model to generate region description information according to the sample control panel image marked with the second identification information, where the region description information is used to describe the operation function of the sample operation region; calling a multi-modal large model to generate sample text operation information according to the region description information, the second identification information, and the sample control panel image; and using a plurality of sample text operation information and a plurality of sample control panel images to fine-tune the vision-language model to obtain a fine-tuned vision-language model.

[0013] Another aspect of the present invention provides a task processing device for a robot, including: an acquisition module, configured to acquire a first control panel image of an object to be operated and multi-modal operation information indicated by task information in response to receiving the task information input by a user, where the multi-modal operation information includes text operation information for operating the object to be operated; a calling module, configured to call a large language model to generate an operation sequence according to the task information and the text operation information, where the operation sequence includes at least one operation region; a generation module, configured to generate an instruction sequence according to the operation sequence, the first control panel image, and the multi-modal operation information, where the instruction sequence includes control instructions corresponding to each operation region; and a control module, configured to control the robot to operate each operation region according to the instruction sequence to process the task corresponding to the task information.

[0014] Another aspect of the present invention provides an electronic device, including: one or more processors; a memory, configured to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method as described above.

[0015] Another aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions that are used to implement the method as described above when executed.

[0016] Another aspect of the present invention provides a computer program product including computer-executable instructions that are used to implement the method as described above when executed.

[0017] Another aspect of the present invention provides a robot for implementing the method as described above.

[0018] In an embodiment of the present invention, by invoking a large language model, an operation sequence is generated according to task information and text operation information. Since the text operation information for operating the object to be operated and the user's task information are input into the large language model, the large language model can understand the operation logic of the first control panel image and flexibly generate an operation sequence corresponding to the task information, thereby ensuring that the robot can well understand the task information. Further, by generating an instruction sequence according to the operation sequence, the first control panel image, and multimodal operation information, where the instruction sequence includes control instructions corresponding to each operation area; controlling the robot to operate each operation area according to the instruction sequence to ensure that the robot can accurately control the operation area to accurately process the task. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0020] Figure 1 An exemplary system architecture to which the task processing method and apparatus for a robot according to the present invention can be applied is shown.

[0021] Figure 2 A flowchart of a task processing method for a robot according to an embodiment of the present invention is shown.

[0022] Figure 3 A schematic diagram of a scenario for generating an operation sequence according to an embodiment of the present invention is shown.

[0023] Figure 4 A flowchart of generating a control instruction according to an embodiment of the present invention is shown.

[0024] Figure 5 A schematic diagram of a scenario for generating a control instruction according to an embodiment of the present invention is shown.

[0025] Figure 6 A schematic diagram of a scenario for generating an instruction sequence according to an embodiment of the present invention is shown.

[0026] Figure 7 A block diagram of a task processing apparatus for a robot according to an embodiment of the present invention is shown.

[0027] Figure 8 A block diagram of an electronic device suitable for implementing task processing for a robot according to an embodiment of the present invention is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.

[0029] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. as used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0030] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0031] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0032] In the embodiments of the present invention, in terms of the collection, update, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the data involved (for example, including but not limited to user personal information), they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and to safeguard the security of user personal information, network security, and national security.

[0033] In the embodiments of the present invention, before obtaining or collecting user personal information, the authorization or consent of the user is obtained.

[0034] In daily life, corresponding control panels are provided for most of the devices that need to be operated. Thus, an embodied intelligent robot can operate the devices by operating the control panels of the devices. For example, when making coffee in the morning, one needs to select the preferred coffee type through the control panel of the coffee machine; when taking the elevator out, select different elevator floors according to the needs; when working, use the mouse and keyboard according to different requirements; when going home, use the air conditioner remote control to set the air conditioner to the appropriate mode, and when doing laundry at night, operate the control panel of the washing machine, etc.

[0035] However, in the prior art, the task learning of an embodied intelligent robot can be achieved based on imitation learning or by using a large language model to generate the action sequence of the robot. However, the above learning regards a certain task of the robot as an end-to-end skill for learning. The robot can complete the learned task, but it cannot flexibly complete other tasks related to the task according to the user's task information. That is to say, the existing robots cannot understand the operation logic of the tasks, and thus cannot flexibly execute multiple tasks on various operation panels according to various task information.

[0036] Therefore, an embodiment of the present invention provides a task processing method for a robot. By inputting the text operation information for operating the object to be operated and the user's task information into a large language model, the large language model can understand the operation logic of the first control panel image and flexibly generate an operation sequence corresponding to the task information, thereby ensuring that the robot can well understand the task information. Then, an instruction sequence is generated through the operation sequence, the first control panel image, and multi-modal operation information, and the robot is controlled to operate each operation area according to the instruction sequence, ensuring that the robot can accurately control the operation area and achieve accurate processing of the task.

[0037] Figure 1 An exemplary system architecture to which the task processing method and device for a robot according to the present invention can be applied is shown. It should be noted that Figure 1 The shown is only an example of the system architecture to which the embodiments of the present invention can be applied, to help those skilled in the art understand the technical content of the present invention, but it does not mean that the embodiments of the present invention cannot be used in other devices, systems, environments or scenarios.

[0038] As Figure 1 shown, the system architecture 100 according to this embodiment may include a user 101, a robot 102, and an object to be operated 103.

[0039] The object to be operated 103 can be a device with a control panel, such as a coffee machine, an elevator, a washing machine, a printer, a shredder, etc. Physical controls such as buttons, keys, knobs, levers, shrapnel, sliders, etc. can exist on the control panel; the control panel can also include a display screen, and various virtual controls are included on the display screen. The above physical controls and virtual controls can both be the operation areas of the robot. For example, the robot can perform operations such as pressing, turning the knob, pushing, pulling, touching, sliding, etc. on the operation area to process the task corresponding to the task information.

[0040] The user 101 can interact with the robot 102 in various forms to obtain task information. For example, it can interact with the robot 102 through touch screen, voice, keyboard, gesture, etc.

[0041] A variety of sensors can be set in the robot 102 to convert the input of the user 101 into task information in text form. For example, the robot obtains a visual image through a camera and extracts the content of the visual image to obtain task information; uses a microphone to receive a sound signal and converts the sound signal into task information, etc.

[0042] It should be noted that the task processing method for the robot provided by the embodiments of the present invention can generally be executed by the robot 102. Correspondingly, the task processing device for the robot provided by the embodiments of the present invention can generally be set in the robot 102. In addition, the task processing method for the robot provided by the embodiments of the present invention can also be implemented by a server, and the task processing device for the robot can also be set in the server.

[0043] For example, after the user 101 interacts with the robot 102 to obtain task information, the robot 102 transmits the task information to the server. The server, in response to receiving the input task information, obtains the first control panel image of the object to be operated indicated by the task information and multi-modal operation information, where the multi-modal operation information includes text operation information for operating the object to be operated; calls a large language model, and generates an operation sequence according to the task information and the text operation information, where the operation sequence includes at least one operation area; generates an instruction sequence according to the operation sequence, the first control panel image, and the multi-modal operation information, where the instruction sequence includes control instructions corresponding to each operation area. The server sends the instruction sequence to the robot to control the robot to operate each operation area according to the instruction sequence to process the task corresponding to the task information.

[0044] The server can be a server that provides various services. For example, the server can analyze and process the received data, etc., and feedback the processing results (such as web pages, information, or data, etc. obtained or generated according to user requests) to the robot.

[0045] It should be understood,Figure 1 The number of robots and the objects to be manipulated in [the description] is merely illustrative. According to the implementation requirements, there can be any number of robots and objects to be manipulated.

[0046] Figure 2 A flowchart of a task processing method for a robot according to an embodiment of the present invention is shown. As Figure 2 shown, the method 200 includes operations S210 to S240.

[0047] In operation S210, in response to receiving task information input by the user, obtain a first control panel image of the object to be manipulated indicated by the task information and multimodal operation information, where the multimodal operation information includes text operation information for operating the object to be manipulated.

[0048] The task information is used to represent the task requirements of the user. For example, the task information may include the object to be manipulated, and thus, the object to be manipulated can be directly determined from the task information. For example, the task information may be "Help me make a cup of coffee in the coffee machine", "Help me put the clothes into the washing machine to wash", etc., and the object to be manipulated may be "coffee machine", "washing machine", etc.

[0049] Alternatively, the task information may not include the object to be manipulated, and thus, the object to be manipulated can be determined by performing content understanding on the task information. For example, the task information may be "Please help me make a cup of coffee", "Help me wash the clothes", etc., and by performing content understanding on the task information, "coffee machine", "washing machine" are determined as the objects to be manipulated. In one embodiment, content understanding tasks can be performed by calling large language models (LLMs), and the LLMs for determining the object to be manipulated from the task information may be the same or different from the LLMs for generating the operation sequence.

[0050] The multimodal operation information can be pre-determined and stored in a database, and can be directly obtained from the database when needed. The multimodal operation information at least includes text operation information in text mode. The text operation information is used to indicate the operation logic for operating the control panel of the object to be manipulated. For example, when the object to be manipulated is a rice cooker, the text operation information may be "Long press the power-on button to adjust the cooking mode".

[0051] The first control panel image can be taken by the robot. The first control panel image refers to an image including the control panel of the object to be manipulated. To ensure the clarity of the text part and the operation area in the control panel, the first control panel image can be an image taken at a predetermined perspective, such as the perspective directly facing the control panel.

[0052] In operation S220, a large language model is called to generate an operation sequence according to the task information and text operation information, where the operation sequence includes at least one operation area.

[0053] In an embodiment of the present invention, large language models (LLMs) can understand the task information and text operation information, and generate an operation sequence for processing the task corresponding to the task information according to the understood text operation information. For example, after understanding the task information and text operation information, the large language model can split the task information and learn to complete the task corresponding to the task information through multiple operations on the control panel. Each operation on the control panel corresponds to an operation area, and the operation area can be a physical control or a virtual control on the control panel, or the position area of the control.

[0054] For example, taking the task information "Make me a cup of hot American coffee" as an example, the operation sequence can include three operation areas A, B, and C, corresponding to the position areas of the controls for selecting coffee type, temperature, and making coffee respectively.

[0055] In operation S230, an instruction sequence is generated according to the operation sequence, the first control panel image, and the multimodal operation information, where the instruction sequence includes control instructions corresponding to each operation area.

[0056] In the implementation of the present invention, the entire task processing process includes: understanding the task and processing the task. The operation sequence generated by the LLMs completes the choreography of the task based on understanding the task. However, for a specific robot, it is also necessary to control the robot to operate on at least one operation area to achieve task processing.

[0057] Specifically, the first control panel image presents the control panel of the object to be operated from a visual perspective. After determining the operation sequence, for each operation area in the operation sequence, it is necessary to locate each operation area in the first control panel image to control the robot to operate on the operation area at the accurate position. Considering that the multimodal operation information is used to indicate how to operate the object to be operated, therefore, when locating the operation area, the multimodal operation information and the first control panel image can be used to locate the operation area, and control instructions for operating each operation area are generated according to the located position information, and multiple control instructions form an instruction sequence.

[0058] In one embodiment, there is an operation order between at least one operation area in the operation sequence, and at least one control instruction in the generated instruction sequence is arranged according to the operation order. For example, still taking the task information "Make me a cup of hot American coffee" as an example, the operation sequence may include three operation areas A, B, and C, and the generated instruction sequence includes control instructions A1, B1, and C1, which are respectively used to control the robot to press the controls for selecting coffee type, temperature, and making coffee.

[0059] In operation S240, control the robot to operate each operation area according to the instruction sequence to process the task corresponding to the task information.

[0060] For example, the robot can be controlled to operate each operation area according to each control instruction in the instruction sequence, and after completing all the control instructions in the instruction sequence, the processing of the task is realized. In one embodiment, the robot is usually a robot with embodied intelligence, and the robot is built-in with a control chip or a processor for processing the above tasks.

[0061] In the embodiment of the present invention, by calling the large language model, an operation sequence is generated according to the task information and the text operation information. Since the text operation information for operating the object to be operated and the user's task information are input into the large language model, the large language model can understand the operation logic of the first control panel image and flexibly generate an operation sequence corresponding to the task information, thus ensuring that the robot can well understand the task information. Further, according to the operation sequence, the first control panel image, and the multimodal operation information, an instruction sequence is generated, where the instruction sequence includes control instructions corresponding to each operation area; controlling the robot to operate each operation area according to the instruction sequence ensures that the robot can accurately control the operation area to achieve accurate task processing.

[0062] According to the embodiment of the present invention, calling the large language model to generate an operation sequence according to the task information and the text operation information includes: obtaining prompt information, where the prompt information includes an example operation sequence, at least one first blank field corresponding to the object to be operated, and a second blank field corresponding to the task information; filling the first blank field and the second blank field with the text operation information and the task information respectively to obtain updated prompt information; and inputting the updated prompt information into the large language model to output an operation sequence in the same format as the example operation sequence.

[0063] The prompt information is pre-constructed and stored, and when calling the large language model, the prompt information can be obtained from a predetermined address. It can be understood that the example operation sequence, the first blank field, and the second blank field are also pre-determined and are all used to guide the large language model to generate an operation sequence in a predetermined format.

[0064] It should be noted that the first and second in the first blank field and the second blank field are only used to distinguish the corresponding relationship with the object to be operated and the task information.

[0065] In one embodiment, the operation sequence not only includes the operation area, but also includes the operation duration for each operation area and / or the waiting duration after operating each operation area. Thus, each control instruction in the instruction sequence can also include a time parameter, which is used to control the robot to operate the above operation area according to the operation duration and / or wait for the above waiting duration after operating the above operation area.

[0066] In this embodiment, for the control instruction including the waiting duration, when controlling the robot to execute the operation based on this control instruction, it can enable the robot to process the task with an operation more in line with the control panel logic, so as to ensure accurate task processing and improve the user experience.

[0067] For example, the example operation sequence can be: <1>0, <4>0, <4>0, <3>20, <1>0. Each line represents an operation, where the serial number in <> corresponds to different operation areas in the device operation function, and the number after <> represents the waiting duration after the operation, with the unit of seconds. The above operation description indicates that after the fourth operation, it waits for 20 seconds before performing the fifth operation.

[0068] The first blank field and the second blank field represent fields with empty field values. There can be multiple first blank fields. For example, the first blank field can include at least one of the following: device name (template['device_name']), device function (template['device_function']), device operation description (template['description']), device operation function module (modules), all possible pages (pages), the current operation page (this_page), the description of the current operation page (page_description), etc.

[0069] The field value of the first blank field can be obtained from the text operation information in the multimodal operation information. For example, for multiple first blank fields, for each first blank field, the field value matching each first blank field can be determined from the text operation information, and this field value can be filled into the corresponding first blank field.

[0070] The second blank field can be one, such as task, and the task information can be directly filled as the field value of the second blank field.

[0071] In one embodiment, the field values of the first missing field and the second missing field in the prompt message can be marked by special front identifiers and back identifiers, so as to fill the field values between the front identifier and the back identifier by using text operation information and task information.

[0072] After filling the prompt message by using text operation information and task information, an updated prompt message can be obtained. After inputting the prompt message into the large language model, an operation sequence can be output.

[0073] In the embodiment of the present invention, by obtaining the prompt message, and respectively filling the first missing field and the second missing field by using text operation information and task information, an updated prompt message is obtained; the updated prompt message is input into the large language model, and an operation sequence consistent with the format of the example operation sequence is output, so that the large language model can generate an operation sequence that conforms to the operation logic of the object to be operated and is consistent with the format of the example operation sequence with the assistance of text operation information, flexibly adapting to different objects to be operated and task information, thereby improving the flexibility of the task and the accuracy of task execution.

[0074] For example, the prompt message can be: "There is the following device instruction manual: Device name: {template['device_name']}, Device function: {template['device_function']}, Device operation description: {template['description']}, Device operation function module: {modules}, All possible pages: {pages}, Current operating page: {this_page}, Description of the current operating page: {page_description}, Please refer to the above instruction manual and generate an operation sequence for the device operation according to the user's operation instruction. The sequence is output in the following format, and do not generate any other extra text: Example operation sequence (as above). The user's task information is: {task}. Please generate an operation sequence. The missing parts in the curly brackets are obtained from the corresponding text operation information. In this prompt message, the front identifier and the back identifier are { and} respectively."

[0075] Figure 3 Shows a schematic diagram of the scenario for generating an operation sequence according to an embodiment of the present invention. As Figure 3As shown, in Embodiment 300, the obtained prompt information 301 includes a first blank field 311, a second blank field 312, and an example operation sequence 313. After filling the first blank field 311 with text operation information T1 and the second blank field 312 with task information T2, the example operation sequence 313, the filled first blank field, and the filled second blank field form the updated prompt information 304. After inputting the updated prompt information 304 into the large language model M1, an operation sequence T3 in the same format as the example operation sequence 313 is obtained.

[0076] For example, the operation sequence can be <2>0, <3>0, <1>0, <1>3, <2>0, which means operating on the second button (such as the power-on button) of the rice cooker to turn on the rice cooker; operating on the third button (such as the mode button) to select the mode; operating on the first button (such as the time button) twice (each operation increases by half an hour) to set the timer for one hour, and after waiting for 3s, the timing is completed; then operating on the second button to confirm the start of cooking.

[0077] According to an embodiment of the present invention, using text operation information and task information to fill the first blank field and the second blank field respectively to obtain updated prompt information, including: determining a target field value matching each first blank field from the text operation information according to the field name of the first blank field in the prompt information; filling the target field value into the first blank field and filling the task information into the second blank field to obtain updated prompt information.

[0078] In an embodiment of the present invention, both the prompt information and the text operation information may include field names. Thus, the target field value matching the first blank field can be determined from the text operation information through the field name. The field name of the first blank field may be the above-mentioned device_name, etc.

[0079] For example, by calculating the similarity between field names, a field name similar to the field name of the first blank field can be determined from the text operation information, and the field value of this field name can be determined as the target field value matching the first blank field.

[0080] In one embodiment, the field name - field value in the text operation information has a specific format, such as the specific format is: "field name": "field value". The field name can be determined by identifying the specific format, and the similarity between this field name and the field name of the first blank field in the prompt information can be calculated.

[0081] For the second blank field, the task information can be directly filled into the default position of the field value of the blank field without matching operations. After filling the first blank field and the second blank field, updated prompt information can be obtained.

[0082] In an embodiment of the present invention, by determining, according to the field name of the first blank field in the prompt information, a target field value matching each first blank field from the text operation information, and using the target field value to fill the first blank field and using the task information to fill the second blank field, it is possible to ensure the correspondence between the first blank field and the target field value in the updated prompt information, so as to avoid generating an incorrect operation sequence according to an incorrect matching relationship subsequently.

[0083] In one embodiment, the text operation information may be: {"image_dir":"0_Color.png","points":[{"label":"1","points":[75.92028985507247,63.956521739130444]},{"label":"2","points":[160.84782608695653,64.10144927536233]},{"label":"3","points":[245.7753623188406,63.521739130434796]},{"label":"4","points":[75.6304347826087,123.37681159420289]},{"label":"5","points":[160.84782608695653,123.37681159420289]},{"label":"6","points":[246.21014492753625,123.81159420289855]}],"modules":{"1":"<Button>Make Espresso","2":"<Button>Make Hot Latte","3":"<Button>Make Hot Mocha","4":"<Button>Make Hot Americano","5":"<Button>Make Hot Cappuccino","6":"<Button>Make Hot Chocolate Milk"},"description":"When using the device, different types of coffee can be made by clicking different buttons on the device control panel. The use of this coffee machine includes three pages: the start page, which is used to select the type of coffee to be made; the detailed settings page, which is used to set the sweetness; and the confirmation page, which is used to confirm the production. The current page is: the start page","device_name":"Coffee Machine","device_function":"Used to make coffee","next_template":{"1":"0_1_template.json","2":"0_2_template.json","3":"0_3_template.json","4":"0_4_template.json","5":"0_5_template.json","6":"0_6_template.json"},"pages":["Start Page","Detailed Settings Page","Confirmation Page"],"this_page":"Start Page","this_page_description":"Used to select the type of coffee you want to make"}。

[0084] Among them, the meanings of the field names are as follows: "image_dir": the address of the second control panel image in the multimodal operation information; "points": a list of all possible operation areas, where each element is a dictionary containing two pieces of information, "label" is the label of the operation area, and "points" are the pixel coordinates of the operation area on the second control panel image, represented by ."modules": the text description of each operation area, which starts with the type of the specific operation area and is marked with '<>'. The types of operation areas include: (1) buttons / touch screen buttons; (2) knobs; (3) levers / metal shims / sliders; (4) others, four categories. "description": the detailed functions and operation logics of the control panel and the operation target. "device_name": the name of the operation target corresponding to the control panel. "device_function": the function of the operation target. "next_template": if there is a page change for the control panel, the next template corresponding to a certain operation needs to be specified in sequence. For example, after a rice cooker selects the cooking function, it will switch to the next control panel logic to select the hardness of the rice. At this time, the physical control panel may not change, but the entire operation logic changes, so a new template needs to be specified. "pages": used to define all possible logic panels of the current control panel. "this_page": used to specify the logic panel to which the current control panel belongs. "this_page_description": used to describe what function the current logic panel is used to implement.

[0085] In a specific embodiment, the text operation information may include a template for each page. Combining the meanings of the field names and the text operation information, it can be determined that the above text operation information belongs to the template of the "start page", and the start page provides six buttons for selecting different coffee types.

[0086] Figure 4 The flowchart of generating a control instruction according to an embodiment of the present invention is shown.

[0087] As Figure 4 shown, the embodiment 400 of generating a control instruction includes operation S431 to operation S433.

[0088] In operation S431, according to the first control panel image and the second control panel image, a position transformation relationship is determined, where the position transformation relationship represents the position transformation relationship between the coordinate system of the second control panel image and the coordinate system of the first control panel image.

[0089] According to an embodiment of the present invention, the multi-modal operation information further includes a pre-acquired second control panel image, and the first control panel image is acquired by the robot in response to receiving the input task information. In this embodiment, the second control panel image may be acquired by a device other than the robot, or may be acquired by the robot before acquiring the first control panel image. For example, the second control panel image is acquired by the robot when generating the multi-modal operation information.

[0090] For example, when the second control panel image is acquired by a device other than the robot, the coordinate systems of the first control panel image and the second control panel image are different, and the coordinates of the same operation area in the first control panel image and the second control panel image are different. In the case where both the first / second control panels are acquired by the robot, the coordinate systems of the above two images may also be different due to different shooting positions.

[0091] Since both the first control panel image and the second control panel image include the same control panel, therefore, the position transformation relationship between the two coordinate systems can be determined through the same control panel.

[0092] In one embodiment, as in the above points field, its field value includes xy two-dimensional positions. Therefore, the position change relationship in the embodiment of the present invention can be the position change relationship in the xy two-dimensional coordinates.

[0093] In operation S432, according to the first control panel image, determine the depth information of each operation area.

[0094] According to an embodiment of the present invention, a depth prediction model can be adopted to predict the depth of each pixel point in the first control panel image, and obtain the depth information of each pixel point in the first control panel image. The depth information of each operation area includes: the depth information of each pixel point within the operation area.

[0095] For example, the depth prediction model can be a network such as Convolutional Neural Networks (CNN), deep learning, U-net, etc.

[0096] In operation S433, generate a control instruction according to the depth information, the position transformation relationship, and the text operation information.

[0097] Based on the above statements, the text operation information can include the two-dimensional position information of each operation area, the depth information can be used as the third dimension, and the three-dimensional position information of the operation area can be determined by combining the two-dimensional position information. Thus, the coordinates of the operation area in the coordinate system of the robot can be obtained, and further, a control instruction for controlling the robot to operate on the operation area can be generated.

[0098] In an embodiment of the present invention, by determining the position transformation relationship based on the first control panel image and the second control panel image, the two-dimensional coordinate offset of the control panel caused by the coordinate system difference can be determined; by the first control panel image, the depth information of each operation area can be determined, and the third-dimensional coordinate based on the current captured image can be determined. Thus, the control instruction generated according to the depth information, the position transformation relationship, and the text operation information is the accurate position information in the robot coordinate system to ensure the accuracy of subsequent task processing.

[0099] According to an embodiment of the present invention, determining the position transformation relationship according to the first control panel image and the second control panel image includes: performing feature registration on the first control panel image and the second control panel image to obtain at least one pair of feature points; and determining the position transformation relationship according to the two-dimensional position information of the two feature points in the pair of feature points.

[0100] For example, through a feature detection algorithm, the first control panel image can be subjected to feature detection to obtain at least one feature point of the first control panel image and the descriptor of each feature point. Similarly, through a feature detection algorithm, the second control panel image can be subjected to feature detection to obtain at least one feature point of the second control panel image and the descriptor of each feature point. Then, through a feature matching algorithm, according to the descriptors of the feature points, the feature points matching each feature point in the first control panel image are determined from at least one feature point in the second control panel image to form a pair of feature points. In one embodiment, there are usually multiple pairs of feature points between the first control panel image and the second control panel image.

[0101] The feature detection algorithm can be a Scale-Invariant Feature Transform (SIFT) algorithm, a Speeded Up Robust Features (SURF) algorithm, an Oriented FAST and Rotated BRIEF (ORB) algorithm. The feature matching algorithm can be a Fast Approximate Nearest Neighbor Search Library (FLANN) algorithm, a Brute-Force Matcher (BFMatcher), etc.

[0102] In an embodiment of the present invention, each feature point in the feature point pair is a pixel point in the first control panel image / the second control panel image. Therefore, the two-dimensional position information of the feature points in the second control panel image can be obtained according to the points field in the text operation information. When the robot captures the first control panel image, a pose detector can be set inside the robot. When capturing the first control panel image, the two-dimensional position information of each feature point can be determined by the pose detector.

[0103] Using the two feature points in each feature point pair, the position transformation relationship between the coordinate system of the second control panel image and the coordinate system of the first control panel image can be calculated. In one embodiment, this position transformation relationship can be represented by a Homography Matrix.

[0104] For example, the random sample consensus (RANSAC) algorithm is used to calculate the Homography Matrix from the second control panel image to the first control panel image using multiple pairs of feature points. As shown, the following formula (1) can be referred to:

[0105] (1)

[0106] Among them, for the two feature points in the feature point pair, x and y are the horizontal and vertical coordinates of the feature point on the first control panel image (the two-dimensional position information in the first control panel image), z is 1, x t 、y t are respectively the horizontal and vertical coordinates of the feature point on the second control panel image (the two-dimensional position information in the second control panel image), z t is 1, is the Homography Matrix, and h 11 ~ h 33 are the 9 parameter values of the solved Homography Matrix.

[0107] In an embodiment of the present invention, at least one feature point pair is obtained by feature registration of the first control panel image and the second control panel image, and the position transformation relationship is determined according to the two-dimensional position information of the two feature points in the feature point pair, so as to accurately determine the position transformation relationship between the second control panel image and the currently captured first control panel image, which helps to accurately locate the operation area and control the robot to accurately execute tasks subsequently.

[0108] According to an embodiment of the present invention, the text operation information includes two-dimensional position information of each operation area in the second control panel image; according to the depth information, the position transformation relationship, and the text operation information, a control instruction is generated, including: for each operation area, according to the two-dimensional position information of the operation area in the second control panel image and the position transformation relationship, determining the two-dimensional position information of the operation area in the robot coordinate system; according to the depth information and the two-dimensional position information of the operation area in the robot coordinate system, determining the three-dimensional position information of the operation area; and generating a control instruction according to the three-dimensional position information.

[0109] It can be understood that since the first control panel image is captured by the robot, the coordinate system of the first control panel image is also referred to as the robot coordinate system.

[0110] For example, after determining the two-dimensional position transformation relationship, the field value of the points field in the text operation information can be used as the two-dimensional position information of the operation area. Thus, for each operation area, according to the homography matrix, the known two-dimensional position information in the second control panel image can be converted into the two-dimensional position information in the robot coordinate system, so as to determine the accurate two-dimensional position without performing complex positioning processing on the second control panel image. On this basis, taking the depth information as the coordinate of the third dimension and combining the two-dimensional position information of the robot coordinate system determined above, the three-dimensional position information of the operation area in the real three-dimensional space, such as three-dimensional coordinates, can be obtained. After determining the three-dimensional position information, a control instruction can be generated based on the control logic of the robot, so that the robot can perform corresponding operations on the operation area at the three-dimensional position information according to the control instruction.

[0111] In the embodiment of the present invention, by using the depth information and the two-dimensional position information of the operation area in the robot coordinate system to determine the three-dimensional position information of the operation area and generating a control instruction according to the three-dimensional position information, the operation area can be simply and accurately located and the robot can be controlled to accurately process tasks according to the control instruction.

[0112] Figure 5 A schematic diagram of a scenario for generating a control instruction according to an embodiment of the present invention is shown.

[0113] Such as Figure 5As shown, in this scenario 500, the multimodal operation information 502 includes the second control panel image P2 and the text operation information T1. Depth prediction is performed on the first control panel image P1 to obtain depth information 503. The position transformation relationship 504 can be determined based on the first control panel image P1 and the second control panel image P2. According to the position transformation relationship 504 and the two-dimensional position information 505 of each operation area in the second control panel image in the text operation information T1, the two-dimensional position information 506 of each operation area in the first control panel image is determined. According to the depth information 503 and the two-dimensional position information 506 of each operation area in the first control panel image, the three-dimensional position information 507 of each operation area can be determined, and then control instructions 508 corresponding to each operation area are generated based on the three-dimensional position information 507.

[0114] According to an embodiment of the present invention, the method further includes: calling a region detection model to perform region detection on the second control panel image to obtain the first identification information of at least one operation area; calling a fine-tuned vision-language model to generate text operation information based on the first identification information and the second control panel image; and forming multimodal operation information with the text operation information and the second control panel image and storing it in a database, so that in response to receiving the input task information, the multimodal operation information is obtained from the database.

[0115] In an embodiment of the present invention, for each object to be operated, only one multimodal operation information needs to be generated based on the second control panel image. When a task needs to be processed, in response to receiving the task information input by the user, the multimodal operation information of the object to be operated can be directly obtained from the database according to the identification information of the object to be operated.

[0116] The region detection model can be pre-trained. The region detection model can be a general model for performing region detection on the second control panel images of multiple objects to be operated. For example, the region detection model can be the yolo (You Only Look Once) model, such as yolo v11.

[0117] When training the region detection model, as many images of objects to be operated including the control panel as possible can be obtained. After screening and annotation, the sample data for training is obtained. In the sample data, the operation areas of each object to be operated are marked, and these operation areas include buttons, knobs, levers, sliders, etc. on the control panel. Then, the initial region detection model is trained using the sample data until the training conditions are met to obtain the region detection model.

[0118] After training, the second control panel image can be input into the region detection model, and the first identification information including at least one operation region is output. The first identification information can be a detection frame and / or the center position of the detection frame. It can be understood that if there are multiple operation regions in the second control panel image, the region detection model can output the first identification information of each of the multiple operation regions simultaneously.

[0119] The Vision-Language Model (VLM) adopted in the present invention is a fine-tuned VLM. The input of the vision-language model can be information in text modality and image modality, such as the first identification information in text modality and the second control panel image in image modality.

[0120] The fine-tuned vision-language model is used to perform content understanding tasks and text generation tasks. The content understanding tasks can be: understanding the operation logic of the control panel in the second control panel image, the functions of each operation region, the operation logic, etc. The text generation tasks can be: generating text descriptions of the operation regions, two-dimensional position information, the types of operation regions, etc.

[0121] For example, for an operated object of the coffee machine type, after inputting the second control panel image and the first identification information for positioning the operation region into the fine-tuned vision-language model, the text operation information as in the above embodiment can be generated.

[0122] In one embodiment, one operated object corresponds to one multi-modal operation information. After generating the text operation information of each operated object, the text operation information and the second control panel image are combined into a data pair and stored in the database as multi-modal operation information. Thus, in response to receiving the input task information, the multi-modal operation information can be obtained from the database according to the identification information of the operated object indicated by the task information.

[0123] In the embodiment of the present invention, by calling the region detection model, region detection is performed on the second control panel image to obtain the first identification information of at least one operation region; calling the fine-tuned vision-language model, according to the first identification information and the second control panel image, generating text operation information; combining the text operation information and the second control panel image into multi-modal operation information and storing it in the database, so that in the process of processing tasks, there is no need to perform region detection on the second control panel image, and the multi-modal operation information can be directly obtained, ensuring the simplicity and convenience of task processing operations.

[0124] According to an embodiment of the present invention, the vision - language model is fine - tuned through the following method: Call a region detection model to perform region detection on a sample control panel image to obtain second identification information of at least one sample operation region; Call a multi - modal large model to generate region description information according to the sample control panel image marked with the second identification information, where the region description information is used to describe the operation function of the sample operation region; Call a multi - modal large model to generate sample text operation information according to the region description information, the second identification information, and the sample control panel image; Use multiple sample text operation information and multiple sample control panel images to fine - tune the vision - language model to obtain a fine - tuned vision - language model.

[0125] The sample control panel image can be an image of the control panels of multiple devices collected through the invention channel. Through the region detection model trained above, region detection can be performed on multiple sample control panel images to obtain second identification information of at least one sample operation region in each sample control panel image. The method of determining the second identification information is similar to the method of obtaining the first identification information and will not be elaborated here.

[0126] Input the sample control panel image marked with the second identification information into the multi - modal large model, and use prompt words to instruct the multi - modal large model to generate a function description for each operation region. Thus, the multi - modal large model can output region description information for each operation region.

[0127] For example, the region description information can be the information in the description field. Taking the coffee machine above as an example, the region description information can be: When using this device, by clicking different buttons on the control panel, different types of coffee can be made. The use of this coffee machine includes three pages: the start page for selecting the type of coffee to be made, the detailed settings page for setting the sweetness, and the confirmation page for confirming the production. The current page is: the start page.

[0128] After obtaining the region description information, if a certain region description information does not match the function description of the operation region on the control panel, it can be readjusted through multiple interactions with the multi - modal large model until the correct region description information is generated; or, for a certain sample control panel image, the region description information that does not match the function description can be screened out. Then, input the region description information, the second identification information, and the sample control panel image into the multi - modal large model again, and use prompt words to instruct the multi - modal large model to output sample text operation information in the same format as the above - mentioned text operation information.

[0129] It can be understood that the sample text operation information generated by the multimodal large model is also the sample data used to fine-tune the vision-language model. In this embodiment, a small number of sample control panel images can be used to enable the multimodal large model to output a sufficient amount of sample data.

[0130] For multiple sample control panel images, multiple sample text operation information can be obtained. Specifically, multiple sample text operation information can be used as examples to fine-tune the vision-language model according to the multiple sample text operation information and the multiple sample control panel images, so that the fine-tuned vision-language model can output the correct text operation information.

[0131] In the embodiment of the present invention, by calling the region detection model for region detection and calling the multimodal large model to generate region description information; and generating sample text operation information according to the region description information, the second identification information, and the sample control panel image, it is possible to generate a relatively large number of sample text operation information with relatively accurate sample text operation information, improve the accuracy of the subsequent fine-tuned vision-language model, thereby improving the accuracy of generating text operation information and improving the accuracy of the robot's task processing.

[0132] Figure 6 Shows a schematic diagram of a scenario for generating an instruction sequence. As Figure 6 shown, in Embodiment 600 of generating an instruction sequence, for a new object to be operated, the second control panel image P2 can be collected, and the region detection model can be used to perform region detection on the second control panel image P2 to obtain the second identification information 602. Inputting the second control panel image P2 and the second identification information 602 into the fine-tuned vision-language model M2, the text operation information T1 can be generated, and the text operation information T1 and the second control panel image P2 can be stored in the database together. Thus, in the application stage, the robot can respond to the received task information T2 from the user, and obtain the multimodal operation information including the text operation information T1 and the second control panel image P2 from the database. Call the large language model M1 to generate the operation sequence T3 according to the text operation information T1 and the task information T2. In addition, by performing feature registration of the operation S607 on the first control panel image P1 and the second control panel image P2, the position transformation relationship can be obtained, and combined with the operation sequence T3, the two-dimensional position information in the text operation information T1, and the depth information of the first control panel image P1, the instruction sequence 608 can be determined.

[0133] In the embodiments of the present invention, for different control panels and different task information of users, corresponding operation arrangements and the positioning of the operation area of each operation can be flexibly output, so that the robot can be guided to complete the operation of the control panel step by step, thereby realizing the processing of task information. For various task information and various control panels, it can ensure that the robot can well understand and accurately process the tasks corresponding to the task information, and the user experience is better.

[0134] Figure 7 FIG. shows a block diagram of a task processing device for a robot according to an embodiment of the present invention. As Figure 7 shown, the task processing device 700 for a robot includes an acquisition module 710, a call module 720, a generation module 730, and a control module 740.

[0135] The acquisition module 710 is configured to, in response to receiving task information input by a user, acquire a first control panel image of an object to be operated indicated by the task information and multimodal operation information, where the multimodal operation information includes text operation information for operating the object to be operated.

[0136] The call module 720 is configured to call a large language model and generate an operation sequence according to the task information and the text operation information, where the operation sequence includes at least one operation area.

[0137] The generation module 730 is configured to generate an instruction sequence according to the operation sequence, the first control panel image, and the multimodal operation information, where the instruction sequence includes control instructions corresponding to each operation area.

[0138] The control module 740 is configured to control the robot to operate each operation area according to the instruction sequence to process the task corresponding to the task information.

[0139] According to an embodiment of the present invention, the call module 720 includes: an acquisition sub-module, an update sub-module, and an output sub-module. The acquisition sub-module is configured to acquire prompt information, where the prompt information includes an example operation sequence, at least one first blank field corresponding to the object to be operated, and a second blank field corresponding to the task information. The update sub-module is configured to fill the first blank field and the second blank field with the text operation information and the task information respectively to obtain updated prompt information. The output sub-module is configured to input the updated prompt information into the large language model and output an operation sequence in the same format as the example operation sequence.

[0140] According to an embodiment of the present invention, the update sub-module includes: a matching unit and a filling unit. The matching unit is configured to determine, according to the field name of the first blank field in the prompt information, a target field value that matches each first blank field from the text operation information. The filling unit is configured to fill the target field value into the first blank field and fill the task information into the second blank field to obtain the updated prompt information.

[0141] According to an embodiment of the present invention, the multi-modal operation information further includes a pre-collected second control panel image, and the first control panel image is collected by the robot in response to receiving the input task information. The generation module 730 includes: a position relationship determination sub-module, a depth determination sub-module, and a generation sub-module. The position relationship determination sub-module is configured to determine a position transformation relationship according to the first control panel image and the second control panel image, where the position transformation relationship represents the position transformation relationship between the coordinate system of the second control panel image and the coordinate system of the first control panel image. The depth determination sub-module is configured to determine the depth information of each operation area according to the first control panel image. The generation sub-module is configured to generate a control instruction according to the depth information, the position transformation relationship, and the text operation information.

[0142] According to an embodiment of the present invention, the position relationship determination sub-module includes: a registration unit configured to perform feature registration on the first control panel image and the second control panel image to obtain at least one pair of feature points. A determination unit configured to determine the position transformation relationship according to the two-dimensional position information of the two feature points in the pair of feature points.

[0143] According to an embodiment of the present invention, the text operation information includes the two-dimensional position information of each operation area in the second control panel image.

[0144] The generation sub-module includes: a first determination unit, a second determination unit, and an instruction generation unit. For each operation area: The first determination unit is configured to determine the two-dimensional position information of the operation area in the robot coordinate system according to the two-dimensional position information of the operation area in the second control panel image and the position transformation relationship. The second determination unit is configured to determine the three-dimensional position information of the operation area according to the depth information and the two-dimensional position information of the operation area in the robot coordinate system. The generation unit is configured to generate a control instruction according to the three-dimensional position information.

[0145] According to an embodiment of the present invention, the task processing device 700 for a robot further includes: a region detection module, an operation information generation module, and a storage module. The region detection module is configured to call a region detection model to perform region detection on the second control panel image to obtain first identification information of at least one operation region. The operation information generation module is configured to call a fine-tuned vision-language model to generate text operation information according to the first identification information and the second control panel image. The storage module is configured to form multimodal operation information by combining the text operation information and the second control panel image and store it in a database, so that in response to receiving input task information, the multimodal operation information is obtained from the database.

[0146] According to an embodiment of the present invention, the vision-language model is fine-tuned in the following manner: calling a region detection model to perform region detection on a sample control panel image to obtain second identification information of at least one sample operation region; calling a multimodal large model to generate region description information according to the sample control panel image marked with the second identification information, where the region description information is used to describe the operation function of the sample operation region; calling a multimodal large model to generate sample text operation information according to the region description information, the second identification information, and the sample control panel image; and using multiple sample text operation information and multiple sample control panel images to fine-tune the vision-language model to obtain a fine-tuned vision-language model.

[0147] According to an embodiment of the present invention, any plurality of modules, sub-modules, units, and sub-units, or at least part of the functions of any of them can be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present invention can be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present invention can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits in hardware or firmware, or implemented in any one of the three implementation ways of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to the embodiments of the present invention can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.

[0148] For example, any combination of the obtaining module 710, the invoking module 720, the generating module 730, and the control module 740 may be integrated and implemented in one module / unit / sub-unit, or any one of the modules / units / sub-units may be split into multiple modules / units / sub-units. Alternatively, at least part of the functions of one or more of these modules / units / sub-units may be combined with at least part of the functions of other modules / units / sub-units and implemented in one module / unit / sub-unit. According to an embodiment of the present invention, at least one of the obtaining module 710, the invoking module 720, the generating module 730, and the control module 740 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or any other reasonable way of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the obtaining module 710, the invoking module 720, the generating module 730, and the control module 740 may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.

[0149] It should be noted that the device part in the embodiments of the present invention corresponds to the method part in the embodiments of the present invention. For the description of the device part, please refer to the method part for details, and will not be elaborated herein.

[0150] Figure 8 The block diagram of an electronic device suitable for implementing task processing for a robot according to an embodiment of the present invention is shown.

[0151] As Figure 8 shown, the electronic device 800 according to an embodiment of the present invention includes a processor 801, which may perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 may also include on-board memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0152] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. The processor 801 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the programs can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.

[0153] According to an embodiment of the present invention, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the input / output (I / O) interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 810 as needed so that a computer program read from it can be installed into the storage portion 808 as needed.

[0154] According to an embodiment of the present invention, the method flow according to the embodiments of the present invention can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication portion 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system according to the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described system, device, apparatus, module, unit, etc. can be implemented by computer program modules.

[0155] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0156] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium. For example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0157] For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories other than the above-described ROM 802 and / or RAM 803 and / or ROM 802 and RAM 803.

[0158] An embodiment of the present invention further includes a computer program product, which includes a computer program that contains program code for executing the method provided by the embodiments of the present invention. When the computer program product runs on an electronic device, the program code is used to cause the electronic device to implement the task processing method for a robot provided by the embodiments of the present invention.

[0159] When the computer program is executed by the processor 801, the above functions defined in the system / apparatus of the embodiments of the present invention are executed. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.

[0160] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and downloaded and installed through the communication part 809, and / or installed from the removable medium 811. The program code included in the computer program may be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0161] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, programming languages such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0162] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or by a combination of dedicated hardware and computer instructions. Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0163] The above describes the embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A task processing method for a robot, characterized in that: The method comprises: In response to receiving task information input by a user, acquiring a first control panel image and multimodal operation information of an operated object indicated by the task information, wherein the multimodal operation information includes text operation information for operating the operated object; Calling a large language model to generate an operation sequence according to the task information and the text operation information, wherein the operation sequence includes at least one operation area; generating an instruction sequence according to the operation sequence, the first control panel image and the multimodal operation information, wherein the instruction sequence includes a control instruction corresponding to each of the operation areas; and The robot is controlled to operate each of the operating areas according to the instruction sequence to process the task corresponding to the task information.

2. The method according to claim 1, characterized in that The calling of the large language model generates an operation sequence according to the task information and the text operation information, including: Acquire prompt information, wherein the prompt information includes an example operation sequence, at least one first vacant field corresponding to the operated object, and a second vacant field corresponding to the task information; Filling the first vacant field and the second vacant field with the text operation information and the task information respectively to obtain updated prompt information; and The updated prompt information is input into the large language model, and the operation sequence consistent with the format of the example operation sequence is output.

3. The method according to claim 2, characterized in that The using the text operation information and the task information to fill the first vacant field and the second vacant field respectively to obtain updated prompt information includes: Determining, from the text operation information, a target field value matching each of the first missing fields according to the field name of the first missing field in the prompt information; and The target field value is filled into the first empty field, and the task information is filled into the second empty field to obtain the updated prompt information.

4. The method according to claim 1, characterized in that: The multimodal operation information further includes a pre-collected second control panel image, wherein the first control panel image is collected by the robot in response to receiving the input task information; The generating an instruction sequence according to the operation sequence, the first control panel image and the multimodal operation information includes: Determining a position transformation relationship according to the first control panel image and the second control panel image, wherein the position transformation relationship represents a position transformation relationship between a coordinate system of the second control panel image and a coordinate system of the first control panel image; determining depth information of each of the operation areas according to the first control panel image; and The control instruction is generated according to the depth information, the position transformation relationship and the text operation information.

5. The method according to claim 4, characterized in that The determining of the position transformation relationship according to the first control panel image and the second control panel image includes: Performing feature registration on the first control panel image and the second control panel image to obtain at least one feature point pair; and The position transformation relationship is determined according to the two-dimensional position information of each of the two feature points in the feature point pair.

6. The method according to claim 4, characterized in that The text operation information includes two-dimensional position information of each operation area in the second control panel image; generating the control instruction according to the depth information, the position transformation relationship and the text operation information includes: For each of the operating areas, Determine the two-dimensional position information of the operation area in the robot coordinate system according to the two-dimensional position information of the operation area in the second control panel image and the position transformation relationship; Determining three-dimensional position information of the operating area according to the depth information and the two-dimensional position information of the operating area in the robot coordinate system; and The control instruction is generated according to the three-dimensional position information.

7. The method according to any one of claims 4 to 6, characterized in that: The method further comprises: Calling a region detection model to perform region detection on the second control panel image to obtain first identification information of at least one operation region; calling the fine-tuned visual language model to generate the text operation information according to the first identification information and the second control panel image; and The text operation information and the second control panel image are combined into the multimodal operation information and stored in a database, so that in response to receiving the input task information, the multimodal operation information is acquired from the database.

8. The method according to claim 7, characterized in that The visual language model is fine-tuned by the following method: Calling the region detection model to perform region detection on the sample control panel image to obtain second identification information of at least one sample operation region; Calling the multimodal large model to generate region description information according to the sample control panel image marked with the second identification information, wherein the region description information is used to describe the operation function of the sample operation region; calling the multimodal large model to generate sample text operation information according to the area description information, the second identification information and the sample control panel image; and The visual language model is fine-tuned using the plurality of sample text operation information and the plurality of sample control panel images to obtain the fine-tuned visual language model.

9. A task processing device for a robot, characterized in that: The device comprises: an acquisition module, configured to acquire, in response to receiving task information input by a user, a first control panel image and multimodal operation information of an operated object indicated by the task information, wherein the multimodal operation information includes text operation information for operating the operated object; A calling module, used for calling a large language model, generating an operation sequence according to the task information and the text operation information, wherein the operation sequence includes at least one operation area; a generating module, configured to generate an instruction sequence according to the operation sequence, the first control panel image and the multimodal operation information, wherein the instruction sequence includes a control instruction corresponding to each of the operation areas; and A control module is used to control the robot to operate each of the operating areas according to the instruction sequence to process the task corresponding to the task information.

10. An electronic device, comprising: one or more processors; a memory for storing one or more programs, It is characterized in that when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having executable instructions stored thereon, characterized in that: When the instruction is executed by a processor, the processor implements the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that The computer program implements the method according to any one of claims 1 to 8 when executed by a processor.

13. A robot, characterized in that: Used to implement the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Intelligent robot control system, method and device based on artificial intelligence

    CN104965426A

  • Accounting information processing method and device based on voice recognition and electronic equipment

    CN110659970A

  • Training method and training device of multi-modal pre-training model and electronic equipment

    CN113283551A

  • Mechanical arm sensing method based on multi-modal data fusion

    CN117103277A

  • Man-machine interaction assembly method and system based on multi-modal large model and reinforcement learning

    CN118744426A

Cited By

  • Robot control method and device based on artificial intelligence, equipment and medium

    CN120735005A