Intelligent agent training method, certificate information acquisition method and device
By training the intelligent control robot to collect certificate information, using multi-task action blocking model and reinforcement strategy, the problem of low efficiency in the collection of certificate information in the existing technology is solved, and efficient and accurate certificate information collection is achieved.
Patent Information
- Application Number
- CN202411987040.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
AI Technical Summary
In the prior art, the collection efficiency of document information is low and it is difficult to accurately traverse different environmental scenarios, resulting in low efficiency and high cost of manual collection methods.
By training the agent, the robot controls the preset certificate information collection task, and uses the multi-task action blocking model combined with the reinforcement strategy to update the agent parameters to realize the automatic execution of the certificate image acquisition action sequence.
It realizes efficient automation of document information collection, can accurately traverse different environmental scenarios, improves collection efficiency and accuracy, and reduces the cost of manual collection.
Smart Images

Figure CN120046647A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an intelligent agent training method, a certificate information collection method and a device. Background Art
[0002] In many business scenarios, it is necessary to identify document information. For this purpose, it is necessary to collect a large amount of document information in different environmental scenarios to train the document information recognition model. At present, the industry usually adopts manual collection, but this method is very inefficient and difficult to accurately traverse the required environmental scenarios. Summary of the invention
[0003] One or more embodiments of the present specification provide an intelligent agent training method, a document information collection method and device, which controls a robot to perform a preset document information collection task by training an intelligent agent, and can accurately collect the required document images with extremely high efficiency.
[0004] In a first aspect, a method for training an intelligent agent is provided, the method comprising:
[0005] Acquire training samples, wherein the training samples include view images, action sequences, joint position data, and task prompt information at each time step of the robot in the process of performing a preset document information collection task;
[0006] Inputting the training sample into the intelligent agent to obtain a document image acquisition action sequence predicted by the intelligent agent;
[0007] Controlling the robot to execute the certificate image acquisition action sequence and extracting certificate information from the acquired certificate image;
[0008] An incentive signal is determined based on the difference between the certificate information and the target information of the certificate information collection task, and a reinforcement strategy is adopted to update the parameters of the agent until a target agent is obtained.
[0009] As an optional implementation of the method of the first aspect, the agent is implemented using a multi-task action block model; the multi-task action block model includes:
[0010] A language encoder for learning feature embedding of the task description information;
[0011] A conditional variational autoencoder for learning a latent encoding of the action sequence;
[0012] A transformer encoder, configured to perform feature fusion on the view image, the feature embedding of the task description information, the latent coding of the action sequence and the joint position data to obtain a fused feature;
[0013] A decoder is used to predict the next set of document image acquisition action sequences of the current time step according to the fusion features.
[0014] As an optional implementation of the method of the first aspect, the robot includes a mechanical arm, an image acquisition device and a perception system;
[0015] The mechanical arm is used to grab the image acquisition device and perform image acquisition operations;
[0016] The perception system is used to extract the certificate information required for the certificate information collection task from the certificate image collected by the image collection device.
[0017] As an optional implementation manner of the method of the first aspect, controlling the robot to execute a set of certificate image acquisition actions corresponding to the action block and extracting certificate information from the acquired certificate image specifically includes:
[0018] The robot's perception system is used to extract the document information required for the document information collection task from the document image.
[0019] As an optional implementation manner of the method described in the first aspect, the task prompt information is obtained by decomposing the certificate information collection task according to the time step.
[0020] In a second aspect, a method for collecting certificate information is provided, the method comprising:
[0021] According to the target certificate information collection task, construct task description information;
[0022] The task description information is input into an intelligent agent, and the intelligent agent controls a robot to perform an operation of collecting certificate information on a target certificate to obtain the certificate information required for the target certificate information collection task; the intelligent agent is pre-trained using the above-mentioned intelligent agent training method.
[0023] In a third aspect, an intelligent agent training device is provided, the device comprising:
[0024] A first data acquisition module is configured to acquire training samples, wherein the training samples include a view image, an action sequence, joint position data, and task prompt information at each time step of the robot in a process of performing a preset document information collection task;
[0025] The training module is configured to input the training samples into the intelligent agent to obtain the ID image acquisition action sequence predicted by the intelligent agent; control the robot to execute the ID image acquisition action sequence and extract ID information from the acquired ID image; determine the incentive signal based on the difference between the ID information and the target information of the ID information acquisition task, and adopt the reinforcement strategy to update the parameters of the intelligent agent until the target intelligent agent is obtained.
[0026] As an optional implementation manner of the device of the third aspect, the agent is implemented by a multi-task action block model; the multi-task action block model includes:
[0027] A language encoder for learning feature embedding of the task description information;
[0028] A conditional variational autoencoder for learning a latent encoding of the action sequence;
[0029] A transformer encoder, configured to perform feature fusion on the view image, the feature embedding of the task description information, the latent coding of the action sequence and the joint position data to obtain a fused feature;
[0030] A decoder is used to predict the next set of document image acquisition action sequences of the current time step according to the fusion features.
[0031] As an optional implementation of the device of the third aspect, the robot includes a mechanical arm, an image acquisition device and a perception system;
[0032] The mechanical arm is used to grab the image acquisition device and perform image acquisition operations;
[0033] The perception system is used to extract the certificate information required for the certificate information collection task from the certificate image collected by the image collection device
[0034] As an optional implementation of the device of the third aspect, the training module is specifically used to utilize the perception system of the robot to extract the document information required for the document information collection task from the document image.
[0035] As an optional implementation manner of the device described in the third aspect, the task prompt information is obtained by decomposing the certificate information collection task according to the time step.
[0036] In a fourth aspect, a device for collecting certificate information is provided, comprising:
[0037] A second data acquisition module is configured to construct task description information according to the target certificate information collection task;
[0038] The certificate information collection module is configured to input the task description information into an intelligent agent, and control the robot through the intelligent agent to perform the certificate information collection operation on the target certificate to obtain the certificate information required for the target certificate information collection task; the intelligent agent is pre-trained using the above-mentioned intelligent agent training method.
[0039] In a fifth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor executes the above-mentioned intelligent agent training method, or executes the above-mentioned document information collection method.
[0040] In a sixth aspect, an electronic device is provided, including:
[0041] One or more processors; and a memory associated with the one or more processors, the memory being used to store program instructions, which, when read and executed by the one or more processors, enable the electronic device to execute the above-mentioned intelligent agent training method, or, to execute the above-mentioned document information collection method.
[0042] The beneficial effect of the intelligent agent training method described in one or more embodiments of this specification is that the method uses a reinforcement learning strategy to train an intelligent agent for controlling a robot to collect document information. The intelligent agent includes multimodal inputs such as vision and text, has strong generalization capabilities, can adapt to the document information collection needs in different business scenarios, and has strong flexibility. Based on the intelligent agent, the robot can be controlled to automatically perform document collection tasks by constructing task description information, with high collection efficiency and the ability to accurately traverse the required collection scenarios. The intelligent agent training device, document information collection method and device described in the embodiments of this specification also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0044] Figure 1 A flowchart of an intelligent agent training method provided for one or more embodiments of this specification.
[0045] Figure 2 A schematic diagram of the structure of a multi-task action block model provided in one or more embodiments of this specification.
[0046] Figure 3A flowchart of a method for collecting certificate information provided in one or more embodiments of this specification.
[0047] Figure 4 A schematic diagram of the structure of an intelligent agent training device provided in one or more embodiments of this specification.
[0048] Figure 5 A schematic diagram of the structure of a certificate information collection device provided for one or more embodiments of this specification.
[0049] Figure 6 A schematic diagram of the structure of an electronic device provided in one or more embodiments of this specification. DETAILED DESCRIPTION
[0050] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0051] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0052] Those skilled in the art will appreciate that the terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0053] In many business scenarios, it is necessary to identify the certificate information, and it is necessary to collect a large amount of certificate information in different environmental scenarios, such as facial images, text information, patterns, etc. in the certificate, to train the certificate information recognition model. For example, in the certificate anti-counterfeiting recognition business scenario, it is necessary to collect a large number of forged certificate images, such as screen-shot certificate images, color-printed certificate images, and tampered certificate images, to train the certificate anti-counterfeiting recognition model. These forged certificate images also need to be collected in different light, collection angles, and other collection scenes.
[0054] Traditional manual collection methods are inefficient and costly, and manual collection is difficult to accurately traverse the above-mentioned collection scenarios.
[0055] In view of this, one or more embodiments of the present specification propose an intelligent agent training method, a document information collection method and a device, which can use the trained intelligent agent to control a robot to perform automated collection of document information required for different document information collection tasks, so as to at least partially overcome the above-mentioned technical problems.
[0056] The following will further describe in detail the intelligent agent training method, certificate information collection method and device described in one or more embodiments of this specification in conjunction with the drawings and specific embodiments of the specification, but this detailed description does not constitute a limitation on the embodiments of this specification.
[0057] Please refer to Figure 1 , Figure 1 Flowchart of an agent training method proposed in one or more embodiments of this specification. Figure 1 As shown, the method may include steps S100 to S106.
[0058] S100: Obtain training samples.
[0059] The training samples mentioned above include the view images, action sequences, joint position data and task prompt information of each time step when the robot performs the preset document information collection task. Among them, the action sequence mentioned above can be directly collected, or calculated according to the control instructions issued to the robot joints and grippers. The robot's joint position data can include parameters such as the robot's joint position, speed, and the position and speed of the grippers. The task prompt information is obtained by decomposing the document information collection task according to the time step, specifically referring to the information to be collected by the robot at each time step, such as the target object to be collected (such as identity card, passport, driver's license, etc.), the reference object in the collection path, the location of the document, and other information.
[0060] In order to obtain the above-mentioned training samples, based on the preset document information collection task, the robot can be controlled to execute the standard document image collection process for the document information collection task, and the robot's view image, action sequence, joint position data and task description information at each time step in the document image collection process can be collected as training samples.
[0061] The above-mentioned certificate information collection task can be set according to needs, and this embodiment does not limit this.
[0062] The number of the above-mentioned perspective images can be set according to the requirements, and this embodiment does not limit this. Taking the acquisition of four perspective images at each time step as an example, the four perspective images can be images of four different perspectives, or perspective images of two perspectives and depth maps of the two perspective images can be selected.
[0063] In some embodiments, the robot may include a mechanical arm, an image acquisition device, and a perception system. The mechanical arm is used to grab the image acquisition device and perform image acquisition operations; the perception system is used to extract the document information required for the document information acquisition task from the document image acquired by the image acquisition device.
[0064] In some implementations, data augmentation may be performed on the above-mentioned training samples to quickly increase the training sample data set with different semantic scene changes, thereby obtaining sufficient training samples to train an intelligent agent with generalization capabilities.
[0065] For example, the pre-collected training sample set can be increased by creating a diverse semantic enhancement set based on existing robot experience. When performing semantic enhancement, the robot's trajectory data can be retained, and new objects or new scene changes can be introduced by patching the view image in each trajectory, thereby achieving semantic enhancement of the trajectory data. Specifically, the semantic segmentation model SAM (Segment Anything Model) can be used to achieve semantic control-based target segmentation, which can segment and replace the background in the view image, and can also segment and replace the certificate in the view image, thereby creating a new view image.
[0066] S102: Input the training samples into the intelligent agent to obtain the document image acquisition action sequence predicted by the intelligent agent.
[0067] The above-mentioned intelligent agent can be implemented using a large language model LLM or a multimodal large model VLM. When a large language model LLM is used as an intelligent agent, the view image can be converted into a natural language description in combination with the image description generation technology, that is, the image features of the view image are extracted through a convolutional neural network, and then the image features are input into the large language model, so that the large language model can process image modality data.
[0068] In some embodiments, the above-mentioned intelligent agent can adopt a general robot intelligent agent RoboAgent, and then based on the RoboAgent framework, use training samples in the document information collection task scenario to generalize RoboAgent to achieve an intelligent agent suitable for document information collection tasks.
[0069] RoboAgent is implemented using the Multi-Task Action Chunking Transformer (MT-ACT), a new language-conditioned policy architecture for training powerful agents that can recover multiple skills from multimodal datasets.
[0070] like Figure 2 As shown in Figure 1, MT-ACT includes a language encoder (Lang Encoder), a conditional variational autoencoder (CVAE), a transformer encoder (Transformer Encoder) and a decoder (Transformer Decoder). The language encoder is used to learn the feature embedding T of the task description information. The conditional variational autoencoder is used to learn the action sequence a 1:T The transformer encoder is used to embed the features of the view image, task description information into T, the latent code z of the action sequence, and the joint position data j t Perform feature fusion to obtain fused features. The decoder is used to predict the next action block of the current time step based on the fused features.
[0071] S104: Control the robot to execute the certificate image acquisition action sequence, and extract the certificate information from the acquired certificate image.
[0072] In some implementations, all overlapping actions predicted in the current time step may be averaged, and the robot may be controlled to execute the averaged action to smooth the robot's motion trajectory.
[0073] In some embodiments, for the document image collected by the robot after executing a set of document image collection actions corresponding to the action block, the perception system can be used to extract the document information in the document image. The perception system is deployed with a recognition algorithm corresponding to the document information collection task, so that the document information required for the document information collection task can be extracted from the document image.
[0074] For example, in some business scenarios, it is necessary to detect certain targets in the ID image to determine the authenticity of the ID image. For this type of target detection task, an open set target detection algorithm (GroundingDINO) can be deployed in the perception system. By inputting the prompt word prompt, the open set target detection algorithm can detect various targets such as the corresponding ID, face, background, and screen of the ID photo, and return the detection box position information of the target.
[0075] In some scenarios, it is necessary to extract text information from the document image. For this type of OCR recognition task, an OCR recognition algorithm can be deployed in the perception system. After locating the document, the text lines in the collected document image of the document are detected by OCR text line detection to detect keywords in the document image, such as name, document number, etc.
[0076] In some scenarios, it is necessary to segment certain objects in the document image for subsequent processing. For such target segmentation tasks, a semantic segmentation model SAM (Segment Anything Model) can be deployed in the perception system to segment the target objects in the document image through semantic control. For example, a mask image of a given detection box can be output based on the given detection box.
[0077] It should be noted that the above-mentioned certificate information collection task can be adaptively set according to needs. Accordingly, the recognition algorithm deployed in the perception system can also be selected and deployed according to the needs of the certificate information collection task. This embodiment does not impose any restrictions on this.
[0078] S106: Determine an incentive signal based on the difference between the certificate information and the target information of the certificate information collection task, and use a reinforcement strategy to update the parameters of the agent until the target agent is obtained.
[0079] The target information of the above-mentioned certificate information collection task is related to the task content of the certificate information collection task. For example, if the certificate information collection task is a certificate text information extraction task, the target information is the text information of the certificate. If the certificate information collection task is a target detection task, the target information is the detection box and mask features of the detection target of the certificate.
[0080] In this step, an incentive function can be constructed based on the difference between the above-mentioned certificate information and the target information of the certificate information collection task. The incentive function gives a positive incentive signal when the above-mentioned certificate information is close to the target information of the certificate information collection task, and gives a negative incentive signal when the above-mentioned certificate information is far from the target information of the certificate information collection task. Through this design, the control agent generates an action block that can make the robot's certificate information collection result gradually approach the target information of the certificate information collection task.
[0081] The above is Figure 1The specific implementation process of the agent training method shown in the figure. This method trains a general agent suitable for the document collection scene for the vertical scene collection task of the document. The agent can process multimodal inputs such as vision and text, and has strong generalization ability, and can realize document image collection under different document information collection requirements. Based on this agent, combined with the document information extraction algorithm, various document information collection tasks can be generated by constructing prompt words. There is no need to retrain the agent, and different document information collection tasks can be completed automatically and accurately, which greatly improves the efficiency and accuracy of document information collection.
[0082] Based on the intelligent agent trained by the above-mentioned intelligent agent training method, one or more embodiments of this specification also propose a method for collecting certificate information. Figure 3 As shown, the method may include steps S300 to S304.
[0083] S300: Constructing task description information according to the target certificate information collection task.
[0084] Specifically, the target document information collection task can be decomposed according to preset time steps to obtain the task description information corresponding to each time step.
[0085] S302: Input the task description information into the intelligent agent, and control the robot through the intelligent agent to perform the document information collection operation on the target document to obtain the document information required for the target document information collection task.
[0086] The above-mentioned intelligent agent is pre-trained by using the above-mentioned intelligent agent training method. The intelligent agent can generate the optimal action block for the current time step based on the input task description information, so that the document information in the document image collected by the robot can gradually approach the target information of the document information collection task.
[0087] After obtaining the task description information, the task description information is constructed as prompt words and input into the agent, which can guide the agent to generate corresponding action blocks. According to the action blocks, the action control instructions of the robot are determined to drive the robot to execute a set of document image acquisition action sequences specified by the action blocks.
[0088] Corresponding to the above-mentioned intelligent agent training method, one or more embodiments of this specification also provide an intelligent agent training device. Figure 4 , Figure 4 1 is a schematic diagram of the structure of an intelligent agent training device proposed in one or more embodiments of this specification. The device can be used to implement the above-mentioned intelligent agent training method. It should be noted that the intelligent agent training method described in one or more embodiments of this application can rely on Figure 4The intelligent agent training device shown is implemented, but not limited to this device.
[0089] like Figure 4 As shown, the intelligent agent training device includes:
[0090] The first data acquisition module 401 is configured to acquire training samples, which include view images, action sequences, joint position data and task prompt information at each time step when the robot performs a preset document information collection task.
[0091] The training module 402 is configured to input the training samples into the intelligent agent to obtain the ID image acquisition action sequence predicted by the intelligent agent; control the robot to execute the ID image acquisition action sequence and extract the ID information from the acquired ID image; determine the incentive signal based on the difference between the ID information and the target information of the ID information acquisition task, and use the reinforcement strategy to update the parameters of the intelligent agent until the target intelligent agent is obtained.
[0092] With respect to the above-mentioned first data acquisition module 401, this module can collect the robot's view image, action sequence, joint position data and task description information at each time step in the standard document image acquisition process when the robot performs a preset document information acquisition task, as training samples.
[0093] Among them, the above-mentioned action sequence can be directly acquired through the first data acquisition module 401, or calculated according to the control instructions issued to the robot joints and grippers. The robot's joint position data can include parameters such as the robot's joint position, speed, and the position and speed of the grippers. The task prompt information is obtained by decomposing the document information collection task according to the time step, specifically referring to the information to be collected by the robot at each time step, such as the target object to be collected (such as an ID card, passport, driver's license, etc.), the reference object in the collection path, the location of the document, and other information.
[0094] The robot may include a mechanical arm, an image acquisition device, and a perception system. The mechanical arm is used to grab the image acquisition device and perform image acquisition operations. The perception system is used to extract the document information required for the document information acquisition task from the document image acquired by the image acquisition device.
[0095] In some embodiments, the first data acquisition module 401 may also perform data augmentation on the above-mentioned training samples to quickly increase the training sample data set with different semantic scene changes, thereby obtaining sufficient training samples to train an intelligent agent with generalization ability.
[0096] For example, the first data acquisition module 401 can increase the pre-collected training sample set by creating a diversified semantic enhancement set based on existing robot experience. When performing semantic enhancement, the robot's trajectory data can be retained, and new objects or new scene changes can be introduced by patching the perspective image in each trajectory, thereby achieving semantic enhancement of the trajectory data. Specifically, the first data acquisition module 401 can achieve target segmentation based on semantic control through the semantic segmentation model SAM (Segment Anything Model), and can segment and replace the background in the perspective image, and can also segment and replace the certificate in the perspective image, thereby creating a new perspective image.
[0097] For the above-mentioned training module 402, the intelligent agent trained by the module can be implemented by using a large language model LLM or a multimodal large model VLM. When the large language model LLM is used as the intelligent agent, the view image can be converted into a natural language description in combination with the image description generation technology, that is, the image features of the view image are extracted through a convolutional neural network, and then the image features are input into the large language model, so that the large language model can process image modality data.
[0098] In some embodiments, the agent deployed in the training module 402 can adopt the basic framework of the general robot agent RoboAgent. RoboAgent is implemented using the multi-task action chunking model MT-ACT (Multi-Task ActionChunking Transformer), including a language encoder (Lang Encoder), a conditional variational autoencoder (CVAE), a transformer encoder (Transformer Encoder) and a decoder (Transformer Decoder). The language encoder is used to learn the feature embedding T of the task description information. The conditional variational autoencoder is used to learn the action sequence a 1:T The transformer encoder is used to embed the features of the view image, task description information into T, the latent code z of the action sequence, and the joint position data j t Perform feature fusion to obtain fused features. The decoder is used to predict the next action block of the current time step based on the fused features.
[0099] In some embodiments, before the training module 402 controls the robot to execute a set of document image acquisition actions corresponding to the action block, it can also average all overlapping actions predicted for the current time step and control the robot to execute the averaged action to smooth the robot's motion trajectory.
[0100] In some embodiments, the training module 402 can use the perception system to extract the document information in the document image collected by the robot after executing a set of document image collection actions corresponding to the action block. The perception system is deployed with a recognition algorithm corresponding to the document information collection task, so that the document information required for the document information collection task can be extracted from the document image.
[0101] In some embodiments, the training module 402 can construct an incentive function based on the difference between the above-mentioned certificate information and the target information of the certificate information collection task, and the incentive function gives a positive incentive signal when the above-mentioned certificate information is close to the target information of the certificate information collection task, and gives a negative incentive signal when the above-mentioned certificate information is far from the target information of the certificate information collection task. Based on this design, through reinforcement learning, it is possible to control the intelligent agent to generate action blocks that make the certificate information collection results of the robot gradually approach the target information of the certificate information collection task.
[0102] For the above-mentioned intelligent body training device, taking a module as an example of a software functional unit, the first data acquisition module 401 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the first data acquisition module 401 may include code running on multiple hosts / virtual machines / containers. The multiple hosts / virtual machines / containers used to run the code may be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.
[0103] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0104] As an example of a hardware functional unit, the first data acquisition module 401 may include at least one computing device, such as a server, etc. Alternatively, the first data acquisition module 401 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0105] The multiple computing devices included in the first data acquisition module 401 can be distributed in the same region or in different regions. The multiple computing devices included in the first data acquisition module 401 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the first data acquisition module 401 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0106] In other embodiments, the first data acquisition module 401 can be used to execute any step in the above-mentioned intelligent agent training method, and the training module 402 can be used to execute any step in the above-mentioned intelligent agent training method. The steps that the first data acquisition module 401 and the training module 402 are responsible for implementing can be specified as needed, and the first data acquisition module 401 and the training module 402 respectively implement different steps in the above-mentioned intelligent agent training method to realize all the functions of the above-mentioned intelligent agent training device.
[0107] In this implementation, the intelligent agent training device can also be applied to computing devices such as computers and servers, or to a computing device cluster including at least one computing device, to realize the specific functions of the intelligent agent training device.
[0108] Corresponding to the above-mentioned document information collection method, one or more embodiments of this specification also propose a document information collection device. Figure 5 , Figure 51 is a schematic diagram of the structure of a certificate information collection device proposed in one or more embodiments of this specification. The device can be used to implement the certificate information collection method described above. It should be noted that the certificate information collection method described in one or more embodiments of this application can rely on Figure 5 The document information collection device shown is implemented, but not limited to this device.
[0109] like Figure 5 As shown, the certificate information collection device includes:
[0110] The second data acquisition module 501 is configured to construct task description information according to the target certificate information collection task.
[0111] The certificate information collection module 502 is configured to input the task description information into the intelligent agent, and control the robot through the intelligent agent to perform the certificate information collection operation on the target certificate to obtain the certificate information required for the target certificate information collection task.
[0112] With respect to the above-mentioned second data acquisition module 501, the module can specifically decompose the target certificate information collection task according to preset time steps, so as to obtain task description information corresponding to each time step.
[0113] For the above-mentioned document information collection module 502, the intelligent agent used in this module is pre-trained using the above-mentioned intelligent agent training method. The intelligent agent can generate the optimal action block for the current time step based on the input task description information, so that the document information in the document image collected by the robot can gradually approach the target information of the document information collection task.
[0114] After obtaining the task description information generated by the second data acquisition module 501, the document information acquisition module 502 constructs the task description information as prompt words and inputs them into the agent, thereby guiding the agent to generate corresponding action blocks. According to the action blocks, the action control instructions of the robot are determined to drive the robot to execute a set of document image acquisition action sequences specified by the action blocks.
[0115] For the above-mentioned certificate information collection device, taking a module as an example of a software functional unit, the second data acquisition module 501 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the second data acquisition module 501 may include code running on multiple hosts / virtual machines / containers. The multiple hosts / virtual machines / containers used to run the code may be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.
[0116] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0117] As an example of a hardware functional unit, the second data acquisition module 501 may include at least one computing device, such as a server, etc. Alternatively, the second data acquisition module 501 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0118] The multiple computing devices included in the second data acquisition module 501 can be distributed in the same region or in different regions. The multiple computing devices included in the second data acquisition module 501 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the second data acquisition module 501 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0119] In other embodiments, the second data acquisition module 501 can be used to execute any step in the above-mentioned certificate information collection method, and the certificate information collection module 502 can be used to execute any step in the above-mentioned certificate information collection method. The steps that the second data acquisition module 501 and the certificate information collection module 502 are responsible for implementing can be specified as needed, and the second data acquisition module 501 and the certificate information collection module 502 respectively implement different steps in the above-mentioned certificate information collection method to realize all the functions of the above-mentioned certificate information collection device.
[0120] In this implementation, the certificate information collection device can also be applied to computing devices such as computers and servers, or to a computing device cluster including at least one computing device, to achieve the specific function of certificate information collection.
[0121] In some embodiments, an electronic device is also provided. Figure 6 , the electronic device includes: a bus 601, a processor 602, a memory 603 and a communication interface 604. The processor 602, the memory 603 and the communication interface 604 communicate with each other through the bus 601. The electronic device can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the electronic device.
[0122] The bus 601 may be a peripheral component interconnect (PCI) bus, or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The bus 601 is represented by only one line, but it does not mean that there is only one bus or one type of bus. The bus 601 may include a path for transmitting information between various components of the electronic device (for example, the processor 602, the memory 603 and the communication interface 604).
[0123] The processor 602 may include any one or more of a processor such as a CPU, a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0124] The memory 603 may include a volatile memory, such as a random access memory (RAM). The memory 603 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0125] The memory 603 stores executable program codes, and the processor 602 executes the executable program codes to implement the aforementioned intelligent agent training method, or to implement the aforementioned document information collection method.
[0126] The communication interface 604 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the electronic device and other devices or a communication network.
[0127] In some embodiments, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the processor executes the above-mentioned intelligent agent training method, or executes the above-mentioned document information collection method.
[0128] The computer-readable storage medium may be any available medium that can be stored in an electronic device or a data storage device such as a data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct the electronic device to execute the above-mentioned agent training method, or to execute the above-mentioned document information collection method.
[0129] It is to be understood that the structure illustrated in the embodiments of this specification does not constitute a specific limitation on the system of the embodiments of this specification. In other embodiments of the specification, the above system may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0130] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0131] The above is a description of a specific embodiment of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0132] It should be noted that the above examples are only specific embodiments of the present invention. Obviously, the present invention is not limited to the above examples, and there are many similar variations. All variations directly derived or associated from the contents disclosed by the technicians in this field should fall within the protection scope of the present invention.
Claims
1. A method for training an intelligent agent, comprising: Acquire training samples, wherein the training samples include view images, action sequences, joint position data, and task prompt information at each time step of the robot in the process of performing a preset document information collection task; Inputting the training sample into the intelligent agent to obtain a document image acquisition action sequence predicted by the intelligent agent; Controlling the robot to execute the certificate image acquisition action sequence and extracting certificate information from the acquired certificate image; An incentive signal is determined based on the difference between the certificate information and the target information of the certificate information collection task, and a reinforcement strategy is adopted to update the parameters of the agent until a target agent is obtained.
2. According to the method of claim 1, the agent is implemented using a multi-task action block model; The multi-task action block model includes: A language encoder for learning feature embedding of the task description information; A conditional variational autoencoder for learning a latent encoding of the action sequence; A transformer encoder, configured to perform feature fusion on the view image, the feature embedding of the task description information, the latent coding of the action sequence and the joint position data to obtain a fused feature; A decoder is used to predict the next set of document image acquisition action sequences of the current time step according to the fusion features.
3. The method according to claim 1, wherein the robot comprises a mechanical arm, an image acquisition device and a perception system; The mechanical arm is used to grab the image acquisition device and perform image acquisition operations; The perception system is used to extract the certificate information required for the certificate information collection task from the certificate image collected by the image collection device.
4. The method according to claim 1, controlling the robot to execute a set of document image acquisition actions corresponding to the action block and extracting document information from the acquired document image, specifically comprising: The robot's perception system is used to extract the document information required for the document information collection task from the document image.
5. According to the method of claim 1, the task prompt information is obtained by decomposing the document information collection task according to the time step.
6. A method for collecting certificate information, comprising: According to the target certificate information collection task, construct task description information; The task description information is input into an intelligent agent, and a robot is controlled by the intelligent agent to perform an operation of collecting certificate information on a target certificate, so as to obtain the certificate information required for the target certificate information collection task; the intelligent agent is pre-trained using the method described in any one of claims 1 to 5.
7. An intelligent agent training device, comprising: A first data acquisition module is configured to acquire training samples, wherein the training samples include a view image, an action sequence, joint position data, and task prompt information at each time step of the robot in a process of performing a preset document information collection task; The training module is configured to input the training samples into the intelligent agent to obtain the ID image acquisition action sequence predicted by the intelligent agent; control the robot to execute the ID image acquisition action sequence and extract ID information from the acquired ID image; determine the incentive signal based on the difference between the ID information and the target information of the ID information acquisition task, and adopt the reinforcement strategy to update the parameters of the intelligent agent until the target intelligent agent is obtained.
8. The device according to claim 7, wherein the agent is implemented using a multi-task action block model; The multi-task action block model includes: A language encoder for learning feature embedding of the task description information; A conditional variational autoencoder for learning a latent encoding of the action sequence; A transformer encoder, configured to perform feature fusion on the view image, the feature embedding of the task description information, the latent coding of the action sequence and the joint position data to obtain a fused feature; A decoder is used to predict the next set of document image acquisition action sequences of the current time step according to the fusion features.
9. The device according to claim 7, wherein the robot comprises a mechanical arm, an image acquisition device and a perception system; The mechanical arm is used to grab the image acquisition device and perform image acquisition operations; The perception system is used to extract the certificate information required for the certificate information collection task from the certificate image collected by the image collection device.
10. The device according to claim 7, wherein the training module is specifically used to utilize the perception system of the robot to extract the document information required for the document information collection task from the document image.
11. The device according to claim 7, wherein the task prompt information is obtained by decomposing the certificate information collection task according to the time step.
12. A certificate information collection device, comprising: A second data acquisition module is configured to construct task description information according to the target certificate information collection task; The certificate information collection module is configured to input the task description information into an intelligent agent, and control the robot through the intelligent agent to perform a certificate information collection operation on the target certificate to obtain the certificate information required for the target certificate information collection task; the intelligent agent is pre-trained using the method described in any one of claims 1 to 5.
13. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor executes the method according to any one of claims 1 to 5, or the method according to claim 6.
14. An electronic device comprising: one or more processors; And a memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the electronic device executes any method as claimed in claims 1 to 5, or executes the method as claimed in claim 6.