Model training method and device, equipment and computer medium
By training a target interaction model and combining word processing and visual encoding techniques, the problem of low accuracy in intelligent robots when performing tasks was solved, and the accuracy and generalization ability of task execution were improved.
Patent Information
- Application Number
- CN202311559285.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2043-11-21
AI Technical Summary
When intelligent robots execute tasks according to user instructions, the accuracy of task execution is low, especially when dealing with semantic ambiguity in visual information and scene analysis.
By acquiring sample task and interaction information, word processing and visual encoding are performed using the initial interaction model to generate word vector sequences and image feature sequences. These are then analyzed in conjunction with a transformation model to train the target interaction model and improve the accuracy of task execution.
It improves the accuracy of intelligent robots in executing tasks according to user instructions, enhances their ability to recognize various task types and interactive objects, and improves their generalization ability.
Smart Images

Figure CN117851552B_ABST
Abstract
Description
Technical Field
[0001] This disclosure pertains to the field of intelligent control, and particularly relates to a model training method, apparatus, device, and computer medium. Background Technology
[0002] With the widespread use of intelligent robots, the accuracy of task execution based on user instructions is becoming increasingly important. In some scenarios, intelligent robots need to perform tasks based on interactions with users, and these interactions may be ambiguous and require resolution. However, current technologies lack effective solutions to this problem, resulting in low accuracy for intelligent robots when executing tasks based on user instructions. Summary of the Invention
[0003] This disclosure provides an implementation scheme different from related technologies to solve the technical problem of low accuracy in task execution by intelligent robots when performing tasks according to user instructions in related technologies.
[0004] Firstly, this disclosure provides a model training method, including:
[0005] The task of obtaining samples;
[0006] Based on the type of the sample task, obtain the corresponding sample interaction information and the sample image corresponding to the sample interaction information; the sample interaction information includes: sample interaction content and sample task execution result;
[0007] The initial word processing unit in the initial interaction model performs word processing on the sample task and the sample interaction content to obtain the corresponding word vector sequence.
[0008] The sample image is encoded by the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image;
[0009] The word vector sequence and the image feature sequence are analyzed using the initial transformation model in the initial interaction model to obtain the corresponding task prediction results;
[0010] The initial interaction model is trained based on the sample task execution results and the task prediction results to obtain a target interaction model. The target interaction model is used to interact with the interaction object based on the first interaction content with the interaction object and the environmental image.
[0011] Secondly, this disclosure provides a model training apparatus, comprising:
[0012] The acquisition unit is used for the task of acquiring samples;
[0013] The acquisition unit is further configured to acquire corresponding sample interaction information and sample images corresponding to the sample interaction information based on the type of the sample task; the sample interaction information includes: sample interaction content and sample task execution result;
[0014] The processing unit is used to perform word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence.
[0015] The encoding unit is used to encode the sample image through the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image;
[0016] The analysis unit is used to analyze the word vector sequence and the image feature sequence through the initial transformation model in the initial interaction model to obtain the corresponding task prediction result;
[0017] The training unit is used to train the initial interaction model based on the sample task execution results and the task prediction results to obtain a target interaction model. The target interaction model is used to interact with the interaction object based on the first interaction content with the interaction object and the environmental image.
[0018] Thirdly, this disclosure provides an electronic device, including:
[0019] Processor; and
[0020] Memory for storing the executable instructions of the processor;
[0021] The processor is configured to execute the first aspect or any of the possible implementations of the first aspect by executing the executable instructions.
[0022] Fourthly, embodiments of this disclosure provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the methods in the first aspect or any of the possible implementations of the first aspect.
[0023] This disclosure provides a sample acquisition task; based on the type of the sample task, it acquires corresponding sample interaction information and sample images corresponding to the sample interaction information; the sample interaction information includes: sample interaction content and sample task execution results; the sample task and sample interaction content are processed by the initial word processing unit in the initial interaction model to obtain corresponding word vector sequences; the sample images are encoded by the initial visual encoding unit in the initial interaction model to obtain image feature sequences corresponding to the sample images; the initial transformation model in the initial interaction model is used to analyze the word vector sequences and image feature sequences to obtain the corresponding task. Prediction results; Based on the sample task execution results and the task prediction results, the initial interaction model is trained to obtain a target interaction model. The target interaction model is used to: determine the interaction scheme with the interaction object based on the first interaction content with the interaction object and the environmental image; the sample interaction information used to train the target interaction model can be determined based on the type of sample task, giving the target interaction model the function of recognizing different task types, thereby enabling the trained target recognition model to have the ability to interact with the interaction object according to multiple task types, thereby improving the generalization ability of the target recognition model and improving the accuracy of the intelligent robot in performing tasks according to user instructions. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments or related technologies of this disclosure, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0025] Figure 1 This is a schematic diagram of the structure of a system provided in one embodiment of the present disclosure;
[0026] Figure 2a A schematic flowchart of a model training method provided in an embodiment of this disclosure;
[0027] Figure 2b A schematic flowchart of a model training method provided in an embodiment of this disclosure;
[0028] Figure 2c A flowchart illustrating an interaction method provided in an embodiment of this disclosure;
[0029] Figure 3 This is a schematic diagram of the structure of a model training device provided in an embodiment of the present disclosure;
[0030] Figure 4This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0031] Embodiments of this disclosure are described in detail below, with examples of these embodiments illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting it.
[0032] The terms "first" and "second," etc., used in the specification, claims, and drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the present disclosure described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] First, some terms used in the embodiments of this disclosure will be explained below to facilitate understanding by those skilled in the art.
[0034] Interactive Visual Semantic Disambiguation (IVSD) dialogue refers to resolving semantic ambiguity in visual information through human-machine interaction. Typically, IVSD dialogue employs a natural language-based interaction method, allowing human users to converse with the machine and provide additional information or explanations in natural language to help the machine better understand the semantic information in images.
[0035] In related technologies, when an intelligent robot performs a task based on user instructions, it locates the target object related to the task in an image based on the interaction with the user, and then performs the task based on the interaction. This method requires a high level of understanding from the intelligent robot; existing intelligent robots cannot accurately locate the target object based on the interaction with the user, resulting in low accuracy in task execution.
[0036] Furthermore, the robot's ability to perform interactive visual semantic disambiguation is weak. The intelligent robot cannot accurately understand the surrounding scene or analyze the target objects in the scene, which will affect the accuracy of the intelligent robot in performing tasks.
[0037] The technical solutions of this disclosure and how they solve the aforementioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this disclosure will now be described with reference to the accompanying drawings.
[0038] Figure 1 A schematic diagram of a system structure provided for an exemplary embodiment of this disclosure, the structure including: a model training device 10 and an intelligent robot 11, wherein the model training device 10 is used for:
[0039] The task of obtaining samples;
[0040] Based on the type of the sample task, obtain the corresponding sample interaction information and the sample image corresponding to the sample interaction information; the sample interaction information includes: sample interaction content and sample task execution result;
[0041] The initial word processing unit in the initial interaction model performs word processing on the sample task and the sample interaction content to obtain the corresponding word vector sequence.
[0042] The sample image is encoded by the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image;
[0043] The word vector sequence and the image feature sequence are analyzed using the initial transformation model in the initial interaction model to obtain the corresponding task prediction results;
[0044] The initial interaction model is trained based on the sample task execution results and the task prediction results to obtain a target interaction model. The target interaction model is used to interact with the interaction object based on the first interaction content with the interaction object and the environmental image.
[0045] After the model training device 10 trains the target interaction model, it can send the target interaction model to the intelligent robot 11, and the intelligent robot 11 can interact with the interaction object based on the target interaction model.
[0046] In some optional embodiments of this application, the process of training the target interaction model and the process of interacting with the interaction object based on the target interaction model can both be performed by the intelligent robot 11.
[0047] Figure 2a This is a flowchart illustrating a model training method provided as an exemplary embodiment of the present disclosure. The method can be executed by the aforementioned intelligent robot or by the aforementioned model training device; this application does not limit the execution of this method. The method includes at least the following steps S201-S206:
[0048] S201, Sample Acquisition Task;
[0049] In some optional embodiments of this application, the sample task is any of the following tasks: answering a question posed by a questioner, making a corresponding question based on the interaction content with the respondent, indicating the first region information of the image area occupied by the first target object corresponding to the interaction content in the sample image based on the interaction content with the questioner or the respondent, answering whether the first target object corresponding to the interaction content can be specified in the sample image based on the interaction content with the questioner or the respondent, or describing the sample image.
[0050] In some optional embodiments of this application, the task of answering the question raised by the questioner may include: answering the question raised by the questioner based on a sample image, specifically, it may include: answering the question raised by the questioner based on a preset area in the sample image.
[0051] In some optional embodiments of this application, when the task type of the sample task is to answer a question raised by the questioner based on a preset region in the sample image, the aforementioned word processing of the sample task and the sample interaction content by the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence means: the initial word processing unit in the initial interaction model processes the sample task, the regional information of the preset region, and the sample interaction content to obtain the corresponding word vector sequence.
[0052] In some optional embodiments, the sample task can be text information or voice information; optionally, the sample task can be any of the following:
[0053] (1)Answer the question based on the region.
[0054] region:<bin_362><bin_655><bin_776><bin_835> .
[0055] (2)Ask a question to guess what I want.
[0056] (3)Which region does the context describe?
[0057] Among them, the task type of sample task (1) is: to answer the question raised by the questioner based on a preset area in the sample image. The task type of sample task (2) is: to make a corresponding question based on the interaction content with the questioner or the answerer. The task type of sample task (3) is: to point out the first area information of the first target object in the image area corresponding to the interaction content in the sample image based on the interaction content with the questioner or the answerer.
[0058] S202. Obtain the corresponding sample interaction information and the sample image corresponding to the sample interaction information based on the type of the sample task; the sample interaction information includes: sample interaction content and sample task execution result;
[0059] In some optional embodiments of this application, the sample interaction information may include interaction information between multiple roles, such as multiple dialogue messages. Specifically, a dialogue message may include the role's name and the dialogue content between the role and its opposing role. The dialogue content may include text information and / or voice information.
[0060] In some embodiments of this application, the sample interaction content precedes the sample task execution result in the sample interaction information, and the sample interaction content is adjacent to the sample task execution result. Both the sample interaction content and the sample task execution result belong to dialogue information.
[0061] In some optional embodiments of this application, the step of obtaining the corresponding sample interaction information based on the type of the sample task in S202 includes the following S2021-S2023:
[0062] S2021. Determine the corresponding interaction role based on the type of the sample task, wherein the interaction role is a questioner or a respondent;
[0063] In some optional embodiments of this application, determining the corresponding interaction role based on the type of the sample task in S2021 may include: determining the interaction role corresponding to the type of the sample task based on a preset correspondence. The preset correspondence may store different correspondences between sample task types and their corresponding interaction roles.
[0064] In some optional embodiments of this application, when the type of the sample task is to answer a question raised by a questioner, the interaction role corresponding to the type of sample task is the answerer.
[0065] In some optional embodiments of this application, when the type of the sample task is a task that asks a question based on the interaction content with the respondent, the interaction role corresponding to the type of sample task is the questioner.
[0066] In some optional embodiments of this application, when the type of the sample task is a task based on the interaction content with the questioner, that is, to point out the first region information of the image area occupied by the first target object corresponding to the interaction content in the sample image, the interaction role corresponding to the type of the sample task is the respondent.
[0067] In some optional embodiments of this application, when the type of sample task is a task based on the interaction content with the respondent, indicating the first region information of the image area occupied by the first target object corresponding to the interaction content in the sample image, the interaction role corresponding to the type of sample task is the questioner.
[0068] In some optional embodiments of this application, when the type of the sample task is a task based on the interaction content with the questioner, answering whether a first target object corresponding to the interaction content can be specified in the sample image, the interaction role corresponding to the type of sample task is the respondent.
[0069] In some optional embodiments of this application, when the type of the sample task is a task based on the interaction content with the respondent, answering whether a first target object corresponding to the interaction content can be specified in the sample image, the interaction role corresponding to the type of the sample task is the questioner.
[0070] In some optional embodiments of this application, when the type of the sample task is a task describing the sample image, the interaction role corresponding to the type of sample task is a responder or a questioner.
[0071] S2022. Based on the interaction role, determine the sample task execution result from the target sample interaction information set;
[0072] Optionally, the sample interaction information base may contain multiple sample interaction information sets. Each sample interaction information set can be considered as a group of dialogue information, and each group of dialogue information may include multiple dialogue messages. The target sample interaction information set is the sample interaction information set in the sample interaction information base that corresponds to the type of sample task.
[0073] Optionally, the above method further includes: acquiring multiple sample interaction information sets; and determining the target sample interaction information set corresponding to the type of the sample task from the multiple sample interaction information sets. Specifically, each type of sample task corresponds to one sample interaction information set.
[0074] In some optional embodiments of this application, in the aforementioned S2022, determining the sample task execution result from the target sample interaction information set based on the interaction role includes: randomly selecting a dialogue content between the interaction role and its opposing role from the target sample interaction information set; and using the dialogue content as the sample task execution result.
[0075] In some other optional embodiments of this application, in the aforementioned S2022, determining the sample task execution result from the target sample interaction information set based on the interaction role includes: determining the sample task execution result from the target sample interaction information set based on the interaction role and the type of the sample task.
[0076] Optionally, the aforementioned determination of the sample task execution result from the target sample interaction information set based on the type of the interaction role and the sample task includes: when the type of the sample task is to answer a question raised by the questioner, the dialogue content of the interaction role in the target sample interaction information set, where the preceding sentence is a question dialogue content of the opposing role, is taken as the sample task execution result, and the dialogue content of the interaction role is a declarative sentence.
[0077] Optionally, based on the type of the interaction role and the sample task, the sample task execution result is determined from the target sample interaction information set, including: when the type of the sample task is a task of making a corresponding question based on the interaction content with the respondent, the dialogue content of the interaction role in the target sample interaction information set, where the preceding sentence is the dialogue content of the opposing role, is taken as the sample task execution result, and the dialogue content of the interaction role is a question.
[0078] Optionally, the aforementioned determination of the sample task execution result from the target sample interaction information set based on the interaction role and the type of the sample task includes: when the type of the sample task is a task based on the interaction content with the questioner or the answerer, indicating the first region information of the image region occupied by the first target object corresponding to the interaction content in the sample image, the dialogue content of the interaction role in the target sample interaction information set, where the preceding sentence is the dialogue content of the opposing role, is taken as the sample task execution result, and the dialogue content of the interaction role contains region information.
[0079] S2023. The target sample interaction information is collected, and the dialogue content before the sample task execution result is taken as the sample interaction content.
[0080] In some optional embodiments of this application, the target sample interaction information set contains a set of dialogue information.
[0081] Optionally, the set of dialogue information referred to in this application refers to the dialogue information that is clear and unambiguous to the target at the end of the dialogue.
[0082] In some other optional embodiments of this application, the target sample interaction information set may include multiple sets of dialogue information. In the aforementioned S2023, the dialogue content before the sample task execution result in the target sample interaction information set is used as the sample interaction content, which includes: using the dialogue content before the sample task execution result in any set of dialogue information in the target sample interaction information set as the sample interaction content.
[0083] When the type of the sample task is based on the interaction content with the questioner or responder, and the task is to answer whether a first target object corresponding to the interaction content can be specified in the sample image, the step of obtaining the corresponding sample interaction information based on the type of the sample task includes the following S01-S03:
[0084] S01. Determine the corresponding interaction role based on the type of the sample task, wherein the interaction role is either a questioner or a respondent;
[0085] When the type of the sample task is based on the interaction content with the questioner, and the task is to answer whether the first target object corresponding to the interaction content can be specified in the sample image, the interaction role corresponding to the type of sample task is the respondent.
[0086] When the type of sample task is based on the interaction content with the respondent, and the task is to answer whether the first target object corresponding to the interaction content can be specified in the sample image, the interaction role corresponding to the type of sample task is the questioner.
[0087] S02. Obtain the preset sample task execution result, wherein the preset sample task execution result is yes or no;
[0088] S03. When the sample task execution result is yes, and the target sample interaction information set includes only one set of dialogue information, then all dialogue content in the target sample interaction information set is taken as the sample interaction content; when the sample task execution result is no, and the target sample interaction information set includes only one set of dialogue information, then part of the dialogue content in the target sample interaction information set is taken as the sample interaction content. When the sample task execution result is yes, and the target sample interaction information set includes multiple sets of dialogue information, then all dialogue content in any set of dialogue information in the target sample interaction information set is taken as the sample interaction content; when the sample task execution result is no, and the target sample interaction information set includes multiple sets of dialogue information, then part of the dialogue content in any set of dialogue information in the target sample interaction information set is taken as the sample interaction content.
[0089] In some optional embodiments of this application, see Figure 2bAs shown, when the sample task is: answer the question based on the region, the preset region information is:<bin_362><bin_655><bin_776><bin_835> The sample interaction content can be:
[0090] agent: Can you help me?
[0091] human:Which one do you want?
[0092] The results of the sample task execution can be:
[0093] I want that banana.
[0094] When the sample task is: ask a question to guess what I want, the sample interaction content can be:
[0095] Human: Can you help me?
[0096] agent:Which one do you want?
[0097] human:Give me the monkey's favorite fruit.
[0098] The results of the sample task execution can be:
[0099] Is it the banana?
[0100] When the sample task is: which region does the context describe?, the sample interaction content can be:
[0101] Human: Can you help me?
[0102] agent:Which one do you want?
[0103] human:Give me the monkey's favorite fruit.
[0104] The results of the sample task execution can be:<bin_362><bin_655><bin_776><bin_835> .
[0105] S203. The initial word processing unit in the initial interaction model performs word processing on the sample task and the sample interaction content to obtain the corresponding word vector sequence.
[0106] In some optional embodiments of this application, see Figure 2b As shown, the initial word processing unit includes a word segmentation unit and an initial embedding layer unit. In the aforementioned S203, the initial word processing unit in the initial interaction model performs word processing on the sample task and the sample interaction content to obtain the corresponding word vector sequence, including the following S2031-S2032:
[0107] S2031. The sample task and the sample interaction content are segmented into words by the word segmentation unit in the initial interaction model to obtain the word segmentation result.
[0108] S2032. The word segmentation result is processed by the initial embedding layer unit in the initial interaction model to obtain the corresponding word vector sequence.
[0109] The parameters in the initial embedded layer unit are adjustable.
[0110] S204. The sample image is encoded by the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image;
[0111] In some optional embodiments of this application, see [link to relevant documentation]. Figure 2b As shown, the initial visual encoding unit includes: an initial residual neural network and an initial linear projection layer. In the aforementioned S204, the process of encoding the sample image through the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image includes the following S2041-S2042:
[0112] S2041. The sample image is encoded using the initial residual neural network to obtain image feature information;
[0113] S2042. The image feature information is dimensionally transformed through the initial linear projection layer to obtain the image feature sequence corresponding to the sample image.
[0114] Optionally, before encoding the sample image through the initial residual neural network, the size of the sample image can be adjusted to a preset size, specifically, the preset size can be 512*512.
[0115] Encoding the sample image using the initial residual neural network to obtain image feature information may include: determining image feature information of a corresponding size for the sample image using the initial residual neural network. Optionally, the corresponding size is related to the parameters of the initial residual neural network. Optionally, the corresponding size can be 32*32.
[0116] In some embodiments, the image feature sequence obtained by performing dimensional transformation on the image feature information through the initial linear projection layer can be an image feature sequence with dimensional information corresponding to the parameters of the initial linear projection layer. This dimensional information can be 1*1024.
[0117] S205. Analyze the word vector sequence and the image feature sequence using the initial transformation model in the initial interaction model to obtain the corresponding task prediction result;
[0118] Optionally, the initial conversion model in this application is a backbone network, and the parameter capacity of the initial conversion model can be 930M. The initial conversion model is a transformer model, which can consist of multiple encoders and decoders. Specifically, the number of encoder layers can be 24, and the number of decoder layers can be 12. The transformer is a sequential-to-sequential architecture, which involves a beam search method, and the beam length is equal to 5.
[0119] Optionally, the aforementioned image feature sequence can be input into the first layer encoder, and the decoder can output the task prediction result corresponding to the autoregressive sample interaction information.
[0120] S206. The initial interaction model is trained based on the sample task execution results and the task prediction results to obtain a target interaction model. The target interaction model is used to interact with the interaction object based on the first interaction content with the interaction object and the environmental image.
[0121] In some optional embodiments of this application, in the aforementioned S206, the initial interaction model is trained based on the sample task execution result and the task prediction result to obtain the target interaction model, including the following S2061-S2062:
[0122] S2061. Determine loss information based on a preset loss function, the sample task execution result, and the task prediction result;
[0123] S2062. If the loss information is less than a preset threshold, the initial interaction model is used as the target interaction model; if the loss information is not less than the preset threshold, the parameters in the initial interaction model are adjusted, and the process returns to the step of performing word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence, until the target interaction model is determined when the loss information is less than the preset threshold.
[0124] In some other optional embodiments of this application, in the aforementioned S206, the initial interaction model is trained based on the sample task execution results and the task prediction results to obtain the target interaction model, including the following S261-S262:
[0125] S261. Determine loss information based on a preset loss function, the sample task execution result, and the task prediction result;
[0126] S262. If the number of times the parameters in the initial interaction model are adjusted exceeds a preset number, then the initial interaction model is taken as the target interaction model; if the number of times the parameters in the initial interaction model are adjusted does not exceed the preset number, then the parameters in the initial interaction model are adjusted based on the loss information, and the process returns to the step of performing word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence, until the target interaction model is determined.
[0127] In some optional embodiments of this application, when the type of the sample task is: based on the interaction content with the questioner or the answerer, the task of indicating the first region information of the image area occupied by the first target object corresponding to the interaction content in the sample image, the first region information of the image area occupied by the first target object in the sample image is used to indicate the coordinate range of the area occupied by the first target object in the sample image.
[0128] Optionally, the first target object corresponding to the interactive content is an object that is associated with the interactive content.
[0129] Optionally, the total range of the sample image can be represented by (a, b, c, d). In some embodiments, a and b represent the coordinates of the upper left corner of the sample image, and c and d represent the coordinates of the lower right corner of the sample image.
[0130] The first region information of the image area occupied by the first target object in the sample image can be represented by (x1, y1, x2, y2). Here, x1 and y1 represent the coordinates of the upper left corner of the image area occupied by the first target object in the sample image, and x2 and y2 represent the coordinates of the lower right corner of the image area occupied by the first target object in the sample image.
[0131] Optionally, the total range of the sample image can be normalized, normalizing each coordinate information in the sample image to the range of 0-1. After normalization, a and b are [0, 0], and c and d are (1, 1).
[0132] Optionally, after normalizing each coordinate information of the sample image to between 0 and 1, each coordinate information can also be mapped to between 0 and a preset value, where the preset value can be 999.
[0133] At this point, each coordinate information in the sample image can be represented as "<bin_i> ", where i∈{0,1,…,999}. For example, the first region information of the image region occupied by the first target object in the sample image can be converted from (0.0,0.12,0.3,0.4) into "".<bin_0> ,<bin_120> ,<bin_300> ,<bin_400> ".
[0134] In some optional embodiments of this application, it may be combined with Figure 2c As shown, the method further includes the following S01-S05:
[0135] S01. Obtain the first interactive content and the environmental image;
[0136] Optionally, the first interaction content is the interaction content between the intelligent robot and the user, obtained by the intelligent robot after the target interaction model has been trained and built into the intelligent robot. The environmental image is an environmental image captured when the first interaction content is obtained.
[0137] Optionally, the acquisition time of the first interactive content and the acquisition time of the environmental image are the same.
[0138] S02. Obtain a first preset task, wherein the task type of the first preset task is a task based on the first interaction content to determine whether a second target object corresponding to the first interaction content can be specified in the environment image;
[0139] S03. Input the first interactive content, the first preset task, and the environmental image into the target interactive model to obtain a first output result. If the first output result is yes, then obtain a second preset task. The task type of the second preset task is: based on the first interactive content, point out the second region information of the image region occupied by the second target object corresponding to the first interactive content in the environmental image.
[0140] S04. Input the second preset task, the first interactive content, and the environmental image into the target interactive model to obtain the second region information;
[0141] The second region information refers to the coordinate range of the image region occupied by the second target image in the environment image.
[0142] S05. Execute the corresponding display task or capture task based on the information of the second area.
[0143] In some optional embodiments of this application, the aforementioned S05, performing the crawling task based on the second region information, may specifically include the following S051-S053:
[0144] S051. Input the environmental image and the second region information into a preset segmentation model to obtain the mask corresponding to the environmental image;
[0145] The segmentation model can be implemented using Segment anything.
[0146] In the aforementioned mask, the partial mask information corresponding to the area occupied by the second target object in the environment image is 1, and the partial mask information corresponding to the area occupied by the remaining objects in the environment image other than the second target object is 0.
[0147] S052. Input the depth image information corresponding to the mask and the environment image into the preset grasping model to obtain the target grasping position of the grasping device;
[0148] S053. Control the gripping device to move to the target gripping position and grip the second target object.
[0149] The crawling model can be implemented using ContactGraspNet.
[0150] Optionally, the gripping device can be a robotic arm, and the target gripping position is the location of the second target object in the real world.
[0151] In some optional embodiments of this application, the above method further includes the following S001-S005:
[0152] S001. If the first output result is negative, then obtain the third preset task. The task type of the third preset task is: a task to ask a question based on the first interactive content.
[0153] S002. Input the third preset task, the first interactive content, and the environmental image into the target interactive model to obtain the corresponding question content;
[0154] S003. Ask the interactive object a question based on the question content;
[0155] This scheme allows for the creation of questions when a second target object corresponding to the first interactive content cannot be specified in the environmental image based on the first interactive content. This enables the acquisition of more interactive content with the interactive object and the disambiguation of the interactive content.
[0156] S004. Obtain the second interactive content of the interactive object based on the question content;
[0157] S005. Take the second interactive content and the question content as the new first interactive content, and return to execute the step of inputting the first interactive content, the first preset task, and the environmental image into the target interactive model until the second region information is determined.
[0158] In this application, the target interaction model includes: a target word processing unit, a target visual encoding unit, and a target transformation model, wherein the target word processing unit includes a word segmentation unit and a target embedding layer unit, and the target visual encoding unit includes a target residual neural network and a target linear projection layer.
[0159] Specifically, when the initial interaction model is trained to become the target interaction model, the initial word processing unit is trained to become the target word processing unit; the initial visual encoding unit is trained to become the target visual encoding unit; the initial conversion model is trained to become the target conversion model; the initial embedding layer unit is trained to become the target embedding layer unit; the initial residual neural network is trained to become the target residual neural network; and the initial linear projection layer is trained to become the target linear projection layer.
[0160] In the aforementioned S03, the first interactive content, the first preset task, and the environmental image are input into the target interactive model to obtain a first output result, specifically including:
[0161] The first preset task and the first interaction content are processed by the target word processing unit in the target interaction model to obtain the corresponding word vector sequence.
[0162] The environmental image is encoded by the target visual encoding unit in the target interaction model to obtain the image feature sequence corresponding to the environmental image;
[0163] The word vector sequence and the image feature sequence are analyzed using the target transformation model in the target interaction model to obtain the corresponding first output result.
[0164] In this application, the datasets involved in training the target interaction model are shown in Table 1:
[0165] Table 1. Data sets involved in this application
[0166]
[0167] As shown in Table 1, when the sample task type is a task describing the sample image, such as image captioning, the dataset used is SBU Cpations. When the sample task type is a task of making corresponding questions based on the interaction content with the respondent, such as visual question answering (VQA) or visual question generation, the dataset used is LLaVA. When the sample task type is Dialog, the dataset used is Visdial, where Dialog represents visual question answering, visual question generation (VQG), image captioning, and visual alignment (VG). Visual alignment is one of the tasks that, based on the interaction content with the questioner or respondent, points out the first region information of the image area occupied by the first target object corresponding to the interaction content in the sample image. When the sample task type is VG+Dialog, the dataset used is GuessWhat?!. When the sample task type is VG+Dialog, the dataset used can also be InViG. When the sample task type is VG+Caption, the dataset used is RefCOCO. When the sample task is of type VG+Caption, the dataset used is RefCOCOg. When the sample task is of type VG+Caption, the dataset used can also be RefCOCO+. When the sample task is of type VG, the dataset used can be OpenImages.
[0168] Additionally, "#image" in Table 1 can refer to the data volume of the corresponding sample image, and "#Sample" in Table 1 can refer to the data volume of the corresponding sample task information. For visual alignment, data can be collected from RefCOCO and OpenImages. Data for visual alignment can also be collected from LLaVA instructions and Visual Dialogs, which contain over 350K high-quality dialogues, very useful for learning multi-turn language modeling. This application used a total of 955K sample images.
[0169] This disclosure provides a sample acquisition task; based on the type of the sample task, it acquires corresponding sample interaction information and sample images corresponding to the sample interaction information; the sample interaction information includes: sample interaction content and sample task execution results; the sample task and sample interaction content are processed by the initial word processing unit in the initial interaction model to obtain corresponding word vector sequences; the sample images are encoded by the initial visual encoding unit in the initial interaction model to obtain image feature sequences corresponding to the sample images; the initial transformation model in the initial interaction model is used to analyze the word vector sequences and image feature sequences to obtain the corresponding task. Prediction results; Based on the sample task execution results and the task prediction results, the initial interaction model is trained to obtain a target interaction model. The target interaction model is used to: determine the interaction scheme with the interaction object based on the first interaction content with the interaction object and the environmental image; the sample interaction information used to train the target interaction model can be determined based on the type of sample task, giving the target interaction model the function of recognizing different task types, thereby enabling the trained target recognition model to have the ability to interact with the interaction object according to multiple task types, thereby improving the generalization ability of the target recognition model and improving the accuracy of the intelligent robot in performing tasks according to user instructions.
[0170] The solution proposed in this application enables intelligent robots to understand complex visual relationships, human state behaviors, and complex user expressions more accurately and robustly in more open scenarios. This facilitates natural and accurate interaction between intelligent robots and humans in most indoor and outdoor scenarios, enabling them to accurately understand and complete human language commands.
[0171] Figure 3 A schematic diagram of the structure of a data processing apparatus provided for an exemplary embodiment of this disclosure;
[0172] The device includes:
[0173] Acquisition unit 31 is used for acquiring sample tasks;
[0174] The acquisition unit 31 is further configured to acquire corresponding sample interaction information and sample images corresponding to the sample interaction information based on the type of the sample task; the sample interaction information includes: sample interaction content and sample task execution result;
[0175] Processing unit 32 is used to perform word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence;
[0176] Encoding unit 33 is used to encode the sample image through the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image;
[0177] Analysis unit 34 is used to analyze the word vector sequence and the image feature sequence through the initial transformation model in the initial interaction model to obtain the corresponding task prediction result;
[0178] Training unit 35 trains the initial interaction model based on the sample task execution results and the task prediction results to obtain a target interaction model. The target interaction model is used to interact with the interaction object based on the first interaction content with the interaction object and the environmental image.
[0179] Optionally, when the aforementioned device is used to obtain corresponding sample interaction information based on the type of the sample task, it is specifically used for:
[0180] The corresponding interaction role is determined based on the type of the sample task, wherein the interaction role is either a questioner or a respondent.
[0181] Based on the interaction roles, the sample task execution results are determined from the target sample interaction information set;
[0182] The dialogue content preceding the execution result of the sample task is taken as the sample interaction content in the target sample interaction information set.
[0183] Optionally, the aforementioned device is also used for:
[0184] Obtain multiple sample interaction information sets;
[0185] The target sample interaction information set corresponding to the type of the sample task is determined from the plurality of sample interaction information sets.
[0186] Optionally, the sample task is any of the following:
[0187] The tasks include answering questions posed by the questioner, making corresponding questions based on the interaction content with the respondent, identifying the first region information of the image area occupied by the first target object corresponding to the interaction content in the sample image based on the interaction content with the questioner or respondent, answering whether the first target object corresponding to the interaction content can be specified in the sample image based on the interaction content with the questioner or respondent, and describing the sample image.
[0188] Optionally, the initial word processing unit includes a word segmentation unit and an initial embedding layer unit. The step of processing the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence includes:
[0189] The word segmentation unit in the initial interaction model is used to segment the sample task and the sample interaction content into words to obtain the word segmentation result.
[0190] The word segmentation results are processed by the initial embedding layer unit in the initial interaction model to obtain the corresponding word vector sequence.
[0191] Optionally, the initial visual encoding unit includes: an initial residual neural network and an initial linear projection layer. When the aforementioned device encodes the sample image using the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image, it is specifically used for:
[0192] The sample image is encoded using the initial residual neural network to obtain image feature information;
[0193] The initial linear projection layer is used to perform dimensional transformation on the image feature information to obtain the image feature sequence corresponding to the sample image.
[0194] Optionally, when the aforementioned apparatus is used to train the initial interaction model based on the sample task execution results and the task prediction results to obtain the target interaction model, it is specifically used for:
[0195] Loss information is determined based on a preset loss function, the sample task execution results, and the task prediction results.
[0196] If the loss information is less than a preset threshold, then the initial interaction model is used as the target interaction model.
[0197] If the loss information is not less than a preset threshold, the parameters in the initial interaction model are adjusted, and the process returns to the step of performing word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence, until the target interaction model is determined when the loss information is less than the preset threshold.
[0198] Optionally, the aforementioned device is also used for:
[0199] Obtain the first interactive content and the environmental image;
[0200] Obtain a first preset task, wherein the task type of the first preset task is a task based on the first interaction content to determine whether a second target object corresponding to the first interaction content can be specified in the environment image;
[0201] The first interactive content, the first preset task, and the environmental image are input into the target interactive model to obtain a first output result. If the first output result is yes, then a second preset task is obtained. The task type of the second preset task is: based on the first interactive content, the task of pointing out the second region information of the image region occupied by the second target object corresponding to the first interactive content in the environmental image.
[0202] The second preset task, the first interactive content, and the environmental image are input into the target interactive model to obtain the second region information;
[0203] Based on the information from the second region, perform the corresponding display or capture task.
[0204] Optionally, the aforementioned device is also used for:
[0205] If the first output result is negative, then a third preset task is obtained. The task type of the third preset task is: a task to ask a question based on the first interaction content.
[0206] The third preset task, the first interactive content, and the environmental image are input into the target interactive model to obtain the corresponding question content;
[0207] Ask the interactive object a question based on the question content;
[0208] Obtain the second interactive content of the interactive object based on the question content;
[0209] The second interactive content and the question content are used as the new first interactive content. The process of inputting the first interactive content, the first preset task, and the environmental image into the target interactive model is repeated until the second region information is determined.
[0210] It should be understood that the device embodiments and method embodiments can correspond to each other, and similar descriptions can be referred to the method embodiments. To avoid repetition, they will not be repeated here. Specifically, the device can execute the above method embodiments, and the foregoing and other operations and / or functions of each module in the device correspond to the corresponding processes in the various methods in the above method embodiments, which will not be repeated here for the sake of brevity.
[0211] The apparatus of this disclosure embodiment has been described above from the perspective of functional modules in conjunction with the accompanying drawings. It should be understood that this functional module can be implemented in hardware, in software instructions, or in a combination of hardware and software modules. Specifically, the steps of the method embodiments in this disclosure can be completed by integrated logic circuits in the processor's hardware and / or by software instructions. The steps of the method disclosed in this disclosure embodiment can be directly embodied as being executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. Optionally, the software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps in the above method embodiments.
[0212] Figure 4 This is a schematic block diagram of an electronic device provided in an embodiment of this disclosure. The electronic device may include:
[0213] The system includes a memory 401 for storing computer programs and a processor 402 for transferring program code to the processor 402. In other words, the processor 402 can retrieve and run the computer program from the memory 401 to implement the methods described in this embodiment.
[0214] For example, the processor 402 can be used to execute the above-described method embodiments according to instructions in the computer program.
[0215] In some embodiments of this disclosure, the processor 402 may include, but is not limited to:
[0216] General-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0217] In some embodiments of this disclosure, the memory 401 includes, but is not limited to:
[0218] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static RAM (SRAM), Dynamic RAM (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0219] In some embodiments of this disclosure, the computer program may be divided into one or more modules, which are stored in the memory 401 and executed by the processor 402 to perform the method provided in this disclosure. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the electronic device.
[0220] like Figure 4 As shown, the electronic device may also include:
[0221] Transceiver 403, which can be connected to processor 402 or memory 401.
[0222] The processor 402 can control the transceiver 403 to communicate with other devices; specifically, it can send information or data to other devices or receive information or data sent by other devices. The transceiver 403 may include a transmitter and a receiver. The transceiver 403 may further include antennas, and the number of antennas may be one or more.
[0223] It should be understood that the various components in the electronic device are connected through a bus system, which includes a data bus, a power bus, a control bus, and a status signal bus.
[0224] This disclosure also provides a computer storage medium storing a computer program thereon, which, when executed by a computer, enables the computer to perform the methods of the above-described method embodiments. Alternatively, embodiments of this disclosure also provide a computer program product containing instructions that, when executed by a computer, cause the computer to perform the methods of the above-described method embodiments.
[0225] When implemented using software, it can be implemented wholly or partially as a computer program product. This computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital video disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0226] According to one or more embodiments of this disclosure, a model training method is provided, comprising:
[0227] The task of obtaining samples;
[0228] Based on the type of the sample task, obtain the corresponding sample interaction information and the sample image corresponding to the sample interaction information; the sample interaction information includes: sample interaction content and sample task execution result;
[0229] The initial word processing unit in the initial interaction model performs word processing on the sample task and the sample interaction content to obtain the corresponding word vector sequence.
[0230] The sample image is encoded by the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image;
[0231] The word vector sequence and the image feature sequence are analyzed using the initial transformation model in the initial interaction model to obtain the corresponding task prediction results;
[0232] The initial interaction model is trained based on the sample task execution results and the task prediction results to obtain a target interaction model. The target interaction model is used to interact with the interaction object based on the first interaction content with the interaction object and the environmental image.
[0233] According to one or more embodiments of this disclosure, obtaining the corresponding sample interaction information based on the type of the sample task includes:
[0234] The corresponding interaction role is determined based on the type of the sample task, wherein the interaction role is either a questioner or a respondent.
[0235] Based on the interaction roles, the sample task execution results are determined from the target sample interaction information set;
[0236] The dialogue content preceding the execution result of the sample task is taken as the sample interaction content in the target sample interaction information set.
[0237] According to one or more embodiments of this disclosure, the method further includes:
[0238] Obtain multiple sample interaction information sets;
[0239] The target sample interaction information set corresponding to the type of the sample task is determined from the plurality of sample interaction information sets.
[0240] According to one or more embodiments of this disclosure, the sample task is any of the following:
[0241] The tasks include answering questions posed by the questioner, making corresponding questions based on the interaction content with the respondent, identifying the first region information of the image area occupied by the first target object corresponding to the interaction content in the sample image based on the interaction content with the questioner or respondent, answering whether the first target object corresponding to the interaction content can be specified in the sample image based on the interaction content with the questioner or respondent, and describing the sample image.
[0242] According to one or more embodiments of this disclosure, the initial word processing unit includes a word segmentation unit and an initial embedding layer unit. The step of performing word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence includes:
[0243] The word segmentation unit in the initial interaction model is used to segment the sample task and the sample interaction content into words to obtain the word segmentation result.
[0244] The word segmentation results are processed by the initial embedding layer unit in the initial interaction model to obtain the corresponding word vector sequence.
[0245] According to one or more embodiments of this disclosure, the initial visual encoding unit includes: an initial residual neural network and an initial linear projection layer. The step of encoding the sample image through the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image includes:
[0246] The sample image is encoded using the initial residual neural network to obtain image feature information;
[0247] The initial linear projection layer is used to perform dimensional transformation on the image feature information to obtain the image feature sequence corresponding to the sample image.
[0248] According to one or more embodiments of this disclosure, the initial interaction model is trained based on the sample task execution results and the task prediction results to obtain a target interaction model, including:
[0249] Loss information is determined based on a preset loss function, the sample task execution results, and the task prediction results.
[0250] If the loss information is less than a preset threshold, then the initial interaction model is used as the target interaction model.
[0251] If the loss information is not less than a preset threshold, the parameters in the initial interaction model are adjusted, and the process returns to the step of performing word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence, until the target interaction model is determined when the loss information is less than the preset threshold.
[0252] According to one or more embodiments of this disclosure, the method further includes:
[0253] Obtain the first interactive content and the environmental image;
[0254] Obtain a first preset task, wherein the task type of the first preset task is a task based on the first interaction content to determine whether a second target object corresponding to the first interaction content can be specified in the environment image;
[0255] The first interactive content, the first preset task, and the environmental image are input into the target interactive model to obtain a first output result. If the first output result is yes, then a second preset task is obtained. The task type of the second preset task is: based on the first interactive content, the task of pointing out the second region information of the image region occupied by the second target object corresponding to the first interactive content in the environmental image.
[0256] The second preset task, the first interactive content, and the environmental image are input into the target interactive model to obtain the second region information;
[0257] Based on the information from the second region, perform the corresponding display or capture task.
[0258] According to one or more embodiments of this disclosure, the method further includes:
[0259] If the first output result is negative, then a third preset task is obtained. The task type of the third preset task is: a task to ask a question based on the first interaction content.
[0260] The third preset task, the first interactive content, and the environmental image are input into the target interactive model to obtain the corresponding question content;
[0261] Ask the interactive object a question based on the question content;
[0262] Obtain the second interactive content of the interactive object based on the question content;
[0263] The second interactive content and the question content are used as the new first interactive content. The process of inputting the first interactive content, the first preset task, and the environmental image into the target interactive model is repeated until the second region information is determined.
[0264] According to one or more embodiments of the present disclosure, a data processing apparatus is provided, the apparatus comprising:
[0265] The acquisition unit is used for the task of acquiring samples;
[0266] The acquisition unit is further configured to acquire corresponding sample interaction information and sample images corresponding to the sample interaction information based on the type of the sample task; the sample interaction information includes: sample interaction content and sample task execution result;
[0267] The processing unit is used to perform word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence.
[0268] The encoding unit is used to encode the sample image through the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image;
[0269] The analysis unit is used to analyze the word vector sequence and the image feature sequence through the initial transformation model in the initial interaction model to obtain the corresponding task prediction result;
[0270] The training unit trains the initial interaction model based on the sample task execution results and the task prediction results to obtain a target interaction model. The target interaction model is used to interact with the interaction object based on the first interaction content with the interaction object and the environmental image.
[0271] According to one or more embodiments of this disclosure, when the aforementioned apparatus is used to obtain corresponding sample interaction information based on the type of the sample task, it is specifically used for:
[0272] The corresponding interaction role is determined based on the type of the sample task, wherein the interaction role is either a questioner or a respondent.
[0273] Based on the interaction roles, the sample task execution results are determined from the target sample interaction information set;
[0274] The dialogue content preceding the execution result of the sample task is taken as the sample interaction content in the target sample interaction information set.
[0275] According to one or more embodiments of this disclosure, the aforementioned apparatus is further used for:
[0276] Obtain multiple sample interaction information sets;
[0277] The target sample interaction information set corresponding to the type of the sample task is determined from the plurality of sample interaction information sets.
[0278] According to one or more embodiments of this disclosure, the sample task is any of the following:
[0279] The tasks include answering questions posed by the questioner, making corresponding questions based on the interaction content with the respondent, identifying the first region information of the image area occupied by the first target object corresponding to the interaction content in the sample image based on the interaction content with the questioner or respondent, answering whether the first target object corresponding to the interaction content can be specified in the sample image based on the interaction content with the questioner or respondent, and describing the sample image.
[0280] According to one or more embodiments of this disclosure, the initial word processing unit includes a word segmentation unit and an initial embedding layer unit. The step of performing word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence includes:
[0281] The word segmentation unit in the initial interaction model is used to segment the sample task and the sample interaction content into words to obtain the word segmentation result.
[0282] The word segmentation results are processed by the initial embedding layer unit in the initial interaction model to obtain the corresponding word vector sequence.
[0283] According to one or more embodiments of this disclosure, the initial visual encoding unit includes: an initial residual neural network and an initial linear projection layer. When the aforementioned apparatus is used to encode the sample image through the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image, it is specifically used for:
[0284] The sample image is encoded using the initial residual neural network to obtain image feature information;
[0285] The initial linear projection layer is used to perform dimensional transformation on the image feature information to obtain the image feature sequence corresponding to the sample image.
[0286] According to one or more embodiments of this disclosure, when the aforementioned apparatus is used to train the initial interaction model based on the sample task execution result and the task prediction result to obtain the target interaction model, it is specifically used for:
[0287] Loss information is determined based on a preset loss function, the sample task execution results, and the task prediction results.
[0288] If the loss information is less than a preset threshold, then the initial interaction model is used as the target interaction model.
[0289] If the loss information is not less than a preset threshold, the parameters in the initial interaction model are adjusted, and the process returns to the step of performing word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence, until the target interaction model is determined when the loss information is less than the preset threshold.
[0290] According to one or more embodiments of this disclosure, the aforementioned apparatus is further used for:
[0291] Obtain the first interactive content and the environmental image;
[0292] Obtain a first preset task, wherein the task type of the first preset task is a task based on the first interaction content to determine whether a second target object corresponding to the first interaction content can be specified in the environment image;
[0293] The first interactive content, the first preset task, and the environmental image are input into the target interactive model to obtain a first output result. If the first output result is yes, then a second preset task is obtained. The task type of the second preset task is: based on the first interactive content, the task of pointing out the second region information of the image region occupied by the second target object corresponding to the first interactive content in the environmental image.
[0294] The second preset task, the first interactive content, and the environmental image are input into the target interactive model to obtain the second region information;
[0295] Based on the information from the second region, perform the corresponding display or capture task.
[0296] According to one or more embodiments of this disclosure, the aforementioned apparatus is further used for:
[0297] If the first output result is negative, then a third preset task is obtained. The task type of the third preset task is: a task to ask a question based on the first interaction content.
[0298] The third preset task, the first interactive content, and the environmental image are input into the target interactive model to obtain the corresponding question content;
[0299] Ask the interactive object a question based on the question content;
[0300] Obtain the second interactive content of the interactive object based on the question content;
[0301] The second interactive content and the question content are used as the new first interactive content. The process of inputting the first interactive content, the first preset task, and the environmental image into the target interactive model is repeated until the second region information is determined.
[0302] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0303] In the embodiments provided in this disclosure, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0304] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. For example, the functional modules in the various embodiments of this disclosure may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0305] The above are merely specific embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A model training method, characterized in that, include: The task of obtaining samples; Based on the type of the sample task, obtain the corresponding sample interaction information and the sample image corresponding to the sample interaction information; The sample interaction information includes: sample interaction content and sample task execution results; The initial word processing unit in the initial interaction model performs word processing on the sample task and the sample interaction content to obtain the corresponding word vector sequence. The sample image is encoded by the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image; The word vector sequence and the image feature sequence are analyzed using the initial transformation model in the initial interaction model to obtain the corresponding task prediction results; The initial interaction model is trained based on the sample task execution results and the task prediction results to obtain a target interaction model. The target interaction model is used to interact with the interaction object based on the first interaction content with the interaction object and the environmental image. The step of obtaining the corresponding sample interaction information based on the type of the sample task includes: The corresponding interaction role is determined based on the type of the sample task, wherein the interaction role is either a questioner or a respondent. Based on the interaction roles, the sample task execution results are determined from the target sample interaction information set; The dialogue content preceding the execution result of the sample task is taken as the sample interaction content in the target sample interaction information set.
2. The method according to claim 1, characterized in that, The method further includes: Obtain multiple sample interaction information sets; The target sample interaction information set corresponding to the type of the sample task is determined from the plurality of sample interaction information sets.
3. The method according to claim 1, characterized in that, The sample task is any of the following: The tasks include answering questions posed by the questioner, making corresponding questions based on the interaction content with the respondent, identifying the first region information of the image area occupied by the first target object corresponding to the interaction content in the sample image based on the interaction content with the questioner or respondent, answering whether the first target object corresponding to the interaction content can be specified in the sample image based on the interaction content with the questioner or respondent, and describing the sample image.
4. The method according to claim 1, characterized in that, The initial word processing unit includes a word segmentation unit and an initial embedding layer unit. The initial word processing unit in the initial interaction model performs word processing on the sample task and the sample interaction content to obtain the corresponding word vector sequence, including: The word segmentation unit in the initial interaction model is used to segment the sample task and the sample interaction content into words to obtain the word segmentation result. The word segmentation results are processed by the initial embedding layer unit in the initial interaction model to obtain the corresponding word vector sequence.
5. The method according to claim 1, characterized in that, The initial visual encoding unit includes an initial residual neural network and an initial linear projection layer. Encoding the sample image using the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image includes: The sample image is encoded using the initial residual neural network to obtain image feature information; The initial linear projection layer is used to perform dimensional transformation on the image feature information to obtain the image feature sequence corresponding to the sample image.
6. The method according to claim 1, characterized in that, The initial interaction model is trained based on the sample task execution results and the task prediction results to obtain the target interaction model, including: Loss information is determined based on a preset loss function, the sample task execution results, and the task prediction results. If the loss information is less than a preset threshold, then the initial interaction model is used as the target interaction model. If the loss information is not less than a preset threshold, the parameters in the initial interaction model are adjusted, and the process returns to the step of performing word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence, until the target interaction model is determined when the loss information is less than the preset threshold.
7. The method according to claim 1, characterized in that, The method further includes: Obtain the first interactive content and the environmental image; Obtain a first preset task, wherein the task type of the first preset task is a task based on the first interaction content to determine whether a second target object corresponding to the first interaction content can be specified in the environment image; The first interactive content, the first preset task, and the environmental image are input into the target interactive model to obtain a first output result. If the first output result is yes, then a second preset task is obtained. The task type of the second preset task is: based on the first interactive content, the task of pointing out the second region information of the image region occupied by the second target object corresponding to the first interactive content in the environmental image. The second preset task, the first interactive content, and the environmental image are input into the target interactive model to obtain the second region information; Based on the information from the second region, perform the corresponding display or capture task.
8. The method according to claim 7, characterized in that, The method further includes: If the first output result is negative, then a third preset task is obtained. The task type of the third preset task is: a task to ask a question based on the first interaction content. The third preset task, the first interactive content, and the environmental image are input into the target interactive model to obtain the corresponding question content; Ask the interactive object a question based on the question content; Obtain the second interactive content of the interactive object based on the question content; The second interactive content and the question content are used as the new first interactive content. The process of inputting the first interactive content, the first preset task, and the environmental image into the target interactive model is repeated until the second region information is determined.
9. A model training device, characterized in that, include: The acquisition unit is used for the task of acquiring samples; The acquisition unit is further configured to acquire corresponding sample interaction information and sample images corresponding to the sample interaction information based on the type of the sample task. The sample interaction information includes: sample interaction content and sample task execution results; The processing unit is used to perform word processing on the sample task and the sample interaction content through the initial word processing unit in the initial interaction model to obtain the corresponding word vector sequence. The encoding unit is used to encode the sample image through the initial visual encoding unit in the initial interaction model to obtain the image feature sequence corresponding to the sample image; The analysis unit is used to analyze the word vector sequence and the image feature sequence through the initial transformation model in the initial interaction model to obtain the corresponding task prediction result; The training unit is used to train the initial interaction model based on the sample task execution results and the task prediction results to obtain a target interaction model. The target interaction model is used to interact with the interaction object based on the first interaction content with the interaction object and the environmental image. The step of obtaining the corresponding sample interaction information based on the type of the sample task includes: The corresponding interaction role is determined based on the type of the sample task, wherein the interaction role is either a questioner or a respondent. Based on the interaction roles, the sample task execution results are determined from the target sample interaction information set; The dialogue content preceding the execution result of the sample task is taken as the sample interaction content in the target sample interaction information set.
10. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the method of any one of claims 1-8 by executing the executable instructions.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-8.
Citation Information
Patent Citations
Visual question and answer and visual question and answer model training method and device, equipment and storage medium
CN113392288A
Robust visual question and answer model training method based on comparative learning
CN116662591A